A method for generating quality metrics of optical coherence tomography (OCT) data

By applying machine learning to generate quality maps in the OCT system, the subjectivity and time-consuming nature of OCT scan quality assessment are resolved, enabling rapid and objective quality identification and optimization, and improving the reliability and efficiency of scan quality.

CN116508060BActive Publication Date: 2026-03-10CARL ZEISS MEDITEC INC +1
View PDF 11 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-11-29
Publication Date
2026-03-10

AI Technical Summary

Technical Problem

Existing OCTA scan quality assessment methods are subjective and time-consuming, failing to provide objective quality metrics. This leads to diagnostic uncertainty and data loss in low-quality data, hindering the rapid identification and improvement of scan quality.

Method used

By applying machine learning techniques, especially deep learning and neural networks, to the OCT system, quantitative OCT/OCTA scan quality maps are generated, low-quality areas are identified, and corrective measures are provided, enabling automatic or semi-automatic scan optimization.

Benefits of technology

It provides objective quality metrics for OCT/OCTA scans, quickly identifies low-quality areas and recommends corrective measures, improving the reliability and efficiency of scan quality and reducing errors from subjective judgment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116508060B_ABST
    Figure CN116508060B_ABST
Patent Text Reader

Abstract

A system, method, and / or apparatus for determining a quality metric for OCT structured data and / or OCTA functional data, using a machine learning model trained to provide a single overall quality metric or quality map distribution for the OCT / OCTA data based on the generation of multiple feature maps extracted from one or more slice views of the OCT / OCTA data. The extracted feature maps may be different texture type maps, and the machine model is trained to determine the quality metric based on the texture maps.
Need to check novelty before this filing date? Find Prior Art

Description

Field of the invention

[0001] The present invention relates generally to optical coherence tomography (OCT) systems. More particularly, the present invention relates to methods of determining quality metrics of OCT and OCT angiography scans and generating quality maps. BACKGROUND

[0002] Optical coherence tomography (OCT) is a non-invasive imaging technique that uses light waves to penetrate tissue and produce image information at different depths within the tissue. Essentially, an OCT system is an interferometric imaging system that determines the scattering profile of a sample along an OCT beam by detecting the interference of light reflected from the sample with a reference beam to create a three-dimensional (3D) representation of the sample. Each scattering profile in the depth direction (e.g., z-axis or axial) can be individually reconstructed as an axial scan or A-scan. Cross-sectional slice images (e.g., two-dimensional (2D) bi- section scans or B-scans) and volumetric images (e.g., three-dimensional (3D) cubic scans or C-scans) can be constructed from a plurality of A-scans acquired as the OCT beam is scanned / moved through a set of transverse (e.g., x-axis and y-axis) positions on the sample. OCT systems also allow for the construction of planar en face (e.g., face) images of selected portions of a tissue volume (e.g., a target tissue slice view (sub-volume) or a target tissue layer, such as the retina of an eye).

[0003] Within the field of ophthalmology, OCT systems were initially developed to provide structural data, such as cross-sectional images of retinal tissue, but can now also provide functional information, such as flow information. While OCT structural data allows one to view different tissue layers of the retina, OCT angiography (OCTA) extends the functionality of OCT systems to also identify (e.g., present in image format) the presence or lack of flow in retinal tissue. For example, OCTA can identify flow by identifying differences (e.g., contrast differences) over time in multiple OCT scans of the same retinal region and designating differences that meet predetermined criteria as flow. While data produced by an OCT system (e.g., OCT data) can include both OCT structural data and OCT flow data, depending on the functionality of the OCT system, for ease of discussion, OCT structural data can be referred to herein as “OCT data” and OCT angiography (or flow) data can be referred to herein as “OCTA data,” unless otherwise stated or understood from context. Thus, OCT can be said to provide structural information, while OCTA provides flow (e.g., functional) information. However, since both OCT data and OCTA data can be extracted from the same OCT scan(s), the term “OCT scan” can be understood to include OCT structural scans (e.g., OCT acquisitions) and / or OCT functional scans (e.g., OCTA acquisitions), unless otherwise stated. A more in-depth discussion of OCT and OCTA is provided below.

[0004] OCTA provides valuable diagnostic information not found in structural OCT, but OCTA scans can have acquisition issues that can make their quality less than desirable. Existing attempts to quantify OCT scan quality have focused on OCT structural data and typically rely on signal intensity measures, such as described in D. M. Stein et al., “A new OCT quality assessment parameter,” Br J Ophthalmol. 2006; 90: 186-190. While signal intensity measures for assessing OCT structural data have been found to be useful, the use of such methods in OCTA data is limited because the quality of the derived flow information depends on many other factors not included in such quantification.

[0005] Accordingly, OCTA scan quality is typically determined subjectively by an observer in order to determine whether a particular OCTA acquisition (e.g., OCTA scan) is usable for diagnosis or inclusion in a broad study. Examples of this approach can be found in: “Determinants of quantitative optical coherence tomography angiography metrics in diabetic patients,” Tang FY et al., Sci Rep. 2018; 8: 7314; “Contact lens-related corneal vascularization scan source optical coherence tomography angiography,” Ang M et al., Ophthalmol J. 2016, 2016, 9685297; and “Effect of eye tracking technology on imaging quality of age-related macular degeneration OCT angiography,” Raunam et al., Graefe’s Arch Clin Exp Ophthalmol. 2017, 255: 1535. However, these approaches are highly subjective and time consuming. Furthermore, the subjective quality is assessed after the patient examination during a-posteriori scan review, when the patient is away from the clinic, making it impossible to attempt and acquire additional scans of better quality to replace the low quality data, and leading to data loss or uncertain diagnosis. Even for operators who are able to actively judge the quality of an OCTA scan during acquisition while the patient is still in the clinic, there is currently no guidance with a quantitative quality score that would help establish an objective quality cutoff measure for rescanning or for improving the quality of subsequent acquisitions.

[0006] It is an object of the present invention to provide a system / apparatus / method for providing an objective quality measure of OCT / OCTA data.

[0007] It is another object of the present invention to provide a quick determination of when an OCT / OCTA scan is of insufficient quality and can need to be redone.

[0008] It is another object of the present invention to provide a quality measure of OCTA data on a per-A-scan basis.

[0009] It is a further object of the present invention to provide a (e.g., 2D or 3D) quality map of OCTA data that visually identifies portions of an OCTA scan that can have poor quality, e.g., as determined by the present system / method / apparatus. SUMMARY

[0010] The above objects are met in a method / system / apparatus for identifying low quality OCT scans (or low quality portions within an OCT scan, e.g., OCT structural scans and / or OCTA functional scans), identifying possible sources of low quality, and recommending (or implementing) corrective actions for improving subsequent OCT scans. In addition, a quality map of an OCT scan can also be provided.

[0011] For example, the present system / method / apparatus can provide one or more (e.g., 2D and / or 3D) quantitative quality maps that describe the quality of the OCT / OCTA acquisition, e.g., at each facial location (e.g., pixel or pixel region / window location). The resulting quality maps are well correlated with subjective quality metrics (e.g., provided by human testers) observed in the slices produced from the acquisition at the corresponding facial locations. Optionally, the values in the quality maps can be averaged to provide an overall quality score for the acquisition, which is also found to be well correlated with subjective quality ratings. This is in contrast to previous quality assessment methods that determine a measure of total signal intensity recorded in the OCT structural component compared to a noise baseline. Such previous methods do not provide reliable quality values for the OCTA flow component or location-specific quality values. The appended claims describe the present invention in more detail.

[0012] The system / method / apparatus can also identify and output one or more possible sources / causes for the low quality acquisition or low quality region. For example, the system / method / apparatus can identify the source of the low quality acquisition as incorrect focus, opacity (e.g., cataract or floaters of opaque media), illumination below a predetermined threshold (e.g., possibly caused by small pupil), tracking problems (e.g., due to blinking), and recommend corrective actions such as correcting / adjusting the focus, suggesting replacing the imaging angle to avoid opacity, identifying the need for pupil dilation, and identifying possible causes of loss of eye tracking. This information can be used to provide recommendations to the system operator (or automatic / semi-automatic subsystem within the OCT system) during data acquisition, which can be used to acquire repeat scans to achieve better image quality. For example, the OCT system can use this information to automatically (or semi-automatically, e.g., in response to an OK input signal from the system operator) take the recommended corrective actions to improve subsequent scan acquisition.

[0013] Other objects and merits of the present invention, as well as a fuller understanding thereof, will become apparent and appreciated upon consideration of the following description and

[0014] Several publications can be cited or referred to herein in order to facilitate a fuller understanding of the application. All publications cited or referred to in this specification are herein incorporated by reference in their entirety.

[0015] The embodiments disclosed herein are only examples and the scope of the application is not limited to them. Any embodiment features mentioned in one claim category, e.g. system, device or method, can also be claimed in another claim category, e.g. system, device or method. The dependencies or references in the appended claims are chosen solely for formal reasons. However, any subject matter resulting from the intentional or accidental combination of any previous claims can also be claimed, such that any combination of features from any of the claims can be claimed, irrespective of the dependencies selected in the appended claims. BRIEF DESCRIPTION OF DRAWINGS

[0016] Priority applications US serial numbers 63 / 119377 and 63 / 233033 contain at least one color drawing and are hereby incorporated by reference.

[0017] In the drawings, like reference numerals / characters designate like parts throughout the various figures:

[0018] Figure 1 A set of Haralick features are shown, extracted from a retinal flow slice in a 250 micron circular neighborhood in a pixel-by-pixel sliding window.

[0019] Figure 2 A representation of an exemplary training protocol workflow according to the present application is provided.

[0020] Figure 3 A representation of an exemplary application (or test) phase workflow according to the present application is provided.

[0021] Figure 4 The effect of reducing the overlap ratio, e.g. increasing the computation speed, as well as the application of the current Gaussian filter is illustrated.

[0022] Figure 5 An exemplary result of applying extrapolation to a quality map with standard filtering to fill non-number (NaN) values (e.g. fill NaN values along the perimeter of the quality map) according to the present application is provided, wherein the quality map is obtained from a retinal angiogram slice (e.g. en face image).

[0023] Figures 6A-6E Various exemplary criteria for grading flow image quality using a 1-5 scale are shown.

[0024] Figure 6F A first example of results from a good quality scan (top row, Gd1) and a second example of results from a generally poor quality scan (bottom row, Pr1) are shown.

[0025] Figure 7 An example of the four slices 71-74 for feature extraction (max_flow, avg_struc, max_struc and min_stuc, respectively) and the labeling of the considered neighborhood (white circle) is provided, as well as the resulting target quality map 75.

[0026] Figure 8A A graph of the results of the predicted values in the training data before applying the adjustment (correction) by using a quadratic polynomial is provided.

[0027] Figure 8B The results after applying the quadratic polynomial adjustment / correction are provided.

[0028] Figure 9 The effect of smoothing the average grader map is illustrated when considering the ground truth.

[0029] Figure 10A An analysis of the average score for all pixels of all images in the training set with respect to the predicted values is shown, where the horizontal axis represents the average given score by the expert grader and the vertical axis represents the predicted score by the training model.

[0030] Figure 10B The percentage of failure cases is provided, as the threshold of failure is considered to vary for images in the training set, where failure is defined as a given ratio (e.g., percentage or fraction) of images having a deviation from the ground truth higher than 1 quality point.

[0031] Figure 11 An example result comparing retinal slices 111 is shown, the ground truth quality scores collected from three expert graders 112, the quality grading results of the current algorithm 113, the difference between the ground truth and the algorithm quality map 114 scaled in 0-5 gray levels, and the areas with a deviation from the ground truth and the algorithm quality map greater than 1.0 quality score 115.

[0032] Figure 12A A plot of the average grader score values for all 6,500,000 data points in the test set with respect to the predicted values is shown.

[0033] Figure 12B The percentage of failure cases encountered is shown, considering the threshold of failure, the failure of images in the training set is considered to vary, failure is defined as a given ratio (e.g., percentage or fraction) of images having a computed quality metric deviation from the ground truth higher than 1 quality point.

[0034] Figure 13A Exemplary results of three images considered to be acceptable are shown, given the 20% pass deviation limit.

[0035] Figure 13B Examples of three images considered suboptimal from 26 analyses giving the same 20% bias limit are shown.

[0036] Figure 14 A first example of quality maps generated / obtained (using the present application) from the right eye (left four images) and left eye (right four images) of a first patient for different 3x3 mm OCTA acquisitions (organized in pairs) is shown.

[0037] Figure 15 A second example of quality maps obtained from the right eye (left four images) and left eye (right four images) of a second patient for different 3x3 mm OCTA acquisitions (organized in pairs) is shown.

[0038] Figure 16 An example of quality maps obtained from the right eye of a patient for different field of view (FOV) OCTA acquisitions: 3x3 mm, 6x6 mm, 9x9 mm and 12x12 mm. The OCTA acquisitions (organized in pairs) are shown in a similar manner to Figure 14

[0039] Figure 17 The percentage of acceptable results of the algorithm providing each human grader and the average of the three human graders as ground truth test data is provided.

[0040] Figure 18 The percentage of acceptable results obtained by comparing the annotations of each grader (individually removed from those establishing the ground truth test data in turn) and the results of the algorithm with the average results of the other two remaining graders as ground truth test data is provided.

[0041] Figure 19 A general purpose spectral domain optical coherence tomography system suitable for use in the present application for collecting 3D image data of an eye is shown.

[0042] Figure 20 An exemplary OCT B-scan image of a normal retina of a human eye is shown, and various canonical retinal layers and boundaries are identified exemplarily.

[0043] Figure 21 An example of an en face vasculature image is shown.

[0044] Figure 22 An image of an exemplary vasculature B-scan (OCTA) is shown.

[0045] Figure 23 An example of a multi-layer perceptron (MLP) neural network is shown.

[0046] ​Figure 24 A simplified neural network is shown consisting of an input layer, a hidden layer, and an output layer.

[0047] Figure 25 An example convolutional neural network architecture is shown.

[0048] Figure 26 An example U-Net architecture is shown.

[0049] Figure 27 An example computer system (or computing device or computer) is shown. DETAILED DESCRIPTION

[0050] Optical coherence tomography (OCT) system scans can suffer from acquisition problems that can adversely affect the acquisition / scanning quality. Such problems include, among others: misfocus, presence of floaters of opaque media, low illumination (e.g., signal intensity less than 6 on a scale from 0 to 10), low light penetration (e.g., less than half of the target penetration value, or less than 5 pm), presence of tracking / motion artifacts, and / or high noise (e.g., root mean square noise value above a predetermined threshold). These problems can adversely affect the quality of the OCT system data, e.g., shown in a B-scan or en face view (e.g., in the form of a slice), and can adversely affect the accuracy of data extraction (or image processing) techniques or algorithms applied to the OCT system data, such as segmentation or blood vessel density quantification techniques. Thus, low quality OCT system data, particularly OCTA data, can potentially make correct diagnosis difficult. Therefore, there is a need to assess the quality of acquired OCT / OCTA scans in a quantitative manner in order to quickly determine the feasibility or effectiveness of the scans.

[0051] The quality of OCT structural data is typically determined based on a total signal intensity value. Typically, if the total signal intensity value is below a predetermined threshold, the entire OCT structural scan is considered to be poor, e.g., the scan fails. Thus, the signal intensity based quality metric only provides (e.g., outputs) a single quality value / measurement for the entire volumetric field of view (FOV), but this approach is not a reliable enough approach to adjust when trying to assess the quality of en face images, e.g., whose structural information varies from one scan location to another. This approach is also particularly unsuitable for assessing the quality of OCTA acquisitions, which provide functional, flow information rather than structural information. Thus, it is the Applicants’ understanding that there is no commercially available method to automatically and quantitatively assess the quality of flow information in OCTA scans, either as a unique value per scan or in the form of a quality map.

[0052] A system and method for generating a quality map of OCT system data is provided, the quality metric of which can vary throughout the acquisition (e.g., over the desired FOV). Some portions of the current discussion can describe the present invention as applying to one or the other of OCT structural data or OCTA flow data, but it should be understood that the description of the invention can also apply to the other of OCT structural data or OCTA flow data unless otherwise noted.

[0053] The present invention provides a system and method for quantitatively measuring the relative quality of OCT system data at each of a plurality of image quality locations / positions in the OCT system data (e.g., at each scan location (e.g., each A-scan location) or each quality metric window location (or pixel neighborhood) that can span multiple A-scan locations). While the present quality assessment method can apply to any OCT system data observation / imaging technique (e.g., en face, A-scan, B-scan, and / or C-scan images), for ease of discussion, the present method will be primarily described herein as applying to en face images (unless otherwise noted), while it should be understood that the same (or substantially similar as would be understood by one of skill in the art) method / technique can apply to any other OCT system data observation / imaging technique (e.g., A-scan, B-scan, and / or C-scan images).

[0054] The present system / method can assess a set of texture attributes of the OCT / OCTA data at each image quality location (e.g., each en face location (e.g., pixel or image quality window or pixel neighborhood)) and assign a quantitative quality score related to the scan quality for that location (and / or neighborhood). In the case of en face images, the result is a two-dimensional quality map that describes the scan quality at each en face location, e.g., by using a color code (or grayscale code) that is indicative of the image quality.

[0055] This quality map can be used to judge / determine / calculate the quality of individual scans across its FOV, quantify the quality differences between multiple acquisitions (e.g., OCT system scan acquisitions) of the same subject at each en face position, and / or provide an overall quality measure (e.g., metric) of each acquisition, e.g., by averaging the quality map values. As discussed in more detail below, OCTA flow data can be determined by identifying contrast differences over time in multiple OCT scans (or acquisitions) of the same tissue (e.g., retinal) region. The quality map techniques of the present disclosure can be determined for individual OCT scans used to define an OCTA flow image, and the quality maps of the individual OCT scans can be averaged to define a quality map for the OCTA flow image defined therefrom. Alternatively or additionally, the present quality map techniques can be applied directly to the defined OCTA flow data or image (which can be based on contrast information or other flow-indicative data from multiple OCT scans). Optionally, this directly determined OCTA quality map can also be combined (e.g., weighted average, e.g., equally weighted or more heavily weighted toward the directly determined OCTA quality map) with the quality maps of the individual OCT scans that define the OCTA flow data / image.

[0056] Regardless, the defined quality map (or overall quality measure of the acquisition) can provide important information to the OCT system operator to determine when the quality of an acquisition is low and needs to be rescanned (e.g., OCT scan or OCTA scan). The present system can further identify one or more possible causes of the low quality and output (e.g., to the system operator or to an automatic / semi-automatic subsystem of the OCT system) a recommendation aimed at obtaining better quality in a subsequent scan. For example, if the quality map indicates that at least a predefined target retinal region (e.g., a predefined region of interest, ROI) in the acquisition is below a predefined threshold quality measure, or if the overall measure of the acquisition is below a predefined threshold overall quality measure, the quality map (or overall measure) can be used to determine in an automated system that another acquisition is needed. The automated system can then initiate another acquisition, either automatically or in response to an approval input signal from the system operator. The present system can also identify one or more corrective measures (actions) for improving the quality of the acquisition, and automatically make one or more of the identified corrective measures prior to initiating another acquisition. Alternatively or additionally, the quality maps (e.g., OCTA quality maps) of multiple acquisitions of the same retinal region can be compared to one another, and the best quality (or higher quality) portions / regions of the multiple acquisitions (as determined from their respective quality maps, e.g., pixel-by-pixel or window-by-window) can be combined to define a composite acquisition having a higher overall quality than each of the individual acquisitions (OCT and / or OCTA acquisitions).

[0057] Certain embodiments of the present application apply to en face level OCTA acquisition. This embodiment generates 2D quantitative maps that describe the OCTA acquisition (scanning) quality at each en face location. This technique first extracts a set of features related to image texture and other characteristics from a neighborhood of pixels in a slice visualization (e.g., en face image) obtained from an OCTA volume. Features are extracted for different pixel neighborhoods and assigned to the neighborhood in a sliding window fashion. For example, the window can be any shape (e.g., rectangular, circular, etc.) and contain a predetermined number of pixels (e.g., 3x3 pixel window). At each window location, information from multiple (e.g., all) pixels within the window can be used to determine features for a target pixel (e.g., center pixel) within the window. Once the features for the target (e.g., center) pixel are determined, the window can be moved one (or more) pixel location and new features determined for another pixel (e.g., new center pixel) in the new window location. The result is a set of two-dimensional feature maps, each describing a different image feature at each en face location. These features can be hand-crafted (e.g.: intensity, energy, entropy) or learned using a deep learning approach (or other machine learning or artificial intelligence techniques) as a result of training. Examples of machine learning techniques can include artificial neural networks, decision trees, support vector machines, regression analysis, Bayesian networks, etc. Generally, machine learning techniques include one or more training phases followed by one or more testing or application phases. A more detailed discussion of, for example, neural networks that can be used in the present application is provided below.

[0058] In the training phase of one or more machine learning methods of the present application, the set of two-dimensional feature maps obtained from a set of training OCTA acquisitions are combined in a machine learning or deep learning approach to produce a model with outputs corresponding to quality scores previously manually provided by a (human) expert grader for the same set of acquisitions. In addition, the model can also be trained to indicate common acquisition problems that can be derived from the images, such as misfocus, low illumination or light penetration, tracking / motion artifacts, etc., that were previously annotated.

[0059] In the testing or application phase, the model learned from the training phase is applied to a set of 2D feature maps obtained from data that has never been seen (e.g., data not used in the training phase) to produce a 2D quality map as output. Individual quality metrics (or a combined quality metric of one or more sub-regions (e.g., partial areas / parts) of the 2D quality map, such as by averaging the individual quality metrics within the respective sub-regions) can be compared to a predetermined minimum quality threshold to identify regions in the scan that are below a desired quality threshold. Alternatively or additionally, the values in the 2D quality map can also be averaged across the map to produce an overall quality score. In addition, if the model is trained to indicate possible acquisition issues in the image, the feature maps can also be used to provide this information in unseen test images. Different parts of this approach are discussed in more detail below.

[0060] Extracting feature maps

[0061] A single front-facing image (or slice) or N multiple front-facing images are produced from the OCTA cube. Each of these front-facing images is analyzed to produce a set of M feature maps. As discussed above, these feature maps can be designed from known hand-crafted image properties (e.g., gradients, entropy, or texture) in a given neighborhood (e.g., window or pixel neighborhood) of each front-facing location, or be the result of an intermediate layer in a deep learning (or other machine learning) scheme. The result is a set of N x M feature maps for each OCTA acquisition.

[0062] For each hand-crafted image property (or abstract property from a deep learning scheme) and the generated slice, a sliding window approach method is considered to generate a single map from the set with the same dimensions as the slice, where a neighborhood of each pixel is considered to generate a unique property value (e.g., a texture value, such as one or more Haralick features). This property value is assigned to the pixel neighborhood in the map. As the sliding window is moved, centered on different pixel locations in the slice, the values computed for each neighborhood are averaged to produce a resulting value for each neighborhood. The neighborhood can be defined differently depending on the application, such as a rectangular or circular neighborhood. In a similar manner, the extent of the neighborhood and overlap in the sliding window approach can be defined differently depending on the application. For example, Figure 1A set of 22 Haralick features H1 to H22 is shown, extracted from a retinal flow slice 13 using a pixel-by-pixel sliding window 11 within a 250-micron circular neighborhood, as indicated by arrow 15. Haralick features, or Haralick texture features, are a well-known mathematical method for extracting texture features from matrices or image regions. A more detailed discussion of Haralick features can be found in Robert M. Haralick's "Statistical and Structural Methods of Texture," IEEE Journal, Vol. 67, No. 5, 1979, pp. 786–804, which is incorporated herein by reference in its entirety.

[0063] Training phase

[0064] Figure 2 A representation of an exemplary training scheme workflow according to the present invention is provided. This example illustrates K training OCTA acquisition (scanning) samples A1 to AK. For each acquisition, N slices (e.g., frontal images) 23 can be generated, and M features can be defined for each of the N slices 23. Figure 25 (For example, such as) Figure 1 (The Haralick feature shown). This example can be trained using a holistic scoring method, where OCTA acquisitions are assigned a single holistic (e.g., quality) score, and / or using a region-based scoring method, where OCTA acquisitions are divided into multiple regions (e.g., P distinct regions) and each region is assigned a corresponding (e.g., quality) score. How the training vector is defined can depend on whether a holistic scoring method or a region-based scoring method is used. In any case, for the purposes of this discussion, the training vector used for each OCTA acquisition is referred to herein as a "case". Therefore, Figure 2 K cases are shown (e.g., cases 1 to K), with one case for each OCTA acquisition A1 to AK. Optionally, additional labels 27 (e.g., overall quality, segment quality, image or physiological or other recognition features, etc.) can be combined with the feature maps to define the corresponding cases, as shown by arrow 29. Cases 1 to K can then be submitted to the model training module 31, which outputs the generated model 33.

[0065] In summary, one can train with overall quality scores and / or information scores given to the overall OCTA scan and / or region-based scores given to specific regions of the OCTA scan. For example, if overall scores are provided (e.g., per OCTA acquisition), the mean value (or any other aggregation function) of each feature map can be used to provide a corresponding single value. In this case, from a single OCTA acquisition (e.g., Al), a single N x M feature vector can be generated for training (e.g., training input), and the overall value provided as the training result (e.g., training target output). Alternatively, if region-based scores are provided for each acquisition, the mean value (or any other aggregation function) of each feature map (e.g., per region) can be used for training, resulting in multiple training instances.

[0066] In this case, if the OCTA acquisition / scans are graded in P different regions, this would illustrate training P number of N x M feature vectors [f 1 to f P ], and P corresponding values as the training result (e.g., training target output). This approach is flexible for different labeling of the training phase, even for the case of training with overall scores of the overall of each acquisition. This accelerates the collection of training data, as images can be graded with scores to the images.

[0067] Adjusting predicted scores with higher order polynomials

[0068] Depending on the model and data used to train the algorithm used according to the present application, additional adjustments can be made to the resulting quality scores. For example, using a linear model to describe the quality based on a combination of features and fitted weights (as in linear regression) can not appropriately follow subjective way scores, and can require adjustment. That is, there is no guarantee that the quantitative difference between scores of 1 and 2 is the same as the difference between scores of 2 and 3. Additionally, using training data where some specific scores are more represented than others can result in an unbalanced model, which can produce better results adjusted to a given score, while having large errors for other scores. One way to mitigate this behavior is by adding an additional adjustment or fit of the model predicted scores to the given target scores using a higher order polynomial than the one initially considered / used to train the model. For example, when a linear model is used to train the algorithm, a quadratic polynomial can be considered to adjust the predicted scores to better represent the target data. An example of such adjustment (e.g., applied to a linear model) is provided below.

[0069] Application phase

[0070] Figure 3 A representation of an exemplary application (or test) phase workflow according to the present application is provided. Similar to the training phase, the application phase can be performed on a per acquisition basis, or on an overall basis (e.g., per OCTA scan). In the case of an overall basis, the application phase can be performed on the overall scores of the overall of each acquisition, or on the region-based scores of each acquisition. Figure 2the application phase, the corresponding slices 41 and feature maps 43 are extracted from a sample OCTA acquisition 45 that was not seen / used in the training phase previously. As shown in block 47, the resulting values for each pixel neighborhood are averaged, taking into account the values computed for each neighborhood centered at different pixel locations of the feature map as the sliding window moves. The result is a map of each en face location of the OCTA acquisition 45, such as a quality map 49 indicating the quality of angiography (or any other property or information that the model was designed to characterize). The values in the resulting map 49 can also be aggregated (e.g., by averaging) to provide an overall value (e.g., a quality score) for the acquisition 45. Figure 2 The resulting model 33 obtained as shown in block 33 can be applied in a sliding window fashion to a given pixel neighborhood of the feature map 43, resulting in a value for each pixel neighborhood. As the sliding window moves, the values computed for each neighborhood are averaged, taking into account the values computed for each neighborhood centered at different pixel locations of the feature map. The result is a map of each en face location of the OCTA acquisition 45, such as a quality map 49 indicating the quality of angiography (or any other property or information that the model was designed to characterize). The values in the resulting map 49 can also be aggregated (e.g., by averaging) to provide an overall value (e.g., a quality score) for the acquisition 45.

[0071] Post-processing

[0072] When using texture neighborhoods to predict the quality of individual pixels in the resulting WxH (width x height) dimensional quality map, a total of WxH neighborhoods need to be evaluated to produce the complete map. This process can be applied in a sliding window fashion, but it can take longer than optimal / expected to slide the window pixel by pixel because the feature extraction process and model prediction can be computationally expensive. To speed up the process, the sliding window can be defined to have a given overlap in the considered neighborhoods, and the average value in the overlapping neighborhoods. While a lower overlap between neighborhoods is desirable for faster computation, this can result in a pixelated image with abrupt transitions when the overlap is too low. To correct for this, a Gaussian filter can be used / applied to the resulting quality map. Since larger neighborhoods with a lower overlap ratio can require more aggressive filtering, this Gaussian filter can be adapted to the defined neighborhood size and overlap ratio between neighborhoods to produce a visually pleasing result with minimal image disruption. Exemplary filter parameters can be defined as follows:

[0073] [σ x ,σ y ] = [2 · (RadFeatPix x / 3) · (1 - overlapR x ), 2 · (RadFeatPix y / 3) · (1 - overlapR y )],

[0074] [Filter_radius x ,Filter_radius y ] = [(2 · RadFeatPix x)+1, (2 · RadFeatPix y )+1,

[0075] where σ x and σ y are the sigma parameter of the Gaussian filter; Filter_radius x and Filter_radius y are the filter radius (extent) of the filter function; RadFeatPix x and RadFeatPix y are the neighborhood (or window) size in pixels of the neighboring regions in horizontal and vertical direction, respectively; and overlapR x and overlapR y are the overlap ratio defined in horizontal and vertical direction, respectively.

[0076] Figure 4 The effect of reducing the overlap ratio, e.g. increasing the computation speed, and the application of the current Gaussian filter is illustrated. Block B1 provides the color and / or grayscale quality map results for a given input image before filtering on the overlap ratio instances from 95% overlap to 50% overlap. Block B2 provides the color and / or grayscale of the same quality maps after applying the Gaussian filter. It can be seen that reducing the overlap ratio creates more pixelated and abrupt behavior between neighboring regions in the image before filtering, but this effect can be reduced by applying the appropriate filter while still producing results substantially similar to the higher overlap ratio employed. For example, after filtering, there is little difference between the quality maps with 50% overlap and those with 85% or 95% overlap.

[0077] Figure 4 It is further shown that one of the limitations of the sliding window approach is that not all image border positions will be evaluated when using a circular neighborhood, as not all circular neighborhoods will be able to assign a neighborhood with a given radius and overlap ratio. This problem can be more pronounced when applying a Gaussian filter, as areas without information (e.g. “not-a-number” or “NaN”) cannot be included in the filtering process. This effect can be eliminated by applying a Gaussian filter such that it ignores the NaN pixels with appropriate weighting, and / or extrapolate to produce values in such NaN pixels. This can be achieved by the following steps:

[0078] 1) Define a map of the same size as the quality map with all pixel values being 1: Img1

[0079] 2) Replace all NaN positions in the unfiltered quality map and in Img1 with the value 0

[0080] 3) Apply a Gaussian filter to Img1

[0081] 4) Apply a Gaussian filter to the unfiltered quality map

[0082] 5) Divide the result of step 4 by the result of step 3

[0083] For comparison, Figure 5 Exemplary results 51 (e.g., color or grayscale) of applying extrapolation of the above NaN values to a quality map having a standard filter 53 (e.g., color or grayscale) obtained from a retinal angiogram slice (en face image) 55 are provided. The result of this operation is a filtered map 51 (e.g., color or grayscale) having values extrapolated correctly at the NaN locations.

[0084] Training protocol

[0085] Exemplary training and testing of a model to characterize subjective angiogram quality in retinal flow OCTA slices is provided herein. The model is trained on retinal flow slices collected from multiple acquisitions (e.g., 72 or 250 OCTA scans, each scan being a 6x6 mm OCTA scan).

[0086] Collected annotations

[0087] A set of (e.g., human) graders independently grade each acquisition at each pixel location in each en face slice, for example, using a numerical grading scale. For example, each grader can delineate different regions within each slice according to the respective quality of each region using a grading scale. The numerical grading scale can consist of quality values ranging from 1 to 5, where 1 can indicate the worst quality (unusable data) and 5 can indicate the best quality (best). Figures 6A-6E Various exemplary criteria (or examples thereof) for grading the quality of the flow maps using a 1-5 scale are shown. The manual grading output can also be used to define one or more weights of the Haralick coefficients in the flow quality algorithm.

[0088] Figure 6Ais a first example of individual quality grading using a 1-5 scale to assign each quality grade to different regions of the positive flow map 61. For ease of illustration, each region is identified by its grade value and corresponding color coding or line pattern or intensity border. For example, a blue or black solid line border corresponds to grade 1, a red or gray solid line border corresponds to grade 2, a green or long dashed line border corresponds to grade 3, a light purple or short dashed line border corresponds to grade 4, and a yellow or dashed line border corresponds to grade 5. In this example, grade 5 (yellow / dashed line perimeter region) identifies the highest quality image regions with excellent brightness and contrast, as determined by a (e.g., human) expert grader. The grade 5 regions show excellent brightness and / or contrast, capillaries are well delineated, and the grader can follow the capillaries. In the grader’s estimation, the grade 5 regions are examples of the best achievable image quality. Grade 4 (light purple / short dashed line region) identifies image regions with reduced brightness and / or contrast (compared to the grade 5 regions), but the grader can still follow the capillaries well. Grade 3 (green / long dashed line region) is of lower quality than grade 4, but the grader can still infer (or guess) the presence of capillaries with some discontinuities. Thus, within the grade 3 regions, the grader can miss some capillaries. Grade 2 regions (red / solid gray region) have lower quality than grade 3 and define regions where the grader can see some signal between large blood vessels, but cannot resolve capillaries. Within the grade 2 regions, some areas between capillaries can appear falsely ischemic. In grade 1 regions (blue / solid black region), the contents are washed out and the grader considers these regions unusable.

[0089] Figure 6B A second example is provided having elements similar to Figure 6A , which have like reference numerals and are defined above. Figure 6B The impact of artifacts, such as due to floaters, on image quality is shown. In this example, the floaters are identified by noting that the artifacts are not present in other scans of the same region.

[0090] To provide more accurate annotations that can be used to train algorithms, specific regions of the retinal slice can need to have specific annotation instructions. For example, since the foveal region is typically avascular, it is more difficult to judge the visibility of capillaries. Below are examples of dividing the positive flow map into three regions of interest, each considering differences in defining grading criteria to account for vascularization features. Figure 6C The use of the central avascular zone (FAZ) for grading is shown, Figure 6D The use of the nasal segment for grading is shown, and Figure 6E The use of the rest of the image for grading is shown.

[0091] In Figure 6C Grade 5 again represents the best image quality with excellent brightness and / or contrast. Within these regions, the grader can follow the contours of the FAZ without difficulty or hesitation. Grade 4 regions have reduced brightness and / or contrast, but the grader can still track the FAZ. In Grade 3 regions, the capillaries are in focus, and the grader can follow the contours of the FAZ with some noticeable discontinuities. Within this region, capillaries can be missed. In Grade 2 regions, poor signal blocks some portions of the FAZ, and some capillaries appear out of focus. As previously mentioned, Grade 1 regions are not usable, and at least some portions of the FAZ content are washed out.

[0092] In Figure 6D the example of the grading nasal segment, Grade 5 represents the best image quality with excellent brightness and / or contrast. In Grade 5 regions, the retinal nerve fiber layer (RNFL) is well delineated, allowing the grader to follow the RNFL without difficulty. Grade 4 regions have reduced contrast, but the RNFL pattern can still be identified. In Grade 3 regions, the grader can infer (guess) the presence of the RNFL, despite some discontinuities. Within this region, the grader can deduce that the RNFL capillaries are present, but it can be difficult to track individual capillaries. Grade 2 represents regions where the grader can see some signal between large vessels but the RNFL cannot be resolved. Within Grade 2 regions, some areas can appear pseudohypoxic. Grade 1 regions are considered unusable. Within Grade 1 regions, medium vessels can appear blurry, and the content generally appears washed out.

[0093] In Figure 6E Grade 5 represents regions of best quality with excellent brightness and / or contrast, with very well delineated capillaries that can be easily followed. Grade 4 is a region of reduced contrast, where capillaries can still be tracked, and have a higher visible density compared to Grade 3, which is a region where the presence of capillaries can be guessed but not followed. Nonetheless, in Grade 3, fine capillary branches are well in focus (arteriolar and venular branches are in focus), but details of capillaries are lost. Grade 2 identifies regions where some signal between large vessels can still be seen but capillaries are not resolved. Within Grade 2 regions, fine capillaries can appear slightly out of focus, and some regions between capillaries can appear pseudohypoxic (e.g., due to the presence of floaters). Grade 1 represents unusable regions, where generally large vessels appear blurry and the content is washed out.

[0094] Using the above grading example, grader annotations can be collected using, for example, free drawing tools such as ImageJ plugins known in the art. Graders can draw and mark regions of the slice according to the quality of the region using any shape with the goal of covering the entire slice field of view with annotated regions. Annotations collected from different graders can be averaged to generate an average manual quality map that can be used as the target outcome when training the quality map model.

[0095] Figure 6F A first example of results from a scan of overall good quality (top row, Gd1) and a second example of results from a scan of overall poor quality (bottom row, Pr1) are shown. The leftmost image (e.g., the image of column C1) is the individual 6x6 angiographic retinal slice; the center image (e.g., the image of column C2) is the quality map value (in grayscale) with a scale from 1 to 5; and the rightmost image (e.g., the image of column C3) shows the quality map overlaid on the retinal slice (e.g., C1) with a color-coded scale or a grayscale-coded scale from 1 to 5, as shown. The present exemplary method evaluates a set of texture attributes of the OCTA data in the neighborhood of each frontal location (e.g., pixel or window / region) and assigns a quantitative score related to the quality of the scan at that location. The result is a two-dimensional quality map that describes the quality of the scan at each frontal location. This map can be used to judge the quality of an individual scan across its FOV, quantify the quality differences between multiple acquisitions of the same subject at each frontal location, or provide an overall quality measure of each acquisition by averaging the map values. In addition, these quality maps can also provide important information to the OCT / OCTA operator about the need to re-do scans due to poor quality and / or suggestions aimed at obtaining better quality scans in subsequent acquisitions. Other possible applications of this algorithm include reliability measures of other algorithms applied to the OCTA (or OCT) data, automatic exclusion of low quality regions of the OCTA scan from quantification by different algorithms, and use as relative weights for mosaicking, as well as averaging or stitching of OCTA slices or OCTA volumes from multiple acquisitions in overlapping frontal locations according to their quality.

[0096] Consideration of extracted features

[0097] For each OCTA scan considered for training, the retinal slice definition can be used to generate four different slice images: a front flow slice (max_flow) generated by averaging the five maximum value pixels at each A-scan position; a front structure slice (avg_struc) generated by averaging the values at each A-scan position; a front structure slice (max_struc) generated by averaging the five maximum value pixels at each A-scan position; and a front structure slice (min_struc) generated by averaging the five minimum value pixels at each A-scan position. Optionally, no further processing or resizing is considered when generating these front projections. For each of the four slice images, a circular neighborhood with a 250-micrometer radius circular sliding window with a given offset (e.g. a pixel-wise offset or a 75% (i.e. 0.75) overlap) can be considered to extract a set of 22 Haralick features indicative of texture properties. For example, if 72 images are staged, this accounts for the extraction of 88 features from 133128 different neighborhoods used in the training process. Figure 7 An example of the labeling of the four slices 71, 72, 73 and 74 for feature extraction (max_flow, avg_struc, max_struc and min_stuc, respectively) and the considered neighborhoods (white circles) is provided, where eighty-eight features (22 per image) are extracted. The resulting target quality map 75 (e.g. in color or grayscale) is also shown.

[0098] It can be noted that the Haralick features extracted from the images can be highly dependent on the characteristics of the specific instrument and software version, such as baseline signal level, internal normalization or possible internal data filtering. For example, in the proof-of-concept implementation of the present invention, the scans used for training underwent an internal process of 3D Gaussian filtering of the collected flow volumes. In order to apply the algorithm in subsequent software versions that do not support this internal Gaussian filtering, it is necessary to apply the same type of filtering beforehand. That is, the scans in the form of flow volumes obtained with an instrument that does not include the internal Gaussian filtering are pre-filtered by the algorithm prior to feature extraction.

[0099] Training the model

[0100] In one example, Lasso regression (a generalized linear regression model) is used to train the computed mean values of the feature maps to predict the overall quality score of the acquisition. For example, Lasso regression is used to train the set of 88 features extracted from 133128 neighborhoods to predict the manual quality grading given to the center pixel of the corresponding neighborhood in the target quality map, as Figure 7The Lasso model is trained to select the regularization coefficient (e.g., lambda) that will produce the least mean squared error in prediction. A more detailed discussion of Lasso regression can be found in "Regression Shrinkage and Selection via the Lasso," Journal of the Royal Statistical Society, Series B (Methodological), Wiley, 1996, 58(1): pp. 267-88, by Robert Tibshirani, which is hereby incorporated by reference in its entirety. The resulting model can be applied to different acquisitions in a sliding window fashion (defined for feature extraction) to produce angiographic quality maps.

[0101] As discussed above in the "Adjusting the prediction scores by a higher order polynomial" section, since a linear model is used in this example, and different amounts of data are considered for all 1-5 ratings, an additional 2ndorder polynomial is used to adjust the results of the training. For comparison purposes, Figure 8A A plot of the results of the predicted values in the training data before this adjustment (correction) is provided, and Figure 8B The results after applying the 2ndorder polynomial (after adjustment / correction) are provided. In both plots, the horizontal axis represents a given score in the target quality map of each training neighborhood (e.g., the aggregate assigned manually by an expert), while the vertical axis represents the score predicted by the training model. Also in both plots, the light gray shaded area represents the standard deviation, while the dark gray shaded area represents the 95% limits of the predicted score for a given target score. From these two plots, it can be observed how the average prediction for each given score is closer to the given score after applying the current adjustment, and how the confidence interval of the prediction is also more stable over different given scores.

[0102] Results

[0103] To evaluate the accuracy of the algorithm, the quality maps produced by the automated algorithm are compared to the ground truth quality maps. As explained above, the ground truth quality maps are constructed from the average manual region ratings on different raters (see "Collected annotations" section). Since the average rater map can present sharp transitions from the region annotation process, and the automated quality map presents a smooth behavior from the moving window analysis of different regions and the post-processing smoothing, the average rater map is smoothed to provide a fair comparison. This smoothing is considered as an expected behavior of the automated algorithm, which uses a smoothing filter with a kernel equal to the region extent used in the moving window process of the automated algorithm (in this case, a circular neighborhood of 250 microns radius).

[0104] Figure 9The effect of smoothing the average grader map when considered as ground truth is illustrated. For illustration purposes, the max_flow map of the OCTA retinal slice 91, the average of the manual grader scores 92 of the max_flow map, the smoothed average grader map 93 (considered as ground truth), and the results (e.g., color or grayscale) of the present algorithm 94 are shown.

[0105] The evaluation is performed in two steps: (1) first analyzing the behavior on the same data used to train the algorithm to understand the best expected behavior; and (2) then analyzing the behavior in a separate test set of images.

[0106] Expected results - analysis of the training data

[0107] While the traditional accuracy in the same data used for training typically does not represent the expected results in independent test data, in the present example, a linear model is used to fit a very large number of instances with a much smaller number of extracted features as predictors, so overfitting is extremely unlikely, and the results are a good illustration of the expectations in independent test data. The results obtained in the training data are analyzed to understand the best expected behavior of the algorithm and to set the limits that can be considered as good results and non-optimal results.

[0108] Back to Figure 9 which shows example results obtained in one of the cases used to train the algorithm, it can be observed how and to what extent the automatic results 94 from the algorithm resemble those collected in the smoothed average grader scores 93 (used as ground truth). After reviewing the results in the training cases, all results are “acceptable” from a subjective point of view.

[0109] To understand how close the results from the algorithm are to the ground truth in the training data, the values predicted by the algorithm in all pixels from all cases (14,948,928 data points in total) from the training set are compared to the values in the ground truth. Figure 10A An analysis of the average grader score values versus the predicted values for all pixels of all images in the training set is shown, where the horizontal axis represents the average given score values by the expert grader and the vertical axis represents the predicted score values by the training model. The light gray shaded area and the dark gray shaded area represent the standard deviation and the 95% limits of the predicted scores for a given target score, respectively. It can be observed how the average predicted scores resemble the average grader scores. Outliers in the predicted scores will cause larger tails, but the standard deviation and 95% limits of the differences fall mostly below 0.5 and 1.0 points of the specified quality at different quality levels.

[0110] The results on the training data are used to determine what can constitute optimal and suboptimal results, ultimately helping to determine pass and fail rates when establishing algorithmic requirements. To do so, as the threshold for what is considered a failure is varied, one can determine what percentage of cases are failures, as shown in Figure 10B As shown, a failure is defined as a given rate (or percentage) of images deviating from the ground truth by more than 1 quality point. That is, Figure 10B The percentage of cases that are failures is provided, as for the images in the training set, the threshold for what is considered a failure is varied, where a failure is defined as a given rate (e.g., percentage or fraction) of images deviating from the ground truth by more than 1 quality point. As a smaller fraction (e.g., percentage or fraction) of images with a deviation greater than 1 is considered a failure, one would see a larger percentage of cases that are failures, as this would be a more restrictive requirement. It is observed that when the rate threshold is set to 0.2, the algorithm does not produce failures on the training data. That is, by setting the requirement for acceptable results to no more than 20% of images can have a deviation from the ground truth of more than 1 quality point, all results would be acceptable on the training set images. This analysis is used to further evaluate the results on independent test data.

[0111] 5.2 Results on Independent Test Data

[0112] As part of the proof of concept implementation, 26 6x6 mm OCTA scans of a different eye than used in the training set were analyzed as independent test data. As indicated in the “Considering Extracted Features” section above, the fluid volume of scans acquired using an instrument version that did not include internal Gaussian filtering was pre-filtered by the algorithm prior to feature extraction. Following the same approach as the training data, the retinal OCTA flow slices of each scan were manually labeled by three different expert graders independently at each pixel location in each en face slice, as discussed in the “Collected Annotations” section above.

[0113] Figure 11exemplary results (e.g., in color or grayscale) of a comparison of the retinal slice (“retinal OCTA front,” 111), the ground truth quality scores collected from the three expert graders (“smoothed average grader plot,” 112), the quality grading results of the current algorithm (“algorithm quality plot,” 113), the difference between the ground truth and the algorithm quality plot with a 0-5 gray scale (“difference between grader and algorithm,” 114), and the areas with a deviation greater than 1.0 quality score from the ground truth and the algorithm quality plot (“areas with difference > 1,” 115). It can be observed that, for this particular example, the results of the algorithm are, on average, similar to those collected from the expert graders and also similar to the quality of the retinal slice, with an area of significantly lower quality on the left side of the image. It can also be seen how, by analyzing the training data, with the help of the algorithm criteria set earlier, it was considered that this case was acceptable: for example, the degree of deviation between the algorithm and the ground truth was less than 20% of the image, greater than 1.0.

[0114] Similarly, the values predicted by the algorithm in all the pixels of all the cases from the test set (6,500,000 data points in total) were compared with the values in the ground truth. Figure 12A is a plot of the average grader score values versus the predicted values for all 6,500,00 data points in the test set. This plot represents that, on average, the predicted scores are similar to the scores given by the average grader. The observed tail is smaller than in the training group, but the standard deviation of the difference (e.g., the light gray area on the plot) and the 95% limit (dark gray area on the plot) seem slightly larger than in the training set. This can be because the data used to train the algorithm provides a better fit and the use of only 3 graders in the test set ground truth can account for a lower reliability. Figure 12B shows the percentage of failed cases encountered, since the threshold of images from the training set that are considered to be failed is varied, a failure being defined as a given rate (e.g., percentage or fraction) of images that have a calculated quality metric that deviates more than 1 quality point from the ground truth. Using a requirement of acceptable results of no more than 20% of the images having a quality metric that deviates more than 1 from the ground truth (determined from the training data), 23 of the 26 images are acceptable, which constitutes 88% of the images having acceptable results.

[0115] Figure 13A shows (e.g., in color or grayscale) an example of the results of the three images considered to be acceptable given this 20% pass deviation limit, and Figure 13B shows (e.g., in color or grayscale) an example of the three images considered to be suboptimal of the 26 analyzed given the same 20% deviation limit. In Figure 13A and Figure 13BFrom left to right, the first column of images is an example of a retinal slice (“retinal OCTA front,” 131); the second column of images is the corresponding ground truth map (“smoothed average grader map,” 132); the third column is the corresponding result of the current algorithm (“algorithm quality map,” 133; the fourth column represents the difference between the corresponding ground truth and the algorithm map result (“grader vs. algorithm difference,” 134, with a scale of 0-5 gray levels; the fifth column represents the areas where the deviation is greater than 1 (“areas where difference > 1,” 135). Ref. Figure 13B It can be observed that, although the areas where the deviation is greater than 1 in the failed cases are greater than 20% of their images, the results produced by the quality map algorithm (e.g., the quality map of column 133) are somewhat similar to the quality of their corresponding retinal slice (e.g., the retinal OCTA front image of column 131).

[0116] Figure 14 A first example of quality maps generated / obtained (e.g., shown in color or gray scale) for different 3x3mm OCTA acquisitions (organized in pairs) from the right eye (four images on the left) and the left eye (four images on the right) of a first patient (using the present application) is shown. Each pair of images shows on the right the acquired retinal flow slice, with the assigned visual score given by the reader on top, and on the left the overlaid angiogram quality map, with the overall average score computed by the current model on top (the overlay also shows the retinal flow slice). It can be seen that the computed overall score correlates well with the given manual score. In addition, the areas of higher quality in each slice are shown with higher area scores compared to the areas of lower quality.

[0117] Figure 15 A second example of quality maps obtained (e.g., shown in color or gray scale) for different 3x3mm OCTA acquisitions (organized in pairs) from the right eye (four images on the left) and the left eye (four images on the right) of a second patient is shown. For ease of discussion, Figure 15 the results are shown in a similar way as Figure 14 the results of Figure 15 In the example of it can be clearly seen how the higher quality acquisitions produce higher overall scores, and how the areas of higher quality within each image also have higher scores.

[0118] Figure 16 An example of quality maps obtained (e.g., shown in color or gray scale) for OCTA acquisitions of different FOVs: 3x3mm, 6x6mm, 9x9mm, and 12x12mm, from the right eye of a patient is shown. The OCTA acquisitions (organized in pairs) are shown in a similar way as Figure 14 the results of Figure 16Examples show how the present method performs flexibly for (e.g., adapts to) different scan sizes.

[0119] Comparison to inter-reader variability

[0120] To better understand the performance of the algorithm (e.g., as compared to human experts), in an exemplary application, the same three graders (readers) that were used to establish the ground truth training set (e.g., the algorithm) that was used to train the model according to the present application are given an evaluation image test set. The ground truth test data is defined (e.g., for each test image in the test image set) by taking the average of the quality ratings provided by the three graders. The individual results of each grader and the results produced by the present algorithm are compared to the ground truth test data and their individual performance is determined using the same quality criteria used to define a failure of the algorithm, as discussed above. More specifically, if the submitted quality data for a test image deviates from the ground truth test data by more than 1 quality point for 20% or more (e.g., not less than 20%) of the test image, the test image result is considered a failure. First, the annotations made by each of the three graders are compared to the average quality map used as the ground truth test data in the same manner. Figure 17 The percentage of acceptable results for each human grader and the algorithm is provided using the average of the three human graders as the ground truth test data. Although this analysis is biased because the annotations from each grader are also included in the average annotations used as the ground truth test data, it can be observed that the percentage of acceptable results from the human graders is similar to the percentage of acceptable results for the algorithm. In fact, one of the human graders actually achieves a lower percentage of acceptable results than the algorithm.

[0121] To remove the bias, the experiment is repeated individually for each human grader by removing the results from one grader from the data used to establish the ground truth test data and comparing the results of the removed grader to the revised ground truth test data. Figure 18 The percentage of acceptable results is provided that is obtained by comparing the annotations of each grader (individually removed from those used to establish the ground truth test data) and the results of the algorithm to the average results of the other two remaining graders as the ground truth test data. As shown, the percentage of acceptable results for the individual graders using this method is lower than the percentage of acceptable results for the algorithm as compared to the ground truth test data. Figure 18 The good performance of the present algorithm and the ability to produce a quality map of quality annotations similar to what would be produced by an average expert grader is highlighted taking into account the task difficulty of grading scan quality in images and the stringent validation criteria used.

[0122] Processing speed

[0123] In an example application, the average execution time for 26 scans (e.g., a combination of scans that require and do not require additional Gaussian filtering) was 2.58 seconds (with a 0.79 standard deviation). Of these 26 scans, 11 scans required additional Gaussian filtering (within the algorithm), while 15 scans did not. For those that required Gaussian filtering, the average processing time was 2.9 seconds (0.96 std), while for those that did not require Gaussian filtering, the average processing time was 2.34 seconds (0.56 std).

[0124] The following provides a description of various hardware and architectures suitable for use with the present application.

[0125] Optical coherence tomography imaging system

[0126] Generally, optical coherence tomography (OCT) uses low-coherence light to produce two-dimensional (2D) and three-dimensional (3D) internal views of biological tissue. OCT is capable of in vivo imaging of retinal structures. OCT angiography (OCTA) produces flow information, such as from blood vessels within the retina. Examples of OCT systems are provided in U.S. Patents 6,741,359 and 9,706,915, and examples of OCTA systems can be found in U.S. Patents 9,700,206 and 9,759,544, all of which are incorporated by reference in their entirety. An example OCT / OCTA system is provided herein.

[0127] Figure 19A generic frequency domain optical coherence tomography (FD-OCT) system suitable for collecting 3D image data of an eye in accordance with the present application is shown. The FD-OCT system OCT l includes a light source LtSrci. Typical light sources include, but are not limited to, a broadband light source with a short temporal coherence length or a swept frequency laser source. The beam of light from the light source LtSrci is typically guided by an optical fiber Fbrl to illuminate a sample, such as an eye E; a typical sample is tissue in a human eye. In the case of spectral domain OCT (SD-OCT), the light source LrSrcl can be, for example, a broadband light source with a short temporal coherence length, or in the case of swept source OCT (SS-OCT) the light source can be a wavelength tunable laser source. The light is typically scanned with a scanner Scnr1 located between the output of the optical fiber Fbrl and the sample E so that the beam of light (dashed line Bm) is scanned laterally over the sample region to be imaged. The beam of light from the scanner Scnr1 can pass through a scan lens SL and an ophthalmic lens OL and be focused onto the sample E being imaged. The scan lens SL can receive the beam of light from the scanner Scnr1 at multiple angles of incidence and produce substantially collimated light, and then the ophthalmic lens OL can focus onto the sample. The present example shows a scanning beam that needs to be scanned in two lateral directions (e.g., x and y directions on a Cartesian plane) to scan a desired field of view (FOV). One example of this is point field OCT, which scans a sample with a point field beam. Thus, the scanner Scnr1 is illustratively shown as including two sub-scanners: a first sub-scanner Xscn to scan a point field beam through the sample in a first direction (e.g., a horizontal x direction); and a second sub-scanner Yscn to scan the point field beam on the sample in a direction transverse to a second direction (e.g., a vertical y direction). If the scanning beam is a line field beam (e.g., line field OCT), which can sample an entire line portion of the sample at a time, then only one scanner can be needed to scan a line field beam through the sample to span a desired FOV. If the scanning beam is a full field of view beam (e.g., full field of view OCT), then no scanner is needed and the full field of view beam can be applied over the entire desired FOV at a time.

[0128] Regardless of the type of beam used, light scattered from the sample (e.g., sample light) is collected. In the present example, the scattered light returned from the sample is collected into the same optical fiber Fbrl that was used to direct the illuminating light. Reference light from the same light source LtSrci is passed through a separate path, in this case involving optical fiber Fbr2 with an adjustable optical delay and a retroreflector RRl. Those skilled in the art will recognize that a transmissive reference path can also be used, and that the adjustable delay can be placed in either the sample or reference arm of the interferometer. The collected sample light is combined with the reference light, for example in fiber coupler Cplr1, to form an optical interference in OCT light detector Dtctr1 (e.g., photodetector array, digital camera, etc.). Although a single fiber port is shown to reach the detector Dtctr1, those skilled in the art will recognize that various designs of interferometers can be used for balanced or unbalanced detection of the interference signal. The output from the detector Dtctr1 is provided to a processor (e.g., internal or external computing device) Cmpl, which converts the observed interference into depth information of the sample. The depth information can be stored in a memory associated with the processor Cmpl and / or displayed on a display (e.g., computer / electronic display / screen) Scnl. The processing and storage functions can be located within the OCT instrument, or the functions can be offloaded to an external processor (e.g., performed on an external computing device) to which the collected data can be transferred. An example of a computing device (or computer system) is shown in Figure 27

[0129] ​The sample and reference arms in an interferometer can be composed of bulk optics, fiber optics, or hybrid optical systems and can have different configurations, such as Michelson, Mach-Zehnder, or common-path based designs known to those skilled in the art. Beam as used herein should be interpreted as any carefully directed path of light. Instead of mechanically scanning the beam, the light field can illuminate a one- or two-dimensional region of the retina to generate OCT data (see, e.g., U.S. Patent 9332902; D. Hillmann et al., "Holographic mirror-holographic optical coherence tomography," Opt. Express, 36(13):2390 2011; Y. Nakamura et al., "High-speed three-dimensional human retina imaging by line-field spectral-domain optical coherence tomography," Opt. Express, 15(12):7103 2007; Blazkiewicz et al., "Signal-to-noise ratio study of full-field Fourier-domain optical coherence tomography," Appl. Opt., 44(36):7772 (2005)). In time-domain systems, the reference arm needs to have an adjustable optical delay to produce interference. Balanced detection systems are typically used in TD-OCT and SS-OCT systems, while a spectrometer is used in the detection port of SD-OCT systems. The present invention described herein can be applied to any type of OCT system. Various aspects of the present invention can be applied to any type of OCT system or other types of ophthalmic diagnostic systems and / or multiple ophthalmic diagnostic systems, including but not limited to fundus imaging systems, visual field testing devices, and scanning laser polarimetry.

[0130] In Fourier-domain optical coherence tomography (FD-OCT), each measurement is a real-valued spectral interferogram (Sj(k)). The real-valued spectral data is typically subjected to several post-processing steps, including background subtraction, dispersion correction, etc. The Fourier transform of the processed interferogram yields a complex-valued OCT signal output The absolute value of this complex OCT signal |Aj| reveals the profile of scattering intensity at different path lengths, and thus as a function of depth (z-direction) in the sample. Similarly, the phase Scattering profiles as a function of depth are referred to as axial scans (A-scans). A set of A-scans measured at adjacent locations in a sample produces a cross-sectional image of the sample (a tomogram or B-scan). A collection of B-scans collected at different lateral locations on the sample constitutes a data volume or cube. For a particular data volume, the term "fast axis" refers to the scan direction along a single B-scan, while the "slow axis" refers to the axis along which multiple B-scans are collected. The term "cluster scan" can refer to a single data unit or block produced by repeated acquisition at the same (or substantially the same) location (or region) for analysis of motion contrast, which can be used to identify flow. A cluster scan can consist of multiple A-scans or B-scans collected at approximately the same location on the sample at relatively short time intervals. Since the scans in a cluster scan are at the same region, static structures remain relatively unchanged from scan to scan in a cluster scan, while motion contrast between scans that meet a predetermined criterion can be identified as blood flow.

[0131] Various ways of creating B-scans are known in the art, including but not limited to: along the horizontal or x-direction, along the vertical or y-direction, along the diagonal of x and y, or in a circular or spiral pattern. B-scans can be in the x-z dimension, but can be any cross-sectional image including the z dimension. In some embodiments, the B-scan is a cross-sectional image of the sample in the x-z dimension. Figure 20 An example OCT B-scan image of a normal retina of a human eye is shown in FIG. 1. The OCT B-scan of the retina provides a view of the retinal tissue structure. For illustrative purposes, Figure 20 Various typical retinal layers and layer boundaries are identified. The determined retinal boundary layers include (from top to bottom): Internal limiting membrane (ILM) Layr 1, Retinal nerve fiber layer (RNFL or NFL) Layr 2, Ganglion cell layer (GCL) Layr 3, Inner plexiform layer (IPL) Layr 4, Inner nuclear layer (INL) Layr 5, Outer plexiform layer (OPL) Layr 6, Outer nuclear layer (ONL) Layr 7, Junction between the outer segment (OS) and inner segment (IS) of photoreceptors (denoted by reference symbol Layr 8), External or outer limiting membrane (ELM or OLM) Layr 9, Retinal pigment epithelium (RPE) Layr 10, and Bruch’s membrane (BM) Layr 11.

[0132] In OCT angiography or functional OCT, analysis algorithms can be applied to OCT data collected at the same or approximately the same sample locations on a sample at different times (e.g., cluster scans) to analyze motion or flow (see, e.g., U.S. Patent Publication Nos. 2005 / 0171438, 2012 / 0307014, 2010 / 0027857, 2012 / 0277579, and U.S. Patent No. 6,549,801, the entireties of which are hereby incorporated by reference). The OCT system can use any of a variety of OCT angiography processing algorithms (e.g., motion contrast algorithms) to identify blood flow. For example, a motion contrast algorithm can be applied to intensity information derived from the image data (intensity-based algorithms), phase information from the image data (phase-based algorithms), or complex image data (complex-based algorithms). An en face image is a 2D projection of 3D OCT data (e.g., by averaging the intensity of each individual A-scan such that each A-scan defines a pixel in the 2D projection). Similarly, an en face vasculature image is an image that displays the motion contrast signal, where the data dimension corresponding to depth (e.g., the z-direction along an A-scan) is displayed as a single representative value (e.g., a pixel in a 2D projection image), typically by summing or integrating over all or an isolated portion of the data (see, e.g., U.S. Patent No. 7,301,644, the entirety of which is hereby incorporated by reference). An OCT system that provides angiography imaging functionality can be referred to as an OCT angiography (OCTA) system.

[0133] Figure 21 An example of an en face vasculature is shown. After processing the data using any motion contrast technique known in the art to highlight motion contrast, a range of pixels corresponding to a given tissue depth of the internal limiting membrane (ILM) surface in the retina can be summed to generate an en face (e.g., front view) image of the vasculature. Figure 22 An example B-scan of a vasculature (OCTA) image is shown. As shown, the structural information can not be well-defined, as the blood flow can pass through multiple retinal layers such that their definition is not as well-defined as in a structural OCT B-scan, as Figure 20OCTA provides a non-invasive technique for imaging microvascular pathology of the retina and choroid, which can be critical for diagnosing and / or monitoring various pathologies. For example, OCTA can be used to differentiate diabetic retinopathy by identifying microaneurysms, neovascular complexes, and quantifying central avascular zones and non-perfusion areas. Further, OCTA has been shown to be well-consistent with fluorescein angiography (FA), a more traditional but more invasive technique that requires injection of dye to visualize blood flow in the retina. Further, in dry age-related macular degeneration, OCTA has been used to monitor the overall reduction in choroidal vascular layer flow. Similarly, in wet age-related macular degeneration, OCTA can provide qualitative and quantitative analysis of choroidal neovascularization and membranes. OCTA is also used to study vascular occlusions, such as evaluating non-perfusion areas as well as the integrity of superficial and deep vascular plexuses.

[0134] Neural network

[0135] As discussed above, the present application can use a neural network (NN) machine learning (ML) model. For completeness, a general discussion of neural networks is provided here. The present application can use any of the neural network structures described below, either individually or in combination. A neural network, or neural network, is a network of interconnected neurons (nodes), where each neuron represents a node in the network. Groups of neurons can be arranged in layers, where the output of one layer is fed forward to the next layer in a multi-layer perceptron (MLP) device. An MLP can be understood as a feedforward neural network model that maps a set of input data to a set of output data.

[0136] Figure 23 An example of a multi-layer perceptron (MLP) neural network is shown. Its structure can include a plurality of hidden (e.g., internal) layers HL1 to HLn that map an input layer InL (which receives a set of inputs (or vector inputs) in_1 to in_3) to an output layer OutL that produces a set of outputs (or vector outputs), such as out_1 and out_2. Each layer can have any given number of nodes, which are exemplarily shown herein as circles within each layer. In this example, the first hidden layer HL1 has two nodes, while the hidden layers HL2, HL3, and HLn each have three nodes. Generally speaking, the deeper the MLP (e.g., the greater the number of hidden layers in the MLP), the greater its learning capacity. The input layer InL receives a vector input (illustratively shown as a three-dimensional vector consisting of in_1, in_2, and in_3), and can apply the received vector input to the hidden layer sequence

[0137] The first hidden layer in the column is HL1. The output layer OutL receives the output from the last hidden layer (e.g., HLn) in the multi-layer model, processes its input, and produces a vector output result (illustratively shown as a two-dimensional vector consisting of out_1 and out_2).

[0138] Typically, each neuron (or node) produces a single output, which is fed forward to neurons in the immediately following layer. However, each neuron in a hidden layer can receive multiple inputs from either the input layer or from the outputs of neurons in the preceding hidden layer. Generally, each node can apply a function to its inputs to produce that node's output. Nodes in hidden layers (e.g., learning layers) can apply the same function to their respective inputs to produce their respective outputs. However, some nodes, such as those in the input layer InL, receive only one input and can be passive, meaning they simply forward the value of their single input to their output; for example, they provide a copy of their input to their output, as indicated by the dashed arrow within the node in the input layer InL.

[0139] For the purpose of explanation, Figure 24 A simplified neural network consisting of an input layer InL', a hidden layer HL1', and an output layer OutL' is shown. The input layer InL' is shown as having two input nodes i1 and i2 that receive inputs Input_1 and Input_2, respectively (e.g., the input nodes of layer InL' receive a two-dimensional input vector). The input layer InL' is fed forward to a hidden layer HL1' with two nodes h1 and h2, which in turn is fed forward to an output layer outL' with two nodes o1 and o2. The interconnections or links between neurons (as indicated by solid arrows) have weights w1 to w8. Typically, in addition to the input layers, nodes (neurons) can receive the output of the node in their immediately preceding layer as input. Each node computes its output by summing the products of its inputs (multiplying each of its inputs by the corresponding interconnection weights for each input), adding (or multiplying by) a constant defined by another weight or bias that can be associated with that particular node (e.g., node weights w9, w10, w11, w12 corresponding to nodes h1, h2, o1, and o2, respectively), and then applying a nonlinear or logarithmic function to the result. The nonlinear function can be called an activation function or a transfer function. Various activation functions are known in the art, and the choice of a particular activation function is not critical to the present discussion. However, it should be noted that the operation of an ML model or the behavior of a neural network depends on the weight values, which can be learned to make the neural network provide the desired output for a given input.

[0140] A neural network learns (e.g., is trained to determine) appropriate weight values to achieve a desired output for a given input during a training or learning phase. Prior to training a neural network, an initial (e.g., random and optionally non-zero) value can be assigned to each weight individually, such as a random number seed. Various methods of assigning initial weights are known in the art. The weights are then trained (optimized) so that, for a given training vector input, the neural network produces an output that is close to the desired (predetermined) training vector output. For example, the weights can be incrementally adjusted over thousands of iteration cycles by a technique known as backpropagation. In each cycle of backpropagation, a training input (e.g., a vector input or training input image / sample) is fed forward through the neural network to determine its actual output (e.g., a vector output). An error is then computed for each output neuron or node based on the actual neuron output and the target training output for that neuron (e.g., a training output image / sample corresponding to the current training input image / sample). The weights are then updated based on how much each weight contributes to the total error by propagating backwards (in a direction from the output layer back to the input layer) through the neural network so that the output of the neural network moves closer to the desired training output. The cycle is then repeated until the actual output of the neural network is within an acceptable error range of the desired training output for a given training input. It will be appreciated that many backpropagation iterations can be needed for each training input before achieving the desired error range. Typically, an epoch refers to one backpropagation iteration (e.g., one forward pass and one backward pass) for all training samples, so that many epochs can be needed to train a neural network. Generally, the larger the training set, the better the performance of the trained ML model, so various data augmentation methods can be used to increase the size of the training set. For example, when the training set includes multiple pairs of corresponding training input images and training output images, the training images can be divided into multiple corresponding image segments (or patches). Corresponding patches from the training input images and the training output images can be paired to define multiple training patch pairs from one input / output image pair, which expands the training set. However, training on large training sets places high demands on computational resources (e.g., memory and data processing resources). The computational demand can be reduced by dividing the large training set into multiple small batches, where the size of the small batches defines the number of training samples in one forward / backward pass. In this case, one epoch can include multiple small batches. Another issue is the potential for the NN to overfit the training set, which reduces its generalization ability from specific inputs to different inputs. The problem of overfitting can be mitigated by creating an ensemble of neural networks or by randomly dropping nodes within the neural network during training, which effectively removes the dropped nodes from the neural network. Various dropout regularization methods, such as dropout, are known in the art.

[0141] It should be noted that the operation of a trained NN machine model is not a direct algorithm of the operational / analytical steps. In fact, when a trained NN machine model receives an input, the input is not analyzed in the traditional sense. Rather, regardless of the subject matter or nature of the input (e.g., vectors defining real-time images / scans or vectors defining other entities such as demographic descriptions or activity records), the input will be subject to the same pre-defined architecture configuration of the trained neural network (e.g., same node / layer arrangement, trained weight and bias values, pre-defined convolution / deconvolution operations, activation functions, pooling operations, etc.) and it can not be clear how the architecture configuration of the trained network produces its output. Moreover, the values of the trained weights and biases are not deterministic and depend on many factors such as the amount of time the neural network is given to train (e.g., number of epochs in training), the random starting values of the weights before training begins, the computer architecture of the machine on which the NN is trained, the selection of training samples, the distribution of training samples among multiple mini-batches, the selection of activation functions, the selection of error functions that modify the weights, and even if the training is interrupted on one machine (e.g., with a first computer architecture) and completed on another machine (e.g., with a different computer architecture). The point is that the reasons why a trained ML model reaches certain outputs are not clear and much research is currently being conducted to try to determine the factors on which ML models base their outputs. Thus, the processing of real-time data by a neural network cannot be reduced to a simple step algorithm. Rather, its operation depends on its training structure, training sample set, training sequence, and various circumstances in the training of the ML model.

[0142] In summary, the construction of a NN machine learning model can include a learning (or training) phase and a classification (or operational) phase. In the learning phase, a neural network is trained for a particular purpose and can be provided with a set of training examples, including training (sample) inputs and training (sample) outputs, and optionally a set of validation examples to test the progress of the training. During this learning process, various weights associated with the nodes and node interconnections in the neural network are incrementally adjusted in order to reduce the error between the actual outputs of the neural network and the desired training outputs. In this way, a multi-layer feed-forward neural network (such as discussed above) can be enabled to approximate any measurable function to any desired degree of accuracy. The result of the learning phase is a (neural network) machine learning (ML) model that has been learned (e.g., trained). In the operational phase, a set of test inputs (or real-time inputs) can be submitted to the learned (trained) ML model, which can apply what it has learned to produce an output prediction based on the test inputs.

[0143] With Figure 23 And Figure 24Similar to regular neural networks, convolutional neural networks (CNNs) are also composed of neurons with learnable weights and biases. Each neuron receives inputs, performs an operation (e.g., dot product), and optionally follows that with a nonlinearity. However, CNNs can receive raw image pixels as input at one end (e.g., the input end) and provide a classification (or class) score at the other end (e.g., the output end). Because CNNs expect images as input, they are optimized to work with volumes (e.g., the height and width of the pixels of an image, plus the depth of the image, e.g., the color depth, such as the RGB depth defined by three colors: red, green, and blue). For example, the layers of a CNN can be optimized for neurons arranged in three dimensions. The neurons in a CNN layer can also be connected to small regions before that layer, rather than all neurons in a fully connected NN. The final output layer of a CNN can reduce the entire image to a single vector (classification) arranged along the depth dimension.

[0144] Figure 25 An example convolutional neural network structure is provided. A convolutional neural network can be defined as a sequence of two or more layers (e.g., layer 1 to layer N), where a layer can include an (image) convolution step, a (resulting) weighted sum step, and a nonlinearity function step. Convolution can be performed on its input data by, for example, applying a filter (or kernel) on a moving window over the input data to produce a feature map. Each layer and the components of a layer can have different predetermined filters (from a filter bank), weights (or weighting parameters), and / or function parameters. In this example, the input data is an image with a given pixel height and width, which can be the raw pixel values of the image. In this example, the input image is shown to have a depth of three color channels RGB (red, green, and blue). Optionally, the input image can undergo various pre-processing and the pre-processed results can be input instead of or in addition to the raw input image. Some examples of image pre-processing can include: retinal vessel map segmentation, color space conversion, adaptive histogram equalization, connected component generation, etc. Within a layer, a dot product can be computed between the given weights and the small region they are connected to in the input volume. Many ways to configure a CNN are known in the art, but as an example, a layer can be configured to apply an element-wise activation function, such as a max(0, x) threshold at zero. A pooling function (e.g., along the x-y direction) can be performed to downsample the volume. A fully connected layer can be used to determine a classification output and produce a one-dimensional output vector, which has been found useful for image recognition and classification. However, for image segmentation, a CNN will need to classify each pixel. Since each CNN layer tends to reduce the resolution of the input image, another stage is needed to upsample the image back to its original resolution. This can be achieved by applying a transpose convolution (or deconvolution) stage TC, which typically does not use any pre-defined interpolation method but has learnable parameters.

[0145] Convolutional neural networks have been successfully applied to many computer vision problems. As explained above, training a CNN typically requires a large training dataset. The U-Net architecture is based on a CNN and can typically be trained on a smaller training dataset than a traditional CNN.

[0146] Figure 26 An example U-Net structure is shown. The example U-Net includes an input module (or input layer or stage) that receives an input U-in of any given size (e.g., an input image or image patch). For illustrative purposes, the image size at any stage or layer is indicated within the box representing the image, e.g., the input module contains the number “128x128” to indicate that the input image U-in is composed of 128x128 pixels. The input image can be a fundus image, an OCT / OCTA en face, a B-scan image, etc. However, it should be understood that the input can be of any size or dimension. For example, the input image can be an RGB color image, a monochrome image, a volumetric image, etc. The input image undergoes a series of processing layers, each of which is shown with example dimensions, but these dimensions are for illustrative purposes only and will depend on, e.g., the size of the image, the convolutional filters, and / or the pooling stages. The architecture includes a contracting path (here exemplarily including four encoding modules), followed by an expansive path (here exemplarily including four decoding modules), and copy-and-cut links (e.g., CC1 to CC4) between the respective modules / stages that copy the output of one of the encoding modules in the contracting path and connect it to the upsample input of the corresponding decoding module in the expansive path (e.g., appended to the back of the upsample input of the corresponding decoding module in the expansive path). This results in the characteristic U-shape from which the architecture derives its name. Optionally, a “bottleneck” module / stage (BN) can be located between the contracting path and the expansive path, e.g., for computational considerations. The bottleneck BN can include two convolutional layers (with batch normalization and optional dropout).

[0147] The contraction path is similar to the encoder, and generally captures contextual (or feature) information by using feature maps. In this example, each encoding module in the contraction path can include two or more convolutional layers, as indicated by the asterisk symbol “*”, and can be followed by a max-pooling layer (e.g., a down-sampling layer). For example, the input image U-in is illustratively shown as undergoing two convolutional layers, each with 32 feature maps. It should be understood that each convolutional kernel produces a feature map (e.g., the output from a convolution operation with a given kernel is an image that is generally referred to as a “feature map”). For example, the input U-in undergoes a first convolution that applies 32 convolutional kernels (not shown) to produce an output consisting of 32 respective feature maps. However, as is known in the art, the number of feature maps produced by a convolutional operation can be adjusted (up or down). For example, the number of feature maps can be reduced by averaging groups of feature maps, discarding some feature maps, or other known feature map simplification methods. In this example, the first convolution is followed by a second convolution, the output of which is limited to 32 feature maps. Another way to envision the feature maps can be to consider the output of a convolutional layer as a 3D image, the 2D dimensions / planes of which are given by the listed X-Y plane pixel dimensions (e.g., 128 x 128 pixels), and the depth of which is given by the number of feature maps (e.g., 32 plane image depth). In this analogy, the output of the second convolution (e.g., the output of the first encoding module in the contraction path) can be described as a 128 x 128 x 32 image. The output from the second convolution is then subjected to a pooling operation, which simplifies the 2D dimensions of each feature map (e.g., the X and Y dimensions can each be halved). The pooling operation can be implemented in a down-sampling operation, as indicated by the downward arrow. Several pooling methods are known in the art, such as max-pooling, and the particular pooling method is not critical to the present invention. The number of feature maps can double at each pooling, starting with 32 feature maps in the first encoding module (or block), 64 feature maps in the second encoding module, and so on. Thus, the contraction path forms a convolutional network consisting of multiple encoding modules (or stages or blocks). As is typical for convolutional networks, each encoding module can provide at least one convolutional stage followed by an activation function (e.g., a rectified linear unit (ReLU) or a sigmoid layer) (not shown) and a max-pooling operation. Generally, the activation function introduces non-linearity into the layer (e.g., to help avoid overfitting issues), receives the results of the layer, and determines whether to “activate” the output (e.g., determines whether the value of a given node meets a pre-defined criterion to forward the output to the next layer / node). In summary, the contraction path generally reduces spatial information while increasing feature information.

[0148] The expansion path is analogous to the decoder, and where localization and spatial information can be provided for the results of the contraction path, despite the down-sampling and any max-pooling performed at the contraction stage. The expansion path comprises a plurality of decoding modules, where each decoding module concatenates its currently up-converted input with the output of the corresponding encoding module. In this way, features and spatial information are combined in the expansion path through a sequence of up-convolutions (e.g., up-sampling or transpose convolutions or de-convolutions) and concatenation with high-resolution features from the contraction path (e.g., via CC1 to CC4). Thus, the output of a de-convolution layer is concatenated with the corresponding (optionally, cropped) feature map from the contraction path, followed by two convolution layers and an activation function (with optional batch normalization).

[0149] The output from the last expansion module in the expansion path can be fed to another processing / training block or layer, such as a classifier block, which can be trained together with the U-Net architecture. Alternatively, or additionally, the output of the last up-sampling block (at the end of the expansion path) can be submitted to another convolution (e.g., output convolution) operation before producing its output U-out, as shown by the dashed arrow. The kernel size of the output convolution can be chosen to reduce the dimensionality of the last up-sampling block to a desired size. For example, the neural network can have multiple features per pixel up to the output convolution, which can provide a 1x1 convolution operation to combine these multiple features into a single output value per pixel at a pixel-wise level.

[0150] Computing device / system

[0151] Figure 27 An example computer system (or computing device or computer device) is shown. In some embodiments, one or more computer systems can provide functionality described or illustrated herein and / or perform one or more steps of one or more methods described or illustrated herein. A computer system can take any suitable physical form. For example, a computer system can be embedded in a device, a system-on-chip (SoC), a single-chip computer system (SBC) (for example, a computer-on-module (COM) or system-on-module (SOM)), a desktop computer system, a laptop or notebook computer system, a computer system grid, a mobile telephone, a personal digital assistant (PDA), a server, a tablet computer system, an augmented / virtual reality device, or a combination of two or more of these. Where appropriate, a computer system can reside in a cloud, which can include one or more cloud components in one or more networks. Where appropriate, one or more computer systems can perform without substantial spatial or temporal limitation one or more steps of one or more methods described or illustrated herein. In some embodiments, one or more computer systems might perform significantly different steps of one or more methods described or illustrated herein, or might perform the same steps in a different order or in parallel.

[0152] In some embodiments, the computer system can include a processor Cpntl, a memory Cpnt2, a storage Cpnt3, an input / output (I / O) interface Cpnt4, a communications interface Cpnt5, and a bus Cpnt6. The computer system can also optionally include a display Cpnt7, such as a computer monitor or screen.

[0153] The processor Cpntl includes hardware for executing instructions, such as those that make up a computer program. For example, the processor Cpntl can be a central processing unit (CPU) or a general-purpose computing on graphics processing units (GPGPU). The processor Cpntl can retrieve (or fetch) instructions from an internal register, an internal cache, the memory Cpnt2, or the storage Cpnt3, decode and execute them, and write one or more results to an internal register, an internal cache, the memory Cpnt2, or the storage Cpnt3. In particular embodiments, the processor Cpntl can include one or more internal caches for data, instructions, or both, as well as one or more arithmetic logic units (ALUs). The processor Cpntl can be a multi-core processor; or include one or more processors Cpntl. Although the present disclosure describes and illustrates a particular processor, the present disclosure encompasses any appropriate processor.

[0154] Memory Cpnt2 can include main memory for storing instructions and data that processor Cpntl executes or maintains during processing. For example, a computer system can load instructions or data (e.g., a data table) from storage Cpnt3 or from another source (e.g., another computer system) into memory Cpnt2. Processor Cpntl can load the instructions and data from memory Cpnt2 into one or more internal registers or internal caches. To execute instructions, processor Cpntl can retrieve and decode the instructions from the internal registers or internal caches. During or after instruction execution, processor Cpntl can write one or more results (which can be intermediate or final results) to the internal registers, internal caches, memory Cpnt2, or storage Cpnt3. Bus Cpnt6 can include one or more memory buses (which can each include an address bus and a data bus) and can couple processor Cpntl to memory Cpnt2 and / or storage Cpnt3. Optionally, one or more memory management units (MMUs) facilitate the transfer of data between processor Cpntl and memory Cpnt2. Memory Cpnt2 (which can be a fast volatile memory) can include random access memory (RAM), such as dynamic RAM (DRAM) or static RAM (SRAM). Storage Cpnt3 can include long- or large-term storage for data or instructions. Storage Cpnt3 can be internal or external to the computer system and include one or more of a disk drive (e.g., a hard disk drive, HDD, or a solid-state drive, SSD), flash memory, ROM, EPROM, an optical disk, a magneto-optic disk, magnetic tape, a universal serial bus (USB)-accessible drive, or other types of non-volatile memory.

[0155] I / O interface Cpnt4 can be software, hardware, or a combination of both, and includes one or more interfaces (e.g., serial or parallel communication ports) for communicating with I / O devices, which can enable communication with a human (e.g., a user). For example, the I / O devices can include a keyboard, a keypad, a microphone, a monitor, a mouse, a printer, a scanner, a speaker, a still camera, a stylus, a tablet, a touchscreen, a trackball, a video camera, another suitable I / O device, or a combination of two or more of these.

[0156] The communication interface Cpnt5 can provide a network interface for communicating with other systems or networks. The communication interface Cpnt5 can include a Bluetooth interface or other type of packet-based communication. For example, the communication interface Cpnt5 can include a network interface controller (NIC) and / or a wireless NIC or wireless adapter for communicating with a wireless network. The communication interface Cpnt5 can provide communication with a WI-FI network, an ad hoc network, a personal area network (PAN), a wireless PAN (e.g., a Bluetooth WPAN), a local area network (LAN), a wide area network (WAN), a metropolitan area network (MAN), a cellular telephone network (e.g., a Global System for Mobile Communications (GSM) network), the Internet, or a combination of two or more of these.

[0157] The bus Cpnt6 can provide a communication link between the above-described components of the computing system. For example, the bus Cpnt6 can include an accelerated graphics port (AGP) or other graphics bus, an enhanced industry standard architecture (EISA) bus, a front-side bus (FSB), a HyperTransport (HT) interconnect, an industry standard architecture (ISA) bus, an Infinity bus, a low pin count (LPC) bus, a memory bus, a MicroChannel Architecture (MCA) bus, a Peripheral Component Interconnect (PCI) bus, a PCI-Express (PCIe) bus, a serial advanced technology attachment (SATA) bus, a Video Electronics Standards Association local (VLB) bus, or another suitable bus or combination of two or more of these.

[0158] Although this disclosure describes and illustrates particular computer systems having particular numbers of particular components in particular arrangements, this disclosure contemplates any suitable computer systems having any suitable numbers of any suitable components in any suitable arrangements.

[0159] Here, a computer-readable non-transitory storage medium can include one or more semiconductor-based or other integrated circuits (ICs) (such as for example, field-programmable gate arrays (FPGAs) or application-specific ICs (ASICs)), hard disk drives (HDDs), hybrid hard drives (HHDs), optical discs, optical disc drives (ODDs), magneto-optical discs, magneto-optical drives, floppy diskes, floppy disk drives (FDD s), magnetic tapes, solid-state drives (SSDs), RAM-drives, SECURE DIGITAL cards or drives, any other suitable computer-readable non-transitory storage media, or any suitable combination of two or more of these. A computer-readable non-transitory storage medium can be volatile, non-volatile, or a combination of volatile and non-volatile, it can be embedded in a computer program product, which can be transported or distributed over a computer network.

[0160] While the application has been described in connection with specific embodiments thereof, it will be understood that many modifications, substitutions and changes will be apparent to those skilled in the art in light of the foregoing description. Accordingly, this application is intended to embrace all changes and modifications that fall within the spirit and scope of the claims.

Claims

1. A method for generating a quality metric for optical coherence tomography (OCT) data, comprising: acquiring a volume of OCT data; defining one or more slice views from the volume OCT data; generating a plurality of feature maps from each slice view, wherein each slice view is a frontal planar representation of a sub-volume of the volume of OCT data; determining the quality metric based on image properties of the plurality of feature maps, wherein the image properties of the feature maps comprise image texture features or Haralick features; displaying the quality metric or storing the quality metric for further processing.

2. The method of claim 1, wherein, determining the quality metric comprises submitting the plurality of feature maps to a machine model trained using a plurality of pre-ranked OCT data volume samples, one or more training slice views defined from each OCT data volume sample, and a plurality of training feature maps generated from each training slice view.

3. The method of claim 2, wherein, the machine model is a deep learning model.

4. The method of claim 2, wherein, the machine model is a neural network model.

5. The method of claim 1, wherein, the OCT data is OCT angiography data.

6. The method of claim 1, wherein, the quality metric is a two-dimensional (2D) quality map identifying quality metric of different regions of a corresponding slice view.

7. The method of claim 6, wherein, acquiring a volume of OCT data comprises: a) collecting a plurality of different OCT volume samples of a same retinal tissue region; b) applying to each of the OCT volume samples the steps of defining one or more slice views, generating the plurality of feature maps, and determining the quality metric based on the plurality of feature maps, thereby defining a plurality of 2D quality map samples corresponding to the plurality of different OCT volume samples; and the method further comprises: defining a composite OCT data based on regions of the plurality of different OCT volume samples having higher quality based on the respective 2D quality map samples than other regions.

8. The method of claim 6, wherein: determining the quality metric comprises submitting the plurality of feature maps to a machine model trained to determine the quality map based on the plurality of feature maps; and the machine model is further trained to identify one or more causes of a quality region in the quality map having a lower quality metric than a predetermined threshold, and to identify one or more corrective actions to improve the lower quality metric in subsequent OCT acquisition.

9. The method of claim 8, wherein, the one or more causes are selected from a predefined list of error sources including one or more of: misfocus, opacity, illumination below a predefined threshold, light penetration less than a predefined threshold, tracking or motion artifacts, and noise above a predefined threshold.

10. The method of claim 8, wherein, the one or more corrective actions include focus adjustment, recommendation for pupil dilation, identification of an alternative imaging angle, or identification of a likely cause of eye tracking absence.

11. The method of claim 8, wherein, the corrective actions are output to an electronic display.

12. The method of claim 8, wherein, the corrective actions are communicated to an automated subsystem that automatically implements the corrective actions prior to subsequent acquisition.

13. The method of claim 6, wherein, further comprising defining an overall quality score for the acquisition based at least in part on an average of individual quality metric distributions of the quality map.

14. The method of claim 6, wherein, The acquisition is an OCTA acquisition defined from a plurality of OCT scans of the same region of the retina, and the method further comprises defining an overall quality score for the OCTA acquisition based at least in part on an average of individual quality metric distributions of the quality maps of the plurality of OCT scans defining the OCTA acquisition.

15. The method of claim 6, further comprising: Identifying a target region within an acquisition, and designating the entire acquisition as good or bad based on a quality map metric corresponding to the target region.

Citation Information

Patent Citations

  • High speed spectral domain functional optical coherence tomography and optical doppler tomography for in vivo blood flow dynamics and tissue structure

    US20050171438A1

  • In vivo structural and flow imaging

    US20100027857A1

  • Inter-frame complex oct data analysis techniques

    US20120277579A1

  • Method and apparatus for ultrahigh sensitive optical microangiography

    US20120307014A1

  • Phase-resolved optical coherence tomography and optical doppler tomography for imaging fluid flow in tissue with fast scanning speed and high velocity sensitivity

    US6549801B1