Machine learning method for creating structure-derived field of view priors

By combining ophthalmic imaging systems and machine learning to generate structure-derived field-of-view priors, the starting point and intensity of field-of-view testing are optimized, solving the problems of long testing time and insufficient individualization, and achieving shorter testing time and higher accuracy.

CN114390907BActive Publication Date: 2026-03-17CARL ZEISS MEDITEC INC +1
View PDF 24 Cites 0 Cited by

Patent Information

Application Number
CN202080062565.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-09-06
Filing Date
2020-09-04
Publication Date
2026-03-17
Estimated Expiration
2040-09-04

AI Technical Summary

Technical Problem

Existing field of vision testing methods are time-consuming, especially for glaucoma patients, affecting testing frequency and accuracy. Furthermore, traditional field of vision prior information lacks individualization, leading to extended testing time.

Method used

By combining ophthalmic imaging systems such as OCT and OCTA, machine learning techniques are used to generate structure-derived field-of-view priors, optimize the starting point and intensity of field-of-view tests, and use biometric information to predict thresholds, thereby reducing the number of iterative adjustments.

Benefits of technology

It significantly shortens the field of vision testing time, improves the accuracy and repeatability of the test, reduces patient fatigue, lowers testing costs, and enhances the efficiency of individualized field of vision testing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114390907B_ABST
    Figure CN114390907B_ABST
Patent Text Reader

Abstract

Systems for customized visual field (VF) testing use machine learning models (15) trained on retinal images (12A, 12C, 12D) including optical coherence tomography (OCT), optical coherence tomography angiography (OCTA), fundus, and / or fluorescein angiography images. In operation, when a particular VF test (13) is to be prepared for a patient, a retinal image of the patient is submitted to the current machine model, which responds by synthesizing a VF for the patient. The synthesized VF can be used to optimize the particular VF test prior to performing the particular VF test on the patient.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention generally relates to the field of field testing. More specifically, this invention relates to a system and method for optimizing field testing to improve accuracy, enhance repeatability, reduce total testing time, and for suggesting / identifying new locations within the field of view to be tested. Background Technology

[0002] Glaucoma is one of the leading causes of blindness worldwide, with 44.7 million people suffering from open-angle glaucoma globally, a number projected to reach 58.6 million by 2020. While optical coherence tomography (OCT) and optical coherence tomography angiography (OCTA) are becoming increasingly common in glaucoma management, field of view (VF) analysis remains the clinical gold standard for glaucoma diagnosis and staging, as well as for monitoring functional vision loss over time.

[0003] A field of vision test is a method of measuring an individual's entire field of vision, such as their central and peripheral (lateral) visual acuity. A field of vision test is a method of mapping the field of vision for each eye individually, which can detect blind spots (scotomas) as well as more subtle areas of blurred vision.

[0004] A rangefinder or "perimeter" is a specialized machine / device / system used to perform field of vision testing on a patient. There are different types of perimeters and different types of field of vision tests, but all field of vision tests are subjective. Therefore, the patient must be able to understand the test instructions, cooperate fully, and remain alert throughout the test to provide useful information. Adding to the complexity, field of vision tests can take a relatively long time, which can cause patient fatigue and impair test results.

[0005] A common type of field-of-view test or algorithm is the Standard Automated Perimeter (SAP) test, which determines how dim light can be and how still perceptible it can be at different points within a single eye's field of view (e.g., a threshold). Various algorithms have been developed to determine the threshold for different individual test points within a single field of view. The Swedish Interactive Thresholding Algorithm (SITA) can be combined with the SAP test to more effectively determine the field of view, for example, with... When used with the Humphrey Field Analyzer (HFA), the SITA algorithm optimizes the determination of the field threshold by continuously estimating the predicted threshold based on the patient's age and neighboring thresholds. For example, the intensity of each subsequent stimulus is modified based on the patient's response to the first stimulus. This iterative process is repeated until the possible threshold measurement error is reduced below a predetermined level, typically with one or more reversals at each test location. In this way, the time required to acquire the field of view can be reduced, patient fatigue can be reduced, and reliability can be increased. Improvements to SITA have resulted in Fast SITA and even faster SITA algorithms, which can further reduce test time. Similar to the SITA testing strategy of the HFA, the Trend-Oriented Perimeter (TOP) algorithm was developed for Octopus. TM Visual acuity measurement serves as an alternative to its lengthy stepped threshold procedure. Nevertheless, even with state-of-the-art testing strategies, such as various versions of SITA, field testing for each eye typically still takes several minutes. Testing time also tends to increase with more damaged or glaucoma-related fields of vision.

[0006] In summary, shorter testing strategies may help increase the frequency of field testing in glaucoma management, bringing clinical glaucoma care closer to current professional recommendations. Patients generally prefer shorter testing times to minimize the impact of patient fatigue, thereby obtaining more reliable test results and reducing testing costs.

[0007] The purpose of this invention is to reduce the total testing time for field of view testing.

[0008] Another objective of this invention is to reduce the duration of threshold field of view testing (e.g., the time required for a patient to reach his / her minimum visible light threshold for a single test point) with minimal or no loss of clinical accuracy.

[0009] Another object of the present invention is to provide a system and method for improving the prediction of expected thresholds for individual patient test points.

[0010] Another objective of this invention is to help reduce the duration of threshold field of vision testing by utilizing structural and / or functional characteristics of the patient's eye obtained through different ophthalmological examination methods. Summary of the Invention

[0011] The aforementioned objectives are achieved in the method / system for customizing field of vision testing. This method / system may have multiple components, including: a data system for selecting a field of vision test for a patient, wherein the selected field of vision test has one or more test points with definable light intensity; acquisition or otherwise access to biometric (e.g., structural or functional) measurements of the patient's retina, for example, from an electronic medical record (EMR). Biometric information can be collected using optical coherence tomography (OCT) systems, OCT angiography systems, fundus imagers, or other ophthalmic examination system modalities for collecting physical / empirical ophthalmic data. For example, the biometric information may be at least partially based on retinal images, which may include 3D or depth-resolved data. A computing system or network (e.g., a computing system or network embodying a machine learning architecture (e.g., an artificial intelligence system and / or a neural network system) can be used to predict corresponding threshold sensitivity values ​​for one or more selected test points of the selected field of vision test, at least partially based on the acquired biometric information. Each predicted threshold sensitivity value may include a light intensity measurement that the patient is expected to see with a predetermined success rate (e.g., 50% success rate), and / or an area measurement that the patient is expected to see at a given brightness level (e.g., a specific size of illuminated point / shape / area), and / or a combination of both. The vision testing system can use the predicted threshold sensitivity values ​​as "prior," for example, as input to a selected field of view test (which can be used to optimize the patient's FA test), and / or as the starting intensity / area value for one or more selected test points when the selected field of view test is applied to the patient. By using starting intensity values ​​that are close to the patient's final test result, the patient can reach his / her threshold more quickly, resulting in a shorter overall test duration.

[0012] Alternatively, the predicted threshold sensitivity can be used as a synthetic VF prior to replace, or supplement, the true VF prior in the VF prediction system. The VF prediction system can use the synthetic VF prior (and optionally any available true VF prior) to predict the patient's future field of vision.

[0013] Other objects and achievements, as well as a fuller understanding of the invention, will become apparent and readily understood by taking into account the accompanying drawings and the following description and claims.

[0014] To facilitate understanding of this invention, several publications may be cited or referenced herein. All publications cited or referenced herein are incorporated herein by reference in their entirety.

[0015] The embodiments disclosed herein are merely examples, and the scope of this disclosure is not limited thereto. Any embodiment feature mentioned in one claim class (e.g., system) may also be claimed in another claim class (e.g., method). Dependencies or references in the appended claims are chosen solely for formal reasons. However, any subject matter arising from intentional references to any prior claim may also be claimed, such that any combination of claims and their features may be disclosed and claimed, regardless of the dependencies chosen in the appended claims. Attached Figure Description

[0016] In the accompanying drawings, the same reference symbols / characters denote the same parts:

[0017] Figure 1 An overview of a system for customized field-of-view testing according to the present invention is provided;

[0018] Figure 2 A training example of the neural network NN-1 according to the present invention is shown;

[0019] Figure 3 It shows Figure 2 Example operation of a trained neural network NN-1, wherein real-time data is input after training, or test data is input during the evaluation phase of the training session;

[0020] Figure 4 An alternative training architecture linking multiple NN levels (Stg1 and Stg2) is shown, with each NN level including its own neural network in a modular neural network configuration;

[0021] Figure 5 This demonstrates how a patient's field of view testing history can help predict (forecast) the patient's current or future field of view threshold for a given test point;

[0022] Figure 6 A field-of-view prediction system according to the present invention is shown;

[0023] Figure 7A and Figure 7B Based on the random forest method ( Figure 7A ) and neural network methods ( Figure 7B A graph of the (derived) OCT-estimated threshold versus the (true) VF threshold for the machine learning model.

[0024] Figure 8 Table 1 shows that the overall mean absolute error (MAE) of ZEST-RF and ZEST-CNN is statistically comparable to that of ZEST (p<0.001);

[0025] Figure 9An example of a field of vision testing instrument (perimeter) used to test a patient's field of vision is shown;

[0026] Figure 10 An example of a slit-scan ophthalmic system for fundus imaging is shown;

[0027] Figure 11 A generalized frequency-domain optical coherence tomography system for collecting 3D image data of the eye, suitable for use in this invention, is shown.

[0028] Figure 12 An example of a frontal vascular system image is shown;

[0029] Figure 13 An example of a multilayer perceptron (MLP) neural network is shown;

[0030] Figure 14 A simplified neural network consisting of an input layer, hidden layers, and an output layer is shown.

[0031] Figure 15 An example convolutional neural network architecture is shown;

[0032] Figure 16 An example U-Net architecture is shown;

[0033] Figure 17 An example computer system (or computing device or computer) is shown. Detailed Implementation

[0034] In a typical field of vision (VF) test, a patient is presented with multiple test points distributed (e.g., sequentially) across the field of vision and is asked to perceive the appearance of each individual test point. The size and / or light intensity of each individual test point can be adjusted until the patient is able to recognize the appearance of each individual test point with a predetermined success rate (e.g., 50%). The final size and / or intensity of the test point defines a threshold for that test point, which can serve as the basis for visual sensitivity measurements incorporated into the field of vision test results. If the initial size and / or intensity of a test point deviates significantly from its final threshold, numerous adjustment iterations (e.g., reaching the patient's threshold for that particular test point) may be required before “thresholding,” resulting in longer test times. Therefore, the goal of an effective thresholding strategy is to select initial size and / or intensity values ​​for the corresponding test points that are close to the final threshold for a particular patient, thereby reducing the field of vision test time.

[0035] Efficient thresholding strategies have been pushing the limits of threshold testing. One approach to improving the threshold is to use visual field "prior," or prior information (e.g., historical data or statistical models derived from historical data) to estimate a patient's future VF test performance. By default, Bayesian strategies incorporate the idea of ​​prior information or data that is updated with each stimulus presentation (e.g., test point) and response. The Swedish Interactive Throttling Algorithm (SITA) and the Fast Estimation via Sequential Tests (ZEST) visual field calculation method are examples of strategies that utilize Bayesian prior techniques. For a discussion of SITA, see Boel Bengtsson et al., “SITA Fast, A New Rapid Perimetric Threshold Test, Description of Methods and Evaluation in Patients with Manifest and Suspect Glaucoma,” 1998: 76: 431-437, and Anders Hejil et al., “A New SITA Perimetric Threshold Testing Algorithm: Construction and a Multicenter Clinical Study,” American Journal of Ophthalmology, Vol. 198, February 2019, pp. 154-165. Similarly, for a discussion of ZEST, see Chong et al., “Targeted Spatial Sampling Using GOANNA Improves Detection of Visual Field Progression,” Ophthalmic Physiol Opt, March 2015, 35(2): 155-69. All of these references are incorporated herein by reference. Priors (e.g., previously collected data and / or population-derived data) are typically based on uniform values ​​(usually exceeding the threshold / brightness), correlated with age-matched data, or even derived from previous visual fields of the same patient. Some limitations of using these priors are that uniform or age-matched data are not individualized for a given patient, meaning additional stimulation may be required at a given location. Visual field priors for the same patient are possible but may often be unavailable (e.g., at the patient's first visit) or expire due to the VF test being less frequent than other tests. That said, because the VF test is subjective and takes more time than other more typical ophthalmological tests (e.g., structural / imaging tests), the VF test may not be performed at the recommended intervals.

[0036] A newer approach to facilitating field-of-view (VF) testing is to construct a structurally derived field of view, which may include one or more of the following: a derived field of view, a derived visual sensitivity measurement, and a derived prior based on one or more quantifiable data sources, such as ophthalmic images, patient-specific physiological characteristics / measurements, medical conditions, medical treatments, (visual) evoked potential tests, and / or other visually relevant tests. Structural imaging, such as optical coherence tomography (OCT), has been used to estimate (e.g., derive) the field of view, and it is often positioned as an “alternative” field of view to functional VF testing (e.g., in lieu of standard / functional field-of-view testing). Because structural data is generally more reproducible than functional VF data, structural data can provide the benefit of a more easily reproducible derived field of view. One limitation of previous structure-derived fields of view was that they were typically generated using custom mathematical models, often instrument-specific, as described by Bogunovic et al. in “Relationships of Retinal Structure and Humphrey 24-2 Visual Field Thresholds in Patients with Glaucoma,” Invest. Ophthalmol. Vis. Sci., 2015; 56(1):259-271, which are incorporated herein by reference in their entirety. This use of instrument-specific custom mathematical models limited the utility of structure-derived fields of view. Another obstacle to previous structure-derived fields of view was that the standard (functional) field of view was still considered the gold standard for assessing visual function. Therefore, the functional field of view was likely to be more trusted by the general clinician than estimated fields of view derived from structural a priori.

[0037] In evoked potential (EP) or evoked response (ER) tests, electrodes are used to record potential responses from specific parts of the patient's nervous system (typically the brain) followed by the presentation of a stimulus (sensory stimulus), such as light, sound, or touch. For example, an evoked potential test can measure the time required for the brain to respond to a sensory stimulus. In a visual evoked potential (VEP) test, electrodes may be placed on the patient's scalp while the patient sits in front of a screen viewing a changing pattern of light (e.g., first with one eye, then with the other). A VEP test can record the response of each eye to the changing pattern. For example, the patient may be asked to gaze at a checkerboard pattern on a screen while the colors of the squares alternate at a predetermined frequency and / or a predetermined pattern, and the VEP test records the changes the patient is able to perceive based on the patient's evoked potential responses.

[0038] However, structural priors are believed not to have been used to facilitate the construction / management of standard functional fields of view. This paper presents a method, system, and / or workflow that generates accurate (realistic / functional) fields of view in a novel manner with reduced testing time.

[0039] This invention combines the use of ophthalmic imaging / examination systems (and / or their outputs) with field-of-view testing systems (e.g., perimeters) to optimize functional field-of-view testing (e.g., optimizing its starting point, such as the initial light intensity and / or size of the test point in a functional field-of-view test). For example, for a given test point, the optimized starting point can be estimated / predicted to be close to the patient's expected threshold (e.g., final) value. In this way, the number of iterative adjustments (intensity and / or size) to reach the patient's threshold for a given test point is reduced, resulting in a reduction in total testing time. Generally, a discussion of common vision testing systems and typical (functional) field-of-view testing is given in the "Field-of-View Testing Systems" section below.

[0040] Various types of ophthalmic imaging / examination systems are known in the art, such as fundus imagers, OCT systems, and OCT angiography (OCTA) systems. Fundus imagers can capture two-dimensional (2D) images of the retinal surface or other parts of the eye. Various structural measurements / observations can be made from fundus images. OCT and OCTA enable non-invasive, depth-resolved (e.g., A-scan), volumetric (e.g., C-scan), and 2D (e.g., frontal or cross-sectional / B-scan) visualization of the retinal vascular system. OCT can provide structural images of the vascular system, while OCTA can provide functional images of the vascular system (e.g., blood flow). For example, OCTA can image vascular flow by using the motion of flowing blood as an intrinsic contrast. These types of ophthalmic imaging systems will be discussed below in “Fundus Imaging Systems” and “Optical Coherence Tomography (OCT) Imaging Systems.” Unless otherwise stated, aspects of the invention can be applied to any or all of these ophthalmic imaging systems. For example, the methods / systems presented herein for optimizing thresholds (e.g., optimizing the initial values ​​of test points in a functional field of vision test, and / or providing a synthetic “prior” for a functional field of vision test) may include structural and / or functional (e.g., motor) ophthalmic information (biometric information) extracted from the eye, and this ophthalmic information (biometric information) may be obtained by using a fundus imager, an OCT system, and / or an OCTA system.

[0041] Some embodiments of the present invention utilize existing Bayesian-type strategies and follow-up measures to add synthetic / derived "priorities" (e.g., synthetic fields of view) derived from structural and / or functional ophthalmic data / imaging (e.g., fundus images, OCT scans / images, OCTA scans / images, patient-specific physiological characteristics / measurements, medical conditions, medical treatments, (visual) evoked potential tests, and / or other visual-related tests) instead of genuine VF priors (e.g., prior functional VF test results obtained using a perimeter). These previously synthesized fields of view can be derived using machine learning (ML) techniques, such as deep learning (DL) and / or artificial intelligence (AI) methods. That is, unlike existing methods that attempt to accelerate functional VF testing using genuine VF priors, this method proposes using synthetic VF priors determined from structural (e.g., OCT and fundus imaging) and / or functional (e.g., OCTA imaging) ophthalmic information, collectively referred to herein as "biometrics" and / or "physical characteristics" (measurements / values). One advantage of this approach is that biometrically derived priors are likely to be more reproducible (e.g., less variable) than generating multiple (real) prior fields of view to establish a subject's VF history, and are generally less cumbersome for the subject to obtain (i.e., can be quickly derived during a patient's visit to the same clinic before the patient undergoes a conventional functional field of view test). Furthermore, biometrically derived priors can be created using artificial intelligence (AI), machine learning (ML), and / or deep learning (DL) methods.

[0042] It should be understood that various types of machine learning models are known in the prior art, such as supervised learning, unsupervised learning, semi-supervised learning, reinforcement learning, etc. While examples of specific machine learning models (e.g., DL and AI) are provided in the aspects of this discussion, it should be understood that other types of machine learning models can be used individually or in combination in this invention. For example, according to the invention, one or more of nearest neighbor, Naive Bayes, decision trees, linear regression, support vector machines (SVM), and neural networks can be used to implement supervised learning models.

[0043] By using machine learning techniques to derive biometric-derived priors, this invention not only has the potential to generate more robust and reproducible input (synthetic / derived) fields of view (VFs), but also leverages features inherent to these methods that can help better understand ophthalmic biometrics associated with VF function models (e.g., identifying the relationship between observed biometric information and VF testing). By using structural priors (e.g., biometric-derived priors) created from biometric (e.g., image) data typically collected as part of standard clinical workflows, instead of true field-of-view priors (which may not be available from previous accesses or are less reproducible), as a rapid VF testing strategy (which may lack other available VF prior data), such as SITA, this method is estimated to potentially reduce current threshold VF testing time in glaucoma eyes by up to 30%. In other words, this invention pushes the limitations of threshold testing beyond what can be achieved using modern thresholding strategy types alone, such as SITA and its variants, which may have reached their optimization limits. For example, this method can extend these limitations by introducing biometric information as an additional source of prior information for optimizing functional VF testing.

[0044] Figure 1An overview of a system for customizing / optimizing conventional (functional) field of vision tests according to the present invention is provided. The system may include a graphical user interface (not shown) and an electronic processor to facilitate various processing steps. For example, the user / technician may begin by specifying / selecting a specific VF test (box 11). The system can be configured to work with any type of field of vision test selected for a given perimeter (VF tester) VF0. The selected field of vision test may be any of the following: static automated perimetry test, kinetic perimetry test, frequency doubling perimetry test, or other known field of vision test types that use thresholds (e.g., visual sensitivity measurements) to determine the patient's field of vision. Examples of known VF tests include Standard Automated Perimetry (SAP), Short Wavelength Automated Perimetry (SWAP), Frequency Doubling Technique (FDT), Swedish Interactive Thresholding Algorithm (SITA), Fast SITA, Faster SITA, Tendency-Guided Perimetry (TOP), Objective Perimetry (VEP, Multifocal Electroretinography (ERG / PERF), Pupil Measurement, etc.). Regardless of the type of VF test selected, the chosen field of view test typically has one or more test points with definable light intensity. Some VF tests also provide points with definable size (e.g., area) at a given light intensity. The purpose of this system is to determine the threshold (e.g., termination intensity and / or size) of individual test points for a specific patient, to which the selected VF test will be administered.

[0045] As shown in box 12, this system acquires one or more biometric information (e.g., body characteristic measurements), such as biometric information of the retina of a patient to be visually tested (optionally including the patient's previous functional tests) to construct a structure-derived field of view. The biometric information can be based on retinal images obtained using any of a variety of imaging modalities and / or images (e.g., photocopies, bitmap / raster / vector or other digital images, printouts, etc.) from previous patient tests. For example, the imaging modality can be grayscale, color, infrared, retinal thickness mapping, fundus photography, optical coherence tomography (OCT), Doppler OCT, OCT angiography, and / or fluorescein angiography. Biometric information 12 can be extracted from (e.g., based on) one or more OCT / OCTA images 12A, previous field of view test results 12B (or the primary sensitivity value of a previous field of view test), fundus images 12C, fluorescein angiography (FA) images 12D, VEP 12E, or other imaging modalities or retinal / visual measurement techniques / devices, or include all (or part) of these. Biometric information can be obtained by using ophthalmic testing systems (e.g., OCT systems or fundus imaging devices, not shown) directly on the patient during their visit, or it can be accessed from the patient's medical record data storage, such as from an electronic medical record (EMR). Examples of biometric information may include one or more A-scans, B-scans, C-scans, or frontal images obtained using an OCT / OCTA system. Biometric information may include the shape, size, color, and / or relative location of individual ophthalmic structures, such as measurements of the optic disc (OHN), fovea, retinal thickness, and thickness of individual retinal layers. Other examples of biometric information may include blood flow measurements and / or tissue motion measurements of specific regions of the retina, discoloration areas deviating from expected criteria, vascular transition areas (e.g., their size, location, and / or number), exudate formation (e.g., their size, location, and / or number), large vessel counts, small vessel counts, and identification of specific structures, some of which may indicate pathology (e.g., associated with it). For example, exudate-related disorders are lesions associated with certain types of "wet" age-related macular degeneration (AMD). Biometric information may also include comparisons of relative measurements of different physiological characteristics, such as distances between specific structures (and / or relative orientations / positions of specific structures) and / or relative size ratios of specific structures.

[0046] The acquired biometric information can be submitted to a machine learning model 15, which may be contained in one or more computing systems (e.g., electronic processors). It should be understood that individual retinal images (e.g., OCT / OCTA, fundus, and / or fluorescence images) can be submitted to the machine model 15 as one or more biometric information, and the machine model 15 can extract individual biometric sub-measurements from the submitted images as needed. Optionally, the machine model 15 may also receive input information regarding the selection of a specific VF test algorithm to be administered to the patient. For example, the machine model 15 may be informed of the type of VF test to be performed on the patient, which allows it to better adapt to the construction of its appropriate bio-derived priors. The machine learning model 15 may determine (e.g., predict / synthesize / derive) a corresponding threshold (e.g., visual sensitivity value) for one or more selected test points of the selected VF test type, at least in part, based on the received biometric information. Each threshold sensitivity value may be based on a light intensity measurement and / or dot size measurement that the patient is expected to see clearly at a predetermined success rate (e.g., 50% success rate) for a single VF test point. In other words, machine learning model 15 outputs a synthesized VF threshold (e.g., VFTh_out), which can constitute one or more VF priors, such as a set of numerical data (schematically shown as the derived VF test output 10), and can be used in conjunction with a selected functional VF test performed on the patient, as shown in box 13. Therefore, this system results in an accelerated functional VF examination 17 (e.g., a short-duration VF examination).

[0047] Optionally, as shown in box 14, the individual threshold sensitivity value VFTh_out can be further based on additional, unstructured, or image-related patient data, such as data accessible from the EMR. For example, the determination of the threshold sensitivity value for one or more selected test points for a selected field of view test can be further based on patient age-specific criteria data associated with a specific imaging device (e.g., OCT and / or fundus imager) that provides one or more biometric information. The prediction of the threshold sensitivity value can also be based on unstructured patient-specific data (e.g., physiological data not extracted from the input retinal image in box 12), such as one or more of the patient's age, ethnicity, and medical history. The determination of the threshold sensitivity value can also be based on prior patient-specific functional tests, such as prior VF test results and / or prior (visual) evoked potential test data.

[0048] To reiterate, the thus determined (e.g., predicted / derived) visual sensitivity value VFTh_out can be submitted to perimeter 13, which, when applying a selected functional VF test to a patient, can base its initial VF test point value (e.g., intensity and / or size input prior) on one or more corresponding selected test points (or otherwise optimize its VF test). That is, the derived sensitivity VFTh_out can be modified in the prior construction. For example, the selected VF test can begin using an input prior that has an offset (e.g., higher or lower intensity) from the derived sensitivity VFTh_out.

[0049] Alternatively or additionally, the determined or estimated threshold sensitivity value can be used as a VF prior and / or for determining the patient's field of vision, which can be used for diagnostic / clinical interpretation or structural-functional analysis. For example, the patient's predicted field of vision can be used as part of a clinical decision support (CDS) system that provides clinicians, staff, patients, or other individuals with knowledge and personalized information, intelligently filtered or presented at appropriate times to enhance healthcare. This system can be included as an additional tool within a CDS system to enhance decision-making in clinical workflows. For example, the system can provide computerized alerts and reminders to care providers and patients, and offer clinical guidelines, condition-specific order sets (e.g., recommendations for field of vision testing or other medical tests), centralized patient data reports and summaries, document templates, diagnostic support, and context-sensitive reference information. For example, the currently derived sensitivity VFTh_out can be compared with one or more previously derived sensitivity results and / or actual field of vision test results (e.g., from previous physician visits), and a warning sign / message can be issued when the currently derived sensitivity VFTh_out indicates that the patient's field of vision has changed beyond a predetermined range and / or predetermined area and / or predetermined rate of change. Warning signs / messages can indicate that a patient should be scheduled for a real field of vision test.

[0050] Machine learning models15 can be based on one or more of the following: linear regression, logistic regression, decision trees, support vector machines, Naive Bayes, k-nearest neighbors, k-means, random forests, dimensionality reduction, gradient boosting, and neural networks. Typically, a machine learning model is a computational system that can be trained to perform a specific function, and the choice of a particular model can depend on the type of problem being solved. For example, a support vector machine (SVM) is a linear machine learning model used for classification and regression problems, and can be used to solve both linear and nonlinear problems. The idea behind an SVM is to create a line, or hyperplane, that separates the data into classes. More formally, an SVM defines one or more hyperplanes in a multidimensional space, where hyperplanes are used for classification, regression, outlier detection, etc. Essentially, an SVM model represents labeled training examples as points in a multidimensional space, mapping them such that labeled training examples of different classes are separated by a hyperplane, which can be thought of as a decision boundary separating different classes. When a new test input sample is submitted to the SVM model, the test input is mapped to the same space, and a prediction is made about which class it belongs to based on which side of the decision boundary (hyperplane) the test input lies on.

[0051] However, in a preferred embodiment of the invention, the machine learning model 15 is at least partially implemented in a computing system that includes / contains a trainable neural network capable of deep learning. Reference is made below. Figures 13 to 16 Various examples of neural networks are discussed, any one of which, individually or in combination, can be used in this invention.

[0052] For the purpose of explanation, Figure 2An example of training a neural network NN-1 according to the present invention is shown. For ease of discussion, each training set is shown as consisting of training pairs TP1 to TPn, and in this example, each training pair may include OCT-based images / scans OCT1 to OCTn (e.g., OCT angiography data and / or structural OCT data) as training input samples, which are paired with corresponding, labeled field-of-view test result data VFTR1 to VFTRn, which are collected from the same “test patient” from whom the training input images were obtained, and preferably on approximately the same date on which the training input images were collected. However, as mentioned above, in addition to, or instead of, OCT-based images, training (data) input samples may include fundus images, fluorescein angiography images, (visual) evoked potential test results and other objective field-of-view examination results (multifocal electroretinography ERG / PERF, pupillary measurements, etc.), individual retinal structural measurements, previously diagnosed pathologies (e.g., medical conditions, medical treatments and / or other medical records), physical characteristics of the test patient (e.g., age, race, medical history), standard structural data of the test patient demographics (retinal nerve fiber layer (RNFL) thickness and ganglion cell-inner plexiform layer (GCIPL) thickness), standard functional data of the test patient demographics (e.g., standardized initialization parameters for a specific field-of-view test), etc. Training (data) input may further include prior field-of-view test results (e.g., real or previously synthesized / derived functional VT test results / visual sensitivity measurements) and the date they were taken to help identify trends in the rate of change of field-of-view results associated with specific characteristics (e.g., pathology) of the test patient. It should be understood that these prior field-of-view test results used for training inputs can be based on SAP and / or they can be based on objective field-of-view examinations (VEP, multifocal ERG / PERF, pupil measurement, etc.). For ease of illustration, the OCT-based data for training inputs is shown as depth-encoded frontal panels / images; however, it should be understood that the OCT data for training inputs can be volume data, B-scan, or A-scan. In this example, a neural network NN-1 is trained to determine (or derive) VF priors, such as field-of-view thresholds (e.g., intensity and / or size thresholds for corresponding field-of-view test points for a given field-of-view test type), and thus its field-of-view training target outputs VFTR1 to VFTRn are illustratively shown as labeled, real, functional field-of-view test results (e.g., dark and bright squares and / or digital field-of-view threshold results for corresponding test points along the test field-of-view distribution).In this example, the neural network NN1 is trained to extract field-of-view threshold data from full OCT-based image information; therefore, the training input in each training pair is shown as including full scan information OCT1 to OCTn. Optionally, data augmentation methods can be used to increase the size of the training dataset, for example, by dividing each test input data (OCT1 to OCTn) into smaller data segments (or image / scan pieces), where the pieces can have similar or different sizes. Generally, a larger training set size provides better training results.

[0053] Figure 3 It shows Figure 2 This describes an exemplary operation of training a neural network NN-1, wherein real-time data input is used after training, or test data input is used during the evaluation phase of the training session. The trained neural network NN-1 can include one or more of fully connected neural networks, convolutional neural networks, feedforward neural networks, recurrent neural networks, modular neural networks, and U-Net, as discussed more fully below. This neural network NN-1 can receive acquired image data (e.g., real-time images from an OCT system or fundus imager, or access previously collected images, such as those from a patient's medical records, which can be remotely stored), as input to OCT-in (which may also optionally specify the type of field of view test to be performed on the patient if multiple VF test types are supported), and predict (e.g., determine / synthesize / generate) a corresponding field of view threshold output VFTh_out with prediction thresholds for one or more test points of the specified field of view test type. Figure 1 As shown, the output VFTh_out can be submitted to box 13 for performing a functional VF test on the patient. Note that the input image OCT-in is not an image used during training, nor is it derived from any images used during training. That is, image data that the network NN-1 has not previously seen (e.g., OCT-in) is selected for the testing / evaluation / operation phase. Optionally, during operation, the network NN-1 does not receive any previous real (functional) field of vision test results from the patient as input.

[0054] Figure 4 An alternative training architecture linking multiple neural network levels (Stg1 and Stg2) is shown, with each level comprising its own neural network in a modular neural network configuration. The first level, Stg1, of this architecture is similar to... Figure 2 The first stage, Stg1, can consist of neural networks optimized for image processing, such as convolutional neural networks and / or U-Net. Figure 4 Zhongyu Figure 2All similar elements have similar reference numerals and have been discussed above. In this example, the output from the first stage Stg1 is fed into a second neural network NN-2, which can be optimized to process individual data units (as opposed to images) and can consist of, for example, fully connected neural networks, feedforward neural networks, and / or recurrent neural networks. The input to the second stage Stg2 may not include images and may include separate datasets (e.g., contextual data), such as standard data, patient medical record data, separate biometric information, previous (real or synthetic) field-of-view thresholding results, etc. In operation (e.g., after training), from Figure 4 The predicted VF threshold (not shown) for the architecture can be submitted to Figure 1 The field of view meter 13 is used to manage functional field of view testing by using the predicted VF threshold as the starting test point value and / or prior.

[0055] Note that previous field-of-view (FOV) testing results can help identify trends in patient FOV changes, potentially leading to more accurate predictions. However, because FOV testing has historically been time-consuming and not always performed at prescribed (e.g., regular) intervals, gaps may exist in patient FOV testing results. Therefore, there may not be sufficient data to determine trends or tendencies in patient FOV changes. This system addresses this problem by providing synthesized / derived FOV tests to fill these gaps. For example, while a patient may have skipped FOV testing during a specific outpatient visit (or a specific month / time), they may have already had retinal images (e.g., OCT, OCTA, fundus images, FA, etc.) taken during an outpatient visit (or within a predetermined timeframe, e.g., month or other set number of weeks / days). In this case, the taken retinal images can be used to extract a derived FOV. This derived FOV can then be used in place of the true functional FOV in VF-related analyses. For example, such a derived FOV can be used to create additional training sets in additional training sessions of a neural network (e.g., used as the VF target output VFTRi in a specific training pair TPi, such as...). Figure 2 (as shown) or train another neural network. That is, in Figure 2 and / or Figure 4 In the training configuration, the exported field of view can be used as training data (instead of the previously captured real functional field of view results or as a supplement to them).

[0056] Figure 5 An example of the derived field of view for VF correlation analysis is provided. Figure 5The example graph illustrates a patient's declining field of vision (VF) sensitivity over time and demonstrates how having a patient's VF testing history can help predict (foreshadow) the patient's current or future field of vision (e.g., predict VF sensitivity measurements, such as based on thresholds at a given test point). The vertical axis can correspond to measurements of the patient's visual sensitivity, and the horizontal axis can correspond to the passage of time, such as a series of prescribed VF testing dates or scheduled outpatient visits. In this example, actual prior VF test results are shown as solid dots, and synthetic (derived) VF results from the patient's previous outpatient visits (e.g., based on biometric information or other non-traditional functional field of vision data) are shown as circles. The graph of prior VF test sensitivity results versus time helps illustrate the patient's predicted threshold at time "x". Such prediction is impossible using only the actual prior VF test results (solid dots), which would indicate linear progression, as shown by the dashed line Ln1. However, adding synthetic VF results (circles) as additional "VF priors" to fill the gaps in the test time reveals a more logarithmic or curvilinear graph (indicated by the dashed line Crv1), which better predicts the future VF value at time "x". That is, a set of derived and actual visual fields can be input into a VF prediction system that uses the input to predict the patient's current or future visual field. Such a VF prediction system can be embodied by a computational system implementing any number of prediction techniques, such as machine learning (e.g., linear regression) and / or deep learning (e.g., recurrent neural networks).

[0057] Figure 6 A VF prediction system 21 according to the present invention is illustrated. In this example, to better predict the field of view VFTh_out for subsequent time slot TS10, nine time slots / intervals TS1 to TS9 require VF priors. In this example, for time slots TS1, TS3, TS4, TS6, ST7, and ST8, real VF test results are available, but due to gaps in the VF history, no real VF test results are available for time slots TS2, TS5, and TS9. Assuming the patient has image data (biometrics / physical measurements) corresponding to the missing time slots (e.g., the patient underwent retinal imaging / scanning but did not have a VF test in the specified time slot), this system can be used to synthesize VF priors for the missing time slots TS2, TS5, and TS9. The set of real and synthesized VF priors can be submitted to the prediction tool 21 (sequentially or in parallel), and the prediction tool can then output the predicted field of view VFTh_out for time slot TS10. The output VFTh_out can be used as... Figure 1 The exported field of view VFTh_out is submitted to box 13.

[0058] A preliminary proof-of-concept study was conducted to evaluate the performance of using a structure-derived field-of-view prior (S-prior) to simulate the field of view (VF). This study used qualified (e.g., retrospective) data from 1399 subjects (monocular) from a Singapore population study. Humphrey Field Analyzer (HFA2i) data were collected during the study visits. (ZEISS, Dublin, CA) SITA Standard 24-2 VF and The HD-OCT (ZEISS, Dublin, CA) data included optical cubes. 70% of the eyes were used to train a regressor (e.g., a random forest regressor) to predict 54-point VF. A random forest (RF) was constructed using 256-point perretinal papillary nerve fiber layer data and age. A simplified mixed-scale dense convolutional neural network (CNN) was constructed using RNFL thickness maps, see, for example, “A Mixed-ScaleDense Convolutional Neural Network for Image Analysis” by Pelt et al., PNAS, 2018, 115(2), 254-259, the entire contents of which are incorporated herein by reference. The remaining 30% of the eyes were used to predict the S prior and provide the input field to the VF simulator.

[0059] The VF simulator implements Bayesian ZEST using a bimodal initial probability distribution (SPD) with no prior (ZEST), as described in “Targeted Spatial Sampling Using GOANNA Improves Detection of Visual Field Progression” (Chong et al., Ophthalmic Physiol Opt, March 2015, 35(2): 155-69), except that the normal mode is centered on age normal values ​​determined from a normal set of 118 eyes, as described in “Exploring the Structure-Function Relationship for Perimetry Stimulus Sizes III, V and VI and OCT in Early Glaucoma” by Flanagan et al., ARVO (Association for Research in Vision and Ophthalmology) Abstract, Investigative Ophthalmology & Visual Science (IOVS), September 2016, Vol. 57, 376, the entire contents of which are incorporated herein by reference.

[0060] ZEST (e.g., ZEST-RF, ZEST-CNN) was also simulated using a single-mode SPD designed for custom priors centered on two S-priors. The slope of the visual response frequency was modeled as described in “Response Variability in the Visual Field: Comparison of Optic Neuritis, Glaucoma, Ocular Hypertension, and Normal Eyes” (Henson et al., IOVS, February 2000, Vol. 41, pp. 417-421), the entire contents of which are incorporated herein by reference. False response rates were set at 0%, 5%, and 20%, respectively, for three types of respondents. Performance between simulated (e.g., synthesized) and real VF was evaluated by observing the mean absolute error (MAE) and the total number of questions between the simulated (e.g., synthetic) and real VF. The two locations closest to the blind spot were excluded from the analysis. Significance tests for inter-policy equivalence and ZEST were performed using a consistency limit of ±5% dB of MAE and ±5% of the total number of questions (two-sided paired t-tests, α = 0.05).

[0061] The results showed that the mean VF MD of the training set and the test set were -1.8±2.4dB and -2.7±2.7dB, respectively (p<0.001). Figure 7A and Figure 7B Is it using random forest ( Figure 7A ) and neural networks Figure 7B This is a graph showing the (derived) OCT estimated threshold versus the (true) VF threshold for an example application of the machine learning model. Because this is a proof-of-concept application, the availability of training data is limited, especially for certain thresholds. In each graph, the vertical line VL provides a visual indicator separating the region RA (e.g., at lower thresholds) with less training data from the region RB (e.g., at more normal thresholds) with more training data. It can be understood that the target line t1 indicates the desired distribution / trend to indicate the equivalence between the true and derived thresholds. Figure 7A and Figure 7B Both studies show that this simple model performs better in region RB (e.g., the plotted data distribution follows the target line TL better), where more training data is available than in region RA (e.g., this simple model performs better at more normal thresholds than at lower thresholds). Providing additional training data, especially at lower thresholds, will likely improve the current model and provide better results. In any case, Figure 7B This shows that (deep learning) neural network (CNN) models can achieve better results than random forest (RF) models (e.g., the plotted data follows the target line TL better).

[0062] However, Figure 8 Table 1 shows that the overall MAE of ZEST-RF and ZEST-CNN is statistically comparable to ZEST (p<0.001). ZEST-CNN reduced the total number of problems by 16-19% compared to ZEST. These findings suggest that even a simple model with limited / imbalanced data, which predicts VF based on biometric / structural data (e.g., OCT data and / or fundus images), can reduce the duration of initial VF examinations in this population with comparable error. As more data representing the clinical population becomes available and the models become more refined, performance is likely to improve further.

[0063] The following provides a description of various hardware and architectures applicable to this invention.

[0064] Field of view testing system

[0065] The improvements described in this article can be used with any type of field of view tester / system, such as a field of view meter. One such system is the "bowl" field of view tester VF0, such as... Figure 9 As shown. The subject (e.g., a patient) VF1 is shown observing a hemispherical projection screen (or other type of display) VF2, typically shaped like a bowl, and the tester VF0 is therefore referred to as the bowl. Typically, the subject is instructed to gaze at a point in the center of the hemispherical screen VF3. The subject rests his / her head on a patient support, which may include a chin support VF12 and / or a forehead support VF14. For example, the subject rests his / her head on the chin support VF12 and his / her forehead on the forehead support VF14. Optionally, the chin support VF12 and the forehead support VF14 may move together or independently of each other to properly fix / position the patient's eyes, for example, relative to the test lens holder VF9, which can hold the lens through which the subject can view the screen VF2. For example, the chin support and the forehead support may move independently in the vertical direction to accommodate different patient head sizes and move together in the horizontal and / or vertical directions to properly position the head. However, this is not limiting, and those skilled in the art will conceive of other arrangements / movements.

[0066] Under the control of processor VF5, a projector or other imaging device VF4 displays a series of test stimuli (e.g., test points of any shape) VF6 on screen VF2. Subject VF1 instructs him / her to see stimulus VF6 by initiating user input VF7 (e.g., pressing an input button). This subject response can be recorded by processor VF5, which can be used to assess the eye's field of vision based on the subject's response, for example, determining the size, location, and / or intensity of test stimulus VF6 that subject VF1 can no longer see, thereby determining the (visibility) threshold of test stimulus VF6. Camera VF8 can be used to capture the patient's gaze (e.g., gaze direction) throughout the test. Gaze direction can be used to align the patient and / or determine whether the patient is following the correct test procedure. In this example, camera VF8 is located on the Z-axis relative to the patient's eye (e.g., relative to the test lens holder VF9) and behind the bowl (of screen VF2) to capture real-time images or videos of the patient's eye. In other embodiments, the camera may be located outside this Z-axis. Images from the gaze camera VF8 can optionally be displayed on a second display VF10 to a clinician (or, interchangeably, a technician) to assist the patient in alignment or test verification. The camera VF8 can record and store one or more images of the eye during each stimulus presentation. Depending on the testing conditions, this may result in the collection of dozens to hundreds of images per field of vision test. Alternatively, the camera VF8 can record and store a full-length film during the test, providing timestamps indicating when each stimulus occurred. Furthermore, images can be collected between stimulus presentations to provide details of the subject's overall attention throughout the entire duration of the VF test.

[0067] The test lens holder VF9 can be placed in front of the patient's eye to correct any refractive errors in the eye. Optionally, the lens holder VF9 can carry or hold a liquid test lens (e.g., see U.S. Patent No. 8,668,338, the entire contents of which are incorporated herein by reference), which can be used to provide variable refractive correction for the patient's VF1. However, it should be noted that the invention is not limited to using a liquid test lens for refractive correction, and other conventional / standard test lenses known in the art can also be used.

[0068] In some embodiments, one or more light sources (not shown) may be located in front of the subject's VF1 eye, reflecting off the surface of the eye (e.g., the cornea). In one variation, the light source may be a light-emitting diode (LED).

[0069] Although Figure 9A projection-type field-of-view tester VF0 is shown, but the invention described herein can be used with other types of devices (field-of-view testers), including those that generate images via liquid crystal displays (LCDs) or other electronic displays (see, for example, U.S. Patent No. 8,132,916, incorporated herein by reference). Other types of field-of-view testers include, for example, flat panel screen testers, miniaturized testers, and binocular field-of-view testers. Examples of these types of testers are found in U.S. Patent Nos. 8,371,696, 5,912,723, 8,931,905, and U.S. Design Patent D472,637, the entire contents of each of which are incorporated herein by reference.

[0070] The field of vision tester VF0 may include an instrument control system (e.g., an algorithm, which may be software, code, and / or routines) that uses hardware signals and a motorized positioning system to automatically position the patient's eyes in a desired location, such as the center of the refractive lens at the lens retainer VF9. For example, stepper motors can move the chin support VF12 and forehead support VF14 under software control. A rocker switch may be provided to allow the attending technician to adjust the patient's head position by operating the chin support and forehead stepper motors. Manually movable refractive lenses may also be placed in front of the patient's eyes on the lens holder VF9, as close to the patient's eyes as possible without adversely affecting patient comfort. Optionally, if such movement would interrupt test execution, the instrument control algorithm may pause the visual field test while the chin support and / or forehead motor movement is in progress.

[0071] Fundus imaging system

[0072] Two types of imaging systems used for fundus imaging are flood illumination imaging systems (or flood illumination imagers) and scanning illumination imaging systems (or scanning imagers). A flood illumination imager, for example, uses a flash lamp to simultaneously flood the entire field of interest (FOV) of the sample with light and captures a full-frame image of the sample (e.g., the fundus) with a full-frame camera (e.g., a camera with a sufficiently large two-dimensional (2D) light sensor array to capture the desired FOV as a whole). For example, a flood illumination fundus imager will flood the fundus and capture a full-frame image of the fundus in a single image capture sequence from the camera. A scanning imager provides a scanning beam that scans across an object (e.g., the eye), and as the scanning beam scans across the object, it images at different scanning locations, producing a series of image fragments that can be reconstructed (e.g., synthesized) to produce a composite image of the desired FOV. The scanning beam can be a point, a line, or a two-dimensional region, such as a slit or a wide line.

[0073] Figure 10An example of a slit-scanning ophthalmic system SLO-1 for imaging the fundus F is shown. The fundus F is the inner surface of the eye E opposite the lens (or optic disc) CL and may include the retina, optic disc, macula, fovea, and posterior pole. In this example, the imaging system is in a so-called “scan-to-de-scan” configuration, wherein a scan line beam SB passes through the optical components of the eye E (including the cornea Crn, iris Irs, pupil Ppl, and lens CL) to scan the fundus F. In the case of a floodlight fundus imager, a scanner is not required, and light is applied immediately across the entire desired field of view (FOV). Other scanning configurations are known in the art, and the specific scanning configuration is not important to the present invention. As shown, the imaging system includes one or more light sources LtSrc, preferably a multicolor LED system or a laser system, wherein the light collection rate has been appropriately adjusted. An optional slit Slt (adjustable or static) is located in front of the light source LtSrc and can be used to adjust the width of the scan line beam SB. Furthermore, the slit Slt can remain stationary during imaging or can be adjusted to different widths to allow for different levels of confocality and different applications, for specific scans, or to suppress reflections during scanning. An optional objective lens ObjL can be placed in front of the slit Slt. The objective lens ObjL can be any existing lens, including but not limited to refractive, diffractive, reflective, or hybrid lenses / systems. Light from the slit Slt passes through the pupil splitter SM and is directed to the scanner LnScn. It is desirable to bring the scanning plane and the pupil plane as close together as possible to reduce vignetting in the system. Optional optics D1 can be included to manipulate the optical distance between the images of the two components. The pupil splitter SM transmits the illumination beam from the light source LtSrc to the scanner LnScn and reflects the detection beam from the scanner LnScn (e.g., reflected light returning from the eye E) toward the camera Cmr. The task of the pupil splitter SM is to split the illumination and detection beams and help suppress system reflections. The scanner LnScn can be a rotating galvanometer scanner or other types of scanners (e.g., piezoelectric or voice coil, microelectromechanical systems (MEMS) scanners, electro-optic deflectors, and / or rotating polygon scanners). Depending on whether pupil splitting occurs before or after the scanner LnScn, the scanning can be divided into two steps, where one scanner is in the illumination path and the other is in the detection path. A specific pupil splitting setup is described in detail in U.S. Patent No. 9,456,746, the entire contents of which are incorporated herein by reference.

[0074] From the scanner LnScn, an illumination beam passes through one or more optics, in this case a scanning lens SL and an ophthalmic or eyepiece OL, which allow the pupil of the eye E to image onto the system's image pupil. Typically, the scanning lens SL receives the scanning illumination beam from the scanner LnScn at any of a plurality of scanning angles (incident angles) and produces a scanning line beam SB with a substantially flat surface focal plane (e.g., a collimated optical path). The ophthalmic lens OL can focus the scanning line beam SB onto the fundus F (or retina) of the eye E and image the fundus. In this way, the scanning line beam SB produces a transverse scanning line across the fundus F. One possible configuration of these optics is a Keplerian telescope, in which the distance between two lenses is chosen to create an approximately telecentric intermediate fundus image (4-f configuration). The ophthalmic lens OL can be a single lens, an achromatic lens, or an arrangement of different lenses. As those skilled in the art will know, all lenses can be refractive, diffractive, reflective, or a combination of these. The focal lengths of the ophthalmic lens (OL), scanning lens (SL), and the dimensions and / or forms of the pupillary divider (SM) and scanner (LnScn) can vary depending on the desired field of view (FOV). Therefore, an arrangement can be envisioned where multiple components can switch in and out of the optical path, for example, by using triggers, motorized wheels, or detachable optical elements, depending on the FOV. Since changes in the FOV result in different beam sizes across the pupil, pupillary splitting can also change with the FOV. For example, a 45° to 60° field of view is the typical or standard FOV for fundus cameras. Higher fields of view (e.g., 60°–120° or higher wide field of view FOVs) are also feasible. Wide field of view FOVs can be ideal for combining wide-line fundus imaging (BLFI) with another imaging modality (e.g., optical coherence tomography (OCT)). The upper limit of the field of view can be determined by the accessible working distance combined with the physiological conditions surrounding the human eye. Because the typical human retina has a field of view (FOV) of 140° horizontally and 80°–100° vertically, it may be desirable to have an asymmetrical field of view for the highest possible FOV on the system.

[0075] The scanning line beam SB passes through the pupil Ppl of the eye E and is directed to the retina or fundus surface F. The scanner LnScn1 adjusts the position of the light on the retina or fundus F to illuminate a series of lateral positions on the eye E. Reflected or scattered light (or emitted light in the case of fluorescence imaging) is guided back along a similar path to the illumination to define the collection beam CB on the detection path to the camera Cmr.

[0076] In the "scan-de-scan" configuration of this exemplary slit-scan ophthalmic system SLO-1, the light returning from the eye E is "de-scanned" by the scanner LnScn on its way to the pupillary segmentation mirror SM. That is, the scanner LnScn scans the illumination beam from the pupillary segmentation mirror SM to define the scan illumination beam SB passing through the eye E; however, since the scanner LnScn also receives the returning light from the eye E at the same scanning position, it has the effect of removing the returning light (e.g., canceling the scan action) to define a non-scanning (e.g., stable or stationary) collected beam from the scanner LnScn to the pupillary segmentation mirror SM, which folds the collected beam toward the camera Cmr. At the pupillary segmentation mirror SM, reflected light (or, in the case of fluorescence imaging, emitted light) is separated from the illumination beam onto a detection path pointing toward the camera Cmr, which may be a digital camera with a light sensor to capture an image. An imaging (e.g., objective) lens ImgL may be located in the detection path to image the fundus onto the camera Cmr. Similar to the objective lens ObjL, the imaging lens ImgL can be any type of lens known in the art (e.g., refractive, diffractive, reflective, or hybrid lens). Additional operational details, particularly methods for reducing artifacts in images, are described in PCT Publication No. WO2016 / 124644, the entire contents of which are incorporated herein by reference. The camera Cmr captures the received images, for example, creating an image file, which can be processed by one or more (electronic) processors or computing devices (e.g., ...). Figure 17 The computer system shown further processes the data. Therefore, the collected beam (returning from all scan positions of the scan line beam SB) is collected by the camera Cmr, and the full-frame image Img can be composed of a synthesis of the individually captured collected beams, for example, through synthesis. However, other scanning configurations are also conceivable, including configurations where the illumination beam scans on the eye E and the collected beam scans on the camera's light sensor array. Several embodiments of a slit scanning ophthalmoscope, including various designs where the returned light sweeps across the camera's light sensor array and where the returned light does not sweep across the camera's light sensor array, are described by reference to PCT Publication WO2012 / 059236 and U.S. Patent Publication No. 2015 / 0131050, which are incorporated herein by reference.

[0077] In this example, the camera Cmr is connected to a processor (e.g., a processing module) Proc and a display (e.g., a display module, computer screen, electronic screen, etc.) Dspl. Both can be part of the imaging system itself, or they can be part of separate dedicated processing and / or display units, such as a computer system, where data is transmitted from the camera Cmr to the computer system via cable or a computer network including wireless networks. The display and processor can be integrated. The display can be a conventional electronic display / screen or a touchscreen type, and can include a user interface for displaying and receiving information to and from the instrument operator or user. The user can interact with the display using any type of user input device known in the art, including but not limited to a mouse, knob, button, pointer, and touchscreen.

[0078] When imaging, it may be desirable to keep the patient's gaze fixed. One way to achieve this is to provide a fixed target that the patient can be guided to look at. The fixed target can be inside or outside the instrument, depending on which area of ​​the eye is being imaged. Figure 10 An embodiment of an internal fixation target is illustrated. In addition to the main light source LtSrc for imaging, a second optional light source FxLtSrc, such as one or more LEDs, can be positioned such that a light pattern is imaged onto the retina using a lens FxL, a scanning element FxScn, and a reflector / mirror FxM. The fixation scanner FxScn can move the position of the light pattern, and the reflector FxM guides the light pattern from the fixation scanner FxScn to the fundus F of the eye E. Preferably, the fixation scanner FxScn is positioned such that it is located in the pupillary plane of the system, allowing the light pattern on the retina / fundus to move according to the desired gaze position.

[0079] By selecting filtering elements based on the light source and wavelength used, the slit-lamp ophthalmoscope system can operate in different imaging modes. When imaging the eye with a series of colored LEDs (red, blue, and green), true-color reflective imaging (similar to the imaging observed by clinicians when examining the eye with a handheld or slit-lamp ophthalmoscope) can be achieved. The image for each color can be built progressively with each LED turned on at each scanning position, or each color image can be captured individually as a whole. These three color images can be combined to display a true-color image or displayed individually to highlight different features of the retina. The red channel best highlights the choroid, the green channel highlights the retina, and the blue channel highlights the anterior retina. Furthermore, light of specific frequencies (e.g., individual colored LEDs or lasers) can be used to excite different fluorophores in the eye (e.g., autofluorescence), and the resulting fluorescence can be detected by filtering out the excitation wavelength.

[0080] Fundus imaging systems can also provide infrared reflection images, for example, by using an infrared laser (or other infrared light source). The advantage of infrared mode is that the eye is not sensitive to infrared wavelengths. This allows users to capture images continuously without disturbing the eye (e.g., in preview / alignment mode) to assist the user during instrument alignment. Furthermore, infrared wavelengths increase penetration into tissues and can provide improved visibility of choroidal structures. Additionally, fluorescein angiography (FA) and indocyanine green (ICG) angiography imaging can be accomplished by collecting images after a fluorescent dye has been injected into the subject's bloodstream. For example, in FA (and / or ICG), a series of time-lapse images can be captured after a photoactive dye (e.g., a fluorescent dye) has been injected into the subject's bloodstream. It is important to note that caution must be exercised because fluorescent dyes can cause life-threatening allergic reactions in certain individuals. High-contrast grayscale images are captured by exciting the dye using a selected specific light frequency. As the dye flows through the eye, the corresponding part of the eye emits bright light (e.g., fluorescence), allowing the progress of the dye to be seen, thus revealing the flow of blood through the eye.

[0081] Optical coherence tomography system

[0082] In addition to fundus photography, fundus autofluorescence (FAF), and fluorescein angiography (FA), ophthalmic images can be created using other imaging modalities, such as optical coherence tomography (OCT), OCT angiography (OCTA), and / or ocular ultrasound. This invention, or at least a portion thereof, as understood in the art, with minor modifications, can be applied to these other ophthalmic imaging modalities. More specifically, this invention can also be applied to ophthalmic images generated by OCT / OCTA systems that produce OCT and / or OCTA images. For example, this invention can be applied to OCT / OCTA images. Examples of fundus imagers are provided in U.S. Patents 8,967,806 and 8,998,411, examples of OCT systems are provided in U.S. Patents 6,741,359 and 9,706,915, and examples of OCTA imaging systems are provided in U.S. Patents 9,700,206 and 9,759,544, the entire contents of which are incorporated herein by reference. For completeness, exemplary OCT / OCTA systems are provided herein.

[0083] Figure 11A generalized frequency-domain optical coherence tomography (FD-OCT) system suitable for collecting three-dimensional image data of the eye is illustrated. The FD-OCT system OCT_1 includes a light source LtSrc1. Typical light sources include, but are not limited to, broadband light sources or scanning laser sources with short time coherence lengths. The beam from the light source LtSrc1 is typically guided by an optical fiber Fbr1 to illuminate a sample (e.g., the eye E); the sample is typically tissue in the human eye. The light source LtSrc1 can be a broadband light source with a short time coherence length in the case of spectral domain OCT (SD-OCT), or a wavelength-tunable laser source in the case of scanning source OCT (SS-OCT). The light can typically be scanned by a scanner Scnr1 located between the output of the optical fiber Fbr1 and the sample E, such that the beam (dashed line Bm) is scanned laterally (in the x and y directions) over the area of ​​the sample to be imaged. In the case of full-field-of-view OCT, a scanner is not required, and the light travels through the entire desired field of view (FOV) at once. The light scattered from the sample is collected and typically enters the same optical fiber Fbr1 used to guide the light for illumination. Reference light from the same light source LtSrc1 propagates along a separate path, in this case including fiber Fbr2 and a retroreflector RR1 with adjustable optical delay. Those skilled in the art will recognize that a transmission reference path can also be used, and the adjustable delay can be placed in the sample or reference arm of the interferometer. The collected sample light is typically combined with the reference light in fiber coupler Cplr1 to form optical interference in the OCT photodetector Dtctr1 (e.g., a photodetector array, digital camera, etc.). Although a single fiber port leading to detector Dtctr1 is shown, those skilled in the art will recognize that various designs of the interferometer can be used for balanced or unbalanced detection of the interference signal. The output from detector Dtctr1 is provided to processor Cmp1 (e.g., a computing device), which converts the observed interference into depth information of the sample. The depth information can be stored in a memory associated with processor Cmp1 and / or displayed on display Scn1 (e.g., a computer / electronic display / screen). The processing and storage functions can be located within the OCT instrument or in an external processing unit (e.g., Figure 17 The functions performed on the computer system shown are transmitted to the external processing unit. This unit can be dedicated to data processing or performing other very general tasks, rather than being dedicated to the OCT device. The processor Cmp1 may include, for example, a field-programmable gate array (FPGA), a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a graphics processing unit (GPU), a system-on-a-chip (SoC), a central processing unit (CPU), a general-purpose graphics processing unit (GPGPU), or a combination thereof, which performs some or all of the data processing steps before being transmitted to the main processor or in parallel.

[0084] The sample and reference arms in the interferometer can be composed of bulk optics, fiber optics, or hybrid bulk optics systems, and can have different architectures, such as Michelson, Mach-Zehnder, or common-path-based designs known to those skilled in the art. The beams used herein should be interpreted as any carefully guided optical path. Instead of a mechanical scanning beam, an optical field can illuminate a one-dimensional or two-dimensional region of the retina to generate OCT data (e.g., see U.S. Patent 9,332,902; D. Hillmann et al., “Holoscopy–Holographic Optical Coherence Tomography”, Optics Letters 36(13): 23902011; Y. Nakamura et al., “High-Speed ​​Three-Dimensional Human Retinal Imaging by Line Field Spectral Domain Optical Coherence Tomography”, Optics Express 115(12): 7Fbr2 2007; Blazkiewicz et al., “Signal-To-Noise Ratio Study of Full-Field Fourier-Domain Optical Coherence Tomography”, Applied Optics 44(36): 7722(2005)). In time-domain systems, the reference arm needs to have an adjustable optical delay to generate interference. Balanced detection systems are typically used in TD-OCT and SS-OCT systems, while spectrometers are used at the detection port of SD-OCT systems. The invention described herein can be applied to any type of OCT system. Various aspects of this invention can be applied to any type of OCT system or other types of ophthalmic diagnostic systems and / or multiple ophthalmic diagnostic systems, including but not limited to fundus imaging systems, field-of-view testing devices, and scanning laser polarimeters.

[0085] In Fourier domain optical coherence tomography (FD-OCT), each measurement is a real-valued spectral interferogram (Sj(k)). The real-valued spectral data typically undergoes several post-processing steps, including background subtraction and dispersion correction. The Fourier transform of the processed interferogram produces a complex-valued OCT signal output. The absolute value |Aj| of this complex OCT signal reveals the scattering intensity distribution for different path lengths, thus scattering is a function of depth (z-direction) in the sample. Similarly, phase... It can also be extracted from complex-valued OCT signals. The scattering profile as a function of depth is called an axial scan (A-scan). A set of A-scans measured at adjacent locations in a sample produces a cross-sectional image (tomogram or B-scan) of the sample. A set of B-scans acquired at different lateral locations on the sample constitutes a data volume or cube. For a given amount of data, the term fast axis refers to the scanning direction along a single B-scan, while slow axis refers to the axis along which multiple B-scans are collected. The term "cluster scan" can refer to a single cell or data block generated by repeated acquisition at the same (or substantially the same) location (or region) for analyzing motion contrast, which can be used to identify blood flow. A cluster scan can consist of multiple A-scans or B-scans acquired at substantially the same location on the sample at relatively short time intervals. Because the scans in a cluster scan belong to the same region, the static structure remains relatively unchanged between scans, while the motion contrast between scans that meet predetermined criteria can be identified as blood flow. Various methods for generating B-scans are known in the art, including but not limited to: along the horizontal or x-direction, along the vertical or y-direction, along the diagonals of x and y, or in circular or spiral patterns. B-scans can be in the xz dimension, but can be any cross-sectional image including the z dimension.

[0086] In OCT angiography or functional OCT, analytical algorithms can be applied to OCT data collected at the same or nearly the same sample location on the sample at different times (e.g., cluster scans) to analyze motion or flow (see, for example, U.S. Patent Publications 2005 / 0171438, 2012 / 0307014, 2010 / 0027857, 2012 / 0277579 and U.S. Patent 6,549,801, the entire contents of which are incorporated herein by reference). OCT systems can use any of a variety of OCT angiography processing algorithms (e.g., motion contrast algorithms) to identify blood flow. For example, motion contrast algorithms can be applied to intensity information derived from image data (intensity-based algorithms), phase information derived from image data (phase-based algorithms), or complex image data (complexity-based algorithms). A frontal image is a 2D projection of 3D OCT data (e.g., by averaging the intensity of each individual A-scan, such that each A-scan defines pixels in the 2D projection). Similarly, a frontal vascular system image is an image displaying motion-contrast signals, where the data dimension corresponding to depth (e.g., along the z-direction of the A-scan) is displayed as a single representative value (e.g., a pixel in a 2D projected image), typically displayed by summing or integrating all or isolated portions of the data (e.g., see U.S. Patent No. 7,301,644, the entire contents of which are incorporated herein by reference). An OCT system providing angiographic imaging capabilities may be referred to as an OCT angiography (OCTA) system.

[0087] Figure 12 An example of a frontal vascular system image is shown. After processing the data using any motion contrast technique known in the art to highlight motion contrast, a range of pixels corresponding to a given tissue depth from the surface of the internal limiting membrane (ILM) in the retina can be summed to generate a frontal (e.g., frontal view) image of the vascular system.

[0088] Neural Networks

[0089] As described above, the present invention can utilize neural network (NN) machine learning (ML) models. For completeness, a general discussion of neural networks is provided herein. The present invention can use any of the neural network architectures described below, either alone or in combination. A neural network or neural network is a network of (nodes) composed of interconnected neurons, where each neuron represents a node in the network. Groups of neurons can be arranged hierarchically, and in a multilayer perceptron (MLP) arrangement, the output of one layer is fed forward to the next layer. An MLP can be understood as a feedforward neural network model that maps a set of input data to a set of output data.

[0090] Figure 13 An example of a multilayer perceptron (MLP) neural network is shown. Its structure can include multiple hidden (e.g., inner) layers HL1 to HLn, which map an input layer InL (receiving a set of inputs (or vector inputs) in_1 to in_3) to an output layer OutL, which produces a set of outputs (or vector outputs), such as out_1 and out_2. Each layer can have any given number of nodes, which are schematically shown as circles within each layer in this document. In this example, the first hidden layer HL1 has two nodes, while hidden layers HL2, HL3, and HLn each have three nodes. Generally, the deeper the MLP (e.g., the more hidden layers an MLP has), the stronger its learning ability. The input layer InL receives vector input (schematically shown as a three-dimensional vector consisting of in_1, in_2, and in_3) and can apply the received vector input to the first hidden layer HL1 in the hidden layer sequence. The output layer OutL receives the output from the last hidden layer (e.g., HLn) in the multi-layer model, processes its input, and produces a vector output (exemplarily shown as a two-dimensional vector consisting of out_1 and out_2).

[0091] Typically, each neuron (or node) produces a single output, which is fed forward to neurons in the immediately preceding layer. However, each neuron in a hidden layer can receive multiple inputs, either from the input layer or from the outputs of neurons in the immediately preceding hidden layer. Typically, each node can apply a function to its inputs to produce its output. Nodes in hidden layers (e.g., learning layers) can apply the same function to their respective inputs to produce their respective outputs. However, some nodes (e.g., nodes in the input layer InL) receive only one input and can be passive, meaning they simply relay the value of their single input to their output, for example, providing a copy of their input to their output, as indicated by the dashed arrow within the node in the input layer InL.

[0092] For illustrative purposes, Figure 14 A simplified neural network consisting of an input layer InL', a hidden layer HL1', and an output layer OutL' is shown. The input layer InL' is shown with two input nodes i1 and i2, which receive inputs Input_1 and Input_2 respectively (e.g., the input nodes of layer InL' receive a two-dimensional input vector). The input layer InL' is fed forward to a hidden layer HL1' with two nodes h1 and h2, and the hidden layer HL1' is in turn fed forward to an output layer OutL' with two nodes o1 and o2. The interconnections or links between neurons (shown as solid arrows in the diagram) have weights w1 to w8. Typically, in addition to the input layer, nodes (neurons) can receive the output of the node in their immediately preceding layer as input. Each node can compute its output by multiplying each of its inputs by the corresponding interconnection weights for each input, summing the products of its inputs, adding (or multiplying) a constant defined by another weight or bias that may be associated with that particular node (e.g., node weights (or biases) w9, w10, w11, w12 corresponding to nodes h1, h2, o1, and o2, respectively), and then applying a nonlinear or logarithmic function to the result. The nonlinear function can be called an activation function or a transfer function. Various activation functions are known in the art, and the choice of a particular activation function is not important for this discussion. However, it should be noted that the operation of an ML model, or the behavior of a neural network, depends on the weights, which can be learned so that the neural network provides the desired output for a given input.

[0093] During the training or learning phase, the neural network learns (e.g., is trained to determine) appropriate weight values ​​to achieve the desired output for a given input. Before training the neural network, each weight can be individually assigned an initial (e.g., random and optional non-zero) value, such as a random number seed. Various methods for assigning initial weights are known in the art. The weights are then trained (optimized) such that, for a given training vector input, the neural network produces an output close to the desired (predetermined) training vector output. For example, the weights can be incrementally adjusted over thousands of iterations using a technique called backpropagation. In each iteration of backpropagation, the training input (e.g., a vector input or training input image / sample) is fed forward through the neural network to determine its actual output (e.g., a vector output). The error of each output neuron or output node is then calculated based on the actual neuron output and the target training output of that neuron (e.g., the training output image / sample corresponding to the current training input image / sample). The weights are then updated based on the degree of influence each weight has on the total error, propagating back through the neural network (in the direction from the output layer back to the input layer), making the output of the neural network closer to the desired training output. This loop is then repeated until the actual output of the neural network falls within an acceptable error range of the expected training output for a given training input. It's understandable that each training input may require multiple backpropagation iterations before reaching the expected error range. Typically, an epoch refers to one backpropagation iteration across all training samples (e.g., one forward pass and one backward pass), making training a neural network potentially require many epochs. Generally, the larger the training set, the better the performance of the trained ML model, so various data augmentation methods can be used to increase the size of the training set. For example, when the training set consists of pairs of corresponding training input and training output images, the training images can be divided into multiple corresponding image segments (or blocks). Corresponding blocks from the training input and training output images can be paired to define multiple training block pairs from one input / output image pair, which expands the training set. However, training on large training sets places high demands on computational resources (e.g., memory and data processing resources). The computational requirements can be reduced by dividing the large training set into multiple mini-batches, where the mini-batch size defines the number of training samples in one forward / backward pass. In this case, one epoch can include multiple mini-batches. Another problem is the possibility that neural networks (NNs) overfit the training set, thereby reducing their ability to generalize from a specific input to different inputs. The overfitting problem can be mitigated by creating an ensemble of neural networks or by randomly dropping nodes from the neural network during training, which effectively removes the dropped nodes from the network. Various dropout modulation methods are known in the art, such as reverse dropout.

[0094] It should be noted that the operation of a trained neural network (NN) is not a direct algorithmic step of the operation / analysis process. In fact, when a trained NN receives an input, it does not analyze that input in the traditional sense. Instead, regardless of the subject or nature of the input (e.g., a vector defining a real-time image / scan or a vector defining some other entity, such as a demographic description or activity record), the input will undergo the same predefined architectural construction of the trained neural network (e.g., the same node / layer arrangement, training weights and biases, predefined convolution / deconvolution operations, activation functions, pooling operations, etc.), and it may be unclear how the architecture of the trained network produces its output. Furthermore, the values ​​of the training weights and biases are not deterministic and depend on many factors, such as the amount of time given to the neural network for training (e.g., the number of epochs in training), the random initial values ​​of the weights before training begins, the computer architecture of the machine training the NN, the selection of training samples, the distribution of training samples across multiple mini-batches, the choice of activation function, the choice of error function, and its modification of the weights, even if training is interrupted on one machine (e.g., with a first computer architecture) and completed on another machine (e.g., with a different computer architecture). The key point is that the reasons why trained ML models achieve certain outputs are still unclear, and extensive research is currently underway to try to determine the factors underlying ML model outputs. Therefore, the processing of real-time data by neural networks cannot be simplified to simple algorithmic steps. Instead, its operation depends on its training architecture, training sample set, training sequence, and various circumstances during ML model training.

[0095] In summary, the construction of a neural network (NN) machine learning model can include a learning (or training) phase and a classification (or operation) phase. In the learning phase, the neural network can be trained for a specific purpose, and a set of training examples (including training (sample) inputs and training (sample) outputs) can be provided to the neural network, optionally including a set of validation examples to test the progress of training. During this learning process, various weights associated with the nodes and node interconnections in the neural network are incrementally adjusted to reduce the error between the actual output of the neural network and the desired training output. In this way, a multi-layer feedforward neural network (e.g., as described above) can be made capable of approximating any measurable function to any desired accuracy. The result of the learning phase is a (neural network) machine learning (ML) model that has been learned (e.g., trained). In the operation phase, a set of test inputs (or real-time inputs) can be submitted to the learned (trained) ML model, which can apply what it has learned to produce output predictions based on the test inputs.

[0096] picture Figure 13 and Figure 14Like regular neural networks, convolutional neural networks (CNNs) consist of neurons with learnable weights and biases. Each neuron receives input, performs an operation (e.g., a dot product), and optionally follows a non-linear path. However, a CNN can take raw image pixels at one end (e.g., the input) and provide a classification (or category) score at the other end (e.g., the output). Because a CNN expects an image as input, it is optimized for processing volumes (e.g., the pixel height and width of the image, plus the image depth, such as color depth, e.g., RGB depth defined by three colors: red, green, and blue). For example, a CNN layer can be optimized for neurons arranged in three dimensions. Neurons in a CNN layer can also be connected to a small region preceding that layer, rather than all neurons in a fully connected NN. The final output layer of a CNN can reduce the entire image to a single vector (classification) arranged along the depth dimension.

[0097] Figure 15An example convolutional neural network architecture is provided. A convolutional neural network can be defined as a sequence of two or more layers (e.g., layers 1 to N), where these layers may include (image) convolution steps, (result) weighted sum steps, and nonlinear function steps. Convolution can be performed on the input data, for example, by applying filters (or kernels) over a moving window on the input data to produce feature maps. Each layer and its components may have different predetermined filters (from a filter bank), weights (or weighting parameters), and / or function parameters. In this example, the input data is an image with a given pixel height and width, which may be the raw pixel values ​​of the image. In this example, the input image is shown as a depth image with three color channels RGB (red, green, and blue). Optionally, the input image may undergo various preprocessing steps, and the preprocessed results may be used in place of the original input image or be input in addition to the original input image. Some examples of image preprocessing may include: retinal angiography segmentation, color space transformation, adaptive histogram equalization, connected component generation, etc. Within a layer, the dot product between a given weight and the small regions connected to it in the input volume can be computed. Many ways to configure a CNN are known in the art, but as an example, layers can be configured to apply element-wise activation functions, such as a maximum (0, x) threshold at zero. Pooling functions (e.g., along the xy direction) can be performed to downsample the volume. Fully connected layers can be used to determine the classification output and produce a one-dimensional output vector, which has been found useful for image recognition and classification. However, for image segmentation, a CNN needs to classify each pixel. Since each CNN layer tends to degrade the resolution of the input image, another stage is needed to upsample the image back to its original resolution. This can be achieved by applying a transposed convolution (or deconvolution) stage (TC), which typically does not use any predefined interpolation methods but instead has learnable parameters.

[0098] Convolutional neural networks have been successfully applied to many computer vision problems. As mentioned above, training CNNs typically requires large training datasets. The U-Net architecture, based on CNNs, can usually be trained on smaller training datasets than traditional CNNs.

[0099] Figure 16An exemplary U-Net architecture is illustrated. This exemplary U-Net includes an input module (or input layer or stage) that receives an input U-in (e.g., an input image or image patch) of any given size. For illustrative purposes, the image size of any stage or layer is indicated within a box representing the image; for example, the input module contains the number "128×128" to indicate that the input image U-in consists of 128×128 pixels. The input image can be a fundus image, an OCT / OCTA frontal image, a B-scan image, etc. However, it should be understood that the input can be of any size or dimension. For example, the input image can be an RGB color image, a monochrome image, a volumetric image, etc. The input image undergoes a series of processing layers, each shown at exemplary dimensions, but these dimensions are for illustrative purposes only and will depend on, for example, the image size, convolutional filters, and / or pooling stages. This architecture consists of a shrinking path (exemplarily comprising four encoding modules in this document), followed by an expanding path (exemplarily comprising four decoding modules in this document), and copy and cut links (e.g., CC1 to CC4) located between corresponding modules / stages that copy the output of an encoding module in the shrinking path and cascade it to the upconverted input of the corresponding decoding module in the expanding path (e.g., append it to its back side). This results in a distinctive U-shape, from which the architecture derives its name. Optionally, for example, for computational reasons, a "bottleneck" module / stage (BN) may be located between the shrinking and expanding paths. The bottleneck BN may consist of two convolutional layers (with batch normalization and optional dropout).

[0100] A shrinking path is analogous to an encoder, typically capturing contextual (or feature) information using feature maps. In this example, each encoding module in the shrinking path may include two or more convolutional layers, exemplarily indicated by an asterisk “*”, followed by a max-pooling layer (e.g., a downsampling layer). For example, an input image U-in is illustratively shown undergoing two convolutional layers, each with 32 feature maps. It can be understood that each convolutional kernel produces a feature map (e.g., the output of a convolution operation with a given kernel is an image commonly referred to as a “feature map”). For example, the input U-in undergoes a first convolution applying 32 convolutional kernels (not shown) to produce an output comprising 32 corresponding feature maps. However, as is known in the art, the number of feature maps produced by the convolutional operation can be adjusted (up or down). For example, the number of feature maps can be reduced by averaging multiple sets of feature maps, deleting some feature maps, or other known feature map reduction methods. In this example, the first convolution is followed by a second convolution, the output of which is limited to 32 feature maps. Another way to envision feature maps is to consider the output of the convolutional layer as a 3D image whose 2D dimension is given by the listed XY plane pixel dimensions (e.g., 128 × 128 pixels) and whose depth is given by the number of feature maps (e.g., 32 planar image depths). Following this analogy, the output of the second convolution (e.g., the output of the first encoding module in the shrinking path) can be described as a 128 × 128 × 32 image. The output of the second convolution then undergoes a pooling operation, which reduces the 2D dimension of each feature map (e.g., the X and Y dimensions can each be halved). As indicated by the down arrow, the pooling operation can be embodied in a downsampling operation. Several pooling methods (e.g., max pooling) are known in the art, and the specific pooling method is not important to this invention. In each pool, the number of feature maps can be doubled, starting with 32 feature maps in the first encoding module (or block), 64 feature maps in the second encoding module, and so on. Thus, the shrinking path forms a convolutional network consisting of multiple encoding modules (or stages or blocks). As is typical in convolutional networks, each encoding module can provide at least one convolutional stage, followed by an activation function (not shown) (e.g., a rectified linear unit (ReLU) or sigmoid layer) and a max-gathering operation. Typically, the activation function introduces non-linearity into the layer (e.g., to help avoid overfitting), receives the layer's results, and determines whether to "activate" the output (e.g., determining whether the value of a given node meets a predefined criterion to forward the output to the next layer / node). In summary, shrinking paths generally reduce spatial information while increasing feature information.

[0101] The expansion path is similar to a decoder, specifically providing localization and spatial information for the results of the contraction path, although downsampling and any maximum pooling are performed during the contraction phase. The expansion path comprises multiple decoding modules, each concatenating its current upconverted value with the output of the corresponding encoding module. In this way, feature and spatial information are combined in the expansion path through a series of upconvolutions (e.g., upsampling, transposed convolutions, or deconvolutions) and concatenations with high-resolution features from the contraction path (e.g., via CC1 to CC4). Thus, the output of the deconvolutional layer is concatenated with the corresponding (optionally cropped) feature map from the contraction path, followed by two convolutional layers and an activation function (optionally batch normalized). The output from the last expansion module in the expansion path can be fed into another processing / training block or layer, such as a classifier block, which can be trained with the U-Net architecture.

[0102] Computing device / system

[0103] Figure 17 An example computer system (or computing device or computer apparatus) is illustrated. In some embodiments, one or more computer systems may provide the functionality described or illustrated herein and / or perform one or more steps of one or more methods described or illustrated herein. The computer system may take any suitable physical form. For example, the computer system may be an embedded computer system, a system-on-a-chip (SOC), a single-board computer system (SBC) (e.g., a computer-on-module (COM) or system-on-module (SOM)), a desktop computer system, a laptop or notebook computer system, a computer system grid, a mobile phone, a personal digital assistant (PDA), a server, a tablet computer system, an augmented / virtual reality device, or a combination of two or more of these. Where appropriate, the computer system may reside in a cloud, which may include one or more cloud components within one or more networks.

[0104] In some embodiments, the computer system may include a processor Cpnt1, a memory Cpnt2, a storage device Cpnt3, an input / output (I / O) interface Cpnt4, a communication interface Cpnt5, and a bus Cpnt6. The computer system may also optionally include a display Cpnt7, such as a computer monitor or screen.

[0105] Processor Cpnt1 includes hardware for executing instructions, such as those that make up a computer program. For example, processor Cpnt1 may be general-purpose computing on a central processing unit (CPU) or a graphics processing unit (GPGPU). Processor Cpnt1 may retrieve (or fetch) instructions from internal registers, internal caches, memory Cpnt2, or storage device Cpnt3, decode and execute the instructions, and write one or more results to internal registers, internal caches, memory Cpnt2, or storage device Cpnt3. In a specific embodiment, processor Cpnt1 may include one or more internal caches for data, instructions, or addresses. Processor Cpnt1 may include one or more instruction caches and one or more data caches, for example, for storing data tables. The instructions in the instruction cache may be copies of instructions in memory Cpnt2 or storage device Cpnt3, and the instruction cache may accelerate the retrieval of those instructions by processor Cpnt1. Processor Cpnt1 may include any suitable number of internal registers and may include one or more arithmetic logic units (ALUs). Processor Cpnt1 may be a multi-core processor; or may include one or more processors Cpnt1. Although this disclosure describes and illustrates a particular processor, this disclosure considers any suitable processor.

[0106] Memory Cpnt2 may include main memory for storing instructions that processor Cpnt1 executes or saves temporary data during processing. For example, a computer system may load instructions or data (e.g., a data table) from storage device Cpnt3 or from another source (e.g., another computer system) into memory Cpnt2. Processor Cpnt1 may load instructions and data from memory Cpnt2 into one or more internal registers or internal caches. To execute instructions, processor Cpnt1 may retrieve and decode instructions from internal registers or internal caches. During or after instruction execution, processor Cpnt1 may write one or more results (which may be intermediate or final results) to internal registers, internal caches, memory Cpnt2, or storage device Cpnt3. Bus Cpnt6 may include one or more memory buses (each bus may include an address bus and a data bus) and may couple processor Cpnt1 to memory Cpnt2 and / or storage device Cpnt3. Optionally, one or more memory management units (MMUs) facilitate data transfer between processor Cpnt1 and memory Cpnt2. Memory Cpnt2 (which may be fast volatile memory) may include random access memory, such as dynamic RAM (DRAM) or static RAM (SRAM). Storage device Cpnt3 may include long-term or high-capacity memory for data or instructions. Storage device Cpnt3 may be internal or external to a computer system and includes one or more of the following: disk drive (e.g., hard disk drive HDD or solid-state drive SSD), flash memory, ROM, EPROM, optical disk, magneto-optical disk, magnetic tape, Universal Serial Bus (USB) accessible drive, or other types of non-volatile memory.

[0107] The I / O interface Cpnt4 can be software, hardware, or a combination of both, and includes one or more interfaces (e.g., serial or parallel communication ports) for communicating with I / O devices, enabling communication with a person (e.g., a user). For example, I / O devices may include a keyboard, keypad, microphone, monitor, mouse, printer, scanner, speaker, still camera, pen, tablet computer, touchscreen, trackball, camera, another suitable I / O device, or a combination of two or more of these devices.

[0108] The communication interface Cpnt5 provides a network interface for communication with other systems or networks. Communication interface Cpnt5 may include a Bluetooth interface or other types of packet-based communication. For example, communication interface Cpnt5 may include a network interface controller (NIC) and / or a wireless NIC or wireless adapter for communication with wireless networks. Communication interface Cpnt5 can provide communication with Wi-Fi networks, ad hoc networks, personal area networks (PANs), wireless PANs (e.g., Bluetooth WPANs), local area networks (LANs), wide area networks (WANs), metropolitan area networks (MANs), cellular telephone networks (e.g., GSM networks), the Internet, or combinations of two or more of these networks.

[0109] The Cpnt6 bus can provide communication links between the aforementioned components of a computing system. For example, the Cpnt6 bus may include an Accelerated Graphics Port (AGP) or other graphics bus, an Enhanced Industry Standard Architecture (EISA) bus, a Front Side Bus (FSB), HyperTransport (HT) interconnect, an Industry Standard Architecture (ISA) bus, an InfiniBand bus, a Low Pin Count (LPC) bus, a memory bus, a Micro Channel Architecture (MCA) bus, a Peripheral Component Interconnect (PCI) bus, a High Speed ​​Serial Computer (PCI-Express, PCIe) bus, a Serial Advanced Technology Attachment (SATA) bus, a Video Electronics Standards Association Local (VLB) bus, or other suitable buses combining two or more of these buses.

[0110] Although this disclosure describes and illustrates a particular computer system having a particular number of particular components in a particular arrangement, this disclosure contemplates any suitable computer system having any suitable number of any suitable components in any suitable arrangement.

[0111] In this document, where appropriate, computer-readable non-transitory storage media may include one or more semiconductor-based or other integrated circuits (ICs) (e.g., field-programmable gate arrays (FPGAs) or application-specific integrated circuits (ASICs)), hard disk drives (HDDs), hybrid hard disk drives (HHDs), optical disks, optical disk drives (ODDs), magneto-optical disks, magneto-optical disk drives, floppy disks, floppy disk drives (FDDs), magnetic tape, solid-state drives (SSDs), RAM drives, secure digital cards or drives, any other suitable computer-readable non-transitory storage media, or any suitable combination of two or more of these media. Computer-readable non-transitory storage media may be volatile, non-volatile, or a combination of volatile and non-volatile.

[0112] Although the invention has been described in conjunction with several specific embodiments, it will be apparent to those skilled in the art that many further substitutions, modifications, and variations will be readily apparent from the foregoing description. Therefore, the invention described herein is intended to encompass all substitutions, modifications, applications, and variations that may fall within the spirit and scope of the appended claims.

Claims

1. A method for customizing a visual field test, comprising: selecting a visual field test for a patient, the selected visual field test having one or more test points of definable light intensity; obtaining biometric information of a retina of the patient; deriving, based at least in part on the biometric information, respective threshold sensitivity values for one or more selected test points of the selected visual field test, each threshold sensitivity value being a measure of light intensity at which the patient is expected to see with a predetermined success rate; and determining, when applying the selected visual field test to the patient, one or more starting intensity values for the one or more selected test points using the derived threshold sensitivity values; wherein the deriving is provided at least in part by a neural network, the neural network comprising a first neural network of a first type (NN-1) in series with a second neural network of a second type (NN-2) different from the first type, and wherein the neural network of the first type receives as input image data and the neural network of the second type receives as input non-image data, wherein the deriving of the respective threshold sensitivity values for one or more selected test points of the selected visual field test is further based on prior real visual field test results of the patient and prior derived visual field test predictions of the patient, the prior derived visual field test predictions themselves based on historical biometric information of the retina of the patient taken at an earlier date than the currently obtained biometric information of the retina of the patient.

2. The method of claim 1, wherein: the biometric information is based at least in part on a retinal image; and the retinal image is captured by a particular imaging device using a particular imaging modality.

3. The method of claim 2, wherein, the imaging modality is one of a gray scale image, a color image, an infrared image, a retinal layer thickness image, an optical coherence tomography (OCT), a Doppler OCT, and a fluorescein angiogram.

4. The method of claim 2, wherein, the imaging modality is fundus imaging and the retinal image is a fundus image, or wherein the retinal image is an optical coherence tomography (OCT) angiogram image.

5. The method of any one of claims 1 to 4, wherein, the training of the neural network comprises: collecting a plurality of training data pairs, each training data pair comprising training input data and corresponding training output data, the training input data comprising biometric information of a retina of a patient and the training output data comprising test results from a particular visual field test administered to the patient; for each training data pair, submitting the training input data of the each training data pair as input to a neural network and providing, from the neural network, the corresponding visual field test results of the each training data pair as target output; wherein the training input data further comprises one or more of physical characteristics of the patient, demographic standard biometric data of the patient, and demographic standard functional data of the patient; and the physical characteristics comprise one or more of age, race, and medical history of the patient. The standard biometry data includes one or more of a retinal nerve fiber layer (RNFL) thickness and a ganglion cell-inner plexiform layer (GCIPL) thickness for the patient's demographics; and The standard function data includes one or more standardized initialization parameters for the particular visual field test for the patient's demographics.

6. The method of claim 5, wherein, The biometry information includes an OCT scan of the patient's retina.

7. The method according to one or more of claims 1 to 6, wherein, The first and second neural networks are of a type selected from a group consisting of a fully connected neural network, a convolutional neural network, a feedforward neural network, a recurrent neural network, a modular neural network, and a U-Net.

8. The method according to one or more of claims 1 to 7, wherein, The selected visual field test is one of a static automated perimetry test, a dynamic perimetry test, and a frequency doubling perimetry test.

9. A system for customizing a functional visual field test, comprising: an electronic processor (Cpntl); a perimeter (13) for applying a visual field test to a patient, the visual field test having one or more test points of definable light intensity; a non-transitory computer readable storage device (Cpnt3) storing software instructions that, when executed by a processor, cause the electronic processor to: obtain biometry information of a retina of the patient; and determine respective threshold sensitivity values for one or more selected test points of the visual field test based at least in part on the biometry information, each threshold sensitivity value being a measure of light intensity that the patient is expected to see at a predetermined success rate; wherein, when applying the visual field test to the patient, the perimeter (13) uses the determined threshold sensitivity values to determine starting intensity values for the one or more selected test points; and wherein the electronic processor (Cpntl) is part of a neural network, the neural network comprising a first neural network (NN-1) of a first type in series with a second neural network (NN-2) of a second type different from the first type, and wherein the neural network of the first type receives image data as input and the neural network of the second type receives non-image data as input, wherein the determination of the respective threshold sensitivity values for the one or more selected test points of the selected visual field test is further based on prior real visual field test results of the patient and prior determined threshold sensitivity values of the patient, the prior determined threshold sensitivity values of the patient themselves being based on historical biometry information of the patient's retina taken at an earlier date than the currently obtained biometry information of the patient's retina.

10. The system of claim 9, wherein, The biometry information is based at least in part on retinal fundus images acquired with fundus photography techniques.

11. The system of claim 9 or 10, wherein, The training of the trained neural network comprises: collecting a plurality of training data pairs, each training data pair comprising training input data and corresponding training output data, the training input data comprising biometry information of a retina of a patient and the training output data comprising test results from a particular visual field test administered to the patient; wherein the training of the trained neural network further comprises: training the trained neural network on the plurality of training data pairs to determine the respective threshold sensitivity values for the one or more selected test points of the selected visual field test based at least in part on the biometry information of the patient's retina and the test results from the particular visual field test administered to the patient. for each training data pair, submitting the training input data for the each training data pair as input to a neural network and providing, as target output, from the neural network, a corresponding visual field test result for the each training data pair; wherein the training input data further comprises one or more of a physical characteristic of the patient, standard biometric data for a demography of the patient, and standard functional data for the demography of the patient; and the physical characteristic comprises one or more of an age, a race, and a medical history of the patient; the standard biometric data comprises one or more of a retinal nerve fiber layer (RNFL) thickness and a ganglion cell-inner plexiform layer (GCIPL) thickness for the demography of the patient; and the standard functional data comprises one or more standardized initialization parameters for the particular visual field test for the demography of the patient.

12. The system of claim 11, wherein, the biometric information of the retina of the patient comprises an OCT angiography scan of the retina of the patient.

13. The system according to one or more of claims 9 to 12, wherein, the first neural network and the second neural network are selected from a group consisting of a fully connected neural network, a convolutional neural network, a feedforward neural network, a recurrent neural network, a modular neural network, and a U-Net.

14. The system according to one or more of claims 11 to 13, wherein, the visual field test is one of a static automated perimetry test, a dynamic perimetry test, and a flicker perimetry test.

Citation Information

Patent Citations

  • High speed spectral domain functional optical coherence tomography and optical doppler tomography for in vivo blood flow dynamics and tissue structure

    US20050171438A1

  • Method and apparatus for ultrahigh sensitive optical microangiography

    US20120307014A1

  • Systems and methods for broad line fundus imaging

    US20150131050A1

  • Method and apparatus for early detection of glaucoma

    US5912723A

  • Phase-resolved optical coherence tomography and optical doppler tomography for imaging fluid flow in tissue with fast scanning speed and high velocity sensitivity

    US6549801B1