Improved detecting, tracking, and imaging of vitreous floaters
An end-to-end deep learning pipeline for floater detection and tracking in the eye addresses the limitations of existing methods by providing real-time, accurate identification and localization of individual floaters, enhancing treatment planning and diagnosis through SLO and OCT imaging.
Patent Information
- Application Number
- PCT/IB2025/055325
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-06-13
- Filing Date
- 2025-05-22
- Publication Date
- 2025-11-27
AI Technical Summary
Existing methods for detecting and treating vitreous opacities (floaters) in the eye are limited by slow imaging techniques, inaccurate targeting, and challenges in distinguishing and tracking individual floaters, especially under conditions of overlap or movement, which affects the reliability of real-time treatment planning and diagnosis.
An end-to-end deep learning pipeline using SLO imaging and machine learning models for floater detection, segmentation, and tracking, incorporating synthetic data to enhance accuracy and real-time tracking, and combining with OCT imaging for precise depth localization.
Enables continuous, real-time detection and tracking of individual floaters with high accuracy, improving clinical workflows by reducing manual processing and enhancing the precision of floater identification and treatment planning.
Smart Images

Figure IB2025055325_27112025_PF_FP_ABST
Abstract
Description
IMPROVED DETECTING, TRACKING, AND IMAGING OF VITREOUSFLOATERSTECHNICAL FIELD
[0001] Embodiments of the disclosure generally relate to methods, devices, and systems for detecting, tracking, imaging, and / or treating eye conditions. Some embodiments relate in particular to detecting, tracking, imaging, and / or treating symptomatic vitreous opacities (SVOs), also known as floaters.BACKGROUND
[0002] Symptomatic vitreous opacities (SVOs), commonly referred to as floaters, in a patient’s eye can impact the patient’s vision and / or comfort. Floaters are microscopic fibers that can tend to clump together within the vitreous of the eye and can cast shadows over the patient’s retina. Some treatments for floaters include removing the vitreous fluid that has the floaters and replacing it with a solution. Some other treatments can use lasers to break up the floaters within the vitreous. The lasers can be targeted at the floaters by an ophthalmologist or other practitioner using a targeting laser. The manual targeting process can risk impacting non-floater elements within the patient’s eye. Further, the manual targeting limits the minimum size of the floaters that can be targeted and treated using existing techniques.
[0003] Given the importance of a patient’s eyes to the patient’s quality of life, additional, improved systems and methods for detecting, tracking, and imaging of these floaters (or one or more other eye conditions) are desirable.SUMMARY
[0004] Some systems image a patient’s eye to detect potential floaters. For example, a system may use optical coherence tomography (OCT) to image a patient’s eye. Oftentimes, image can be used to diagnose the patient with a floater, but a full volume scan from OCT imaging can be too slow, in some cases taking approximatelyone second or longer, to be used during treatment procedures. For example, by the time the volume scan is complete, the floater may have moved to another position on the patient’s eye, limiting the volume scan’s diagnostic value and making the volume scan unreliable for real-time treatment planning. Furthermore, once the OCT volume scan is complete, another system (e.g., a manually guided laser device) is required to treat the floater, resulting in fragmented workflows, alignment challenges, and increased risk of inaccurate targeting.
[0005] In another example, a system for detecting and tracking retinal floaters in medical imaging relies on heatmap-based semantic segmentation applied to single Scanning Laser Ophthalmoscopy (SLO) frames. While such systems may offer some ability to localize regions where floaters are present, these systems suffer from low specificity in distinguishing individual floaters, merging of overlapping floaters, and the lack of real-time or near real-time tracking and quantification. Moreover, these limitations reduce the clinician’s ability to accurately monitor floater movement, morphology, and volume over time, which can be important for diagnosis, treatment planning, and interventions (e.g., real-time laser treatment). These inefficiencies make some floater tracking approaches unsuitable for dynamic clinical assessment, real-time intraoperative guidance, or robust longitudinal monitoring.
[0006] To address these challenges, systems, methods, and devices are disclosed herein for detecting, segmenting, identifying, and tracking of retinal floaters in SLO image sequences using an end-to-end deep learning pipeline. A system can include an SLO imaging module, a preprocessing module, a deep learning-based detection and tracking module, and integrated data output and visualization interfaces configured to extract, quantify, and / or track floaters individually with high accuracy and temporal fidelity. In some embodiments, the system leverages synthetic data, such as simulated floaters projected onto retinal images, to improve the reliability of the system’s modeling. For example, such synthetic data can be used in training a machine learning model used for floater identification, tracking, etc.
[0007] The SLO imaging module can be configured to acquire sequential high- frequency retinal images. The preprocessing module can be programmed to standardize and improve the input images. The deep learning model (e.g., a modified instance segmentation neural network) can be trained on both real and synthetic floater data tooutput, for each floater in each image frame, a bounding box, an instance segmentation mask, a unique floater identification number, and motion trajectory data. The tracking module can maintain and update a mapping between floaters’ unique identifications (IDs) and observe each floater with a unique ID across image frames to provide constant identity and motion tracking. The visualization interface can overlay bounding boxes, segmentation masks, and unique IDs onto SLO frames, allowing clinicians to monitor and interact with the output in real time.
[0008] In some embodiments, the system can use projection of digitally modeled floaters onto SLO or OCT image backgrounds to generate diversity in a training set. This enables the model to distinguish, segment, and track floaters even in cases of severe overlap or temporary occlusion, replicating challenging real-world imaging conditions. The data augmentation process simulates realistic floater movement and appearance variability, allowing for improved generalization to unseen clinical cases.
[0009] During operation, the system can detect the presence and trajectory of one or more floaters within an SLO image sequence. For example, the system can identify each floater’s bounding box, generate an instance-specific segmentation mask, assign or update a unique tracking identifier, compute motion vectors or paths for one or more floaters in real time or nearly real time, and so forth. When a floater temporarily exits the imaging field and later returns, the system can re-associate the correct unique ID for the floater to provide longitudinal tracking (e.g., monitor and collect data from the same floater over an extended period of time to observe any changes). The system can provide an interface that visualizes the output directly overlaid on a live SLO stream to allow users (e.g., clinicians) to measure floater size, opacity, displacement, monitor floater movement during examinations, and record temporal and morphological data. In some implementations, the system is configured to analyze floaters, floater movement, etc. The system can integrate with clinical workflow software to generate reports or guide targeted laser therapy in response to floater motion during patient treatment.
[0010] In some embodiments, the system can generate synthetic datasets by simulating floater appearance, movement, overlap, and interactions across modeled SLO sequences. These synthetic datasets, when combined with human-annotated real-world images, can be used to train an instance segmentation and tracking neural network capable of generalizing to complex clinical presentations. For example, the model canlearn to track two or more floaters as they cross paths or occlude one another, maintaining accurate and separate identification and trajectory records for each. This training approach can reduce the need for manual annotation and / or real world imaging, and can adapt the system to a wide spectrum of floater presentations.
[0011] The system can provide several technical advantages for detecting, segmenting, identifying, and tracking of retinal floaters in SLO image sequences using the end-to-end deep learning pipeline. For example, the system can provide continuous, real-time (or nearly real-time), and accurate detection, segmentation, and tracking of individual retinal floaters, overcoming limitations in specificity and speed of some other techniques. The system can reduce the need for manual post-processing, improve the accuracy for identifying overlapping or intersecting floaters, and improve clinical workflows by providing actionable morphological and dynamical information for each floater. The system can support automated tracking and quantification for clinical research or intervention planning, improve the accuracy of volume scans and motion assessment, and improve clinician confidence during diagnosis or treatment, among other advantages that may be realized alone or in any combination.
[0012] In some embodiments, the approaches herein can be used for improved analysis of SLO images. Multiple floaters can be mistakenly grouped as one when using a machine learning (ML) model. The overlapping multiple floaters can be difficult for some ML segmentation models to distinguish or accurately delineate each floater as a unique object. To overcome this problem, the approaches can use a heatmap and contour discretization method that operates in conjunction with the ML model’s segmentation output. Specifically, the method processes the model-generated heatmap to yield more refined and separated contours, enabling the resolution of clustered or merged floater regions into their constituent components. For example, the system can implement the Fast Iterative Shrinkage-Thresholding Algorithm (FIST A) to reliably identify and split the primary subcomponents within the heatmap, thereby improving the precision of floater boundary detection and size estimation. Using algorithms such as FISTA can help the system more accurately identify and track individual floaters for subsequent clinical review or intervention, overcoming the limitations of prior segmentation techniques when floaters are adjacent or overlapping. In some embodiments, other algorithms are usedfor image processing. For example, in some embodiments, a rolling ball algorithm is used to smooth an uneven background in an image.
[0013] In some embodiments, the system can address the difficulty in accurately identifying and tracking floaters with OCT imaging, where determining the correct focal depth is required for capturing optimal floater objects, especially when floaters are faint, clustered, or move unpredictably. The system can use ML to analyze multiple OCT B- scans and sensor metadata to identify the appropriate focal depth for maximizing the visibility of floaters. A B-Scan (e.g., short for Brightness Scan) is a two-dimensional cross- sectional image composed by acquiring a series of one-dimensional depth scans. In particular, OCT B-scans can be generated by sequentially gathering A-scans (e.g., measure reflectivity as a function of depth at a point) across a straight or curved line, creating a 2D slice or section through the retina or vitreous of the eye. For example, the system can create one B-scan showing tissue structure in the cross-section of the retina by collecting 1000 A-scans along a horizontal axis.
[0014] In some embodiments, the system can address the challenge of rapidly and accurately localizing the depth (e.g., z-position) of floaters in OCT imaging. The system can automate the localization of the z-position by systematically sweeping (e.g., panning) the OCT focal plane through the depth of the visible region in discrete increments, collecting a series of B-scans at each focal depth. These B-scans are aligned and assembled into an artificial volumetric dataset in which the focus increments stand in for the physical Y-dimension. The system can then deploy algorithms (e.g., object detection) on the artificial volumetric dataset to determine the precise depth at which a floater is most in focus.
[0015] In some embodiments, the system can combine both the ML OCT focal depth finder, and the panned focus z-position finder described in the two previous paragraphs. The combined approach can offer an integrated ML based system for improved ophthalmic imaging of vitreous floaters using OCT and SLO modalities. The system can first use a ML model that pans the OCT focus across multiple depths, rapidly collecting B-scans to form a volumetric dataset. The system can use a focal depth selector driven by ML to then identify the appropriate imaging plane(s) for maximizing floater signal. The system can use downstream algorithms (e.g., segmentation and heatmap contour discretization) to separate clustered or merged floater regions withinthe B-scans into distinct, individually trackable components. The result is a fully automated pipeline that provides precise, real-time detection, separation, and tracking of floaters with unique IDs and object-level analytics that can improve clinical workflow and the quality of diagnostic information for eye health professionals.
[0016] In some embodiments, the system can process volumetric OCT data of a floater to generate a synthetic two-dimensional shadow representation simulating how the floater would appear in a SLO image. The system can collapse a segmented 3D floater volume into a 2D image that can be realistically modified by altering orientation, opacity, and perspective to reflect different positions within the vitreous. These synthetic SLO frames can be used for model training, validation, patient education, or device simulation, and can reduce the need for real patient data in developing and testing tracking or Z-estimation systems.
[0017] In some embodiments, a series of two-dimensional images captured at different focal planes — such as a sequence of SLO images with panned focal depth or OCT B-scans with adjusted Z-focus — may be analyzed to estimate the Z-axis location of one or more floaters. A machine learning model, such as a convolutional neural network with an attention mechanism, may process the focal sweep data and output an estimated depth corresponding to the optimal focal plane. In certain implementations, synthetic SLO streams may also be generated from the OCT volume to simulate various depth configurations. The estimated depth may be used to guide focus control or targeting in real-time during laser-based floater treatment.
[0018] In some embodiments, the system can quantify the unique identifier metrics for floater tracking in a sequence of eye imaging frames, such as SLO images. The system can use a tracking algorithm that assigns a unique identifier to each detected floater in the initial frame. The system can maintain unique identifier assignments as floaters are tracked across subsequent frames. For each frame, the predicted floater contours and their associated unique identifiers are recorded. To evaluate ID consistency, the system can correlate each tracked floater with its corresponding ground truth annotation, thus excluding false positive trajectories. The system then calculates the Continuous T racking Ratio (CTR) for each floater as the proportion of frames in which the floater retains its consensus unique identifier (e.g., defined by majority presence) over all frames the floater appears. The system uses the average CTR values across floatersto provide an objective measure of the tracking algorithm’s ability to maintain consistent identification. The system further generates CTR analysis plots as a function of floater visibility duration and incorporates Kalman filter-based motion estimation, reporting Kalman filter confidence scores based on prediction error. These combined metrics can offer reliable, quantitative evaluation of tracking performance, support algorithm optimization, and provide informed settings of imaging requirements for volume scans. The approach can help reduce unique identifier switching between floaters due to detection errors, occlusions, or abrupt floater movements, thereby improving overall accuracy in floater quantification, motion modeling, and treatment planning.
[0019] For purposes of this summary, certain aspects, advantages, and novel features are described herein. It is to be understood that not necessarily all such advantages may be achieved in accordance with any particular embodiment. Thus, for example, those skilled in the art will recognize the disclosures herein may be embodied or carried out in a manner that achieves one or more advantages taught herein without necessarily achieving other advantages as may be taught or suggested herein.
[0020] All of the embodiments described herein are intended to be within the scope of the present disclosure. These and other embodiments will be readily apparent to those skilled in the art from the following detailed description, having reference to the attached figures. The invention is not intended to be limited to any particular disclosed embodiment or embodiments.BRIEF DESCRIPTION OF THE DRAWINGS
[0021] These and other features, aspects, and advantages of the disclosure are described with reference to drawings of certain embodiments, which are intended to illustrate, but not to limit, the present disclosure. It is to be understood that the accompanying drawings, which are incorporated in and constitute a part of this specification, are for the purpose of illustrating concepts disclosed herein and may not be to scale.
[0022] Figure 1A shows an example semantic segmentation and tracking pipeline process.
[0023] Figure 1 B shows an illustrative system for detecting and tracking specific floaters according to some embodiments.
[0024] Figure 2 presents illustrative examples of floater detection and tracking using an end-to-end tracking pipeline as described herein.
[0025] Figure 3 shows panel depicting an example of two floaters that are tracked as their visual shadows intersect within a retinal image, in accordance with one or more embodiments of this disclosure.
[0026] Figure 4A illustrates an example of system identifying and discretizing multiple floaters within a single merged segmentation region, in accordance with one or more embodiments of this disclosure.
[0027] Figure 4B presents illustrative panels that demonstrate the generation and positioning of Gaussian kernels used for sparse representation and floater separation in accordance with one or more embodiments of this disclosure.
[0028] Figure 4G illustrates an example of the sparse representation generated by a system from an input heatmap, following the application of the Gaussian kernel decomposition and thresholding process, in accordance with one or more embodiments of this disclosure.
[0029] Figure 4D illustrates the process of the floater identification method that can be carried out by a system in accordance with one or more embodiments of this disclosure.
[0030] Figure 5A illustrates a base model training process according to some embodiments.
[0031] Figure 5B illustrates an iterative model training process according to some embodiments.
[0032] Figure 5C illustrates an example of the application of a base model and iterative model according to some embodiments.
[0033] Figure 6 shows an example diagram depicting a system using an automated approach that includes using machine learning and metadata to find an optimal focus depth.
[0034] Figure 7 shows an example diagram in which a system uses machine learning to find an optimal focal depth range by analyzing acquired B-scans informationat different depth settings, in accordance with one or more embodiments of this disclosure.
[0035] Figure 8A shows an example of the input data that can be provided to a panned focus z-finder algorithm.
[0036] Figure 8B shows an example of the input volume that can be provided to the end-to-end Z-finder model.
[0037] Figure 9 shows example results associated with a trained Z-finder model.
[0038] Figure 10 is a diagram that graphically illustrates a process for generating an SLO image from an OCT image.
[0039] Figure 11A shows a sample of 3 frames selected from the generated SLO stream.
[0040] Figure 11 B shows a sample of a single slice of the volume scan being cropped.
[0041] Figure 11C shows a 3D floater volume being rotated about the x-axis.
[0042] Figure 11D shows simulated 2D representations of a floater added to 2 of the sample frames from the SLO stream with a force in the Z direction.
[0043] Figure 12 illustrates a flow chart for processing a floater volume to add a collapsed floater to an SLO frame.
[0044] Figure 13 illustrates an example of how CTR is computed, in accordance with embodiments of the present disclosure.
[0045] Figure 14 illustrates an example floater consistent ID assignment plot, in accordance with embodiments of the present disclosure.
[0046] Figure 15 illustrates X-Z (left) and Y-Z (right) max projections of a vitreous containing a floater, in accordance with embodiments of the present disclosure.
[0047] Figure 16 illustrates an example streaming 2D visualization of floater position and movement using reconstructed 3D data, in accordance with embodiments of the present disclosure.
[0048] Figure 17 is a block diagram depicting an embodiment of a computer hardware system configured to run software for implementing the approaches for floateridentification, tracking, and imaging and any systems, methods, and devices disclosed herein.DETAILED DESCRIPTION
[0049] Although several embodiments, examples, and illustrations are disclosed below, it will be understood by those of ordinary skill in the art that the approaches described herein extend beyond the specifically disclosed embodiments, examples, and illustrations and include other uses and obvious modifications and equivalents thereof. Embodiments are described with reference to the accompanying figures, wherein like numerals refer to like elements throughout. The terminology used in the description presented herein is not intended to be interpreted in any limited or restrictive manner simply because it is being used in conjunction with a detailed description of certain specific embodiments of the inventions. In addition, embodiments can comprise several novel features and no single feature is solely responsible for its desirable attributes or is essential to practicing the approaches herein described.
[0050] Symptomatic vitreous opacities (SVOs), commonly referred to as floaters, in an eye of a patient (also referred to herein as a subject) can be detected using optical imaging and processing techniques. In some embodiments, SVOs are detected and tracked in real time (e.g., in substantially real time). The detection of the SVOs can be used in evaluating a patient’s eye condition, determining treatment options, and / or treating the SVOs using a therapeutic laser. The treatment can include the ablation or removal or evaporation or liquification of the SVO, or portion thereof through a process of photo-ionization caused by one or more laser pulses. While described largely in terms of SVO detection, tracking, and treatment, it will be appreciated that the approaches described herein can be readily adapted to other diagnostic and / or treatment applications.
[0051] Some imaging and targeting technology does not provide direct feedback telling the clinician if the floater is within a safe treatment zone (e.g., if it’s too close to the retina or the lens. Therefore, there is a need for a system that can image the floater within the eye and determine if it is safe to treat. Additionally, since the floater can move at least partially independently of the eye, delivering a large number (e.g., hundreds, thousands,or more) of laser pulses onto the floater quickly is important. The shockwaves resulting from the laser pulses can result in the floater moving; thus, delivering pulses quickly before the floater has the chance to move is desirable. A significant challenge with treating floaters is the time associated with acquiring 3D images (e.g., OCT images). For example, using OCT technology, it is possible to image a volume, however, acquiring a volume scan can take tens, hundreds, or even thousands of milliseconds, during which time a floater may move. As such, having a methodology to image, detect, and track, floaters quickly enough to determine treatment locations before a floater moves significantly is important to ensure its position at all times and to determine if it is located within a safe treatment zone, and finally deliver laser pulses quickly accurately and effectively to remove the floater, break up the floater, or reduce the size of the floater.
[0052] The detection and tracking of SVOs can be done in various ways as described further below using one or more different imaging devices. For example, a first imaging device, such as a scanning laser ophthalmoscopy (SLO) imaging device, can capture an image of the eye or portion of the eye within which a floater is visible. It will be appreciated that an SLO image may not necessarily capture an image of the actual SVO, but rather a shadow of the SVO on the retina. The image from the first imaging device can provide an X-Y image that allows the position of the floater to be partially determined, although the depth information about the position of the floater may not be determined by the first imaging device. The X-Y position information can be used to select an imaging location of a second imaging device capable of capturing depth information, such as an optical coherence tomography (OCT) imaging device. The images from the first and second imaging devices allow for the 3D location of floaters within the eye to be determined. The combination of multiple imaging devices can allow the 3D tracking of floaters to be done in real-time (e.g., on time scales that allow for treatment to be carried out before a floater has undergone significant movement). The tracking information can be used for various purposes including, for example, measuring details of the floater(s) , treating the floater(s) with a laser, etc. In some embodiments, a treatment laser can be a femtosecond laser. In some embodiments, the treatment laser is a pulsed laser configured to provide a pulse energy of 1-20 pJ, a central wavelength of 1030 nm or about 1030 nm, a pulse rate of from 1 kHz or about 1 kHz to 2 MHz or about 2 MHz. In some embodiments, a pulse duration can be from 100 fs or about 100 fs to 300 fs or about 300 fs.End-to-End Floater Tracker Pipeline
[0053] Figure 1A illustrates an example process for floater detection and tracking. In Figure 1A, detection, tracking, and separation are divided into distinct steps. At operation 105, a system can access an SLO image (or sequence of SLO images). At operation 110, the system can process each accessed SLO image, for example to remove background noise, perform normalization procedures, and so forth. At operation 115, the system can detect floaters in each SLO image, for example using a machine learning model trained to detect floaters. At operation 120, the system can generate a heatmap of the floaters. At operation 125, the system can apply an algorithm to the heatmap to separate larger objects into multiple constituent floaters, as described in more detail herein. At operation 130, the system can track floaters across image frames. For example, the system can assign identifiers to each floater and can track the individual floaters across multiple frames.
[0054] Figure 1 B shows an illustrative system 150 for detecting and tracking specific floaters in a patient’s eye by assigning unique identifiers in accordance with one or more embodiments of this disclosure. For example, a patient can suffer from the shadows of floaters on their retina, which obstructs their ability to see clearly. Accordingly, there is a need to track these floater shadows in real-time. For example, the system 150 can use an end-to-end floater tracking pipeline (e.g., SLO imaging module 160, preprocessing module 170, and localization module 190) to address the need to detect and track specific floaters from a sequence of SLO images 180.
[0055] In some embodiments, system 150 can use the SLO imaging module to capture a temporal sequence of SLO images 180 of the patient’s retina. Each SLO image may contain the shadow of one or more vitreous floaters. System 150 can transmit each SLO image to the preprocessing module 170. Preprocessing module 170 can improve image quality to improve floater detection and tracking. The operations performed by the preprocessing module 170 can include denoising, contrast normalization, and background subtraction, among others. For each SLO image from the sequence of SLO images 180, the system 150 can apply a localization module 190 to detect and track floaters. The localization module 190 can detect individual floaters within each image frame and output both the bounding boxes and segmentation masks that define each floater’s position and shape. For every detected floater, system 150 can assign a uniqueidentifier and match that unique identifier to detections from previous frames. System 150 can consistently track each floater even when floaters move, overlap, or temporarily exit and re-enter the field of view. System 150 can determine motion data for each floater by analyzing changes in detection between consecutive frames. Moreover, system 150 can output detailed floater-specific data comprising segmentation masks, bounding boxes, tracking unique identifiers, and motion information for each frame of the sequence of SLO images 180. The output of data can support real-time clinical floater visualization, longitudinal monitoring, quantitative floater analysis, and intervention guidance.
[0056] System 150 can provide accurate, real-time or near real-time detection and tracking of vitreous floaters under challenging imaging conditions, such as overlapping floaters or abrupt movement (e.g., float moves out of frame and returns). By combining detection, segmentation, and tracking in a unified deep learning model and matching detections in real time from frame to frame, system 150 can provide a user with temporally consistent tracking and reduce the occurrence of identification errors and the need for post-processing association steps. The end-to-end tracking pipeline can be compatible with both real and synthetic SLO data for improved performance.
[0057] In some embodiments, the system can detect and track floaters by assigning a unique identifier for each floater with the appropriate inference speed (15fps, 24fps, 30fps, 60fps, etc.) using a deep learning model. The system can use the deep learning model as a tool for quantifying the floaters, tracking the floaters in real-time, and capturing volume scans as part of diagnosis and / or treatment procedures. For example, the deep learning model may receive a sequence of SLO images 180 of the patient’s retina at each point in time, extract the bounding boxes and segmentation masks for each floater, match the detections in the real-time with the previous detections, and track each floater by assigning a unique identifier. Moreover, the deep learning model can be trained using a combination of real-world data as well as synthetic data. Furthermore, the deep learning model can provide bounding box, segmentation mask, unique identifier, and floater motion data for each floater in a real-time imaging session.
[0058] With the end-to-end tracking pipeline, system 150 can perform detection and tracking can in the same model. Unlike some other approaches, the deep learning model can output the floater objects bounding boxes and instance segmentation mask. The output from system 150 can allow the floater objects to be separated at the deep learningmodel level instead of having to separate the floater objects during post-processing. Several advantages can be realized. The floater objects can still be tracked if they get close, as in the instance segmentation method the objects are not competing on the segmentation mask and can be easily separated. The deep learning model is able to incorporate the floater’s unique identifiers in the ground truth object as the tracking performance can be optimized. The deep learning model can be trained using a combination of real-world data as well as synthetic data. The model outputs a different heatmap / bounding box and tracking identification information for different floaters as opposed to the previous method, in which the model would only output a semantic segmentation heatmap for floaters. The model can potentially incorporate a reidentification method for tracking the floaters that go outside field of view and come back again.
[0059] In some embodiments, the end-to-end floater tracking pipeline can assist ophthalmologists and optometrists in quantifying and tracking floaters of patients in realtime, enabling the collection of data on the size, opacity and motion of the floaters. In some embodiments, the end-to-end floater tracking pipeline would facilitate the volume scanning of moving floaters. In some embodiments, the end-to-end floater tracking pipeline would assist in the treatment of floaters in real-time by providing information on the location of moving floaters.
[0060] Figure 2 presents illustrative examples of floater detection and tracking using an end-to-end tracking pipeline as described herein. Figure 2 includes four panels 200- 1 , 200-2, 200-3, and 200-4, each demonstrating a distinct stage in the floater tracking workflow on SLO images of a patient’s eye.
[0061] In some embodiments, a system can receive, (e.g., via an imaging module such as imaging module 160 of Figure 1 B), an SLO image of an eye. The system can detect the shadow of a floater within the image. The system can apply a localization module (e.g., detecting and tracking module 190 of Figure 1 B) to detect the floater and outputs a bounding box surrounding the identified floater. The system can visually overlay the bounding box onto the input image, to provide the user with visual confirmation of the detected shadow of the floater. For example, panel 200-1 shows a real floater in the eye, with the detected floater enclosed within a bounding box 205 (e.g., displayed as a rectangle) and optionally labeled with a unique identifier.
[0062] In some embodiments, the system applies a heatmap visualization to display the region of highest detection probability for the floater, as seen in panel 200-2. Panel 200-2 demonstrates determining and highlighting the spatial location where a deep learning model infers the floater is most likely present. The heatmap or colored overlay is concentrated around the vicinity of the real floater, corresponding to the bounding box shown in panel 200-1.
[0063] The system can proceed with segmentation and feature extraction, refining the earlier detection (e.g., using preprocessing module 170). The system can overlay segmentation results, which may include pixel-wise or contour-based masks, over the original eye image. At this stage, the system can mark the boundaries and shape of the floater with greater precision, supporting both localization and size quantification as seen in panel 200-3. The bounding box can remain overlaid, while the floater’s contour or other features may be accentuated with color, lines, or other annotation.
[0064] Panel 200-4 shows an example of multi-object tracking as may be carried out according to some embodiments of the present disclosure. A system can track the trajectory of one or more floaters. The system can re-identify multiple floaters across frames. For example, panel 200-4 shows additional bounding boxes around different floaters which can have labels representing unique identifiers, indicating the assignment of identifiers to each detected floater. In some embodiments, each panel can display the synthetic output or simulated detection derived for benchmarking model performance or visualizing a floater’s path in the image plane. The axes, labels, or plot elements can allow for frame-to-frame trajectory plotting, size tracking, or comparison between real and synthetic data.
[0065] A system can be configured to use these visual steps as part of an end-to- end tracking pipeline, facilitating precise, real-time detection, quantification, and identity maintenance for floaters. The panels (e.g., 200-1 to 200-4) illustrate the system’s technical ability to locate, segment, track, and record floater movement using SLO image sequences, supporting quantitative analysis and clinical application.
[0066] Figure 3 shows panel 300 depicting an example of two floaters that are tracked as their visual shadows intersect within a retinal image, in accordance with one or more embodiments of this disclosure. Panel 300 depicts the floaters with clear overlapping regions. Panel 300 illustrates an example output of an end-to-end floaterdetection pipeline capable of distinguishing overlapping floaters by appropriately placing individual bounding boxes or segmentation masks. In some embodiments, different floaters can be assigned different identifiers, annotations, color overlays, etc.
[0067] Many approaches to floater detection and tracking struggle to differentiate between overlapping or intersecting floaters. For example, a semantic segmentation pipeline can compete for assignment of shared pixels in an overlap region, often merging multiple floaters into a single connected mask in the output. This makes it difficult or impossible to separately identify, track, and quantify each individual floater. Furthermore, because classic semantic segmentation models are trained with a ground truth mask that only provides class-level pixel assignment, these approaches cannot incorporate explicit object definitions or unique identifiers during model training or inference. As a result, many systems are limited in their ability to resolve distinct floaters in clustered or intersecting configurations, leading to identification errors and unreliable tracking under real-world clinical conditions.
[0068] By contrast, an end-to-end tracking pipeline can utilize instance segmentation and object-level detection. A system can access SLO image frames and process each SLO image frame to identify, segment, and / or assign a unique identifier to each floater, regardless of their proximity or overlap. In the overlapping region shown in panel 300, a system according to the present disclosure can accurately delineate the boundaries of floaters 305 and 310, even though they overlap, preserving each floater’s separate identities to provide consistent tracking of floaters as individual objects across sequential frames. The end-to-end tracking pipeline can resolve the ambiguity caused by merged masks and supports reliable longitudinal analysis, motion quantification, and intervention planning for individual floaters.Floater Heatmap / Contour Discretization
[0069] As described above, one problem experienced by many systems for detecting floaters observed by imaging of a patient’s eye is not being able to distinguish between close floaters. Many systems may output one large contour instead of outputting / V separate contours, where / V is the number of overlapping floaters in an image being analyzed. This occurs since floaters in SLO images can be close to each other or can overlap in 2D space. As a result, multiple floaters can be mistakenly grouped as one. Asystem according to the present disclosure can overcome the limitations of many other systems. For example, a system can distinguish between floaters that are classified as a single floater using some other approaches, thereby enabling better floater identification and tracking. In some embodiments, a system can improve safety during treatment because having a finer contour helps the treatment to be more accurate and therefore safe. For example, better tracking and / or identification of floaters can help ensure that the correct floaters are targeted.
[0070] In some embodiments, a method of predicting and discretizing the individual floater components of a cluster of floaters may involve the use of a machine learning model. The code for the model can be added to a floater tracker model for easier integration. In some embodiments, a fast iterative shrinkage thresholder algorithm (FISTA) may be used to account for the major subcomponents of the heatmap. In some embodiments, FISTA can be used as a computational method to generate training examples for machine learning. In some embodiments, two different models may be trained and deployed consecutively to separate floaters, yielding more accuracy and time efficiency.
[0071] Figure 4A illustrates an example of system identifying and discretizing multiple floaters within a single merged segmentation region, in accordance with one or more embodiments of this disclosure. Figure 4A includes three panels (e.g., 400-1 , 400- 2, and 400-3), each panel can demonstrate a distinct stage in the floater segmentation and contour discretization workflow.
[0072] Panel 400-1 depicts a heatmap. In some embodiments, a system can generate the heatmap using a segmentation model. The heatmap can highlight a region of interest within the SLO image, corresponding to the presence of possible floaters. The merged heatmap region can indicate that multiple floaters are closely positioned or partially overlapping, resulting in the initial detection representing a single contiguous area.
[0073] Panel 400-2 depicts a contour associated with the heatmap, where a system can overlay the heatmap on the initial SLO image. The SLO image with the overlaid heatmap can show one or more floater shadows in the contour. For example, the system can overlay the contiguous contour produced by the segmentation model onto the original SLO image. Within the outlined contour, the system can provide the user with visuals thatdepict three distinct floater shadows — representing three discrete floaters that fall inside the boundary defined by the model output as seen in panel 400-2. However, at this stage, the segmentation mask does not yet differentiate between the individual floaters; instead, it groups them together as one large object.
[0074] In panel 400-3, the system further processes the merged contour by applying one or more contour discretization algorithms. The system can identify and mark each separate floater inside the region using circular annotations 405A, 405B, 405C, and, in some embodiments, assigns a unique identifier for every distinct floater visible within the original segmentation boundary. For example, the system can use three circles to designate the individual floaters, demonstrating successful separation of floater components that were previously grouped by the segmentation model.
[0075] In some embodiments, to separate out the floaters, a sparse representation is generated from the input heatmap (e.g., panel 400-1) by mimicking a floater with gaussian kernels with different sizes in different locations of the image. The sample gaussian kernels are shown in Figure 4B.
[0076] Figure 4B presents illustrative panels that demonstrate the generation and positioning of Gaussian kernels used for sparse representation and floater separation in accordance with one or more embodiments of this disclosure. Each panel (e.g., 410-1 , 410-2, 410-3) shows a Gaussian kernel applied at different spatial locations on a sampled image grid.
[0077] A system can use Gaussian kernels to construct a sparse representation of the input heatmap (e.g., the heatmap 400-1 of Figure 4A). In each panel of Figure 4B, the system overlays a point corresponding to a Gaussian function at a distinct center position on a discrete grid. The gaussian kernels can mimic the effect of a floater’s visual shadow, with the intensity and spread (controlled by the Gaussian kernel’s standard deviation o) reflecting local floater characteristics. The system can position the Gaussian kernels at different coordinates within the panels to illustrate translation across the image domain.
[0078] To achieve floater separation, a system can formulate and solve an optimization problem, reconstructing the input heatmap as a sparse linear combination of multiple Gaussian kernels of varying size and position. This process is mathematicallygrounded in compressed sensing principles, and can be expressed as minimizing the following cost function: F = | | x - (.aiGi ) 112 +el laI L where x is the input heatmap that is first cropped and then resized to match with the kernel sizes. G is the set of gaussian kernels and a is the set of coefficients for each kernel. £ is a weighting term to help the coefficients to be sparse. Then with a threshold only a few of the largest elements of coefficients are kept.
[0079] Figure 4C illustrates an example of the sparse representation generated by a system from an input heatmap, following the application of the Gaussian kernel decomposition and thresholding process, in accordance with one or more embodiments of this disclosure. The panel in Fig. 40 displays a reconstructed output image comprised of a regular grid with three distinct, spatially separated, points. Each point can be circular or elliptical, with intensity highest at the center and smoothly decreasing toward the periphery, representing the localized contribution of an individual Gaussian kernel to the overall sparse map.
[0080] A system can use the sparse representation marks to estimate positions of separate floaters (e.g., shadows of vitreous opacities) within a detection region. Using the optimization and thresholding methodology described herein, a system can suppress the background (e.g., using a background reduction algorithm such as a rolling ball algorithm) and minimize extraneous activations, highlighting only those regions in the input heatmap most likely to correspond to distinct floater objects. As a result, the system can provide a clear, interpretable visualization of the location and separation of floaters, resolving the ambiguity found in merged or overlapping heatmap outputs. As seen in Fig. 4G., a sparse representation can provide a more helpful image depicting floater locations. For example, the system can reconstruct the image that can then be generated as an output that gives a better image on where the floaters are located (e.g., identifying the separate floaters in the heatmap). In some embodiments, the number of Gaussian kernels is not limited. In some embodiments, the number of Gaussian kernels is restricted, for example based on a maximum number of floaters expected to be encountered within a field of view.
[0081] Figure 4D illustrates the process of the floater identification method that can be carried out by a system in accordance with one or more embodiments of this disclosure. Figure 4D includes three key components: heatmap 432 (e.g., an input imagewith merged floater regions), method 434 (e.g., a process that works on the heatmap 432) and output image 436 (e.g., the processed image with separated floater regions). The heatmap 432 (e.g., generated by the segmentation model, a separate ML model, etc.) is fed to the method 434 as an input, and the generated output image 436 can be a reconstructed image (e.g., a heatmap with the separated floaters identified). Output image 436 can then be used as an input for another model and used to produce finer and more detailed contours or for other purposes, such as generating a treatment plan or controlling a treatment laser or associated components such as steering optics.
[0082] In some embodiments, a system can receive an image (e.g., heatmap 432) representing the initial output of a segmentation model. Heatmap 432 can contain several irregular shapes corresponding to regions with detected floaters. The floaters can touch or overlap, resulting in merged floater regions that obscure the identities of the individual floaters. The system can apply method 434 to process heatmap-432 — such as the optimization-based discriminative approach described herein. Method 434 can include the FISTA algorithm, in which the system uses the FISTA algorithm to resolve and discretize the underlying floater components. The system can process the merged regions by decomposing the heatmap-432 into sparse, well-defined elements. Process 434 can provide output image 436 after the method 434 has been applied. Output image 436 can show the components that make up merged or overlapping shapes. Output image 436 can provide the user with distinctly separated floaters having clear boundaries. The output image 436 can, additionally or alternatively, be used as an input for additional processing, treatment planning, etc. The approach shown in Figure 4D can increase the number of visible floater regions, each region now corresponding to an individual floater component, allowing a system to support precise downstream tracking, segmentation, and contouring of floaters.
[0083] While generating a reconstructed image with finer SVO counters is beneficial for identifying individual floaters in an image, there can be errors in the reconstruction. In some implementations, the reconstruction accuracy in identifying contours when there is more than one SVO can be estimated. In some embodiments, a system can determine the k largest basis components contributing to the reconstruction, k can be an integer, for example 3. The system can determine the largest components (xrs) and the largest basis component (xri).
[0084] The probability of having more than one floater in a contour is defined: 0 < xrl< 255, and c =( m x xrl) / Q]m x 255).
[0085] The mask ‘m’ highlights the region where the most important component exists and in the probability formula derives the xr3 intersection with xn. The constant ‘c’ is a modifier factor, as some parts of the xriwere masked and the values are changed to 1 , and the changes should be considered in the final value. The accuracy can measured based on the probability ‘p’ with threshold of 0.65 on a set of 30 images in some embodiments, although other thresholds and other numbers of images can also be used.
[0086] In some embodiments, the approach of generating a reconstructed image can be implemented into a machine-learning based pipeline. In some embodiments, a FISTA algorithm can be used as a computational method to generate additional training examples for training the machine learning model.
[0087] Figure 5A is a flowchart that illustrates an example process for identifying floaters from a heatmap according to some embodiments. In Figure 5A, a heatmap 505 is provided to a base model 510, and the base model generates a prediction 515. In some embodiments, the base model 510 is a convolutional neural network (CNN). The heatmap 505 can be a heatmap generating using a tracker model for tracking floaters in images (e.g., in SLO images or OCT images). The heatmap can also be provided to a detection algorithm 520 which, when applied to the heatmap 505, produces a ground truth output 525. The detection algorithm can be, for example, FISTA, and can determine locations of different floaters in the heatmap 505. A system can be configured to compare the prediction 515 and the ground truth 525 to determine a loss 530 indicative of the difference between the prediction 515 and the ground truth 525. One challenge with using a base model 510 as shown in Figure 5A is that the base model may fail to detect small floaters. Another problem can be that floaters without clear boundaries may not be reliably separated using the base model 510 even after training.
[0088] Figure 5B is a flowchart showing an example process for identifying floaters from a heatmap using an iterative model according to some embodiments. In Figure 5B, a heatmap 505 is provided to an iterative model 535, which is used by a system to produce a first prediction 540. The iterative model 535 can be a CNN. The system canapply the detection algorithm 520 (e.g., FISTA) to the heatmap 505 to generate the ground truth output 525. The system can compare the first prediction 540 and the ground truth 525 to determine a first loss 545. The prediction 540 can be provided to the iterative model 535 and the system can generate a second prediction 550. The system can compare the second prediction 550 to the ground truth 525 to generate a second loss 555.
[0089] A process such as the one shown in Figure 5B can offer certain advantages. For example, if a prediction is fed back into the iterative model, it can be expected that the output of the iterative model will be the same as the prediction. That is, it can be expected that prediction 540 and prediction 550 are the same. Assuming the same loss function is applied, it can be assumed that the loss 545 is the same as the loss 555. Such an approach can help ensure that small floaters are retained.
[0090] Using both a base model and an iterative model, small floaters can be retained while adjacent or overlapping floaters can be separated. Figure 5C shows an example process for applying an iterative model and a base model to a heatmap to identify individual floaters according to some embodiments. A system can access a heatmap 560 and apply an iterative model 535 to generate a first output 565. The system can apply the base model 510 to the output of the iterative model 535 to produce a second output 570. As shown in Figure 5C, the first output 565 retains small features, while the second output 570 shows only the locations of large features in the heatmap 560. The system can merge the first output 565 and the second output 570 at operation 575 to generate a final output 580. In some implementations, the merge operation 575 can be implemented using a library such as OpenCV. In the final output 580, smaller features (e.g., small floaters) can be depicted as well as the locations of larger floaters. In some embodiments, the second output 570 is used to determine circles or other shapes or outlines to draw on the heatmap 560 to show the locations of individual floaters.
[0091] While Figures 5A, 5B, and 5C are described in terms of receiving a heatmap as a starting point for analysis, any suitable input can be provided. For example, an input can be an SLO image in some embodiments. Moreover, the approaches described with respect to Figures 5A, 5B, and 5C, are not limited to the detection and identification of SVOs, but can be applied in a wide variety of contexts, such as detecting certain featuresin a tissue sample or for other applications, which may be medical applications or nonmedical applications.OCT Focal Depth Finder
[0092] Figure 6 shows an example diagram depicting a system using an automated approach that includes using machine learning and metadata to find the best focal depth (e.g., for floater objects in vitreous). When using conventional OCT imaging devices, the focus is usually set manually, and the procedure of changing the focus from the initial point in depth to any secondary point is also a manual process that requires the user to give input. When imaging floaters in 3D using OCT, it can be especially difficult to decide at which depth to place the OCT focus throughout the vitreous. Accordingly, the process of setting and adjusting the OCT focus depth should ideally be automated. A system can be configured to use ML techniques in conjunction with frame-specific metadata to identify the image frame in which floaters are most clearly visualized, automating what is otherwise a manual, operator-dependent process.
[0093] For example, Figure 6 presents three imaging frames labeled frame 600, frame 602, and frame 604. Each panel can represent a distinct OCT B-scan captured by the system at different focal depths within the vitreous body. Each frame can contain data (e.g., patterns, characteristics, light intensity) near the bottom, corresponding to the tissue interface or region of clinical interest. The set of frames depicts a typical OCT B- scan acquisition in which the focal plane is incrementally swept through depth to capture multiple images.
[0094] Furthermore, Figure 6 displays a table 606 with metadata from each frame. Table 606 can list the sequence of frames as seen on the column labeled "Frame" (e.g., 1 , 2, ... , N) and the column "Focus Index" (e.g., Fi, F2, ... , FN). The focus index indicates the focal depth associated with each frame and can be an important parameter for downstream machine learning analysis. In table 606, the row corresponding to "Frame 2" / "F2" is boxed, denoting the frame selected by the system as optimally focused for floater visualization.
[0095] A system can analyze each frame using a machine learning model trained to evaluate the clarity, prominence, signal strength, or any combination thereof of floater features within the OCT data. For example, edge detection can be used to evaluate thesharpness of floater features in a frame. The system can use both the image content and associated metadata — such as the focus index, acquisition timestamp, and scan position — to determine which frame contains the most distinct floater signal. In some embodiments, the system operates in two modes: (1) utilizing metadata in conjunction with ML predictions to refine focus selection or (2) relying on ML signal analysis alone. That is, while in some embodiments, metadata is used, in other embodiments, an ML model does not use metadata when determining an optimal focus.
[0096] In operation, a system can receive a set of B-scan frames together with their respective metadata. The system can rank or score each frame based on the content, select the appropriate frame for floater assessment, and automatically reference the associated focus index from the metadata in table 606. This process can streamline the identification of the best focal depth and allow automated or operator-independent 3D imaging workflows for vitreous floaters.
[0097] In some embodiments, a machine learning approach can be used to automate the process of selecting or changing the focus point in depth. For example, machine learning can be used alongside available metadata or by itself to find the best focal depth. For this purpose, a set of B-scans with different depth focus settings may be used. In some embodiments, machine learning can be used as part of a fully automated pipeline for setting focal depth for OCT imaging. In some embodiments, the automated setting of focal depth may be used in the characterization of floaters (based on B-scan analysis). In some embodiments, the automated setting of focal depth may be used in any B-scan imaging in sparse environments (e.g., ultrasound B-scan).
[0098] In some embodiments, the system may involve using machine learning and metadata to find an optimized focal depth by selecting the best frame and matching its focal depth. More specifically, the target frame which has the most desirable signal may be chosen (e.g., the target frame with the most well-defined floaters). Metadata can be used to extract the focus information for the selected frame and set the OCT focus or other parameters. This approach can involve processing the images using ML techniques and further extracting more information by relating ML output to the metadata.
[0099] Figure 7 shows an example diagram in which the system can use machine learning to find the best focal depth range by incorporating all the acquired B-scans information at different depth settings, in accordance with one or more embodiments ofthis disclosure. The ML model can select the best focal depth by looking at a set of B- scans without any metadata. The ML model can be configured to determine a desired range of focus by analyzing the B-scans. Each B-scan panel can represent a cross- sectional OCT image of the vitreous and retinal tissue captured at a distinct focal depth. The B-scans are organized side by side, demonstrating how the focal plane is incrementally adjusted throughout the imaging volume during an OCT examination. Overlaid across the three B-scan panels, the system can mark a rectangle that spans horizontally across the upper-middle portion of each scan. The rectangle can designate the “desired range of focus” within the sample, corresponding to the region in which optimal image clarity and structural detail are achieved. For example, a system can automatically identify and select the optimal focal depth range, guiding the acquisition or subsequent analysis of high-quality OCT data for vitreous floater detection, characterization, and clinical intervention.
[0100] In some embodiments, the approach seen in Figure 7 may involve using a machine learning method that has been trained to choose the best focal depth by processing the information in an entire set of B-scans. The information in the entire set of B-scans can be used to detect the target signal and yield a desired depth focus.
[0101] In order to capture floaters at the right focus, a system can be configured to adjust one or more optical elements (e.g., a focus-tunable lens) to focus on desired depth (z). In some embodiments, the approach seen in Figure 7 may utilize an ML model trained to address this challenge. In some embodiments, the machine learning model is referred to as an end-to-end Z-finder (e.g., depth-finder) model.Panned Focus Z-Find
[0102] Figure 8A shows an example of the input data that can be provided to a panned focus z-finder algorithm. In some embodiments, the input data is 3D data with a shape of / V x H x W, where / V is the number of focus steps, and H and W are height and width of the B-scans in each focus. If there is a lack of real data containing real floaters with panned focus, synthetic floater data may be generated to simulate the real data.
[0103] Figure 8B shows an example of the input volume that can be provided to the end-to-end Z-finder model. From left to right, the images are associated with focus sweep from top of the acquired B-scan to its bottom.
[0104] In some embodiments, the architecture of the end-to-end Z-finder model may be a CNN backbone, such as ResNet. In some embodiments, the CNN backbone may be used to derive semantic features from the input data followed by an attention mechanism to give a score to each focus image. The image having the maximum score will be chosen as the best focus point.
[0105] Figures 8A and 8B both illustrate an automated Z-positioning process for accurately locating and imaging faint SVOs (shadows of vitreous opacities, or floaters) within the vitreous using volumetric OCT scanning, in accordance with one or more embodiments of this disclosure. When a potential SVO is identified in an SLO image, a system can initiate a rapid volumetric OCT workflow by selecting the scan location and acquiring a central B-scan. The OCT focus is then panned incrementally from the top to the bottom of the visible region, with each discrete focus setting resulting in the acquisition of a dedicated B-scan. Rather than generating a physical volumetric scan with Y-axis translation, a system can construct an artificial volume in which the “Y” dimension is replaced with a series of B-scans acquired at varying focal depths.A system can align and aggregate these B-scans into a focus-volume object, ensuring focus increments are chosen to be sufficiently small to resolve floaters while large enough to permit rapid acquisition. By analyzing the resulting artificial volume, a system can apply advanced object detection and clustering algorithms — such as previously developed object finders or pulsefinding approaches — to identify the floater’s Z-position, distinguishing true floaters from noise by correlating signals across neighboring B-scans. This approach can allow a system to rapidly, consistently, and automatically determine the optimal focal depth for subsequent targeted OCT imaging or volumetric data collection.
[0106] In some cases, it may be desirable to perform a volumetric OCT scan as quickly as possible (e.g., immediately after a floater is found in an SLO image). However, there are many difficulties associated with imaging floaters in 3D using OCT. Not only are there difficulties in deciding at which depth to place the OCT focus throughout the vitreous (as previously discussed), but floaters can be very faint objects in OCT scans - even when they are directly in focus. The cross-correlation of signals in multiple B-scans is often used to accurately identify a floater from noise.
[0107] Accordingly, an automated way to find the Z location of the floater is needed, which would allow for automation in floater data collection. More specifically, an improved method is needed to identify the Z location of the floaters so that the correct focus in the OCT can be selected for the volumetric scan. In some embodiments, a method of performing the Z-Find uses a previously developed object finder for volumetric objects. This method builds an artificial volume where instead of a physical Y-direction, a focus dimension is used. Panning through all the possible focuses and quickly collecting B- Scans, this volume is generated. An object finder can then be run to discover the location of the object and set the location.
[0108] In some embodiments, when a scan location is selected, a central B-scan is selected and scanned. The OCT focus is panned from the top of the visible region down to the bottom in increments. The individual B-Scans are collected and aligned into a volume object. With the correct increment of the focus, a fast acquisition of the volume can be performed while having the steps close enough that the correlations between subsequent B-Scans can still be used to identify faint floaters above the background. This artificial volume can then be processed by a machine learning algorithm (e.g., a clustering algorithm) to identify the floater and set the Z location for the OCT focus.
[0109] In some embodiments, there may be other possible Z-Finder scanning configurations for identifying a Z depth for a floater. After finding a floater on the SLO (manually or with an ML tracker) and finding the bounding-box / center, various scanning configurations can be used for finding an optimal Z depth. For example, a system can be configured to use panned-focus B-scans, single focus B-scans, single focus cross- sectional B-scans, floater-centered A-scans (e.g., guided by SLO), or a spiral scan pattern. In a panned focus approach, a model can be used to select a B-scan with the best floater representation and output the associated Z-depth, or a model can read a set of panned focus B-scans and output a predicted Z-depth. When single focus B-scans are used, a system can set a static focus in the middle of (visible) vitreous. In some implementations, the system can be configured to read a floater and output a predicted Z-depth. In the context of cross-Osection B-scans, a system can be configured to read two orthogonal B-scans and output a predicted depth. In a floater-centered A-scan approach, a system can be configured to read one or more A scans and output a depth. In some embodiments, a system is configured to provide a user interface to a user, andthe user can define an ML search depth which can constrain the depths used for optimizing z focus.
[0110] Figure 9 shows example results associated with a trained Z-finder model. In some embodiments, the model may be trained on multiple different versions of synthetic data, with each set containing different floater shapes and sizes. Given the input, the output may be the best focus depth number. In Figure 9, the left images are ground truth images, while the images on the right correspond to focus depths determined using the Z-finder model.3D OCT Floater Volume to 2D SLO Shadow
[0111] Obtaining sufficient floater data from SLO images can be difficult, especially when certain movements, timeframes, and rotations are needed for testing or training purposes. To remedy this, methods can be used to synthesize a realistic interpretation of what a floater would look like in an SLO image rather than needing to obtain specific SLO images to extract floater data from. Thus, these methods can be used to generate synthetic SLO floater data, which can be used as training data or to simulate dynamic movement of the floater relative to forces within the eye. These methods can also be used to simulate the SLO environment for device testing purposes or to cross-reference a patient’s description of their floaters with a simulated version.
[0112] In some embodiments, synthesizing an interpretation of what a floater would look like in an SLO image may involve taking a segmented section of a floater found in a volume OCT scan and then collapsing it into a 2D rendition of the floater’s shadow that would be seen in the corresponding SLO image. This 2D shadow can also be realistically manipulated, rotated, or translated to create different views of the same floater. This creates a visualization of what the floater would look like during an SLO imaging session. In some embodiments, the simulated floater environment can be parameterized based on real, physical characteristics such as: 3D morphology, opacity, and movement / motility patterns in a fluid simulation.
[0113] In some embodiments, these methods can be used to generate synthetic floater data that mimics an SLO stream of a floater. This removes the need to collect real patient data for training, testing, or visualization needs. In some embodiments, an OCT volume scan of a floater (usually in the form of an H5 file) can be collapsed into a 2Dform that represents the floater’s shadow projected onto the back of the retina. The shadow can be added to an SLO background to visualize the shadow. A floater volume can be used to manipulate the floater in a 3D space, allowing for changes in opacity, blur, perspective skew, etc. In some embodiments, forces can be applied to the floater, allowing for realistic movement patterns in a fluid simulation (e.g., lateral translation in X, Y, and Z directions, and rotations along each axis). In some embodiments, the simulation can be used to add the 2D form of the floater to a series of frames of the SLO stream, mimicking the floater according to the volume loaded and forces applied over time.
[0114] Figure 10 is a diagram that graphically illustrates a process for taking a 3D volume (1010), segmenting the floater into a floater volume (1020) , collapsing the floater into a 2D representation (1030), and then adding the 2D representation to an SLO image (1040).
[0115] More specifically, volume scan images can be segmented to remove the background (and, in some embodiments, noise) from the volume scan, leaving only the floater or the floater with a limited amount of surrounding vitreous or other material. This new volume can be manipulated using image processing to collapse into a single 2D image that represents the floater from any specific angle. The collapsing process can account for the opacity of the floater, such that the 2D image has an opacity that mimics what would be seen in an actual shadow of the floater in an SLO image. The flattened image can then be inputted into the patient’s SLO image to simulate the corresponding SLO image that would have been taken of that floater.
[0116] For example, in the first step, a system can receive or acquire a three- dimensional OCT volume scan of the eye, capturing the morphology and spatial position of a floater within the vitreous. From this volume during the second step, the system can digitally segment the floater to isolate only the floater structure — removing the surrounding background or non-relevant tissue. In the third step, the system can process the segmented floater by collapsing the 3D data into a 2D projection, effectively generating a shadow or flattened image representative of how the floater would appear from a particular angle in an SLO image. The third step depicts this reduction to a two- dimensional view. The fourth step demonstrates integration of the simulated floater shadow into an SLO frame, including the possibility of programmatically adjusting the floater’s orientation, position, or apparent motion within the generated image.
[0117] A more concrete example of this process is shown in figures 11 A, 11 B, 11 C, and 11 D. Figure 11A shows a sample of 3 frames selected from the generated SLO stream. The simulated 2D form of the floater will be added to one or more of these frames.
[0118] An OCT volume scan of a floater in the eye is obtained. In some embodiments, the volume from the OCT scan is cropped to only include the floater, removing the retina and most of the background. After that, a simple threshold is applied to remove the background (e.g., to set it to black) and any noise surrounding the floater. Figure 11 B shows a sample of a single slice of the volume scan being cropped.
[0119] In some embodiments, the forces inputted by the user are applied to the volume scan. Linear forces can be calculated based on the starting pixel position of the floater, while rotations are done by manipulating the full volume scan to shift the orientation of the values in the 3D array. Figure 11C shows a 3D floater volume being rotated about the x-axis.
[0120] In some embodiments, to set the z (depth) dimension of the floater, the opacity, blurriness and size of the floater is altered based on its distance from the retina. In some embodiments, a force can also be applied in the z dimension, which will alter these 3 parameters to mimic the shadow growing or shrinking as it changes distances from the retina.
[0121] In some embodiments, to mimic the floater’s relative shadow based on its distance from the lens, further changes can be made. For example, slices of the floater that are closer to the retina are set to be more opaque, clearer, and smaller. On the other hand, slices that are closer to the lens are set to be more translucent, blurrier and larger. This allows for a sense of depth perception within the floater. In some embodiments, three linear functions can be used to calculate the corresponding opacity, blur, and size values for the floater based on its distance from the retina.
[0122] In some embodiments, the slices of the floater can then be summed together to collapse into a 2D shape that can be added to the SLO. Parts that overlap more with themselves will leave darker shadows, while those that are thinner will leave less of an impact on the SLO. The new coordinate position and new volume scan are then added to the provided SLO background.
[0123] Figure 11 D shows simulated 2D representations of a floater added to 2 of the sample frames from the SLO stream (e.g., from Figure 11 A) with a force in the Z direction. In reality, this would increase the distance of the floater from the retina. This is reflected by the differences in the appearance of the simulated 2D representations that were added to the 2 frames.
[0124] In some embodiments, a system can receive a three-dimensional OCT volume scan of the eye containing one or more vitreous floaters. The system can segment the floater(s) from the volume using image processing or machine learning techniques, removing background noise and irrelevant anatomical structures. The system can then project the segmented 3D floater onto a two-dimensional plane, simulating the SLO shadow that would be captured in an actual SLO image. The projection algorithm can include adjustments for perspective, illumination, and optical characteristics specific to SLO imaging. The system can allow parameterization and manipulation of the floater in virtual space, such as translating, rotating, or scaling the floater prior to projection. The resulting 2D SLO shadow images can be exported as synthetic training data for machine learning applications related to floater detection, classification, or focus optimization.
[0125] In some embodiments, a system can generate a diverse set of simulated SLO images from a single segmented 3D floater volume. The system can manipulate the segmented floater’s morphology, opacity, and orientation using geometric and photometric transformations. The manipulation may include applying movement patterns, simulating natural floatation, or adjusting blur and size to reflect changes in depth or focus. The manipulated volume is then projected onto a 2D image plane, producing a shadow with attributes that can mimic clinical variability. Multiple frames can be generated in sequence to simulate motion across time, further improving the synthetic dataset and aiding in the development or testing of algorithms that require temporal floater dynamic.
[0126] In some embodiments, a system can use an automated pipeline to process a library of 3D OCT floater volumes. For each input volume, the segmentation, manipulation (movement, rotation, translation, scaling, or application of simulated forces), projection, and post-processing steps are performed in batch, generating a corresponding set of 2D SLO shadow images. The system can randomize andparameterize key physical and optical characteristics — such as opacity, size, and blur — based on the floater’s simulated distance from the retina, and output single images or sequences for use in machine learning training, algorithm benchmarking, or device validation.
[0127] The various technical approaches describe herein from the one or more embodiments can enable a system to segment floater structures from 3D OCT volumes and manipulate them through geometric and optical transformations before projecting them onto a 2D image plane using a model that simulates SLO imaging characteristics. The resulting synthetic SLO shadow images, which can be generated in large quantities and with varying properties, provide a scalable source of diverse, annotated training data for developing and validating machine learning models in ophthalmic applications.
[0128] Figure 12 illustrates a flow chart for processing a floater volume to add a collapsed floater to an SLO frame. In some embodiments, a floater volume 1210 is first preprocessed at operation 1220. The preprocessing may involve cropping the volume to only include the floater or to otherwise substantially reduce the volume, removing the retina and at least a portion or most of the background. In some embodiments, the crop size, threshold, and opacity needed to extract the floater from the background OCT may be configurable parameters.
[0129] In a main loop 1320, forces and rotations on the floater may be applied to adjust the position and orientation of the 2D floater at operation 1240. Depth perception may also be applied to modify the appearance of the 2D floater. In some embodiments, there may be floater parameter ranges (opacity, blur, perspective) representing the ranges of values to be applied to the floater. In some embodiments, the depth (e.g., z position) of the floater can be adjusted at operation 1250. The 2D representation can undergo postprocessing at operation 1260and can be added to an SLO frame, resulting in the SLO frame containing the floater 1270. The opacity, blur, etc., of the floater can depend upon its z position, size, orientation, etc.Quantifying ID Consistency Metrics
[0130] Floater tracking is an essential part in floater diagnosis and treatment. Assigning a unique ID to each floater is essential for detecting and identifying floaters and for matching floaters with their treatment plans during the treatment stage. It will alsohelp to quantify the floaters, estimate the motion of the floaters, measure the opacity of the floaters, and distinguish symptomatic floaters.
[0131] Consistent tracking of floaters using floater ID can aid floater volume scanning by providing information for one or more floaters. For example, a clinician can select a specific floater and perform a volume scan for diagnosis and treatment.
[0132] However, tracking algorithms are often prone to ID switch errors because of a variety of factors such as detection errors, low frame rate, fast floater movements, occlusion, and so forth. Therefore, it is important to have a metric to measure the consistency of tracking for each floater. This metric can be used to help optimize the floater tracking algorithm as well as for setting requirements for volume scan imaging.
[0133] In some embodiments, there may be a set of accuracy metrics and measurement techniques to quantify the performance of floater tracking accuracy. These metrics can be used to evaluate floater detection, tracking, reidentification, and so forth. These metrics can be used to provide a measure of quantifying a floater tracking algorithm’s success in assigning the correct ID to a given floater in an SLO imaging session or other imaging session, or across different imaging sessions. These metrics may also be applied to any application where visually distinct shapes are being tracked and identified. Within these metrics, the variables tp, fp, and fn may stand for true positive, false positive, and false negative, respectively.
[0134] In some embodiments, an F1 score can be used to evaluate performance of a floater detection model. For example, an F1 score can be defined as F1=2 - - = - - - , with values ranging from 0 to 1 , with 1 representing the precision + recall 2tp + fn + fp best performance.
[0135] In some embodiments, the metric for SLO floater tracking accuracy may be a pixel- or object-based accuracy. In some embodiments, this metric may put more emphasis on size of the floaters. For example, larger floaters can have a greater weight in the final evaluation metric.
[0136] In an example when a positive integer number / V SVOs are being tracked,the metric can be defined as: s„P= — x - — -x wi , where w is the pixel- ^=1wtjl- tpi+ fni+ fpil’Kwise size for the / thSVO.
[0137] In some embodiments, Continuous Tracking Ratio (CTR) can be used as a metric for evaluating floater tracking consistency. In some embodiments, the CTR associated with a floater is used to measure the consistency of the floater tracking ID for the floater. In some embodiments, CTR represents the ratio of the number of frames with the same ID assigned to the same floater. In some cases, multiple IDs are assigned to the same floater across different frames. The ID used for determining the CTR can be selected as the ID that is most commonly assigned over the number of frames that the floater is present in the field of view.# frames with CID assigned
[0138] The CTR metric may be defined as: CTR = # frames while SVO is in FOV
[0139] In the formula, the consensus ID (CID) is defined as the ID assigned to a floater for the greatest number of frames in which the floater is within the field of view.
[0140] CTR may have a value between 0 and 1. The highest value (1) means that the floater did not have any ID switch, keeping a consistent ID during the time it was in the field of view (FOV). In some embodiments, CTR may be averaged across all floaters in a validation set.
[0141] In some embodiments, the measurement technique associated with CTR may be as follows. Given an image, filter one or more predicted contours (floaters) with initial IDs assigned. For each subsequent frame, the contour predictions are recorded along with their respective IDs. All tracked contours are associated with their corresponding ground truth floaters. It will remove the false positives tracks. Then, the CTR is calculated for each tracked contour. The average CTR of all floaters is used as a measure of the floater tracking algorithm success. The plot of the average CTR for floaters with different appearances can be as a measure of successful tracking for longstanding floaters in the field of view.
[0142] In some embodiments, Kalman filter confidence score may be a metric for measuring the confidence in the motion estimation of a floater, which can be used in the volume scan process of moving floaters. The Kalman filter confidence score is a measure of tracker accuracy in capturing the motion information of each floater. For each floater, the Kalman filter may be used to estimate the next location of the floater. The Kalman filter is updated using the measurements coming from the detection / tracker at each timepoint. The difference between the Kalman Filter prediction and the measurement for each point is the Kalman filter prediction error. The exponential moving average can be used to measure the Kalman filter prediction error over time.
[0143] The Kalman filter confidence score may be calculated as follows:
[0144] In this formula, MPE is the maximum prediction error set by the user, and the PE is the prediction error. The resulting value is ranged between 0 and 1 , indicating minimum and maximum confidence score respectively.
[0145] In some embodiments, the Kalman filter confidence score may be another important metric to measure and evaluate the model performance (e.g., success in predicting floater movement).
[0146] In some embodiments, a floater velocity measurement metric is used. The metric may be defined as: ev= 2^=1 (FEE) - v;)
[0147] In some embodiments, the metric may represent the mean squared error between the Kalman filter of the longest period of track and the actual values.
[0148] In some embodiments, a metric is used to quantify a switching tracking ID penalty. In some embodiments, these metrics incorporate the model’s ability to successfully track the SVO across frames.
[0149] For each SVO, a likelihood vector can be obtained containing the likelihood probability with previous SVOs. Assuming there has been / V SVOs so far, and identity of / \ / +1th SVO is desired and it is M: PN+1= [pi, p2, - - - , Piv] wherept= 1
[0150] argmax(PN+1) = K
[0151] sw+1= 1 if K = M otherwise 0 yT_s.
[0152] If there are T SVOs in total, the accuracy will be calculated as: s, = —l-
[0153] Overall the two types of metrics can be combined to yield one metric: m =Sp + a S[1 + a
[0154] The value of parameter a controls the contribution of identification score in the final metric m such that value of 1 has equal contribution.
[0155] In some embodiments, there may be a SLO floater re-identification accuracy. The same accuracy that was used for the tracking metric can be used, and it may be used to quantify switching tracking ID penalty.
[0156] In some embodiments, a metric can be used to characterize SLO registration accuracy, in accordance with embodiments of the present disclosure.
[0157] In some embodiments, the keypoint error (artificial transform) may be a metric calculated based on the keypoints that a registration metric chooses (e.g., superpoint). As the affine metric is not able to match more than a limited number of keypoints, some of the keypoints will not be perfectly registered.
[0158] The keypoint error can be calculated as: ea= i=i
[0159] dt is the euclidean distance between / thkeypoint in image 1 (p,1) and image2 registered to image 1 (p;21), calculated as: dt = l l Pt1- Pt21I I2
[0160] In some embodiments, the true transform error (manually select keypoints) may be defined similarly to ea, with the difference that keypoints are manually handpicked instead of a registration algorithm picking them.
[0161] Figure 13 illustrates an example of how CTR is computed, in accordance with embodiments of the present disclosure. Figure 13 presents an exemplary workflow and analysis method for evaluating object tracking performance across a sequence of image frames, in accordance with one or more embodiments of this disclosure. Figure 13 is divided into two principal sections: a visual representation of object tracking through multiple frames, and an analytics component for tracking performance evaluation.
[0162] For example, a system can track and assign each individual floater objects with assigned unique identifiers as they appear in a sequence of frames. In some embodiments, three objects are shown, each associated with a specific object ID (e.g., ID 1 , ID 2, ID 3), and their respective paths or trajectories are tracked across a series of frame numbers (e.g., frames 1 through 13). Each object is visually mapped along a dashed trajectory as it moves through the sequence of frames, with explicit labels connecting objects to their assigned IDs and the frames in which they are visible. The process includes annotations such as “Pre-detection (frame 1)” to denote the onset of tracking, and “Frame #” to indicate temporal progression through the series.
[0163] In some embodiments, a system can calculate a tracking performance metric, referred to as the "CTR" (Correct Tracking Ratio or Continuous Tracking Rate). This analytic is defined by taking the number of frames in which a particular object ID — for example, ID 2 — is successfully tracked, divided by the total number of frames within the sequence under consideration ( e.g., CTR = 5 / 13). This standardized metric can allow the system to quantitatively assess how effectively an object is tracked across the sequence, supporting algorithm validation and system optimization.
[0164] In the example shown, there are 13 total visible frames. Over the course of these frames, three different IDs were assigned to the floater. Of those three IDs, the one associated with the most frames is ID 2, which was assigned to the floater for 5 frames. Accordingly, the CTR would be calculated as 5 / 13.
[0165] Figure 14 illustrates an example floater consistent ID assignment plot, in accordance with embodiments of the present disclosure.
[0166] Across the validation set, the percentage of floaters that have their CID assigned for at least ‘n’ frames can be determined - this can be repeated for a range of ‘n’ (e.g., 30, 45, 60...). In each case of ‘n’, only floaters that are present in the FOV for at least ‘n’ frames are considered. The results can be plotted, with the plot illustrating the relationship of CTR compared to how long the floater is visible in the SLO field-of-view.Method for Vitreous 2D Visualization
[0167] Visualizing floaters can be challenging because the floaters are in a 3D space but they are visualized on a 2D screen. Furthermore, floaters are constantly moving (e.g., within the eye).
[0168] Accordingly, there is a need for providing better visualization of floaters (e.g., on floater diagnostic / treatment device user interfaces). More specifically, there needs to be a way to present floaters and their characteristics on a screen, so that it is easier to visualize where the floaters are and their characteristics.
[0169] In some embodiments, a method for visualizing floaters may involve “flattening” all the 3D floater information into a 2D map encompassing the entire vitreous. This 2D map would have a similar layout to an OCT B-scan, but it would contain the projected information from all floaters in the 3D vitreous. In some embodiments, the method may involve collapsing the 3+1 D information of floaters (e.g., from an OCT scanand / or other modalities) to a 2D visualization without losing key relevant information. In some embodiments, the technique may be applied to the reconstructed 3D data (or multiple volumes if this is a time series of data) to create a streaming 2D visualization for the user.
[0170] In some embodiments, the method may have potential applications in floater screening, diagnostics, and treatment. In some embodiments, the method may have potential applications in providing a simple visualization for patients to view the state of their vitreous health.
[0171] Figure 15 illustrates X-Z (left) and Y-Z (right) max projections of a vitreous containing a floater, in accordance with embodiments of the present disclosure.
[0172] More specifically, Figure. 15 shows an image where the “x-axis” is the horizontal axis of the eye (retina) but with the natural curvature flattened for clarity. The “y-axis” of the image would be the eye depth, normally referred to as the z-axis in imaging. The image depth axis would span from the retina surface to the lens.
[0173] Figure 16 illustrates an example streaming 2D visualization of floater position and movement using reconstructed 3D data, in accordance with embodiments of the present disclosure.
[0174] In some embodiments, each floater may be represented by its projected 2D shape / contour which could be based on the floater’s 3D geometry, a maximum projection in the Y-axis (or whichever dimension is collapsed), the floater’s estimated shape based on its shadow, or a simple non-specific icon. The position of the projected floater is based on the position during imaging.
[0175] In some embodiments, assuming the visualization plane is X-Z, the Y- position of the floater (collapsed dimension) could be represented by the color or outline pattern of the contour.
[0176] In some embodiments, the floater’s trajectory (also flattened in 2D) could also be shown, indicating the speed and path of the floater’s movement. This could be relevant since visualizing the mobility could help distinguish between low-mobility floaters (e.g. Weiss Rings) and high-mobility floaters. Furthermore, some mobile floaters may be unlikely to enter the center field of view.
[0177] In some embodiments, the fovea area is highlighted in the middle so the clinician / patient can see which floaters are currently in the center FOV and have the biggest impact on vision.Example Embodiments
[0178] Embodiment 1. A computer-implemented method for floater detection in an eye of a subject, the computer-implemented method comprising: accessing a sequence of images depicting one or more floaters in the eye of the subject, wherein each image of the sequence of images is a scanning laser ophthalmoscopy (SLO) image; accessing a set of floater identifiers for at least a subset of the one or more floaters; generating an input for a detection and tracking algorithm, wherein the input comprises at least one image of the sequence of images and at least one floater identifier from the set of floater identifiers for at least one of the one or more floaters; detecting, by applying the detection and tracking algorithm to the input, the one or more floaters, wherein the detection and tracking algorithm is configured to decompose a contour representing a plurality of spatially proximate floaters into a plurality of distinct floaters; and generating an output, wherein the output comprises a heatmap depicting locations of the one or more floaters.
[0179] Embodiment 2. The computer-implemented method of embodiment 1 , further comprising determining a depth focus for optical coherence tomography (OCT) imaging of a floater of the one or more floaters, wherein determining the depth focus comprises: accessing a set of OCT scans and a set of focus metadata, the set of focus metadata indicating a focus depth associated with each OCT scan of the set of OCT scans; analyzing the set of OCT scans to determine an optimal OCT scan, wherein the optimal OCT scan is selected based on at least one of: clarity, prominence, or signal strength; and determining, based on the optimal OCT scan and the set of focus metadata, an optimal focus depth.
[0180] Embodiment 3. The computer-implemented method of embodiment 1 , wherein the detection and tracking algorithm is trained using synthetic training data, wherein the synthetic training data is generated by: accessing an optical coherence tomography (OCT) scan depicting a floater; selecting a region of the OCT scan, the region inclusive of the floater and representing a volume that is smaller than a volume ofthe OCT scan; generating a cross-section of the floater; and projecting the cross-section of the floater onto a scanning laser ophthalmoscopy (SLO) image.
[0181] Embodiment 4. The computer-implemented method of embodiment 1, wherein decomposing the contour comprises: providing the input to a first machine learning model, wherein the first machine learning model is trained to separate floaters by comparing an output of the first machine learning model to an output of a detection algorithm and minimizing a loss function characterizing a difference between the output of the first machine learning model and the output of the detection algorithm.
[0182] Embodiment 5. The computer-implemented method of embodiment 1, wherein decomposing the contour comprises: applying an iterative model to the input to generate a first output; applying a base model to the first output to generate a second output; and merging the first output and the second output, wherein the first output comprises an image depicting the locations of floaters, and wherein the second output comprises an image depicting the locations of floaters above a threshold size, wherein the locations determined by the base model are determined with higher precision than the locations determined by the iterative model, wherein the base model is trained to separate floaters by comparing outputs of the base model to outputs a detection algorithm and minimizing a loss function characterizing a difference between the outputs of the base model and the outputs of the detection algorithm, and wherein the iterative model is trained to separate floaters by: comparing outputs of the iterative model to outputs of the detection algorithm; inputting the outputs into the iterative model to generate second outputs; and comparing the second outputs to the outputs of the detection algorithm.
[0183] Embodiment 6. The computer-implemented method of embodiment 4, wherein the detection algorithm is a fast iterative shrinkage-thresholding algorithm (FISTA).
[0184] Embodiment 7. The computer-implemented method of embodiment 1, further comprising: generating an output image depicting the detected floaters, wherein the output image includes at least one of a bounding box or a segment mask indicating locations of the detected floaters.
[0185] Embodiment 8. The computer-implemented method of embodiment 7, wherein the output image further includes at least one label indicating an identifier of at least one of the detected floaters.
[0186] Embodiment 9. The computer-implemented method of embodiment 1 , further comprising: analyzing the sequence of images to determine a motion of at least one floater, the motion comprising one or more of a translation or a rotation.
[0187] Embodiment 10. The computer-implemented method of embodiment 1 , further comprising: determining a continuous tracking ratio for a floater of the one or more floaters, wherein the continuous tracking ratio is determined as a ratio of a number of times a most frequently assigned identifier is assigned to the floater and a total number of images in which the floater appears.
[0188] Embodiment 11. A system for floater detection in an eye of a subject, the system comprising: at least one processor; and a non-transitory, computer-readable medium having instructions stored thereon that, when executed by the at least one processor, cause the system to execute operations comprising: accessing, a sequence of images depicting one or more floaters in the eye of the subject, wherein each image of the sequence of images is a scanning laser ophthalmoscopy (SLO) image; accessing a set of floater identifiers for at least a subset of the one or more floaters; generating an input for a detection and tracking algorithm, wherein the input comprises at least one image of the sequence of images and at least one floater identifier from the set of floater identifiers for at least one of the one or more floaters; detecting, by applying the detection and tracking algorithm to the input, the one or more floaters, wherein the detection and tracking algorithm is configured to decompose a contour representing a plurality of spatially proximate floaters into a plurality of distinct floaters; and generating an output, wherein the output comprises a heatmap depicting locations of the one or more floaters.
[0189] Embodiment 12. The system of embodiment 11 , herein the instructions are further configured to cause the system to determine a depth focus for optical coherence tomography (OCT) imaging of a floater of the one or more floaters, wherein determining the depth focus comprises: accessing a set of OCT scans and a set of focus metadata, the set of focus metadata indicating a focus depth associated with each OCT scan of the set of OCT scans; analyzing the set of OCT scans to determine an optimal OCT scan, wherein the optimal OCT scan is selected based on at least one of: clarity, prominence,or signal strength; and determining, based on the optimal OCT scan and the set of focus metadata, an optimal focus depth.
[0190] Embodiment 13. The system of embodiment 11 , wherein the detection and tracking algorithm is trained using synthetic training data, wherein the synthetic training data is generated by: accessing an optical coherence tomography (OCT) scan depicting a floater; selecting a region of the OCT scan, the region inclusive of the floater and representing a volume that is smaller than a volume of the OCT scan; generating a crosssection of the floater; and projecting the cross-section of the floater onto a scanning laser ophthalmoscopy (SLO) image.
[0191] Embodiment 14. The system of embodiment 11 , wherein decomposing the contour comprises: providing the input to a first machine learning model, wherein the first machine learning model is trained to separate floaters by comparing an output of the first machine learning model to an output of a detection algorithm and minimizing a loss function characterizing a difference between the output of the first machine learning model and the output of the detection algorithm.
[0192] Embodiment 15. The system of embodiment 11 , wherein decomposing the contour comprises: applying an iterative model to the input to generate a first output; applying a base model to the first output to generate a second output; and merging the first output and the second output, wherein the first output comprises an image depicting the locations of floaters, and wherein the second output comprises an image depicting the locations of floaters above a threshold size, wherein the locations determined by the base model are determined with higher precision than the locations determined by the iterative model, wherein the base model is trained to separate floaters by comparing outputs of the base model to outputs a detection algorithm and minimizing a loss function characterizing a difference between the outputs of the base model and the outputs of the detection algorithm, and wherein the iterative model is trained to separate floaters by: comparing outputs of the iterative model to outputs of the detection algorithm; inputting the outputs into the iterative model to generate second outputs; and comparing the second outputs to the outputs of the detection algorithm.
[0193] Embodiment 16. The system of embodiment 15, wherein the detection algorithm is a fast iterative shrinkage-thresholding algorithm (FIST A).
[0194] Embodiment 17. The system of embodiment 11 , herein the instructions are further configured to cause the system to: generate an output image depicting the detected floaters, wherein the output image includes at least one of a bounding box or a segment mask indicating locations of the detected floaters.
[0195] Embodiment 18. The system of embodiment 17, wherein the output image further includes at least one label indicating an identifier of at least one of the detected floaters.
[0196] Embodiment 19. The system of embodiment 11 , herein the instructions are further configured to cause the system to: analyze the sequence of images to determine a motion of at least one floater, the motion comprising one or more of a translation or a rotation.
[0197] Embodiment 20. The system of embodiment 11 , wherein the instructions are further configured to cause the system to: determine a continuous tracking ratio for a floater of the one or more floaters, wherein the continuous tracking ratio is determined as a ratio of a number of times a most frequently assigned identifier is assigned to the floater and a total number of images in which the floater appears.Computer Systems
[0198] Figure 17 is a block diagram depicting an embodiment of a computer hardware system configured to run software for implementing the approaches for floater identification, tracking, and imaging and any systems, methods, and devices disclosed herein. The example computer system 1702 is in communication with one or more computing systems 1720 and / or one or more data sources 1722 via one or more networks 1718. While Figure 17 illustrates an embodiment of a computing system 1702, it is recognized that the functionality provided for in the components and modules of computer system 1702 may be combined into fewer components and modules, or further separated into additional components and modules.
[0199] The computer system 1702 can comprise a module 1714 that carries out the functions, methods, acts, and / or processes described herein. The module 1714 is executed on the computer system 1702 by a central processing unit 1706 discussed further below.
[0200] In general, the word “module,” as used herein, refers to logic embodied in hardware or firmware or to a collection of software instructions, having entry and exit points. Modules are written in a program language, such as Java, C or C++, Python or the like. Software modules may be compiled or linked into an executable program, installed in a dynamic link library, or may be written in an interpreted language such as BASIC, PERL, Lua, or Python. Software modules may be called from other modules or from themselves, and / or may be invoked in response to detected events or interruptions. Modules implemented in hardware include connected logic units such as gates and flipflops, and / or may include programmable units, such as programmable gate arrays or processors.
[0201] Generally, the modules described herein refer to logical modules that may be combined with other modules or divided into sub-modules despite their physical organization or storage. The modules are executed by one or more computing systems and may be stored on or within any suitable computer readable medium or implemented in-whole or in-part within specially designed hardware or firmware. Not all calculations, analysis, and / or optimization require the use of computer systems, though any of the above-described methods, calculations, processes, or analyses may be facilitated through the use of computers. Further, in some embodiments, process blocks described herein may be altered, rearranged, combined, and / or omitted.
[0202] The computer system 1702 includes one or more processing units (CPU) 1706, which may comprise a microprocessor. The computer system 1702 further includes a physical memory 1710, such as random-access memory (RAM) for temporary storage of information, a read only memory (ROM) for permanent storage of information, and a mass storage device 1704, such as a backing store, hard drive, rotating magnetic disks, solid state disks (SSD), flash memory, phase-change memory (PCM), 3D XPoint memory, diskette, or optical media storage device. Alternatively, the mass storage device may be implemented in an array of servers. Typically, the components of the computer system 1702 are connected to the computer using a standards-based bus system. The bus system can be implemented using various protocols, such as Peripheral Component Interconnect (PCI), Micro Channel, SCSI, Industrial Standard Architecture (ISA) and Extended ISA (EISA) architectures.
[0203] The computer system 1702 includes one or more input / output (I / O) devices and interfaces 1712, such as a keyboard, mouse, touch pad, and printer. The I / O devices and interfaces 1712 can include one or more display devices, such as a monitor, which allows the visual presentation of data to a user. More particularly, a display device provides for the presentation of GUIs as application software data, and multi-media presentations, for example. The I / O devices and interfaces 1712 can also provide a communications interface to various external devices. The computer system 1702 may comprise one or more multi-media devices 1708, such as speakers, video cards, graphics accelerators, and microphones, for example.
[0204] The computer system 1702 may run on a variety of computing devices, such as a server, a Windows server, a Structure Query Language server, a Unix Server, a personal computer, a laptop computer, and so forth. In other embodiments, the computer system 1702 may run on a cluster computer system, a mainframe computer system and / or other computing system suitable for controlling and / or communicating with large databases, performing high volume transaction processing, and generating reports from large databases. The computing system 1702 is generally controlled and coordinated by an operating system software, such as z / OS, Windows, Linux, UNIX, BSD, SunOS, Solaris, MacOS, or other compatible operating systems, including proprietary operating systems. Operating systems control and schedule computer processes for execution, perform memory management, provide file system, networking, and I / O services, and provide a user interface, such as a graphical user interface (GUI), among other things.
[0205] The computer system 1702 illustrated in Figure 17 is coupled to a network 1718, such as a LAN, WAN, or the Internet via a communication link 1716 (wired, wireless, or a combination thereof). Network 1718 communicates with various computing devices and / or other electronic devices. Network 1718 is communicating with one or more computing systems 1720 and one or more data sources 1722. The module 1714 may access or may be accessed by computing systems 1720 and / or data sources 1722 through a web-enabled user access point. Connections may be a direct physical connection, a virtual connection, and other connection type. The web-enabled user access point may comprise a browser module that uses text, graphics, audio, video, and other media to present data and to allow interaction with data via the network 1718.
[0206] Access to the module 1714 of the computer system 1702 by computing systems 1720 and / or by data sources 1722 may be through a web-enabled user access point such as the computing systems’ 1720 or data source’s 1722 personal computer, cellular phone, smartphone, laptop, tablet computer, e-reader device, audio player, or another device capable of connecting to the network 1718. Such a device may have a browser module that is implemented as a module that uses text, graphics, audio, video, and other media to present data and to allow interaction with data via the network 1718.
[0207] The output module may be implemented as a combination of an all-points addressable display such as a cathode ray tube (CRT), a liquid crystal display (LCD), a plasma display, or other types and / or combinations of displays. The output module may be implemented to communicate with input devices 1712 and they also include software with the appropriate interfaces which allow a user to access data through the use of stylized screen elements, such as menus, windows, dialogue boxes, tool bars, and controls (for example, radio buttons, check boxes, sliding scales, and so forth). Furthermore, the output module may communicate with a set of input and output devices to receive signals from the user.
[0208] The input device(s) may comprise a keyboard, roller ball, pen and stylus, mouse, trackball, voice recognition system, or pre-designated switches or buttons. The output device(s) may comprise a speaker, a display screen, a printer, or a voice synthesizer. In addition, a touch screen may act as a hybrid input / output device. In another embodiment, a user may interact with the system more directly such as through a system terminal connected to the score generator without communication over the Internet, a WAN, or LAN, or similar network.
[0209] In some embodiments, the system 1702 may comprise a physical or logical connection established between a remote microprocessor and a mainframe host computer for the express purpose of uploading, downloading, or viewing interactive data and databases on-line in real time. The remote microprocessor may be operated by an entity operating the computer system 1702, including the client server systems or the main server system, and / or may be operated by one or more of the data sources 1722 and / or one or more of the computing systems 1720. In some embodiments, terminal emulation software may be used on the microprocessor for participating in the micromainframe link.
[0210] In some embodiments, computing systems 1720 who are internal to an entity operating the computer system 1702 may access the module 1714 internally as an application or process run by the CPU 1706.
[0211] In some embodiments, one or more features of the systems, methods, and devices described herein can utilize a URL and / or cookies, for example for storing and / or transmitting data or user information. A Uniform Resource Locator (URL) can include a web address and / or a reference to a web resource that is stored on a database and / or a server. The URL can specify the location of the resource on a computer and / or a computer network. The URL can include a mechanism to retrieve the network resource. The source of the network resource can receive a URL, identify the location of the web resource, and transmit the web resource back to the requestor. A URL can be converted to an IP address, and a Domain Name System (DNS) can look up the URL and its corresponding IP address. URLs can be references to web pages, file transfers, emails, database accesses, and other applications. The URLs can include a sequence of characters that identify a path, domain name, a file extension, a host name, a query, a fragment, scheme, a protocol identifier, a port number, a username, a password, a flag, an object, a resource name and / or the like. The systems disclosed herein can generate, receive, transmit, apply, parse, serialize, render, and / or perform an action on a URL.
[0212] A cookie, also referred to as an HTTP cookie, a web cookie, an internet cookie, and a browser cookie, can include data sent from a website and / or stored on a user’s computer. This data can be stored by a user’s web browser while the user is browsing. The cookies can include useful information for websites to remember prior browsing information, such as a shopping cart on an online store, clicking of buttons, login information, and / or records of web pages or network resources visited in the past. Cookies can also include information that the user enters, such as names, addresses, passwords, credit card information, etc. Cookies can also perform computer functions. For example, authentication cookies can be used by applications (for example, a web browser) to identify whether the user is already logged in (for example, to a web site). The cookie data can be encrypted to provide security for the consumer. Tracking cookies can be used to compile historical browsing histories of individuals. Systems disclosed herein can generate and use cookies to access data of an individual. Systems can also generate and use JSON web tokens to store authenticity information, HTTPauthentication as authentication protocols, IP addresses to track session or identity information, URLs, and the like.
[0213] The computing system 1702 may include one or more internal and / or external data sources (for example, data sources 1722). In some embodiments, one or more of the data repositories and the data sources described above may be implemented using a relational database, such as DB2, Sybase, Oracle, CodeBase, and Microsoft® SQL Server as well as other types of databases such as a flat-file database, an entity relationship database, and object-oriented database, and / or a record-based database.
[0214] The computer system 1702 may also access one or more databases 1722. The databases 1722 may be stored in a database or data repository. The computer system 1702 may access the one or more databases 1722 through a network 1718 or may directly access the database or data repository through I / O devices and interfaces 1712. The data repository storing the one or more databases 1722 may reside within the computer system 1702.Additional Embodiments
[0215] In the foregoing specification, the systems and processes have been described with reference to specific embodiments thereof. It will, however, be evident that various modifications and changes may be made thereto without departing from the broader spirit and scope of the embodiments disclosed herein. The specification and drawings are, accordingly, to be regarded in an illustrative rather than restrictive sense.
[0216] Indeed, although the systems and processes have been disclosed in the context of certain embodiments and examples, it will be understood by those skilled in the art that the various embodiments of the systems and processes extend beyond the specifically disclosed embodiments to other alternative embodiments and / or uses of the systems and processes and obvious modifications and equivalents thereof. In addition, while several variations of the embodiments of the systems and processes have been shown and described in detail, other modifications, which are within the scope of this disclosure, will be readily apparent to those of skill in the art based upon this disclosure. It is also contemplated that various combinations or sub-combinations of the specific features and aspects of the embodiments may be made and still fall within the scope of the disclosure. It should be understood that various features and aspects of the disclosedembodiments can be combined with, or substituted for, one another in order to form varying modes of the embodiments of the disclosed systems and processes. Any methods disclosed herein need not be performed in the order recited. Thus, it is intended that the scope of the systems and processes herein disclosed should not be limited by the particular embodiments described above.
[0217] It will be appreciated that the systems and methods of the disclosure each have several innovative aspects, no single one of which is solely responsible or required for the desirable attributes disclosed herein. The various features and processes described above may be used independently of one another or may be combined in various ways. All possible combinations and sub-combinations are intended to fall within the scope of this disclosure.
[0218] Certain features that are described in this specification in the context of separate embodiments also may be implemented in combination in a single embodiment. Conversely, various features that are described in the context of a single embodiment also may be implemented in multiple embodiments separately or in any suitable subcombination. Moreover, although features may be described above as acting in certain combinations and even initially claimed as such, one or more features from a claimed combination may in some cases be excised from the combination, and the claimed combination may be directed to a sub-combination or variation of a sub-combination. No single feature or group of features is necessary or indispensable to each and every embodiment.
[0219] It will also be appreciated that conditional language used herein, such as, among others, “can,” “could,” “might,” “may,” “for example,” and the like, unless specifically stated otherwise, or otherwise understood within the context as used, is generally intended to convey that certain embodiments include, while other embodiments do not include, certain features, elements and / or steps. Thus, such conditional language is not generally intended to imply that features, elements and / or steps are in any way required for one or more embodiments or that one or more embodiments necessarily include logic for deciding, with or without author input or prompting, whether these features, elements and / or steps are included or are to be performed in any particular embodiment. The terms “comprising,” “including,” “having,” and the like are synonymous and are used inclusively, in an open- ended fashion, and do not exclude additionalelements, features, acts, operations, and so forth. In addition, the term “or” is used in its inclusive sense (and not in its exclusive sense) so that when used, for example, to connect a list of elements, the term “or” means one, some, or all of the elements in the list. In addition, the articles “a,” “an,” and “the” as used in this application and the appended claims are to be construed to mean “one or more” or “at least one” unless specified otherwise. Similarly, while operations may be depicted in the drawings in a particular order, it is to be recognized that such operations need not be performed in the particular order shown or in sequential order, or that all illustrated operations be performed, to achieve desirable results. Further, the drawings may schematically depict one or more example processes in the form of a flowchart. However, other operations that are not depicted may be incorporated in the example methods and processes that are schematically illustrated. For example, one or more additional operations may be performed before, after, simultaneously, or between any of the illustrated operations. Additionally, the operations may be rearranged or reordered in other embodiments. In certain circumstances, multitasking and parallel processing may be advantageous. Moreover, the separation of various system components in the embodiments described above should not be understood as requiring such separation in all embodiments, and it should be understood that the described program components and systems may generally be integrated together in a single software product or packaged into multiple software products. Additionally, other embodiments are within the scope of the following claims. In some cases, the actions recited in the claims may be performed in a different order and still achieve desirable results.
[0220] Further, while the methods and devices described herein may be susceptible to various modifications and alternative forms, specific examples thereof have been shown in the drawings and are herein described in detail. It should be understood, however, that the embodiments are not to be limited to the particular forms or methods disclosed, but, to the contrary, the embodiments are to cover all modifications, equivalents, and alternatives falling within the spirit and scope of the various implementations described and the appended claims. Further, the disclosure herein of any particular feature, aspect, method, property, characteristic, quality, attribute, element, or the like in connection with an implementation or embodiment can be used in all other implementations or embodiments set forth herein. Any methods disclosed herein need not be performed in the order recited. The methods disclosed herein may includecertain actions taken by a practitioner; however, the methods can also include any third- party instruction of those actions, either expressly or by implication. The ranges disclosed herein also encompass any and all overlap, sub-ranges, and combinations thereof. Language such as “up to,” “at least,” “greater than,” “less than,” “between,” and the like includes the number recited. Numbers preceded by a term such as “about” or “approximately” include the recited numbers and should be interpreted based on the circumstances (for example, as accurate as reasonably possible under the circumstances, for example ±5%, ±10%, ±15%, etc.). For example, “about 3.5 mm” includes “3.5 mm.” Phrases preceded by a term such as “substantially” include the recited phrase and should be interpreted based on the circumstances (for example, as much as reasonably possible under the circumstances). For example, “substantially constant” includes “constant.” Unless stated otherwise, all measurements are at standard conditions including temperature and pressure.
[0221] As used herein, a phrase referring to “at least one of’ a list of items refers to any combination of those items, including single members. As an example, “at least one of: A, B, or C” is intended to cover: A, B, C, A and B, A and C, B and C, and A, B, and C. Conjunctive language such as the phrase “at least one of X, Y and Z,” unless specifically stated otherwise, is otherwise understood with the context as used in general to convey that an item, term, etc. may be at least one of X, Y or Z. Thus, such conjunctive language is not generally intended to imply that certain embodiments require at least one of X, at least one of Y, and at least one of Z to each be present. The headings provided herein, if any, are for convenience only and do not necessarily affect the scope or meaning of the devices and methods disclosed herein.
[0222] Accordingly, the claims are not intended to be limited to the embodiments shown herein but are to be accorded the widest scope consistent with this disclosure, the principles and the novel features disclosed herein.
Claims
CLAIMSWhat is claimed is:
1. A computer-implemented method for floater detection in an eye of a subject, the computer-implemented method comprising: accessing a sequence of images depicting one or more floaters in the eye of the subject, wherein each image of the sequence of images is a scanning laser ophthalmoscopy (SLO) image; accessing a set of floater identifiers for at least a subset of the one or more floaters; generating an input for a detection and tracking algorithm, wherein the input comprises at least one image of the sequence of images and at least one floater identifier from the set of floater identifiers for at least one of the one or more floaters; detecting, by applying the detection and tracking algorithm to the input, the one or more floaters, wherein the detection and tracking algorithm is configured to decompose a contour representing a plurality of spatially proximate floaters into a plurality of distinct floaters; and generating an output, wherein the output comprises a heatmap depicting locations of the one or more floaters.
2. The computer-implemented method of claim 1 , further comprising determining a depth focus for optical coherence tomography (OCT) imaging of a floater of the one or more floaters, wherein determining the depth focus comprises: accessing a set of OCT scans and a set of focus metadata, the set of focus metadata indicating a focus depth associated with each OCT scan of the set of OCT scans; analyzing the set of OCT scans to determine an optimal OCT scan, wherein the optimal OCT scan is selected based on at least one of: clarity, prominence, or signal strength; anddetermining, based on the optimal OCT scan and the set of focus metadata, an optimal focus depth.
3. The computer-implemented method of claim 1 , wherein the detection and tracking algorithm is trained using synthetic training data, wherein the synthetic training data is generated by: accessing an optical coherence tomography (OCT) scan depicting a floater; selecting a region of the OCT scan, the region inclusive of the floater and representing a volume that is smaller than a volume of the OCT scan; generating a cross-section of the floater; and projecting the cross-section of the floater onto a scanning laser ophthalmoscopy (SLO) image.
4. The computer-implemented method of claim 1 , wherein decomposing the contour comprises: providing the input to a first machine learning model, wherein the first machine learning model is trained to separate floaters by comparing an output of the first machine learning model to an output of a detection algorithm and minimizing a loss function characterizing a difference between the output of the first machine learning model and the output of the detection algorithm.
5. The computer-implemented method of claim 1 , wherein decomposing the contour comprises: applying an iterative model to the input to generate a first output; applying a base model to the first output to generate a second output; and merging the first output and the second output, wherein the first output comprises an image depicting the locations of floaters, and wherein the second output comprises an image depicting the locations of floaters above a threshold size, wherein the locations determined by the base model are determined with higher precision than the locations determined by the iterative model,wherein the base model is trained to separate floaters by comparing outputs of the base model to outputs a detection algorithm and minimizing a loss function characterizing a difference between the outputs of the base model and the outputs of the detection algorithm, and wherein the iterative model is trained to separate floaters by: comparing outputs of the iterative model to outputs of the detection algorithm; inputting the outputs into the iterative model to generate second outputs; and comparing the second outputs to the outputs of the detection algorithm.
6. The computer-implemented method of claim 4, wherein the detection algorithm is a fast iterative shrinkage-thresholding algorithm (FIST A).
7. The computer-implemented method of claim 1 , further comprising: generating an output image depicting the detected floaters, wherein the output image includes at least one of a bounding box or a segment mask indicating locations of the detected floaters.
8. The computer-implemented method of claim 7, wherein the output image further includes at least one label indicating an identifier of at least one of the detected floaters.
9. The computer-implemented method of claim 1 , further comprising: analyzing the sequence of images to determine a motion of at least one floater, the motion comprising one or more of a translation or a rotation.
10. The computer-implemented method of claim 1 , further comprising: determining a continuous tracking ratio for a floater of the one or more floaters, wherein the continuous tracking ratio is determined as a ratio of a number of times a most frequently assigned identifier is assigned to the floater and a total number of images in which the floater appears.11 . A system for floater detection in an eye of a subject, the system comprising: at least one processor; and a non-transitory, computer-readable medium having instructions stored thereon that, when executed by the at least one processor, cause the system to execute operations comprising: accessing, a sequence of images depicting one or more floaters in the eye of the subject, wherein each image of the sequence of images is a scanning laser ophthalmoscopy (SLO) image; accessing a set of floater identifiers for at least a subset of the one or more floaters; generating an input for a detection and tracking algorithm, wherein the input comprises at least one image of the sequence of images and at least one floater identifier from the set of floater identifiers for at least one of the one or more floaters; detecting, by applying the detection and tracking algorithm to the input, the one or more floaters, wherein the detection and tracking algorithm is configured to decompose a contour representing a plurality of spatially proximate floaters into a plurality of distinct floaters; and generating an output, wherein the output comprises a heatmap depicting locations of the one or more floaters.
12. The system of claim 11 , herein the instructions are further configured to cause the system to determine a depth focus for optical coherence tomography (OCT) imaging of a floater of the one or more floaters, wherein determining the depth focus comprises: accessing a set of OCT scans and a set of focus metadata, the set of focus metadata indicating a focus depth associated with each OCT scan of the set of OCT scans; analyzing the set of OCT scans to determine an optimal OCT scan, wherein the optimal OCT scan is selected based on at least one of: clarity, prominence, or signal strength; anddetermining, based on the optimal OCT scan and the set of focus metadata, an optimal focus depth.
13. The system of claim 11 , wherein the detection and tracking algorithm is trained using synthetic training data, wherein the synthetic training data is generated by: accessing an optical coherence tomography (OCT) scan depicting a floater; selecting a region of the OCT scan, the region inclusive of the floater and representing a volume that is smaller than a volume of the OCT scan; generating a cross-section of the floater; and projecting the cross-section of the floater onto a scanning laser ophthalmoscopy (SLO) image.
14. The system of claim 11 , wherein decomposing the contour comprises: providing the input to a first machine learning model, wherein the first machine learning model is trained to separate floaters by comparing an output of the first machine learning model to an output of a detection algorithm and minimizing a loss function characterizing a difference between the output of the first machine learning model and the output of the detection algorithm.
15. The system of claim 11 , wherein decomposing the contour comprises: applying an iterative model to the input to generate a first output; applying a base model to the first output to generate a second output; and merging the first output and the second output, wherein the first output comprises an image depicting the locations of floaters, and wherein the second output comprises an image depicting the locations of floaters above a threshold size, wherein the locations determined by the base model are determined with higher precision than the locations determined by the iterative model, wherein the base model is trained to separate floaters by comparing outputs of the base model to outputs a detection algorithm and minimizing a lossfunction characterizing a difference between the outputs of the base model and the outputs of the detection algorithm, and wherein the iterative model is trained to separate floaters by: comparing outputs of the iterative model to outputs of the detection algorithm; inputting the outputs into the iterative model to generate second outputs; and comparing the second outputs to the outputs of the detection algorithm.
16. The system of claim 15, wherein the detection algorithm is a fast iterative shrinkage-thresholding algorithm (FISTA).
17. The system of claim 11 , herein the instructions are further configured to cause the system to: generate an output image depicting the detected floaters, wherein the output image includes at least one of a bounding box or a segment mask indicating locations of the detected floaters.
18. The system of claim 17, wherein the output image further includes at least one label indicating an identifier of at least one of the detected floaters.
19. The system of claim 11 , herein the instructions are further configured to cause the system to: analyze the sequence of images to determine a motion of at least one floater, the motion comprising one or more of a translation or a rotation.
20. The system of claim 11 , wherein the instructions are further configured to cause the system to: determine a continuous tracking ratio for a floater of the one or more floaters, wherein the continuous tracking ratio is determined as a ratio of a number of times a most frequently assigned identifier is assigned to the floater and a total number of images in which the floater appears.
Citation Information
Patent Citations
System and method for detection of floaters
CA3140678A1
System and method for detection of floaters
WO2023097391A1