Unified GAN-feature map-based framework for comprehensive oct image analysis and biomarker segmentation

WO2026206702A1PCT designated stage Publication Date: 2026-10-01UNIV OF PITTSBURGH OF THE COMMONWEALTH SYST OF HIGHER EDUCATION
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/US2026/019704
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-04-14
Filing Date
2026-03-18
Publication Date
2026-10-01

Smart Images

  • Figure US2026019704_01102026_PF_FP_ABST
    Figure US2026019704_01102026_PF_FP_ABST
Patent Text Reader

Abstract

Embodiments described herein provide for a system that is configured to segmenting optical coherence tomography (OCT) images to analyze structures represented by the OCT images. In examples, volumetric data is obtained from an OCT imaging device, the volumetric data representing a three-dimensional image of an eyeball. The system can then extract one or more feature maps to determine segmentations, indicating delineations between structures of the eyeball. The system can then identify a first and second boundary of a structure based on these feature maps and segment the structure. In some examples, the system can also be configured to execute morphological operations when analyzing the structure, or the components of the structure.
Need to check novelty before this filing date? Find Prior Art

Description

076333-1065 / 07025 PATENTUNIFIED GAN-FEATURE MAP-BASED FRAMEWORK FOR COMPREHENSIVE OCT IMAGE ANALYSIS AND BIOMARKER SEGMENTATIONSTATEMENT REGARDING FEDERALLY SPONSORED RESEARCH OR DEVELOPMENT

[0001] This invention was made with government support under EY008098 awarded by the National Institutes of Health. The government has certain rights in the invention.CROSS-REFERENCE TO RELATED APPLICATIONS

[0002] This application claims the benefit of and priority to U.S. Provisional Application No. 63 / 788,646, entitled “Unified Gan-Feature Map-Based Framework for Comprehensive OCT Image Analysis and Biomarker Segmentation,” filed April 14, 2025, and U.S. Provisional Patent Application No. US 63 / 777,528, entitled “AI-Driven Analysis for Drusen Segmentation, Geographic Atrophy Quantification, and Ellipsoidal Zone Loss Detection,” filed March 25, 2025, each of which is incorporated by reference in its entirety.TECHNICAL FIELD

[0003] This application generally relates to techniques for analyzing structures included in a retina and, in some embodiments, to techniques for segmenting optical coherence tomography (OCT) images to analyze structures represented by the OCT images.BACKGROUND

[0004] When analyzing OCT images to identify regions of a retina that are indicative of disease, clinicians can review OCT images to detect retinal pathologies associated with macular degeneration, diabetic retinopathy, retinal vein occlusion, and the like. This manual analysis can involve a trained clinician interpreting OCT images to identify biomarkers like retinal thickness changes, fluid accumulations, or structural abnormalities. But relying solely on a clinician’s expertise can be time-intensive and highly subjective.SUMMARY

[0005] Techniques can be implemented to introduce machine learning models (referred to generally as “models”) to streamline portions of this process, achieving high accuracy in identifying specific pathological signs and improving diagnostic efficiency. These systems can leverage labeled datasets to train such models to be capable of distinguishing healthy retinas from076333-1065 / 07025 PATENTdiseased ones, as well as identifying key abnormalities with increased precision relative to clinicians. But several technical challenges remain. For example, when implementing certain machine learning-based techniques, significant amounts of computing resources can still be involved in processing the OCT images. And while these techniques can streamline certain aspects of the analysis, others that cannot be addressed through the use of machine learning-based techniques can remain prone to variability due to differences in expertise among clinicians, leading to possibly inconsistent diagnoses. Further, as the resolution and complexity of OCT imaging improves, the volume of data generated can increase exponentially, further straining memory and processing resources.

[0006] Embodiments of the present disclosure address certain technical difficulties using techniques that involve segmenting OCT images to analyze structures represented by the OCT images (e.g., indicative of the presence of one or more diseases). For example, a system, such as a server, can be configured to obtain volumetric data, including a plurality of cross-sectional images received from an OCT imaging device, the volumetric data representing three-dimensional image of an eyeball. The system can be configured to extract a plurality of feature maps, based on the volumetric data, to determine a plurality of segmentations across the plurality of cross-sectional images. In some examples, the system can be configured to determine a first boundary of a structure and a second boundary of the structure based on at least one feature map of the plurality of feature maps. In examples, the system can be configured to segment a plurality of substructures within a region established by the first boundary and the second boundary. In some examples, the system can be configured to generate an image of at least a portion of the structure of the eyeball based on the volumetric data and the plurality of substructures to indicate locations of the plurality of substructures within the structure.

[0007] By implementing the techniques described, a variety of efficiencies can be realized. First, system can be configured to more efficiently process and analyze volumetric data by segmenting and focusing on relevant regions of the retina of a patient. This can significantly reduce processing requirements and memory consumption that would otherwise be expended analyzing the entirety of the volumetric data. And by storing detailed data only for segmented portions, rather than the entire dataset, memory resources are used more efficiently. The system can also analyze higher resolution images that might otherwise be unprocessable due to computing constraints as a076333-1065 / 07025 PATENTsmaller portion of the image is analyzed. Additionally, machine learning models, as described herein, can be executed in accordance with the segmented data to improve accuracy, while minimizing computational demands.

[0008] Furthermore, embodiments of the present disclosure address certain technical difficulties using techniques that involve segmenting and analyzing OCT images to identify and quantify specific features within the OCT images (e g., indicative of the presence of one or more diseases). For example, a system such as a server can be configured to obtain volumetric data including a plurality of cross-sectional images received from an OCT imaging device. The server can then be configured to determine a plurality of points associated with a boundary of a structure of the eyeball represented in at least a subset of these cross-sectional images. In some examples, the server can extract a plurality of patches from the volumetric data that corresponds to the boundary of the structure. The server can then determine a first set of patches having a first structural characteristic and a second set of patches having a second structural characteristic. In some examples, the server can generate a graphical user interface based on the plurality of patches to indicate a first region within the eyeball where the first structural characteristic is present and a second region within the eyeball where the second structural characteristic is present.

[0009] By implementing the techniques described, a variety of efficiencies can be realized. First, computing resource consumption can be reduced as the systems described can extract and process patches from the OCT images using these boundaries. In turn, the computing resources that would otherwise be consumed during analysis of the OCT images can be reduced based on the corresponding reduction in the amount of data to be processed. This targeted approach can allow for a reduced overall computational load on the systems when operating at inference.

[0010] Memory consumption can be reduced as certain operations are executed using only the extracted patches instead of the entire volumetric dataset during portions of the described process. By identifying and retaining only the patches that are associated with the boundary of the structure(s) being targeted, a server can significantly decrease the amount of data that needs to be stored in memory. This selective storage approach can allow for memory resources to be used more efficiently, resulting in the handling of larger datasets without exceeding memory capacity.076333-1065 / 07025 PATENT

[0011] When configured as described herein, a server can reduce the amount of network communications that would otherwise be consumed (either internally or externally) by minimizing the amount of data that needs to be transmitted between the OCT imaging device and the server. By processing and extracting the relevant patches locally before transmitting them to the server, the invention can reduce the volume of data sent over the network. This reduction in data transmission not only decreases network bandwidth usage but also speeds up the overall processing time.

[0012] Moreover, a server programmed as described can generate more accurate predictions of whether a disease (e.g., geographic atrophy (GA)) is or is not indicated by a set of OCT images is present in the patient by focusing on the structural characteristics of the extracted patches. By analyzing the first and second sets of patches with distinct structural characteristics, the server can identify regions within the eyeball that exhibit signs of the disease and subsequently focus on these regions in the OCT images. This targeted analysis can allow for more precise detection and diagnosis, leading to improved accuracy in predicting the presence of GA and other ocular diseases.

[0013] In some aspects, a system for segmenting and analyzing OCT images to identify and quantify specific features within the OCT images is disclosed. The system can include or more processors that can be configured to: obtain volumetric data, including a plurality of cross-sectional images received from an OCT imaging device, the volumetric data representing three-dimensional image of an eyeball. The one or more processors can be configured to extract a plurality of feature maps, based on the volumetric data, to determine a plurality of segmentations across the plurality of cross-sectional images. In examples, the one or more processors can be configured to determine a first boundary of a structure and a second boundary of the structure based on at least one feature map of the plurality of feature maps. In some examples, the one or more processors can be configured to segment a plurality of substructures within a region established by the first boundary and the second boundary. In at least some examples, the one or more processors can be configured to generate an image of at least a portion of the structure of the eyeball, based on the volumetric data and the plurality of substructures, to indicate locations of the plurality of substructures within the structure.076333-1065 / 07025 PATENT

[0014] In some aspects, the one or more processors can be configured to determine the first boundary and the second boundary can be configured to: update the at least one feature map by executing one or more morphological operations to adjust a representation of the first boundary and the second boundary. In aspects, the one or more processors configured to determine the first boundary and the second boundary can be configured to compare a relative position of the first boundary and the second boundary within the structure to determine a first type corresponding to the first boundary and a second type corresponding to the second boundary. In at least some aspects, the one or more processors can be further configured to determine one or more aspects of the plurality of substructures of the eyeball to augment in the image of the at least a portion of the structure. The one or more processors configured to determine the first boundary of the structure and the second boundary of the structure can be configured to determine the first boundary of a structure and the second boundary of the structure based on at least one feature map corresponding to the one or more aspects.

[0015] In aspects, the one or more processors can be further configured to: determine one or more aspects of the plurality of substructures of the eyeball to augment in the image of the at least a portion of the structure. The one or more processors configured to segment the plurality of substructures within the region can be configured to segment the plurality of substructures within the region, based on the one or more aspects of the plurality of substructures of the eyeball, to augment in the image. In some aspects, the one or more processors configured to segment the plurality of substructures within the region can be configured to generate a first updated feature map by executing one or more first morphological operations to adjust a representation of the first boundary and the second boundary, and generate a second updated feature map by executing one or more second morphological operations to adjust the representation of the first boundary and the second boundary. The one or more processors can be configured configured to segment the plurality of substructures within the region based on the first updated feature map and the second updated feature map.

[0016] In some aspects, the at least one feature map can include a first feature map, and the plurality of substructures can include a first plurality of substructures. The one or more processors can be further configured to determine a first boundary of a second structure and a second boundary of the second structure based on a second feature map of the plurality of feature076333-1065 / 07025 PATENTmaps. Tn examples, the one or more processors can be configured to segment a second plurality of substructures within the region established by the first boundary and the second boundary. The one or more processors configured to generate the image of the at least a portion of the structure of the eyeball can be configured to generate the image of the at least a portion of the structure of the eyeball based on segmenting the plurality of substructures within the region.

[0017] In some aspects, the techniques described herein relate to a computer-implemented method, including: obtaining volumetric data, including a plurality of cross-sectional images received from an OCT imaging device, the volumetric data representing three-dimensional image of an eyeball. The computer-implemented method can include extracting a plurality of feature maps based on the volumetric data to determine a plurality of segmentations across the plurality of cross-sectional images, and determining a first boundary of a structure and a second boundary of the structure based on at least one feature map of the plurality of feature maps. In examples, the computer method can include segmenting a plurality of substructures within a region established by the first boundary and the second boundary. The computer-implemented method can include generating an image of at least a portion of the structure of the eyeball based on the volumetric data and the plurality of substructures to indicate locations of the plurality of substructures within the structure.

[0018] In some aspects, the determining the first boundary and the second boundary can include updating the at least one feature map by executing one or more morphological operations to adjust a representation of the first boundary and the second boundary. In aspects, determining the first boundary and the second boundary can include comparing a relative position of the first boundary and the second boundary within the structure to determine a first type corresponding to the first boundary and a second type corresponding to the second boundary. In at least some aspects, the computer-implemented method can further include determining one or more aspects of the plurality of substructures of the eyeball to augment in the image of the at least a portion of the structure. Determining the first boundary of the structure and the second boundary of the structure can include determining the first boundary of a structure and the second boundary of the structure based on at least one feature map corresponding to the one or more aspects.

[0019] In some aspects, the computer-implemented method can further include determining one or more aspects of the plurality of substructures of the eyeball to augment in the076333-1065 / 07025 PATENTimage of the at least a portion of the structure. Segmenting the plurality of substructures within the region can include segmenting the plurality of substructures within the region, based on the one or more aspects of the plurality of substructures of the eyeball, to augment in the image. In some aspects, segmenting the plurality of substructures within the region can include generating a first updated feature map by executing one or more first morphological operations to adjust a representation of the first boundary and the second boundary, and generating a second updated feature map by executing one or more second morphological operations to adjust the representation of the first boundary and the second boundary. The computer-implemented method can include segmenting the plurality of substructures within the region based on the first updated feature map and the second updated feature map.

[0020] In aspects, the at least one feature map can include a first feature map and the plurality of substructures can include a first plurality of substructures. The computer-implemented method can include determining a first boundary of a second structure and a second boundary of the second structure based on a second feature map of the plurality of feature maps, and segmenting a second plurality of substructures within the region established by the first boundary and the second boundary. In some aspects, generating the image of the at least a portion of the structure of the eyeball can include generating the image of the at least a portion of the structure of the eyeball based on segmenting the plurality of substructures within the region.

[0021] In some aspects, a non-transitory computer-readable medium is disclosed. The non-transitory computer-readable medium can store instructions thereon that, when executed by one or more processors, cause the one or more processors to obtain volumetric data, including a plurality of cross-sectional images received from an OCT imaging device. The volumetric data can represent a three-dimensional image of an eyeball. The instructions can cause the one or more processors to extract a plurality of feature maps, based on the volumetric data, to determine a plurality of segmentations across the plurality of cross-sectional images. In examples, the instructions can cause the one or more processors to determine a first boundary of a structure and a second boundary of the structure based on at least one feature map of the plurality of feature maps. In some examples, the instructions can cause the one or more processors to segment a plurality of substructures within a region established by the first boundary and the second boundary. In examples, the instructions can cause the one or more processors to generate an image076333-1065 / 07025 PATENTof at least a portion of the structure of the eyeball, based on the volumetric data and the plurality of substructures, to indicate locations of the plurality of substructures within the structure.

[0022] In some aspects, the instructions that cause the one or more processors to determine the first boundary and the second boundary can cause the one or more processors to update the at least one feature map by executing one or more morphological operations to adjust a representation of the first boundary and the second boundary. Tn aspects, the instructions that cause the one or more processors to determine the first boundary and the second boundary can cause the one or more processors to compare a relative position of the first boundary and the second boundary within the structure to determine a first type corresponding to the first boundary and a second type corresponding to the second boundary. In at least some aspects, the instructions can further cause the one or more processors to determine one or more aspects of the plurality of substructures of the eyeball to augment in the image of the at least a portion of the structure. The instructions that cause the one or more processors to determine the first boundary of the structure and the second boundary of the structure can cause the one or more processors to determine the first boundary of a structure and the second boundary of the structure based on at least one feature map corresponding to the one or more aspects.

[0023] In some aspects, a system for segmenting and analyzing OCT images to identify and quantify specific features within the OCT images is disclosed. The system can include one more processors. The one or more processors can be configured to: obtain volumetric data including a plurality of cross-sectional images received from an OCT imaging device, the volumetric data representing three-dimensional image of an eyeball. The one or more processors can be configured to determine a plurality of points associated with a boundary of a structure of the eyeball based on the cross-sectional images, and extract a plurality of patches from the volumetric data that correspond to the boundary of the structure. The one or more processors can be configured to determine a first set of patches having a first structural characteristic and a second set of patches having a second structural characteristic based on the plurality of patches. In some aspects, the one or more processors can be configured to generate a graphical user interface based on the plurality of patches indicating a first region within the eyeball where the first structural characteristic is present and a second region within the eyeball where the second structural characteristic is present.076333-1065 / 07025 PATENT

[0024] In some aspects, the one or more processors configured to extract the plurality of patches from the volumetric data can be configured to: determine a patch size to be used when extracting the plurality of patches; and extract the plurality of patches from the volumetric data that correspond to the boundary of the structure based on the boundary of the structure and the patch size. In aspects, the one or more processors configured to determine the patch size can be configured to: determine the patch size based on a location of one or more adjacent layers to a target layer within the eyeball and a location of a choroid within the eyeball relative to the target layer. In some aspects, the one or more processors configured to determine the first set of patches and the second set of patches can be configured to: execute an artificial neural network based on the plurality of patches to generate a plurality of annotations corresponding to the plurality of patches; and segment the plurality of patches to generate the first set of patches and the second set of patches based on the plurality of annotations.

[0025] In aspects, the plurality of patches can include a first plurality of patches, and the one or more processors configured to determine the first set of patches having the first structural characteristic and the second set of patches having the second structural characteristic can be configured to: execute an artificial neural network by providing at least one patch of patches as an input to the artificial neural network to cause the artificial neural network to execute one or more operations to extract features from the at least one patch at a first resolution and a second resolution to generate at least one feature map. The one or more processors can be configured to determine the first set of patches having the first structural characteristic and the second set of patches having the second structural characteristic based on the at least one feature map.

[0026] In at least some aspects, the one or more processors configured to determine the first set of patches having the first structural characteristic and the second set of patches having the second structural characteristic can be configured to: generate a set of annotations based on determining the first set of patches and the second set of patches. The one or more processors configured to generate the graphical user interface can be configured to: determine the first region within the eyeball where the first structural characteristic is present and the second region within the eyeball where the second structural characteristic is present based on the at least one feature map. The one or more processors can be configured to generate the graphical user interface based on determining the first region within the eyeball and the second region within the eyeball.076333-1065 / 07025 PATENT

[0027] In some aspects, the artificial neural network can include a first artificial neural network, and the one or more processors configured to generate the set of annotations can be configured to: provide the first set of patches and the second set of patches as inputs to a second artificial neural network in accordance with a sequence established by the first set of patches and the second set of patches to cause the second artificial neural network to generate the set of annotations. In aspects, the one or more processors can be further configured to: generate an en face image of at least a portion of the structure of the eyeball based on the volumetric data; execute one or more operations to determine a segmentation mask based on the en face image; and compare the segmentation mask to the first region within the eyeball to determine a third region. The second region can include a subset of the second region. The one or more processors can be configured to generate the graphical user interface can be configured to: generate the graphical user interface to indicate the third region based on determining the third region.

[0028] In at least some aspects, the one or more processors configured to execute the one or more operations to determine the segmentation mask can be configured to: execute an artificial neural network that is configured to receive the en face image as an input and generate the segmentation mask as an output. In some aspects, the one or more processors configured to execute the artificial neural network can be configured to: provide the en face image as an input to the artificial neural network to cause the artificial neural network to execute one or more operations to: iteratively introduce noise to the en face image and, in response to introducing the noise, remove at least a portion of the noise from the en face image to generate the segmentation mask.

[0029] In some aspects, the one or more processors configured to compare the segmentation mask to the first region within the eyeball can be configured to: update the en face image to indicate the first region based on the plurality of patches indicating a first region within the eyeball where the first structural characteristic is present to generate a first annotated image. The one or more processors can be configured to overlay the segmentation mask onto the first annotated image to generate a second annotated image, and determine the third region based on the second annotated image. In some aspects, the one or more processors can be further configured to determine an expected direction for disease progression based on one or more attributes of the third region, and generate the graphical user interface to the expected direction.076333-1065 / 07025 PATENT

[0030] In some aspects, the techniques described herein relate to a computer-implemented method, including: obtaining volumetric data including a plurality of cross-sectional images received from an OCT imaging device, the volumetric data representing three-dimensional image of an eyeball. The method can include determining a plurality of points associated with a boundary of a structure of the eyeball. In examples, the method can include extracting a plurality of patches from the volumetric data that correspond to the boundary of the structure and determining a first set of patches having a first structural characteristic and a second set of patches having a second structural characteristic based on the plurality of patches. The method can include generating a graphical user interface based on the plurality of patches indicating a first region within the eyeball where the first structural characteristic is present and a second region within the eyeball where the second structural characteristic is present.

[0031] In at least some aspects, extracting the plurality of patches from the volumetric data can include determining a patch size to be used when extracting the plurality of patches; and extracting the plurality of patches from the volumetric data that correspond to the boundary of the structure based on the boundary of the structure and the patch size. In aspects, determining the patch size can include determining the patch size based on a location of one or more adjacent layers to a target layer within the eyeball and a location of a choroid within the eyeball relative to the target layer. Determining the first set of patches and the second set of patches can include executing an artificial neural network based on the plurality of patches to generate a plurality of annotations corresponding to the plurality of patches; and segmenting the plurality of patches to generate the first set of patches and the second set of patches based on the plurality of annotations.

[0032] In some aspects, the plurality of patches includes a first plurality of patches and determining the first set of patches having the first structural characteristic and the second set of patches having the second structural characteristic can include: executing an artificial neural network by providing at least one patch of patches as an input to the artificial neural network to cause the artificial neural network to execute one or more operations. Execution of the one or more operations can include extracting features from the at least one patch at a first resolution and a second resolution to generate at least one feature map, and determining the first set of patches having the first structural characteristic and the second set of patches having the second structural characteristic based on the at least one feature map.076333-1065 / 07025 PATENT

[0033] In some aspects, determining the first set of patches having the first structural characteristic and the second set of patches having the second structural characteristic can include generating a set of annotations based on determining the first set of patches and the second set of patches. Generating the graphical user interface can include determining the first region within the eyeball where the first structural characteristic is present and the second region within the eyeball where the second structural characteristic is present based on the at least one feature map; and generating the graphical user interface based on determining the first region within the eyeball and the second region within the eyeball. In aspects, the artificial neural network can include a first artificial neural network, and generating the set of annotations can include providing the first set of patches and the second set of patches as inputs to a second artificial neural network in accordance with a sequence established by the first set of patches and the second set of patches to cause the second artificial neural network to generate the set of annotations.

[0034] In some aspects, a non-transitory computer-readable medium is disclosed. The non-transitory computer-readable medium can store instructions thereon that, when executed by one or more processors, cause the one or more processors to: obtain volumetric data including a plurality of cross-sectional images received from an OCT imaging device. The volumetric data can represent a three-dimensional image of an eyeball. The instructions can cause the one or more processors to determine a plurality of points associated with a boundary of a structure of the eyeball and extract a plurality of patches from the volumetric data that correspond to the boundary of the structure. The instructions can cause the one or more processors to determine a first set of patches having a first structural characteristic and a second set of patches having a second structural characteristic based on the plurality of patches. In examples, the instructions can cause the one or more processors to generate a graphical user interface based on the plurality of patches indicating a first region within the eyeball where the first structural characteristic is present and a second region within the eyeball where the second structural characteristic is present.

[0035] It is to be understood that both the foregoing general description and the following detailed description are exemplary and explanatory and are intended to provide further explanation of the embodiments described herein.076333-1065 / 07025 PATENTBRIEF DESCRIPTION OF THE DRAWINGS

[0036] The accompanying drawings constitute a part of this specification, illustrate one or more embodiments and, together with the specification, explain the subject matter of the disclosure.

[0037] FIG. 1 is a block diagram of an environment for segmenting optical coherence tomography (OCT) images to analyze structures represented by the OCT images, in accordance with one or more embodiments.

[0038] FIG. 2 is a flow chart, illustrating operations of a method for segmenting OCT images to analyze structures represented by the OCT images, in accordance with one or more embodiments.

[0039] FIG. 3 illustrates a non-limiting example of an implementation of a model architecture for generating feature maps, in accordance with one or more embodiments.

[0040] FIG. 4 illustrates non-limiting examples of the generation and post-processing of feature maps, in accordance with one or more embodiments.

[0041] FIG. 5 illustrates a non-limiting example of one or more operations executed to determine the boundaries of one or more structures of a retina, in accordance with one or more embodiments.

[0042] FIG. 6 illustrates a non-limiting example of a process for identifying a boundary of a sclera of a retina, in accordance with one or more embodiments.

[0043] FIG. 7 illustrates a non-limiting example of a process for identifying segmented choroid vasculature of a retina, in accordance with one or more embodiments.

[0044] FIG. 8 illustrates a non-limiting example of a process for identifying segmented portions where fluid is building in a retina, in accordance with one or more embodiments.

[0045] FIG. 9 depicts portions of the retina where pigment epithelial detachment (PED) is identified, in accordance with one or more embodiments.076333-1065 / 07025 PATENT

[0046] FIG. 10 is a flowchart illustrating operations of a processor-implemented or computer-implemented method for segmenting and analyzing OCT images to identify and quantify specific features within the OCT images, in accordance with one or more embodiments.

[0047] FIG. 11 illustrates a non-limiting example of a dataflow for an implementation process of techniques for segmenting and analyzing OCT images to identify and quantify specific features within the OCT images, in accordance with one or more embodiments.

[0048] FIG. 12 illustrates a non-limiting example of dataflow of a process for segmenting and analyzing OCT images to identify and quantify specific features within the OCT images, in accordance with one or more embodiments.DETAILED DESCRIPTION

[0049] Reference will now be made to the embodiments illustrated in the drawings, and specific language will be used here to describe the same. It will nevertheless be understood that no limitation of the scope of the disclosure is thereby intended. Alterations and further modifications of the features illustrated here, and additional applications of the principles as illustrated here, which would occur to a person skilled in the relevant art and having possession of this disclosure, are to be considered within the scope of the disclosure.

[0050] Conventional techniques and technological solutions available today do not include systems, tools, or platforms that segment and analyze OCT images to allow for the identification and / or quantification of specific features within the OCT images. To address the need for techniques that are capable of predicting the onset of an adverse condition, techniques described herein involve the processing of portions of OCT images. More specifically, embodiments described herein include a system for segmenting OCT images to analyze structures represented by the OCT images, the system comprising: at least one processor configured to: obtain volumetric data comprising a plurality of cross-sectional images received from an OCT imaging device, the volumetric data representing three-dimensional image of an eyeball; extract a plurality of feature maps based on the volumetric data to determine a plurality of segmentations across the plurality of cross-sectional images; determine a first boundary of a structure and a second boundary of the structure based on at least one feature map of the plurality of feature maps; segment a plurality of substructures within a region established by the first boundary and the second boundary. In some076333-1065 / 07025 PATENTexamples, the one or more processors can be configured to generate an image of at least a portion of the structure of the eyeball, based on the volumetric data and the plurality of substructures, to indicate locations of the plurality of substructures within the structure.

[0051] By implementing the techniques described herein, systems (e.g., computing devices) can be configured to efficiently process and analyze the volumetric data obtained for diagnosis of a patient by segmenting and focusing on relevant regions and minimizing the need to process the entire dataset. These systems can also reduce memory consumption by only storing highly-detailed data associated with the portions of the retina segmented for analysis instead of the entire volumetric dataset generated during a scan of the retina, allowing for efficient use of memory resources. Additionally, they can reduce network communications by minimizing the amount of data transmitted between the OCT imaging device and the server, or between various systems associated with the server, decreasing external and / or internal network bandwidth usage and speeding up processing time. Furthermore, at least some of the systems described herein can be configured to analyze a subset of volumetric data based on the segmented portions of the retina and utilize deep learning algorithms (e.g., multilayered neural networks) to process the segmented portions of the data more accurately. This can allow for the analysis of higher resolution images that otherwise may not be processable due to computing device constraints, such as memory or processing constraints.

[0052] Moreover, to address the need for techniques that are capable of predicting the onset of an adverse condition, techniques described herein involve the processing of portions of OCT images to determine whether one or more diseases are indicated by the OCT images. More specifically, embodiments described herein include a system for segmenting and analyzing OCT images to identify and quantify specific features within the OCT images, the system comprising: at least one processor configured to: obtain volumetric data comprising a plurality of cross-sectional images received from an OCT imaging device, the volumetric data representing three-dimensional image of an eyeball; determine a plurality of points associated with a boundary of a structure of the eyeball; extract a plurality of patches from the volumetric data that correspond to the boundary of the structure; determine a first set of patches having a first structural characteristic and a second set of patches having a second structural characteristic based on the plurality of patches; and generate a graphical user interface based on the plurality of patches indicating a first076333-1065 / 07025 PATENTregion within the eyeball where the first structural characteristic is present and a second region within the eyeball where the second structural characteristic is present.

[0053] By implementing the techniques described herein, systems (e.g., computing devices) to reduce computational resource consumption by efficiently processing and analyzing the volumetric data, focusing on relevant regions and minimizing the need to process the entire dataset. They can also reduce memory consumption by storing only the extracted patches instead of the entire volumetric dataset, ensuring efficient use of memory resources. Additionally, they can reduce network communications by minimizing the amount of data transmitted between the OCT imaging device and the server, decreasing network bandwidth usage and speeding up processing time. Furthermore, they can generate more accurate predictions of diseases such as geographic atrophy (GA) by analyzing the structural characteristics of the extracted patches, leading to improved detection and diagnosis accuracy.

[0054] FIG. 1 is a block diagram of an environment 100 for segmenting OCT images to analyze structures represented by the OCT images, described in accordance with one or more embodiments herein. The environment 100 can include OCT device 110a, OCT device 110b, and OCT device 110c (referred to collectively as OCT devices 110 and individually as OCT device 110, where contextually appropriate), client device 130a, client device 130b, and client device 130c (referred to collectively as client devices 130 and individually as client device 130, where contextually appropriate), a server 140, and / or a database 150. The OCT devices 110, client devices 130, server 140, and database 150 can all interconnect (e.g., establish a connection to communicate with one another) via a network 120. The network 120 can be a wide area network (such as the Internet), a local area network (LAN), or any other kind of network.

[0055] The OCT devices 110 can include one or more devices (including one or more computing devices) having hardware and software components capable of performing the various processes described herein. For example, the OCT devices 110 can be any device, including a memory and a processor capable of communicating, via the network 120 with one or more other devices of FIG. 1. Tn some implementations, the OCT devices 110 include one or more of a light source, an interferometer, a sample arm, a reference arm, a detector, a signal processor, a scanner, a processor, memory, and a display. The OCT devices 110 can be configured to be in communication with the client devices 130, the server 140 and the database 150 via network 120.076333-1065 / 07025 PATENTIn some embodiments, the OCT devices 110 can be configured to provide (e.g., transmit) data to the client devices 130 and the server 140 as described herein. While the server 140 is described as managing communication between the OCT devices 110 and the database 150, it will be understood that the OCT devices 110 can communicate directly with the database 150 to upload data to the database 150 to be stored. In some embodiments, the OCT devices 110 can be associated with (e.g., controlled or managed by) clinicians as described herein.

[0056] The client devices 130 can include any computing device having hardware and software components capable of performing the various processes described herein. For example, the client devices 130 can be any device, including a memory and a processor capable of communicating, via the network 120, with one or more other devices of FIG. 1. Non-limiting examples of the client devices 130 include desktop computers, mobile devices (e.g., cellular phones and tablets), and / or the like. The client devices 130 can be configured to be in communication with OCT devices 110, server 140, and database 150 via network 120. In some embodiments, the client devices 130 can be associated with clinicians and / or patients as described herein.

[0057] The server 140 can include any computing device comprising hardware and software components capable of performing the various processes described herein. For example, the server 140 can be any device including a memory and processor capable of communicating, via the network 120, with one or more other devices of FIG. 1. Non-limiting examples of the server 140 includes data centers, server computers, workstation computers and / or the like. The server 140 is configured to be in communication with the OCT devices 110 and the client devices 130 via network 120. In some embodiments, the server 140 is associated with one or more clinicians and / or one or more healthcare providers (e.g., hospitals, public and / or private research organizations, and / or the like).

[0058] The database 150 can include any computing device comprising hardware and software components capable of performing the various processes described herein. For example, the database 150 can be any device, including a memory and processor capable of communicating, via the network 120 with one or more other devices of FIG. 1. Non-limiting examples of the database 150 include data centers, server computers, workstation computers and / or the like. The database 150 is configured to be in communication with the OCT devices 110, the client devices076333-1065 / 07025 PATENT130, and the server 140 via network 120. The database 150 can obtain (e.g., receive) data from the OCT devices 110, the client devices 130, and / or the server 140. In some embodiments, the database 150 is associated with (e.g., under the control of) the client devices 130, the server 140 and / or one or more clinicians as described herein.

[0059] The data described herein can include image data associated with one or more images of a portion of an eyeball of a patient. For example, the image data can include data generated by OCT devices 110, while imaging the eyeballs (e.g., the posterior portion of the eyeballs) of patients. Images can be generated by various OCT devices 110 at instant points in time or at points in time over a period of time. And the images can be correlated with one or more patients, one or more periods of time, and / or the like. The database 150 can contain data related to users or research projects (e.g., patient data, researcher data, clinician data, research data), which can be stored in database records (e.g., patient records, research records). Non-limiting examples of the types of data in the database records can include a patient identifier (ID), research ID, timestamps of data inputs, laterality (e.g., vitals, imaging, social determinants of health), data modality (e.g., CF, OCT, FAF), operating mode (e.g., single scan, volume scan), raw image data, and image analysis data, among others.

[0060] FIG. 2 is a flowchart, illustrating operations of a processor-implemented or computer-implemented method 200 for segmenting optical coherence tomography (OCT) images to analyze structures represented by the OCT images, in accordance with one or more embodiments. In some implementations, one or more of the functions described with respect to method 200 can be performed (e.g., completely, partially, and / or the like) by any computing device that is the same as, or similar to, server 140. For instance, the computing device may include a non-transitory computer-readable medium for storing machine-readable processor-executable instructions that, when executed by one or more processors, cause the one or more processors to perform various operations described herein. In some implementations, one or more of the functions described with respect to method 200 can be performed (e.g., completely, partially, and / or the like) by another device or group of devices separate from and / or including the server, such as by one or more OCT devices that are the same as, or similar to, OCT devices 110 of FIG. 1, one or more client devices that are the same as, or similar to, the client devices 130 of FIG. 1, and / or a database that is the same as, or similar to, the database 150 of FIG. 1. As an example,076333-1065 / 07025 PATENToperations 210-250 provide for implementing machine-learning models (e.g., ResNet, LSTM) for feature extraction and identification of EZ loss. In some cases, the machine-learning models may include a diffusion model to augment an output of the method 200 (as in operation 240).

[0061] At operation 210, the method 200 includes obtaining, by a server, data associated with a three-dimensional representation of an eyeball. For example, the server can receive volumetric data, including a plurality of cross-sectional images generated by an OCT imaging device. The volumetric data can include an image of at least a portion of an eyeball of a patient. For example, the server can receive volumetric data represented as a B-scan that represents a two-dimensional cross-section of a retina of a patient (e.g., an anterior or posterior of the retina).

[0062] In some implementations, the volumetric data can represent multiple layers of the retina of the patient (sometimes referred to as “structures” of the eyeball, or as bounding one or more structures of the eyeball). For example, volumetric data can represent the retina (e.g., one or more layers of the retina), the choroid, and / or the sclera of the patient. In some implementations, the one or more layers of the retina represented include the retinal nerve fiber layer (RNFL), the ganglion cell layer (GCL), the inner plexiform layer (IPL), the inner nuclear layer (INL), the outer plexiform layer (OPL), the outer nuclear layer (ONL), the inner segment / outer segment junction (IS / OS), the retinal pigment epithelium (RPE). In some implementations, the first image also represents the vitreous chamber of the eyeball of the patient.

[0063] In some implementations, the server can receive the volumetric data from one or more OCT devices. For example, the server can receive the data associated with the first image from the one or more OCT devices based on the one or more OCT devices imaging eyeballs of patients (e g., generating Cirrus OCT B-scans, standard OCT B-scans, EDI OCT B-scans, OCT volume scans, and / or the like of eyeballs of patients). In some implementations, the one or more images can each be associated with a point in time. For example, the one or more images generated by the OCT devices, while imaging the eyeballs of patients can be associated with a timestamp. The timestamp can indicate the date and time at which the OCT devices performed the imaging of the eyeballs of the patients.

[0064] When generating the images of the patients’ eyeballs, the OCT devices can generate images consecutively starting at the point in time at which the first image is generated. For076333-1065 / 07025 PATENTexample, to generate volumetric data associated with an OCT volume scan (e.g., a three-dimensional image of at least a portion of the eyeball of the patient) the OCT devices can generate multiple B-scans and align (e.g., register) the portions of the B-scans with one another. The position of the multiple portions of the B-scans can be registered with one another as a reference to generate the OCT volume scan (the three-dimensional representation) of the eyeball of the patient. When registering against a reference, the server can align the multiple B-scans of an OCT volume scan, relative to a B-scan, that is transverse to the multiple B-scans (referred to as an en face image) that are being aligned. In some implementations, the OCT volume scan can be generated at the same time the one or more individual B-scans are generated.

[0065] At operation 220, the method 200 includes extracting, by the server, a plurality of feature maps based on the volumetric data. For example, the server can extract a plurality of feature maps based on the volumetric data by processing at least a portion of the volumetric data. In this example, the server can process at least a portion of the volumetric data, such as one or more OCT B-scans, to determine a plurality of segmentations of layers or structures that are represented in a plurality of cross-sectional images included in the volumetric data.

[0066] In an example, the server can extract the plurality of feature maps by executing at least a portion of a model having a model architecture (e.g., that is the same as, or similar to, the model architecture 300 of FIG.3). For example, the server can extract the plurality of feature maps by executing one or more encoder blocks of a model that was trained to receive as an input, either an initial image (e.g., one or more OCT B-scans), or an output from a preceding encoder block. In an example, the server can extract a plurality of feature masks by executing a first encoder in accordance with an input image. The plurality of feature maps can then be generated as an output of the first encoder and can be selected from, based on one or more structures of the retina being analyzed, for further analysis as described herein. For purposes of clarity, the method 200 is described with respect to the generation of a plurality of feature maps and subsequent selection of a feature map from among the plurality of feature maps, however, it will be understood that a plurality of feature maps can be processed, at least, in part, separately to execute one of the more of the operations described herein. An example of separate processing of feature maps is described below with respect to FIG. 7.076333-1065 / 07025 PATENT

[0067] In some embodiments, the server can determine a plurality of segmentations based on one or more of the feature maps generated by executing at least a portion of the model. For example, when segmenting portion of the input image associated with the sclera of a patient, the server can identify feature map no. 2 from among 64 total feature maps generated by the model and execute one or more operations as described below to refine the feature map.

[0068] At operation 230, the method 200 includes determining, by the server, at least one boundary of a structure based on at least one feature map of the plurality of feature maps. For example, the server can determine a first boundary of a structure and a second boundary of the structure based on the feature map selected by the server. The server can determine the first boundary and the second boundary by identifying the initial positions of these boundaries within the feature map selected by the server. In an example, this can involve executing operations to implement morphological operations, such as grey thresholding, adaptive thresholding, etc., to remove noise from the feature map. In examples, the server can additionally, or alternatively, execute one or more operations to implement dilation, erosion, opening, and closing, which can refine the boundaries to better match the features targeted by the server. In some environments, the server can then update the feature map based on executing these operations. For example, the server can adjust the representation of the first boundary and / or the second boundary as included in the feature map by increasing the intensity of each pixel corresponding to the first boundary and / or the second boundary. Additionally, or alternatively, the server can adjust the representation up portions of the feature map that are different from the first boundary or the second boundary by decreasing the intensity of each pixel corresponding to these portions of the feature map.

[0069] In some examples, to determine the first boundary of the structure having a first type and the second boundary of the structure having a second type, the server can compare a relative position of the first boundary and the second boundary to each other within the feature map, and / or to one or more other structures identified in the feature map. For example, in the context of a retina of a patient, the server can identify a plurality of boundaries and compare the plurality of boundaries to known positions of boundaries for retinas of patients. In one example, the server can identify an ILM boundary and determine the location of one or more other boundaries based on the relative position of those other boundaries to the ILM boundary. It will076333-1065 / 07025 PATENTbe understood that the server can determine the location of one or more boundaries based on one or more other boundaries that are different from, or in part involve, the ILM boundary.

[0070] In some embodiments, the server can determine one or more aspects of a plurality of substructures of the eyeball, such as one or more substructures of the retina, to be used when augmenting the input image in accordance with method 200. For example, the server can receive input generated by a client device in response to a clinician providing input to the client device to indicate one or more structures or substructures of the retina of the patient to be analyzed. In this example, the server can determine one or more aspects of the plurality of structures or substructures of the retina of the patient to be used when augmenting the input image. These aspects can include one or more layers, one or more substructures within the layers, etc. that are to be augmented in an output generated by the server. The server can then determine the first boundary of the structure and / or second boundary of the structure, as described above, within the feature map based on the one or more aspects identified by the server.

[0071] At operation 240, the method 200 includes segmenting, by the server, a plurality of substructures within a region established by the at least one boundary. For example, the server can identify and isolate specific layers and structures within the retina or other parts of the eyeball based on the first boundary and the second boundary. In one example, the server can identify specific layer boundaries, layers, and / or structures within the retina of the eyeball of the patient based on the first boundary and / or the second boundary of the retina. In examples, the server can use the first boundary and / or the second boundary as reference points to segment the retina into distinct layers, such as the internal limiting membrane (ILM), retinal pigment epithelium (RPE), etc. The server can segment these structures or substructures in response to receiving input from the clinician at a client device that indicates the server is to identify and augment these structures or substructures. Additionally, or alternatively, the server can segment these structures in response to input from the clinician at the client device requesting that the server segment, and subsequently quantify the structures or substructures described herein.

[0072] In some embodiments, the server can generate an updated feature map based on segmenting plurality of structures and / or substructures. For example, the server can generate the updated feature map by executing one or more morphological operations to adjust a representation of one or more aspects of the structures or substructures represented by the feature map. In one076333-1065 / 07025 PATENTexample, the server can implement grey thresholding, adaptive thresholding, etc., to target these structures or substructures being segmented by the server. This can allow for the reduction of noise in the feature map and augment the structures and / or substructures that are being analyzed by the server. In some embodiments, the server can generate a second updated feature map that is based on the updated feature map. For example, the server can generate the second updated feature map by executing one or more second morphological operations to adjust the representation of the structures and or substructures represented in the feature map (e.g., the updated feature map). These second morphological operations can be at least in part different from the first set of morphological operations. The server can then segment the plurality of structures or substructures within at least a region of the retina, as represented by the input image and the feature maps described herein. Once segmented, the server can update the input image, in accordance with the segmented structures or substructures. In some embodiments, the server can determine the first boundary and or the second boundary of one or more structures or sub structures represented by the second feature map and can then segment the structures or substructures. In this way, the server can, for example, isolate information in accordance with one or more layers of the retina of the patient, corresponding to a first structure of the retina of the patient, and analyze the substructures within those one or more layers (e.g., a second structure that can be in part included in the first structure) when analyzing the input image.

[0073] In some embodiments, the server can segment one or more of the feature maps described herein to determine a first set of boundaries and a second set of boundaries. For example, the server can segment one or more of the featured maps described herein to determine a first boundary of a first structure, such as one or more layers of the retina. In this example, the server can segment the one or more feature maps to identify one or more second boundaries. For example, the one more second boundaries can relate to a different structure, such as a different layer of the retina. In another example, the one or more second boundaries can relate to different substructures within the structure identified by the server. This can include one or more vessels, etc., associated with the first structure.

[0074] At operation 250, the method 200 includes generating, by the server, an image of at least a portion of a structure based on the volumetric data and the plurality of substructures. For example, the server can generate a representation of the targeted retinal structures based on the076333-1065 / 07025 PATENTserver determining and / or generating the segmented boundaries and feature maps as described herein. In one example, the server can generate the image of at least a portion of the structure by updating the input image received by the server, in accordance with the feature maps as processed, to identify the one or more structures and or substructures described above. In examples, the server can generate graphical user interface data associated with (e.g., representing) the image of at least the portion of the structure that was updated using the feature maps as described, where the graphical user interface data is configured to cause a display device of a client device accessible by a clinician to display the image. This can allow the clinician to diagnose or confirm a diagnosis of one or more ocular conditions associated with the given patient.

[0075] In another example, the server can generate an image of at least a portion of a structure, which includes a representation of the structures that is based on the feature maps as described herein. For example, the server can generate the image of at least a portion of the structure as a separate image apart from the input image. In this example, the server can adjust one or more aspects of the portions of the structures and / or substructures included in the image to highlight one or more aspects of the structures and or substructures. This can allow clinician to quickly identify certain aspects of one or more structures or sub structures of the retina of the patient when diagnosing the patient as having or not having one or more diseases that are affecting the retina of the patient.

[0076] FIG. 3 illustrates a non-limiting example of an implementation of a model architecture 300 for generating feature maps, in accordance with one or more embodiments. In some implementations, one or more of the computing devices described can be the same as, or similar to, one or more of the computing devices of FIG. 1. For example, a server that is the same as, or similar to, the server 140 of FIG. 1 can be configured to train and / or execute the model architecture to process at least a portion of volumetric data obtained by the server as described with respect to FIG. 1.

[0077] In some embodiments, the model architecture 300 can be configured to receive an input image as input. For example, the model architecture 300 can be configured to receive an input image, such as a portion of volumetric data, including one or more OCT B-scans generated by an OCT device (e.g., an OCT device that is the same as, or similar to, the OCT device 110 of076333-1065 / 07025 PATENTFIG. 1). The input images can illustrate different layers and structures of a retina of a patient, including layer boundaries, vessel boundaries, etc.

[0078] In some embodiments, the model architecture 300 can be implemented as a machine learning model, including a convolutional neural network, a specialized convolutional neural network (e.g., a U-net), etc., that is configured to generate one or more feature maps based on the input image. For example, the model architecture 300 can include a machine learning model having an encoder-decoder structure with a bottleneck (e.g., a residual block 304) in between, forming a U-shaped design. The encoder-decoder structure can include a plurality of encoders 302a-302d that extend along a contracting path 302 of the model architecture 300 and are configured to extract hierarchical features by progressively downsampling the input image and / or feature maps through convolutional and max-pooling layers. In this example, each first encoder block 302a-302n can have a corresponding decoder 306a-306n for each stage of the model architecture 300. One example of a stage can include first encoder block 302a and corresponding decoder 306a. In some embodiments, each decoder 306a-306d can be configured to reconstruct the segmentation map by upsampling and combining features from corresponding encoder layers via skip connections. This model architecture 300 can allow for the processing of both low-level spatial details and high-level contextual information.

[0079] As described above, the inputs to model architecture 300 can include images, such as OCT B-scans, portions of OCT B-scans (e.g., OCT A-scans), etc. For example, the input to the model architecture 300 can include OCT B-scans that are represented as grayscale images of cross-sectional views of the retina of a patient. The server implementing the model architecture 300 can provide the images into a first encoder block 302a in the contracting path 302 of the model architecture 300. In this example, each encoder block 302a-302d can include (e.g., implement) convolutional layers that execute operations to extract one or more feature maps at various scales (e.g., corresponding to each stage). As an example, a feature map 310 can be generated by the first encoder block (the first encoder block 302a) of the contracting path 302 of the model architecture 300. While not explicitly illustrated, it will be understood that one or more feature maps can be generated at each encoder block 302a-302d of the model architecture 300 and can be used to generate segmented images as described herein. The output from the model architecture 300 can include a segmented image, where each pixel is assigned a label corresponding to specific retinal076333-1065 / 07025 PATENTlayers or regions of interest corresponding to the input image (e.g., the input OCT B-scan). This output can be represented as a multi-channel probability map, with each channel indicating the likelihood of a pixel belonging to a particular class corresponding to a particular layer or region of the retina.

[0080] As described herein, one or more outputs can be obtained by the server based on the execution of the model architecture 300. For example, one or more feature maps that are the same as, or similar to, the feature maps 310 that are generated by the first encoder block 302a along the contracting path 302 can be obtained by the server as an output from the model architecture 300. As a result, the server can be configured to extract one or more feature maps based on execution of one or more of the encoder blocks 302a-302d when processing the input image. This can, in turn, allow the server to extract the plurality of feature maps at one or more resolutions. For example, feature maps can be extracted from the first encoder block 302a at a first resolution associated with the first encoder block 302a. In other examples, feature maps can be extracted from any set or subset of the encoder blocks 302a-302d to be used when segmenting one or more substructures within the retina of the patient, as represented by the input image. And in some examples, a server can implement one or more encoder blocks 302a-302d, without executing the residual block 304 and / or the decoder blocks 306a-306d, to conserve processing and memory resources when analyzing the input image as described, for example, in FIGS. 4-7.

[0081] In some embodiments, the encoder blocks 302a-302d in the contracting path 302 can implement multiple convolutional layers followed by rectified linear unit (ReLU) activations and max-pooling operations for downsampling (not explicitly illustrated). The pool map output by the encoder blocks 302a-302d can be transferred via skip connections to corresponding decoders 306a-306d. These pool maps can represent spatially compressed but contextually rich features that allow for the reconstruction of fine details by the decoders 306a-306d during upsampling. The encoder blocks 302a-302d can also transfer the feature map generated by the encoder blocks 302a-302d to the next encoder block or, in the case of last encoder block 302d, to the residual block 304. In these examples, the feature map can represent progressively abstracted features that capture higher-level patterns represented by the input image.

[0082] In some embodiments, the residual block 304 in model architecture 300 can include two paths: a forward path and a skip connection. In the forward path, the feature map from the last076333-1065 / 07025 PATENTencoder block 302d in the contracting path 302 can be processed. This can include the feature map undergoing two sequential 3*3 convolutional layers, each followed by batch normalization and ReLU activation. In examples, the skip connection can bypass these operations and directly add the input to the output of the forward path. This can allow the residual block 304 to learn residual mappings, simplifying the optimization process and mitigating issues, such as vanishing gradients. If the input and output dimensions differ, a 1 x 1 convolution can be applied to the skip connection to match dimensions before addition. In some embodiments, the output of the residual block 304 can represent a combination of learned features from the forward path and preserved original information from the skip connection. This output can be provided to the first decoder 306d in the expansion path 306, which can be configured to receive the output from the residual block 304 as an input for upsampling operations. By retaining both high-level abstracted features and spatially detailed information, the server implementing the model architecture 300 can more accurately reconstruct the segmented images during decoding.

[0083] In some embodiments, the decoders 306a-306d in the expansion path 306 can execute concatenate operations to combine upsampled feature maps with skip-connected pool maps obtained by the decoders 306a-306n. In some examples, each decoder 306a-306d can then apply transposed convolutions to increase spatial resolution and reduce channel depth, followed by convolutional layers to merge contextual and localized information. The decoders 306d-306a can provide refined feature maps to subsequent decoders along the expansion path 306 until reaching the output resolution. In some embodiments, the decoder 306a (e.g., the final decoder) can output segmentation probabilities through a 1x1 convolution layer, producing masks that align with the input dimensions. The output of the decoder 306a can represent an image (e.g., a segmented image) that corresponds to the input image, where pixels representing various layers or regions are annotated using one or more channels, one or more visual indicators, including predetermined colors for various layers, etc.

[0084] During training, the server can provide OCT B-scans to the model architecture 300 as input to cause the components of the model architecture 300 to generate the output image as described herein. The server can then compare the output image to a ground truth image, including one or more segmentation masks established for the OCT B-scans to calculate a difference between the output image and the ground truth image. In some embodiments, the server can implement a076333-1065 / 07025 PATENTloss function, such as cross-entropy, to quantify the difference to establish a degree of error between the output of the model architecture 300 and the ground truth. In some embodiments, the server can implement backpropagation when updating (e.g., adjusting) the weights of the various components of the model architecture 300 by calculating gradients of the loss with respect to each weight. These gradients can be scaled by a learning rate and used to iteratively adjust weights of portions of the model architecture 300, minimizing segmentation errors over successive training iterations. This process of providing input images to the model architecture 300, comparing the input images to ground truth images, calculating corresponding losses, and updating the weights of the components of the model architecture 300 can be iteratively repeated until a threshold difference representing convergence is satisfied.

[0085] In some embodiments, the encoders 302a-302d (discussed generally with reference to first encoder block 302a in the contracting path 302 the model architecture 300) can be designed with various configurations such as convolutional blocks or residual blocks for feature extraction and downsampling. In an example, the first encoder block 302a can include two consecutive 3x3 convolutions followed by activation layers like ReLU, a max-pooling operation to reduce spatial resolution, and optional batch normalization. In some examples, one or more residual blocks can be inserted between the convolution and / or activation layers of the first encoder block 302a to preserve gradient flow and mitigate the vanishing gradient issue.

[0086] In some embodiments, the first encoder block 302a can receive an image at its input, such as an OCT B-scan, and output feature maps (e.g., feature maps, pooled feature maps, etc.) for subsequent encoders in the contracting path 302 or for a residual block 304 at the end of the contracting path. In an example, the first encoder block 302a can execute operations, as described, based on (e g., in accordance with) the input image and its output that can be passed on to the next encoder (e.g., encoder 302b of FIG. 3) for feature extraction. In some examples, the output of the first encoder block 302a can be processed using a residual block before transitioning to the expansion path. In other examples, intermediate outputs can also be provided to corresponding decoders using skip connections as described herein.

[0087] FIG. 4 illustrates a non-limiting example of a process 400 for the generation and post-processing of feature maps (e.g., that are the same as, or similar to, the feature maps generated as outputs by the encoders 302a-302d of FIG. 3), in accordance with one or more embodiments.076333-1065 / 07025 PATENTIn some implementations, one or more of the computing devices, operations executed, and / or models and / or model architectures implemented as described can be the same as, or similar to, those of FIG. 1 or FIG.3.

[0088] Referring to FIG. 4, initially, a server (e.g., that is the same as, or similar to, the server 140 of FIG. 1) can be configured to receive volumetric data in the form of OCT B-scans from a device, such as an OCT device. The server can then provide at least a portion of the volumetric data, such as one or more OCT B-scans, or portions thereof, as an input to at least a portion of a model, such as one or more encoder blocks. In this example, the model can have a model architecture that is the same as, or similar to, the model architecture 300 of FIG. 3. For example, the model can be associated with a U-Net architecture, or a similar encoder-decoder architecture, and can be configured to receive an input image and generate a feature map as an output at each of the encoders of the model. In this example, the model can provide the image as an input to the first encoder along the contracting path of the model.

[0089] At operation 402, the input image can be provided as an input to a series of convolutional layers associated with the encoder block, where each convolutional layer applies a set of learnable filters to extract spatial and contextual features from the input image. In some examples, the first encoder can be configured to execute operations to implement one or more activation functions, such as ReLU, to introduce non-linearity, and pooling layers, such as max pooling, to reduce the spatial dimensions, while retaining the most salient features represented by the input image. In examples, the output of the first encoder can include a feature map that represents indications of whether or not these features are present in the input image. While the present disclosure discusses the generation of feature maps using a first encoder along the contracting path of the model, it will be understood that feature maps can be used at any depth of the model where appropriate.

[0090] At operation 404, the server can be configured to execute operations when applying one or more transformations (e.g., grey thresholding, etc.) to the feature map to remove noise and allow for segmentation of visual features associated with the feature map. In some examples, grey thresholding can be implemented to isolate certain layers within the feature map by setting predetermined lower and upper threshold values according to which the one or more transformations are performed, allowing regions that exhibit a specific range of grey intensities to076333-1065 / 07025 PATENTbe retained, while suppressing regions that do not satisfy the specific range of grey intensities. In some examples, thresholding operators, such as global thresholding, adaptive thresholding, etc., can be implemented to determine the predetermined threshold values based on histogram analysis of the grey-scale intensities of the feature map. In some examples, additional post-processing operations, including morphological filtering, layer fusion, etc., can be applied to the feature map (e.g., after thresholding) to refine segmentation boundaries and remove noise prior to continued analysis of the feature map.

[0091] At operation 406, the server can be configured to determine one or more boundaries of the retina of the patient based on the execution of the transformations to the feature map. For example, the server can determine one or more boundaries, such as the ILM boundary of the retina, by determining a first boundary represented by the feature map and determining one or more second boundaries based on a known relationship (e.g., a relative position) between that first boundary and the other boundaries of the retina of the patient. In this way, the server can identify one or more boundaries of the retina that allowed the server to update the input image and segment the structures of the retina based on these boundaries.

[0092] At operation 408, the server can be configured to update the input image based on the server determining the one or more boundaries of the retina from the feature map. For example, the server can determine the ILM boundary as described above and can update the pixels in the input image corresponding to the pixels of the feature map that are identified as being associated with the ILM boundary. In some examples, the server can be configured to update the input image to identify a plurality of boundaries. For example, the server can receive an indication that one or more predetermined boundaries are to be used when segmenting one or more layers of the retina. In this example, the server can then determine the one or more predetermined boundaries based on their relative position to one or more other boundaries, such as the ILM boundary. In this way, the server can update the input image (e.g., by adding an additional channel to each pixel of the input image) to indicate one or more boundaries of one or more layers of the retina of the patient.

[0093] FIG. 5 illustrates a non-limiting example of one or more sets of operations (referred to collectively as sets of operations 500) executed to determine the boundaries of one or more structures of a retina, in accordance with one or more embodiments. In some implementations, one or more of the computing devices, operations executed, and / or models implemented, as described,076333-1065 / 07025 PATENTcan be the same as, or similar to, those of FIG. 1, FIG. 3, or FIG. 4. As will be understood, each operation set 502-512 of the sets of operations 500 can involve the execution of operations that are the same as, or similar to, those described herein, for example with respect to the operations discussed in FIG. 4.

[0094] In some embodiments, operation set 502 can include processing and input image extracted from volumetric data as described herein. For example, a server can be configured to process the input image by providing the image input image to a model to cause the model to execute one or more encoder blocks having convolutional layers. In this example, the model can be configured to cause the encoder to extract a plurality of feature maps (e.g., 16, 32, 64, etc.) and select a feature map from among the plurality of feature maps. The server can select the feature map based on the structure(s) to be segmented from the retina for further analysis and / or processing by the server. The server can then apply transformations like grey thresholding to the feature map to remove noise and segment visual features. In this way, the server can isolate layers represented by the feature map in the image of the retina by setting specific grey intensity ranges. In some examples, the server can implement additional post-processing (e.g., morphological filtering, etc.) to update (e.g., refine) the feature map. The server can continue and determine retinal boundaries from the feature map. For example, the server can identify the ILM boundary and other related boundaries of the retina, updating the input image to segment retinal structures. Finally, the server can update the input image based on the determined retinal boundaries by adjusting pixels corresponding to the ILM boundary and, in some instances, identifying multiple boundaries.

[0095] In some embodiments, operation set 504 can include processing and input image extracted from volumetric data as described herein. For example, the server can be configured to process the input image by providing the image input image to a model to cause the model to execute one or more encoder blocks having convolutional layers (similar to above). The server can then apply transformations like grey thresholding to the feature map to remove noise and segment visual features. In some examples, the server can continue and determine retinal boundaries from the feature map. For example, the server can identify the ELM boundary and other related boundaries of the retina. Finaly, the server can update the input image based on the determined retinal boundaries by adjusting pixels corresponding to the ELM boundary and, in some instances, can identify multiple boundaries to update the input image to reflect these layers.076333-1065 / 07025 PATENT

[0096] In some embodiments, operation set 506 can include processing an input image extracted from volumetric data as described herein. For example, the server can process the input image by providing it to a model to extract multiple feature maps (e.g., using one or more encoder blocks), and select one feature map from among the multiple feature maps. In some embodiments, the server can apply transformations, like adaptive thresholding, to the feature map to remove noise and segment visual features. For example, the server can implement adaptive thresholding and dynamically set threshold values based on local pixel intensities, allowing for more precise isolation of layers within the feature map. The server can then determine retinal boundaries, such as the RPE boundary and other related boundaries of the retina and update the input image to segment retinal structures. Finally, the server can update the input image based on the determined retinal boundaries by updating pixels (e.g., by adding a channel) corresponding to the RPE boundary (and other related boundaries) to indicate whether a given pixel is associated with the RPE boundary.

[0097] In some embodiments, operation set 508 can include processing an input image extracted from volumetric data as described herein. For example, the server can process the input image by providing it to a model to extract multiple feature maps and select one feature map from among them. In some embodiments, the server can apply transformations, like adaptive thresholding, to the feature map to remove noise and segment visual features. The server can then determine retinal boundaries from the feature map, such as the sclera boundary and other related boundaries, and update the input image to segment retinal structures. Finally, the server can update the input image based on the determined retinal boundaries by updating pixels (e.g., by adding a channel) corresponding to the sclera boundary and other related boundaries to indicate whether a given pixel is associated with the sclera boundary.

[0098] In some embodiments, operation set 510 can include processing an input image extracted from volumetric data as described herein. For example, the server can process the input image by providing it to a model to extract multiple feature maps and select one feature map from among the multiple feature maps. In some embodiments, the server can apply transformations, like adaptive thresholding, to the feature map to remove noise and segment visual features. Adaptive thresholding can be applied specifically to a region of interest between the RPE and sclera, dynamically setting threshold values based on local pixel intensities to isolate layers within this076333-1065 / 07025 PATENTregion. The server can then determine structures within retinal boundaries from the feature map. For example, the server can identify the choroid vessels and other related structures of the retina and update the input image to segment these structures by updating pixels (e.g., by adding a channel) corresponding to the choroid vessels and other related structures to indicate whether a given pixel is associated with a given choroid vessel. The server can then determine the choroid layer based on one or more boundaries that are in proximity to (e.g., closest to) the choroid vessels relative to one or more other boundaries of the retina.

[0099] In some embodiments, operation set 512 can include processing an input image extracted from volumetric data as described herein. For example, the server can process the input image by providing it to a model to extract multiple feature maps and select one feature map from among the multiple feature maps. In some embodiments, the server can apply transformations, like adaptive thresholding, to the feature map to remove noise and segment visual features. Adaptive thresholding can dynamically set threshold values based on local pixel intensities, allowing for more precise isolation of layers within the feature map. The server can then determine retinal boundaries from the feature map. For example, the server can identify the choroid vessels and other related structures of the retina and update the input image to segment these structures. Finally, the server can update the input image based on the determined retinal boundaries by updating pixels (e.g., by adding a channel) corresponding to the choroid vessels and other related structures to indicate whether a given pixel is associated with a choroid vessel.

[0100] FIG. 6 illustrates a non-limiting example of a process for identifying a boundary of a sclera of a retina, in accordance with one or more embodiments. In some implementations, one or more of the computing devices, operations executed, and / or models implemented as described can be the same as, or similar to, those of FIG. 1 or FIGS.3-5.

[0101] Referring to FIG. 6, initially, a server (e.g., that is the same as, or similar to, the server 140 of FIG. 1) can be configured to receive volumetric data in the form of OCT B-scans from a device, such as an OCT device. The server can then provide at least a portion of the volumetric data, such as one or more OCT B-scans or portions thereof) as an input to a model. In this example, the model can have a model architecture that is the same as, or similar to, the model architecture 300 of FIG. 3. For example, the model can be associated with a U-Net architecture, or a similar encoder-decoder architecture, and can be configured to receive an input image and076333-1065 / 07025 PATENTgenerate a feature map as an output at each of the encoder blocks of the model. For example, the model can provide the image as an input to the first encoder along the contracting path of the model.

[0102] At operation 602, the input image can be provided as an input to the model to cause the model (e.g., one or more encoder blocks of the model) to generate a plurality of feature maps. For example, the server can provide the input image as an input to the model to cause the model to generate a plurality of feature maps using one or more encoder blocks along contracting path of the model. In an example, the server can provide the image as an input to at least one encoder block, such as the first encoder block of the model, to cause the first encoder block to execute a series of convolutional layers associated with the first encoder block. In examples, the output of the first encoder block can include a plurality of feature maps (e.g., 64 feature maps) that represent indications of whether or not features are present in the input image. These features can correspond to, for example, one or more structures of the retina of the patient. In some examples, the server can provide the input to the model to cause the model to generate the plurality of feature maps, and the server can subsequently select a feature map from among the plurality of feature maps. In these examples, the server can select the feature map based on one or more aspects of the retina beings analyzed by the server. In the example shown in FIG.6, the server can select a feature map no. 2 from among 64 total feature maps when analyzing and segmenting a sclera of the retina of a patient.

[0103] At operation 604, the server can be configured to execute operations to one or more transformations, such as adaptive thresholding, to the feature map to remove noise from the feature map and allow for segmentation of visual features. In some examples, the adaptive thresholding can be configured to isolate certain layers within the feature map by setting predetermined lower and upper threshold values according to which the one or more transformations are performed, allowing regions that exhibit a specific range of grey intensities to be retained, while suppressing regions that do not satisfy the specific range of grey intensities. In some examples, thresholding operators, such as global thresholding, adaptive thresholding, etc., can be implemented to determine the predetermined threshold values based on histogram analysis of the grey-scale intensities of the feature map. Additional post-processing operations, including morphological076333-1065 / 07025 PATENTfiltering, layer fusion, etc., can be applied to the thresholded feature map to refine segmentation boundaries and remove noise prior to continued analysis of the feature map.

[0104] At operation 606, after applying adaptive thresholding to segment the image, the server can identify groups of connected pixels (connected components) that share similar properties within the feature map. For example, the server can identify the groups of connected pixels within the feature map that isolate and analyze specific structures within the image, such as the sclera of the retina. In examples, the server can identify the groups of connected pixels along a predetermined portion of the feature map. For example, the server can identify the groups of connected pixels that are located along an upper portion, a middle portion, a lower portion, etc., of the feature map.

[0105] At operation 608, the server can be configured to determine one or more boundaries of the retina of the patient based on the execution of the transformations to the feature map. For example, the server can determine one or more boundaries such as the sclera boundary of the retina by determining a first boundary and determine one or more second boundaries based on a known relationship (e.g., relative positions of the layers of the retina compared to one another) between that first boundary and the other boundaries of the retina of the patient. In this way, the server can identify one or more boundaries of the retina that allowed the server to update the input image and segment the structures of the retina based on these boundaries.

[0106] At operation 610, the server can execute one or more operations to smooth one or more lines of a feature (e.g., the boundary of the sclera) in an image. For example, the server can apply filtering techniques that reduce noise while preserving edge continuity. In some examples, the server can execute convolutional operations using Gaussian or median filters to attenuate high-frequency variations and produce smoother transitions along the lines of the feature in the image. In some examples, the server can execute iterative smoothing algorithms to selectively smooth regions within the feature map, while maintaining sharpness in areas with significant gradient changes.

[0107] At operation 612, the server can be configured to update the input image based on the server determining the one or more boundaries of the retina. For example, the server can determine the sclera boundary as described above and can update the pixels in the input image076333-1065 / 07025 PATENTcorresponding to the pixels of the feature map that are identified as being associated with the sclera boundary. In this way, the server can update the input image (e.g., by adding an additional channel to each pixel of the input image) to indicate one or more boundaries of the sclera of the retina of the patient.

[0108] FIG. 7 illustrates a non-limiting example of a process for identifying segmented choroid vasculature of a retina, in accordance with one or more embodiments. Tn some implementations, one or more of the computing devices, operations executed, and / or models implemented as described can be the same as, or similar to, those of FIG. 1 or FIGS. 3-6.

[0109] Referring to FIG. 7, initially, a server (e.g., that is the same as, or similar to, the server 140 of FIG. 1) can be configured to receive volumetric data in the form of OCT B-scans from a device such as an OCT device. The server can then provide at least a portion of the volumetric data, such as one or more OCT B-scans or portions thereof) as an input to a model. In this example, the model can have a model architecture that is the same as, or similar to, the model architecture 300 of FIG. 3. In some embodiments, the model can be associated with a U-Net architecture, or a similar encoder-decoder architecture, and can be configured to receive an input image and generate a feature map as an output at each of the encoder blocks of the model. For example, the model can provide the image as an input to the first encoder block along the contracting path of the model.

[0110] In some embodiments, the server can provide the input image as an input to at least a portion of the model to cause the model to generate a plurality of feature maps (illustrated as feature maps 2 and 26). For example, the server can provide the input image as an input to the model to cause the first encoder block of the model to generate a plurality of feature maps. The output of the first encoder block can include a plurality of feature maps that represent indications of whether or not features are present in the input image. These features can correspond to, for example, one or more structures of the retina of the patient. In some examples, the server can subsequently select a subset of feature maps (e.g., feature maps 2 and 26) from among the plurality of feature maps based on one or more aspects of the retina beings analyzed by the server. In the example shown in FIG. 7, the server can select a feature map nos. 2 and 26 from among 64 total feature maps when analyzing and segmenting a sclera of the retina of a patient.076333-1065 / 07025 PATENT

[0111] At operations 702 and 704, the server can be configured to execute operations to segment the RPE layer and the sclera layer from feature map 26 and feature map 2, respectively. For example, at operation 702 the server can select feature map 26 and identify one or more boundaries of the retina when segmenting the RPE layer from the other layers represented in the input image. In this example, the server can apply transformations, like adaptive thresholding, to the feature map to remove noise and segment visual features. The server can then determine retinal boundaries, such as the RPE boundary and other related boundaries of the retina and update the input image to segment retinal structures. Finally, the server can update the input image based on the determined retinal boundaries by updating pixels (e.g., by adding a channel) corresponding to the RPE boundary (and other related boundaries) to indicate whether a given pixel is associated with the RPE boundary. Similarly, at operation 704, the server can select feature map 2 and identify one or more boundaries of the retina when segmenting the sclera layer from the other layers represented in the input image. Again, the server can apply transformations, like adaptive thresholding, to the feature map to remove noise and segment visual features. The server can then determine retinal boundaries, such as the sclera boundary and other related boundaries of the retina and update the input image to segment retinal structures. Finally, the server can update the input image based on the determined retinal boundaries by updating pixels (e.g., by adding a channel) corresponding to the sclera boundary (and other related boundaries) to indicate whether a given pixel is associated with the sclera boundary.

[0112] At operation 706, the server can again select feature map 2 and execute operations to implement one or more transformations, such as adaptive thresholding, to the feature map to remove noise from the feature map and allow for segmentation of visual features. In some examples, the operations involved in executing adaptive thresholding can be configured to isolate certain layers within the feature map by setting predetermined lower and upper threshold values according to which the one or more transformations are performed, allowing regions that exhibit a specific range of grey intensities to be retained, while suppressing regions that do not satisfy the specific range of grey intensities. In some examples, thresholding operators, such as global thresholding, adaptive thresholding, etc., can be implemented to determine the predetermined threshold values based on histogram analysis of the grey-scale intensities of the feature map. In some examples, additional post-processing operations, including morphological filtering, layer076333-1065 / 07025 PATENTfusion, etc., can be applied to the thresholded feature map to refine segmentation boundaries and remove noise prior to continued analysis of the feature map.

[0113] At operations 708a-708c, the server can use the output generated by the server at operations 702-706 to generate an updated feature map. For example, the server can combine a first feature map (output at operation 702) where the RPE layer of the retina is segmented, a second feature map (output at operation 704) where the sclera layer is segmented, and a third feature map (output at operation 706) where adaptive thresholding was applied, through a series of operations. For example, the server can integrate these feature maps by aligning and merging at least portions of the feature maps associated with the segmented layers to generate a combined output.

[0114] At operation 710, the server can determine retinal boundaries from the feature map generated as a result of the execution of operations 708a-708c. For example, the server can identify the choroid vessels and other related structures of the retina based on the feature map and update the input image to segment these structures. This can include outlining the choroid vessels, filling in the choroid vessels, etc. At operation 712, the server can update the input image based on the determined retinal boundaries by updating pixels (e.g., by adding a channel) corresponding to the choroid vessels and other related structures to indicate whether a given pixel is associated with the choroid vessels.

[0115] At operation 714, the server can select a different feature map (e.g., feature map 20) generated as a result of the server executing the model to obtain the plurality of feature maps. The server can then apply transformations, like adaptive thresholding, to the feature map (e.g., feature map 20) to remove noise and segment visual features. In this example, the server can apply adaptive thresholding to dynamically set threshold values based on local pixel intensities, allowing for more precise isolation of layers within the feature map. The server can then determine retinal boundaries from the feature map. For example, the server can identify the choroid vessels and other related structures of the retina, updating the input image to segment these structures.

[0116] At operation 716, the server can update the feature map based on the determined retinal boundaries. For example, the server can update the input image by updating pixels (e.g., by adding a channel) corresponding to the choroid vessels and other related structures (indicated by the feature map 20 after the application of adaptive thresholding) to indicate whether a given pixel076333-1065 / 07025 PATENTis associated with the choroid vessels. Similar to operation 710, the server can outline, fdl, etc., the choroid vessels as identified in the feature map.

[0117] At operation 718, the server can generate an image of structures of the retina that were segmented by combining a feature map segmenting choroid vessels (output of operation 716) with an updated image that includes channels indicating the boundaries of the choroid layer (output of operation 712). For example, first, the server can align the feature map with the updated image to confirm the segmented choroid vessels correspond to the choroid layer boundaries. The server can then combine these data sources to highlight the choroid vasculature within the retina of the patient (as illustrated in the input image originally provided as an input by the server to the model) with greater precision than may be achieved through the execution of operations 702-712, and 714 and 716, separately. By merging the segmented choroid vessels with the boundary channels, the server can create a detailed and accurate representation of the choroid vasculature, allowing for improved analysis and visualization of the retinal structures within the choroid layer.

[0118] FIG. 8 illustrates a non-limiting example of a process for identifying segmented portions where fluid is building in a retina, in accordance with one or more embodiments. In some implementations, one or more of the computing devices, operations executed, and / or models implemented as described can be the same as, or similar to, those of FIG. 1 or FIGS. 3-7.

[0119] Referring to FIG. 8, initially, a server (e.g., that is the same as, or similar to, the server 140 of FIG. 1) can be configured to receive volumetric data in the form of OCT B-scans from a device such as an OCT device. The server can then provide at least a portion of the volumetric data, such as one or more OCT B-scans or portions thereof) as an input to a model. In this example, the model can have a model architecture that is the same as, or similar to, the model architecture 300 of FIG. 3. In some embodiments, the model can be associated with a U-Net architecture, or a similar encoder-decoder architecture, and can be configured to receive an input image and generate a feature map as an output at each of the encoder blocks of the model. For example, the model can provide the image as an input to the first encoder block along the contracting path of the model.

[0120] At operation 802, the input image can be provided as an input to the model to cause the model (e.g., one or more encoder blocks of the model) to generate a plurality of feature maps.076333-1065 / 07025 PATENTFor example, the server can provide the input image as an input to the model to cause the model to generate a plurality of feature maps using one or more encoder blocks along contracting path of the model. In an example, the server can provide the image as an input to at least one encoder block, such as the first encoder block of the model, to cause the first encoder block to execute a series of convolutional layers associated with the first encoder block. In examples, the output of the first encoder block can include a plurality of feature maps (e.g., 64 feature maps) that represent indications of whether or not features are present in the input image. These features can correspond to, for example, one or more structures of the retina of the patient. In some examples, the server can provide the input to the model to cause the model to generate the plurality of feature maps, and the server can subsequently select a feature map from among the plurality of feature maps. In these examples, the server can select the feature map based on one or more aspects of the retina beings analyzed by the server. In the example shown in FIG.8, the server can select a feature map no. 2 from among 64 total feature maps when analyzing and segmenting a sclera of the retina of a patient.

[0121] At operation 804, the server can be configured to execute operations to perform one or more transformations, such as adaptive thresholding, to the feature map to remove noise from the feature map. For example, the server can be configured to execute operations to perform one or more transformations and remove noise to allow for more precise segmentation of visual features. In some examples, the adaptive thresholding can be configured to isolate certain layers within the feature map by setting predetermined lower and upper threshold values according to which the one or more transformations are performed, allowing regions that exhibit a specific range of grey intensities to be retained, while suppressing regions that do not satisfy the specific range of grey intensities. In some examples, thresholding operators, such as global thresholding, adaptive thresholding, etc., can be implemented to determine the predetermined threshold values based on histogram analysis of the grey-scale intensities of the feature map. Additional postprocessing operations, including morphological filtering, layer fusion, etc., can be applied to the thresholded feature map to refine segmentation boundaries and remove noise prior to continued analysis of the feature map.

[0122] At operation 806, after applying adaptive thresholding to segment the image (e.g., the feature map), the server can identify groups of connected pixels (connected components) that076333-1065 / 07025 PATENTshare similar properties within the feature map. For example, the server can identify the groups of connected pixels within the feature map that isolate and analyze specific structures within the image, such as the portions of the retina where intraretinal fluid (IRF) and / or subretinal fluid (SRF) is located. In examples, the server can identify the groups of connected pixels along a predetermined portion of the feature map. For example, the server can identify the groups of connected pixels that are located along an upper portion, a middle portion, a lower portion, etc., of the feature map.

[0123] At operation 808, the server can be configured to determine one or more boundaries of the retina of the patient based on the execution of the transformations to the feature map. For example, the server can determine one or more boundaries such as the IRF and / or SRF boundary of the retina by determining a first boundary and determining one or more second boundaries based on a known relationship (e.g., relative positions of the layers of the retina compared to one another) between that first boundary and the other boundaries of the retina of the patient. In the illustrated example, the server can determine that a region between a first boundary and a second boundary represents a portion of the retina where IRF and / or SRF are located by virtue of the position of the region between the first boundary, the second boundary, and the known relationships between the relative position of the layers of the retina and the server can identify the portions of the retina as corresponding to the IRF and / or SRF. In this way, the server can identify one or more boundaries of the retina that allowed the server to update the input image and segment the structures of the retina based on these boundaries as described herein.

[0124] FIG. 9 depicts portions of the retina where pigment epithelial detachment (PED) is identified, in accordance with one or more embodiments, and similar to that of FIG. 8. The portions of the retina where pigment epithelial detachment (PED) is identified can also be segmented using the above-described techniques, such as the operations of the process 800 depicted in FIG. 8.

[0125] FIG. 10 is a flowchart illustrating operations of a processor-implemented or computer-implemented method 1000 for segmenting and analyzing OCT images to identify and quantify specific features within the OCT images, in accordance with one or more embodiments. In some implementations, one or more of the functions described with respect to method 1000 can be performed (e.g., completely, partially, and / or the like) by a server that is the same as, or similar076333-1065 / 07025 PATENTto, server 140. For instance, the computing device may include a non-transitory computer-readable medium for storing machine-readable processor-executable instructions that, when executed by one or more processors, cause the one or more processors to perform the various operations described herein. In some implementations, one or more of the functions described with respect to method 1000 can be performed (e.g., completely, partially, and / or the like) by another device or group of devices separate from and / or including the server, such as by one or more OCT devices that are the same as, or similar to, OCT devices 110 of FIG. 1, one or more client devices that are the same as, or similar to, the client devices 130 of FIG. 1, and / or a database that is the same as, or similar to, the database 150 of FIG. 1.

[0126] At operation 1010, the method 1000 includes obtaining, by a server, volumetric data associated with a three-dimensional representation of an eyeball. For example, the server can receive volumetric data including a plurality of cross-sectional images generated by an OCT imaging device. The volumetric data can include an image of at least a portion of an eyeball of a patient. For example, the server can receive volumetric data represented as a B-scan that represents a two-dimensional cross-section of a retina of a patient (e.g., an anterior or posterior of the retina).

[0127] In some implementations, the volumetric data can represent multiple layers of the retina of the patient (sometimes referred to as “structures” of the eyeball, or as bounding one or more structures of the eyeball). For example, volumetric data can represent the retina (e.g., one or more layers of the retina), the choroid, and / or the sclera of the patient. In some implementations, the one or more layers of the retina represented include the retinal nerve fiber layer (RNFL), the ganglion cell layer (GCL), the inner plexiform layer (IPL), the inner nuclear layer (INL), the outer plexiform layer (OPL), the outer nuclear layer (ONL), the inner segment / outer segment junction (IS / OS), the retinal pigment epithelium (RPE). In some implementations, the first image also represents the vitreous chamber of the eyeball of the patient.

[0128] In some implementations, the server can receive the volumetric data from one or more OCT devices. For example, the server can receive the data associated with the first image from the one or more OCT devices based on the one or more OCT devices imaging eyeballs of patients (e.g., generating Cirrus OCT B-scans, standard OCT B-scans, EDI OCT B-scans, OCT volume scans, and / or the like of eyeballs of patients). In some implementations, the one or more images can each be associated with a point in time. For example, the one or more images generated076333-1065 / 07025 PATENTby the OCT devices while imaging the eyeballs of patients can be associated with a timestamp. The timestamp can indicate the date and time at which the OCT devices performed the imaging of the eyeballs of the patients.

[0129] When generating the images of the patients’ eyeballs, the OCT devices can generate images consecutively starting at the point in time at which the first image is generated. For example, to generate volumetric data associated with an OCT volume scan (e.g., a three-dimensional image of at least a portion of the eyeball of the patient), the OCT devices can generate multiple B-scans and align (e.g., register) the portions of the B-scans with one another. The position of the multiple portions of the B-scans can be registered with one another as a reference to generate the OCT volume scan (the three-dimensional representation) of the eyeball of the patient. When registering against a reference, the server can align the multiple B-scans of an OCT volume scan relative to a B-scan that is transverse to the multiple B-scans (referred to as an en face image) that are being aligned. In some implementations, the OCT volume scan can be generated at the same time the one or more individual B-scans are generated.

[0130] At operation 1020, the method 1000 includes determining a plurality of points associated with a boundary of a structure (e.g., of an eyeball of a patient). For example, the server can determine a plurality of points associated with the boundary of one or more layers of the retina of the patient. In this example, the server can determine the plurality of points across the OCT B-scan to generate one or more segmented images based on (e.g., using) a machine-learning model. In some embodiments, the machine-learning model can include a generative adversarial network (GAN), such as an image-to-image (Pix2Pix) GAN, a U-shaped convolutional neural network (U-net), and / or similar machine-learning models. While certain aspects of the present disclosure are discussed with respect to GANs, it will be understood that any suitable segmentation model can be implemented to perform one or more of the operations described with reference to the GANs as described herein.

[0131] In some embodiments, to train the GAN, the server can provide images (e.g., nonsegmented images) associated with the OCT B-scan (such as one or more A-scans) to the generator network of the GAN, and the server can provide corresponding target images (e.g., segmented images) to the discriminator network. The output of the generator network can then be provided as an input to the discriminator network to allow the discriminator network to generate an output076333-1065 / 07025 PATENTindicating whether the output of the generator network is or is not the target image. Based on whether the discriminator network correctly identifies the output of the generator network as not being the target image, the server can then cause the model to be updated (e.g., by updating one or more weights associated with the generator network or the discriminator network) to improve the probability that subsequent outputs based on the same input will more closely resemble the target image, and iteratively repeat this process until the GAN converges (e.g., until the discriminator network cannot detect generated images from target images so as to satisfy a threshold level of certainty). The server repeats this process for multiple types of images (e.g., EDI OCT B-scans, Cirrus OCT B-scans, and / or the like) and corresponding segmented images (e.g., segmented EDI OCT B-scans, segmented Cirrus OCT B-Scans, and / or the like). In this way, the server can train the model to generate segmented images for multiple image types.

[0132] In some embodiments, the server can use a GAN (Such as the GAN discussed above) to generate the segmented images. For example, the server can provide portions of the volumetric data, such as portions of one or more OCT B-scans, as an input to the generator network of the GAN. In this example, the generator network of the GAN can generate an output representing a segmentation mask, where pixels are identified in accordance with a different channel other than those of the original image to indicate one or more pixels that are associated with edges within the image. These images can include edges that are associated with boundary lines as described herein.

[0133] In some embodiments, the server can generate segmented images, where the segmented images are associated with (e.g., represent) at least one boundary line. For example, in the case of an OCT image, the server can generate the segmented image such that the boundaries between portions of the structures of the eyeballs are denoted. In one example, the server can generate the segmented image such that certain layers of the retina (described herein) are separated from one another. The segmented images can then be reviewed (e.g., by an individual) to confirm that the segmentations (e.g., indicating the separation of differentiated layers of the retina) are correct. For example, the server can communicate with a client device (e.g., a client device that is the same as, or similar to, client devices 130) and transmit data associated with the one or more segmented images to the client devices. The client devices can then display (e.g., via a display device of the client devices, not explicitly shown in FIG. 1) the segmented images and a clinician076333-1065 / 07025 PATENTcan provide input confirming that the segmentation is appropriate. In the case where the segmentation is not appropriate, the clinician can provide an indication via the corresponding client device that the segmentation is not appropriate and / or an indication regarding how to update the segmented image to correct the segmentation. In this case, the clinician can provide input via the client device that causes one or more points along the segmentation lines to move in accordance with the boundary between the different structures of the eyeball of the patient. The server can then update the segmented image and store the segmented image in the database.

[0134] In some embodiments, the server can initially generate the target images (the segmented images) based on one or more segmentation algorithms. For example, the server can generate the target images based on one or more threshold-based techniques. In these examples, the server can implement thresholding that is based on pixel intensity values by comparing adjacent pixels and determining whether the adjacent pixels satisfy a threshold difference in intensity values. The server can also implement adaptive thresholding whereby the threshold difference is determined (e.g., adjusted) based on one or more values within a predetermined distance from a given pixel. In some implementations, the server can generate the segmented images based on one or more graph-based techniques such as by fitting active contours to the images. And in some examples, the server can smooth the one or more segmentation lines by executing operations to perform one or more transformations.

[0135] At operation 1030, the method 1000 includes extracting, by the server, a plurality of patches from the volumetric data that corresponds to the boundary of the structure. For example, the server can extract the plurality of patches from the volumetric data based on one or more patch sizes to be used when extracting the patches and / or one or more boundaries of the structures represented in the volumetric data. The patch sizes can correspond to a predetermined patch size and / or a dynamically selected patch size. In some examples, the patch size can be configured to encompass at least a portion of one or more layers represented by the images included in the volumetric data. In some of these examples, the patch size can correspond to a particular disease or set of diseases that are being targeted for analysis by the server.

[0136] In some embodiments, the server can be configured to extract the plurality of patches from volumetric data by determining a patch size and extracting, in accordance with the patch size, the plurality of patches that correspond to one or more boundaries of one or more076333-1065 / 07025 PATENTstructures of the retina. In an example, the server can determine the patch size based on the resolution of the volumetric data and the relative size of each of the layers when compared to the resolution of the images included in the volumetric data. Additionally, or alternatively, the server can determine the patch size based on the characteristics of the boundaries of the one or more structures as represented by one or more layers adjacent to a target layer and / or the location of one or more adjacent layers to a target layer within the eyeball, such as the location of a choroid within the eyeball relative to the target layer.

[0137] In some embodiments, the server can use techniques such as sliding windows or overlapping patch sampling to allow for coverage of the boundary regions associated with the patches. For example, the server can align the plurality of patches with the boundaries of the structures being targeted for examination. In these examples, the server can determine the boundaries based on the server segmenting one or more portions of the volumetric data (as described above) and / or based on input receives at the server (e g., from a client device and generated in response to input from an individual controlling the client device) indicating one or more regions to be analyzed using the techniques described herein.

[0138] At operation 1040, the method 1000 includes determining, by the server, a first set of patches and a second set of patches from the plurality of patches extracted from the volumetric data. In some embodiments, the first set of patches can have a first structural characteristic and the second set of patches can have a second structural characteristic that is at least in part different from the first structural characteristic. For example the first set of patches can have a first structural characteristics indicated by the structure of the retina that is indicative of a disease, such as ellipsoidal zone (EZ) loss (as described below), etc., and the second set of patches can have a second structural characteristic indicated by the structure of the retina that is at least in part not indicative of the disease.

[0139] In some embodiments, the server can be configured to determine the first set of patches and the second set of patches by executing an artificial neural network (ANN) based on the plurality of patches to generate corresponding annotations usable to segment the plurality of patches into the first set of patches indicative of the presence of a disease and the second set of patches that are indicative of the absence of the disease. For example, the server can execute the ANN by providing at least one patch from the plurality of patches as input to the ANN, causing076333-1065 / 07025 PATENTthe ANN to perform operations that extract features from the patch at one or more resolutions (e.g., a first resolution, a second resolution, etc.). In this example, the extracted features can be used to generate at least one feature map that allows the server to identify structural characteristics of the retina of the patient. In some examples, the server can execute the ANN by providing each patch of the plurality of patches sequentially to the ANN, causing the ANN to extract the features that can be used to generate the at least one feature map as described herein for each of the patches.

[0140] In some examples, the server can use the feature map to determine the first set of patches with the first structural characteristic and the second set of patches with the second structural characteristic. For example, these structural characteristics can correspond to specific regions or features of interest within the volumetric data, such as portions of the retina of the patient, that can be used to identify the presence or non-presence of a disease, such as EZ loss (e.g., degradation or loss of the ellipsoidal zone in the retina). In this case, the server can segment the plurality of patches into the first set and second set based on the annotations generated by the ANN. The server can then separate the plurality of patches into the first set of patches and the second set of patches in accordance with the first structural characteristic or the second structural characteristic being associated with each respective patch.

[0141] In examples involving multi-resolution analysis (e.g., the extraction of features at multiple resolutions), the server can implement hierarchical feature extraction by processing patches at different resolutions. This approach can improve accuracy in detecting complex structures or subtle variations within volumetric data. In some examples described herein, to execute this multi-resolution analysis, the server can implement an ANN having at least a portion associated with an architecture such as a ResNet-based architecture.

[0142] In some examples, the server can implement an ANN having a first architecture and a second architecture. In some examples, the first architecture can correspond to a ResNet-based architecture as indicated above. In examples, the second architecture can correspond to a recurrent neural network (RNN) architecture such as, for example, a Long Short-Term Memory (LSTM) architecture, etc., that is configured to process data in accordance with a sequence. In these examples, the server can implement the ANN such that the plurality of patches are sequentially provided to the ResNet-based architecture portion of the ANN to generate one or more feature maps, and the feature maps are subsequently provided in sequence to the RNN-based076333-1065 / 07025 PATENTarchitecture portion of the ANN to generate a set of annotations indicating whether the first structural characteristic or the second structural characteristic is represented by each patch of the plurality of patches. In this way, the server can implement an ANN to process patches sequentially and allow for consistency when generating the annotations within a continuous three-dimensional space associated with the retina of the patient.

[0143] In some embodiments, ResNet-based architecture can include an ANN that implements residual blocks, which include connections that bypass one or more layers. These connections that bypass the one or more layers can allow the network to learn residual functions instead of directly mapping inputs to outputs. In some embodiments, the ResNet architecture can be organized in a sequential manner that processes input data through a series of layers. For example, initially, the portion of the ANN associated with the ResNet architecture can be configured to receive image data as an input, which can be followed by a convolutional preprocessing stage that includes a convolution layer with a stride that reduces spatial dimensions, followed by batch normalization and a rectified linear unit (ReLU) activation. In this example, the ResNet can include a max pooling layer that further compresses the spatial resolution. Multiple groups of residual blocks can then be arranged in order; each group beginning with a convolutional block that performs down sampling and expands the filter count and is followed by one or more identity blocks that preserve the feature map dimensions through shortcut connections. Finally, the portion of the ANN associated with the ResNet architecture can implement a global average pooling layer that summarizes the learned features into a compact representation and a fully connected layer that outputs predictions (feature maps) via softmax activation.

[0144] To train the portion of the ANN associated with the ResNet architecture, the server can be configured to implement a loss function (e.g., cross-entropy for classification) using backpropagation and gradient descent. During training, the input data can be provided to the portion of the ANN associated with the ResNet architecture to generate predictions, which are compared against ground-truth labels to calculate the loss. The server can then determine gradients of the loss with respect to model parameters computed via backpropagation, leveraging the residual connections to ensure stable gradient flow even in very deep networks. These gradients can be used to update weights of the ResNet architecture, allowing the server to configure the076333-1065 / 07025 PATENTportion of the ANN associated with the ResNet architecture to operate with a predetermined degree of precision (until convergence).

[0145] In some embodiments, the portion of the ANN associated with the LSTM-based architecture can implement LSTM cells, which incorporate input, forget, and output gates to manage the flow of information over a sequence. This can allow the portion of the ANN associated with the LSTM-based architecture to learn spatial dependencies. The portion of the ANN associated with the LSTM-based architecture can be arranged sequentially to process data corresponding to multiple images ordered in accordance with a two-dimensional or three-dimensional space. For example, initially, the portion of the ANN associated with the LSTM architecture can be configured to receive a sequence of feature maps as an input, which can be followed by an embedding or preprocessing stage that transforms aspects of each of the feature maps into vector representations. In this example, the vector representations can then be processed through one or more layers of LSTM cells that maintain internal cell states and leverage gating mechanisms to regulate information retention over time, eventually leading to a fully connected layer that outputs predictions via softmax activation. The output of the LSTM can include indications (e.g., annotations, etc.) that one or more patches extracted from the volumetric data are (or are not) indicative of the presence of a disease such as EZ loss.

[0146] To train the portion of the ANN associated with the LSTM architecture, the server can be configured to implement a loss function (e.g., cross-entropy for classification) using backpropagation through time and gradient descent. During training, sequential input data can be provided to the portion of the ANN associated with the LSTM architecture to generate predictions, which are compared against ground-truth labels to calculate the loss. The server can then determine gradients of the loss with respect to model parameters computed using backpropagation through time, leveraging the gating mechanisms within the LSTM cells to maintain stable gradient flow across time steps. These gradients can then be used to update the weights of the LSTM architecture.

[0147] In some embodiments, the server can generate an en face image of at least a portion of the structure of the eyeball based on volumetric data. This en face image can then be used to determine one or more portions of the volumetric data that are associated with GA as described below. For example, the server can process the volumetric data to generate the en face image. In this example, the server can apply techniques such as maximum intensity projection or averaging076333-1065 / 07025 PATENTalong specific axes to create a 2D representation of the 3D structure of the eyeball. The server can implement preprocessing operations that involve alignment, segmentation, and / or flattening relative to a reference layer (e.g., the retinal pigment epithelium) to allow for improved image coherence.

[0148] In some examples, the server can execute one or more operations to determine a segmentation mask from the en face image. For example, the server can execute operations to implement one or more image processing techniques such as thresholding or edge detection to identify regions of interest within the en face image. In an example, the server can execute a different ANN (such as a diffusion model as described herein) that is configured to receive the en face image and generate a segmentation mask as output. In this example, the ANN can iteratively introduce noise to the en face image and then remove at least a portion of the noise to refine and produce the segmentation mask.

[0149] In some embodiments, the input to the ANN (e.g., the diffusion model) can include an en face image, which is a two-dimensional representation of the retinal layers captured by the OCT device that is substantially transverse to the OCT A-scans included in the OCT B-scan. The output from the ANN can include a segmentation mask that indicates areas associated within the en face image indicative of Geographic Atrophy (GA) within the retina of the patient. This segmentation mask can include a binary or multi-class image where each pixel is labeled to indicate whether it belongs to the GA region. The segmentation mask can then be used to generate a user interface for a clinician to review when assessing an extent and progression of GA in the patient.

[0150] In some embodiments, the server can implement the ANN (e.g., the diffusion model) to process the en face image generated based on the volumetric data. For example, the server can implement a diffusion model that is configured to generate segmentation masks by reversing a diffusion process, which gradually adds noise to the data. In an example, the server can implement a diffusion model that iteratively adds and removes noise to the en face image to generate a segmentation mask that indicate areas associated within the retina (represented by the volumetric data) that are associated with GA.076333-1065 / 07025 PATENT

[0151] In some embodiments, the server can train the ANN (e.g., the diffusion model) using a dataset of en face images paired with corresponding ground truth segmentation masks. During training, the ANN can be configured to map the input images to the correct segmentation masks by minimizing a loss function, such as cross-entropy loss, which measures the difference between the predicted and actual masks. The weights of the ANN can be updated iteratively using backpropagation, where the gradients of the loss function with respect to the weights are computed and used to adjust the weights in the direction that reduces the loss. This process can continue until the ANN achieves satisfactory performance on a validation set (e.g., achieves convergence).

[0152] In some embodiments, the server can compare the segmentation mask to a first region established by the server as being associated with EZ loss (as established by the ANN having the ResNet architecture and the LSTM architecture) within the eyeball to determine a third region. For example, the first region can correspond to areas within the retina of the patient where specific structural characteristics such as EZ loss are detected (as described above). In this example, the server can update the en face image by overlaying patches that indicate the first region, thereby generating a first annotated image. Additionally, or alternatively, the server can overlay the segmentation mask onto this first annotated image to create a second annotated image. In this second annotated image, the region between the first region and the second region can be referred to as a third region, which shows a gap between the portion of the retina associated with EZ loss and a portion of the retina associated with GA.

[0153] In some embodiments, the server can determine the third region based on the second annotated image. For example, the server can determine areas where the segmentation mask associated with GA within the regina overlaps with or diverges from the first region associated with EZ loss. In some embodiments, these operations executed by the server can allow for detailed analysis and visualization of structural features within the eyeball for diagnostic or research applications.

[0154] At operation 1050, the method 1000 includes generating, by the server, a graphical user interface indicating a first region and a second region. For example, the server can generate a graphical user interface based on the first region and the second region to indicate one or more aspects of the structures of the retina of the patients, such as whether EZ loss is or is not present at certain points or portions of the retina of the patient. In some examples, the graphical user interface076333-1065 / 07025 PATENTcan be based on the plurality of patches indicating a first region within the eyeball where the first structural characteristic is present (e.g., where EZ loss is detected) and a second region within the eyeball where the second structural characteristic is present (e.g., where EZ loss is not present). For example, a set of pixels can be reserved along any portion of the graphical user interface (e.g., as a bar located along an upper portion, a middle portion, a lower portion, or combinations thereof) where the set of pixels includes subsets that are colored to indicate regions where the first structural characteristic or the second structural characteristic are identified. It will be understood that any suitable indications can be implemented to identify the different regions as described herein.

[0155] In some examples, the graphical user interface can be based on the plurality of patches as well as the segmentation mask described herein. For example, the graphical user interface can indicate the third region within the reserved set of pixels based on determining the third region. This third region can be displayed on the reserved set of pixels as described above as an additional color that separates the regions where the first characteristic is identified and or the second characteristic is identified. As a result, the third region can indicate the area where specific attributes associated with the disease are identified, such as a region within the retina where GA is shown with respect to detected portions of the retina where EZ loss is detected.

[0156] Additionally, in some examples, the server can determine an expected direction for disease progression based on one or more attributes of the third region when compared to the first and / or second region and generate the graphical user interface to show the expected direction. For example, where the third region is associated with a wider portion along one side of the graphical user interface corresponding to a particular side of the retina of the patient, the server can generate a visual indication of progression of the disease in the direction in which the third region is greater. This can allow for visualization of the structural characteristics within the eyeball and aid in predicting disease progression.

[0157] FIG. 11 illustrates a non-limiting example of a dataflow for an implementation process 1100 of techniques for segmenting and analyzing OCT images to identify and quantify specific features within the OCT images, in accordance with one or more embodiments. In some implementations, one or more of the computing devices described can be the same as, or similar to, one or more of the computing devices of FIG. 1.076333-1065 / 07025 PATENT

[0158] In some embodiments, the process 1100 can include a first stage 1100a and a second stage 1100b. For example, the implementation process 1100 can include a first stage 1100a where a server (e.g., that is the same as, or similar to, the server 140 of FIG. 1) is configured to implement a first ANN to generate one or more indications of whether or not EZ loss is indicated by portions of the structure of a retina of a patient. In some implementations, for both the first stage 1100a and the second stage 1100b, an OCT device that is the same as, or similar to, the OCT devices 110 of FIG. 1 can be configured to generate volumetric data and provide the volumetric data to a server. One or more operations associated with the first stage 1100a and / or the second stage 1100b can be implemented by the server as described herein.

[0159] The first stage 1100a can involve the server obtaining the volumetric data from the OCT device or OCT devices. For example, the server can obtain the volumetric data, where the volumetric data includes one or more OCT B-scans. In this example, the one or more OCT B-scans can include a plurality of OCT A-scans that form of sequence is a form each OCT B-scan. In some embodiments, the server can then extract patches from the volumetric data where the patches overlap around the RPE layer of the retina of the patient involved in the OCT B-scan. The patches can be extracted in accordance with the patch size and or in accordance with one or more portions of the retina structure that are being analyzed by the server as described herein.

[0160] At operation 1102, the server can provide the patches that were extracted from the volumetric data to at least a first portion of an ANN. For example, the server can provide the patches that were extracted from the volumetric data in sequence to the first portion of the ANN. In this example, the first portion of the ANN can include and / or have an architecture associated with a ResNet. In some embodiments, the first portion of the ANN can then be executed by the server to process each patch in accordance with one or more layers to establish feature maps in accordance with one or more resolutions corresponding to the one or more layers.

[0161] At operation 1104, the server can provide the feature maps generated by the first portion of the ANN to a second portion of the ANN. For example, the server can provide the feature maps generated by the first portion of the ANN in sequence to the second portion of the ANN. In this example, the second portion of the ANN can include and / or have an architecture associated with an RNN, such as an LSTM as described herein. In some embodiments, the second portion of the ANN can then be executed by the server to process each feature map to generate076333-1065 / 07025 PATENToutputs indicative of whether or not one or more diseases are present and represented by the structure of the retina such as, for example, EZ loss.

[0162] At operation 1106, the server can generate a graphical user interface to indicate portions of the retina associated with the disease. For example, the server can generate the graphical user interface such that portions of the graphical user interface associated with more and more patches are illustrated along with a set of pixels reserves to indicate where portions of each patch are indicative of, or are not indicative of, a disease such as EZ loss. In the illustrated example, the originally obtained OCT B-scans can be used to generate the graphical user interface and a bar extending along an upper portion of the images can be overlaid onto (or above) the respective OCT B-scans, where sub portions of the bar are colored to indicate the presence or non-presence of the disease as represented by the structure of the retina. It will be understood that any suitable indication can be implemented.

[0163] The second stage 1100b can involve the server obtaining the volumetric data from the OCT device or OCT devices. In some embodiments, the server can then generate an en face image based on the volumetric data. For example, the server can receive the OCT B-scan and generate the en face image by processing multiple B-scans acquired in a raster pattern to create a volumetric dataset, then extracting a transverse slice at a specific depth of the retina of the patient. This can involve the server analyzing the structure of the retina represented by the volumetric data, applying surface segmentation to delineate anatomical boundaries, and using these anatomical boundaries to generate the en face image, which represents a planar view of the retina at a chosen depth.

[0164] At operation 1108, the server can execute a second ANN by providing the en face image as an input to the second ANN. In this example, the second ANN can include and / or have an architecture associated with a diffusion model as described herein. The diffusion model can be configured to receive the en face image, execute one or more operations to process the en face image, and generate a segmentation mask as described here in.

[0165] At operation 1110, the server can generate a segmentation mask using the diffusion model, where the segmentation mask represents a segmented region of the en face image indicative of GA. For example, the diffusion model can iteratively introduce and remove noise to and from076333-1065 / 07025 PATENTthe en face image to generate the segmentation mask. Tn this example, the segmentation mask can include a binary mask indicating portions where GA is and is not present.

[0166] At operation 1112, the server can then update the en face image based on the segmentation mask. For example, the server can update the en face image by subtracting intensity values established as being associated with EZ loss as indicated by the output of the first stage 300a from the en face image generated by the server based on the volumetric data.

[0167] At operation 1114, the server can then overlay the segmentation mask onto the updated en face image to indicate the area according to which GA is detected relative to the area at which EZ loss is detected in the en face image.

[0168] At operation 1116, the server can execute one or more operations to transform the en face image indicative of the areas according to which EZ loss and GA our presence in the retina of the patient such that only the region where junctional EZ loss (the area between the regions where EZ loss and GA are identified) is highlighted.

[0169] At operation 1118, the server can then generate a graphical user interface indicative of the areas where the junctional EZ loss is present. For example, the server can generate a graphical user interface by updating the graphical user interface output as a result of the execution of the first stage 1100a to identify one or more regions along the bar extending along an upper portion of the images where junctional EZ loss is present. The server can then generate graphical user interface data associated with any of the graphical user interfaces described herein and provide the graphical user interface data to a client device (e.g., that is the same as, or similar to, the client devices 130 of FIG. 1) to cause the client device to generate a graphical user interface as an output at a display device of the client device.

[0170] FIG. 12 illustrates a non-limiting example of dataflow of a process 1200 for segmenting and analyzing OCT images to identify and quantify specific features within the OCT images in accordance with one or more embodiments. In some implementations, one or more of the computing devices, operations executed, and / or models implemented as described can be the same as, or similar to, those of FIG. 1 or FIG. 11.076333-1065 / 07025 PATENT

[0171] Referring to FIG. 12, initially, a server (e.g., that is the same as, or similar to, the server 140 of FIG. 1) can be configured to receive volumetric data in the form of OCT B-scans from a device such as an OCT device. For example, the server can receive the volumetric data along with an indication from a clinician operating a client device that the volumetric data is to be processed to identify whether or not one or more diseases are represented by the structures as indicated in the volumetric data of the retina patient.

[0172] At operation 1202, the server can be configured to extract en face images from the volumetric data to be used when segmenting portions of the en face images to indicate whether or not one or more diseases are present.

[0173] At operation 1204, the server can be configured to segment a portion of the retinal pigment epithelium (RPE) from the volumetric data using a GAN (such as a Pix2Pix GAN). In this example, the GAN can be configured to receive the volumetric data as an input and generate segmentation maps that delineate the RPE layer, averaging out data from 100 microns below the RPE to improve accuracy and robustness against noise. In some examples, the generator within the GAN can implement a U-Net structure that extracts features from the volumetric data at one or more resolutions and outputs one or more feature maps. In some examples, the feature maps can then be averaged over a specified depth (e.g., 100 microns below the RPE) to generate an en face image indicating a view of the underlying structures.

[0174] At operation 1206, the server can execute a diffusion model to generate GA segmentation maps corresponding to the en face image generated at operation 1204. For example, the server can implement a diffusion model to iteratively execute noise introduction and denoising operations in accordance with the en face image. In this example, the server can cause the diffusion model to generate a segmentation mask that indicates the areas according to which GA is present or not present as indicated by the en face image.

[0175] At operation 1208, the server can project the GA segmentation maps onto the B-scans. For example, the server can project the GA segmentation maps onto the B-scans to indicate, as an additional channel within each pixel of the B-scans, the presence or absence of GA.

[0176] At operation 1210, the server can classify portions of the volumetric data as being associated with EZ loss. For example, the server can process the volumetric data, which includes076333-1065 / 07025 PATENTa series of cross-sectional images of the eyeball of a patient, to identify and quantify specific features indicative of diseases. In this example, the server can determine points associated with the boundary of the retina of the patient and extracts patches from the volumetric data that correspond to these boundaries. These extracted patches can then be processed and classified as described herein.

[0177] At operation 1212, the server can classify the patches extracted from the volumetric data (e.g., from the A-scans of the respective OCT B-scans) into sets (e.g., a first set, a second set, etc.) based on their structural characteristics using an ANN having one or more architectures. For example, the server can execute an ANN having a first architecture portion that is associated with a ResNet and a second architecture portion that is associated with an LSTM. The server can execute the ANN to generate feature maps using the first architecture portion that identify regions within the eyeball exhibiting signs of EZ loss. The server can then provide the feature maps as an input to the second architecture portion to generate indications of whether or not EZ loss is indicated based on the structure of the retina associated with the initial patch being processed by the ANN.

[0178] At operation 1214, the server can combine the classifications generated as described herein to indicate a region according to which EZ junction loss is detected. The server can project GA into B-scans as a different channel and combine them with A-scan classifications to generate an en face image indicative of EZ junction loss. The server can enhance the visualization of the EZ junction loss through this combination. In some examples, the server can provide a more comprehensive assessment of the retinal structure by executing the operations described with respect to the process 1200 and improve diagnostic accuracy by integrating B-scans and A-scan classifications. In some examples, the server can generate a graphical user interface as described herein to indicate the areas according to which easy junctional loss is identified.

[0179] The foregoing method descriptions and the process flow diagrams are provided merely as illustrative examples and are not intended to require or imply that the steps of the various embodiments must be performed in the order presented. The steps in the foregoing embodiments can be performed in any order. Words such as “then,” “next,” etc., are not intended to limit the order of the steps; these words are simply used to guide the reader through the description of the methods. Although process flow diagrams can describe the operations as a sequential process,076333-1065 / 07025 PATENTmany of the operations can be performed in parallel or concurrently. In addition, the order of the operations can be re-arranged. A process can correspond to a method, a function, a procedure, a subroutine, a subprogram, and the like. When a process corresponds to a function, the process termination can correspond to a return of the function to a calling function or a main function.

[0180] Some non-limiting embodiments of the present disclosure are described herein in connection with a threshold. As described herein, satisfying a threshold can refer to a value being greater than the threshold, more than the threshold, higher than the threshold, greater than or equal to the threshold, less than the threshold, fewer than the threshold, lower than the threshold, less than or equal to the threshold, equal to the threshold, and / or the like.

[0181] No aspect, component, element, structure, act, step, function, instruction, and / or the like used herein should be construed as critical or essential, unless explicitly described as such. In addition, as used herein, the articles “a” and “an” are intended to include one or more items and can be used interchangeably with “one or more” and “at least one.” Furthermore, as used herein, the term “set” is intended to include one or more items (e.g., related items, unrelated items, a combination of related and unrelated items, etc.) and can be used interchangeably with “one or more” or “at least one.” Where only one item is intended, the term “one” or similar language is used. Also, as used herein, the terms “has,” “have,” “having,” or the like are intended to be open ended terms. Further, the phrase “based on” is intended to mean “based at least partially on” unless explicitly stated otherwise.

[0182] The various illustrative logical blocks, modules, circuits, and algorithm steps described in connection with the embodiments disclosed herein can be implemented as electronic hardware, computer software, or combinations of both. To clearly illustrate this interchangeability of hardware and software, various illustrative components, blocks, modules, circuits, and steps have been described above generally in terms of their functionality. Whether such functionality is implemented as hardware or software depends upon the particular application and design constraints imposed on the overall system. Skilled artisans can implement the described functionality in varying ways for each particular application, but such implementation decisions should not be interpreted as causing a departure from the scope of the present disclosure.076333-1065 / 07025 PATENT

[0183] Embodiments implemented in computer software can be implemented in software, firmware, middleware, microcode, hardware description languages, or any combination thereof. A code segment or machine-executable instructions can represent a procedure, a function, a subprogram, a program, a routine, a subroutine, a module, a software package, a class, or any combination of instructions, data structures, or program statements. A code segment can be coupled to another code segment or a hardware circuit by passing and / or receiving information, data, arguments, parameters, or memory contents. Information, arguments, parameters, data, etc. can be passed, forwarded, or transmitted via any suitable means, including memory sharing, message passing, token passing, network transmission, etc.

[0184] The actual software code or specialized control hardware used to implement these systems and methods is not limiting. Thus, the operation and behavior of the systems and methods were described without reference to the specific software code being understood that software and control hardware can be designed to implement the systems and methods based on the description herein.

[0185] When implemented in software, the functions can be stored as one or more instructions or code on a non-transitory computer-readable or processor-readable storage medium. The steps of a method or algorithm disclosed herein can be embodied in a processor-executable software module, which can reside on a computer-readable or processor-readable storage medium. A non-transitory computer-readable or processor-readable media includes both computer storage media and tangible storage media that facilitate transfer of a computer program from one place to another. A non-transitory processor-readable storage media can be any available media that can be accessed by a computer. By way of example, and not limitation, such non-transitory processor-readable media can comprise RAM, ROM, EEPROM, CD-ROM or other optical disk storage, magnetic disk storage or other magnetic storage devices, or any other tangible storage medium that can be used to store desired program code in the form of instructions or data structures and that can be accessed by a computer or processor. Disk and disc, as used herein, include compact disc (CD), laser disc, optical disc, digital versatile disc (DVD), floppy disk, and Blu-ray disc, where disks usually reproduce data magnetically, while discs reproduce data optically with lasers. Combinations of the above should also be included within the scope of computer-readable media. Additionally, the operations of a method or algorithm can reside as one or any combination or set076333-1065 / 07025 PATENTof codes and / or instructions on a non-transitory processor-readable medium and / or computer-readable medium, which can be incorporated into a computer program product.

[0186] The preceding description of the disclosed embodiments is provided to enable any person skilled in the art to make or use the present disclosure. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the generic principles defined herein can be applied to other embodiments without departing from the spirit or scope of the disclosure. Thus, the present disclosure is not intended to be limited to the embodiments shown herein but is to be accorded the widest scope consistent with the following claims and the principles and novel features disclosed herein.

[0187] While various aspects and embodiments have been disclosed, other aspects and embodiments are contemplated. The various aspects and embodiments disclosed are for purposes of illustration and are not intended to be limiting, with the true scope and spirit being indicated by the following claims.

Claims

076333-1065 / 07025 PATENTCLAIMSWhat is claimed is:

1. A system for segmenting optical coherence tomography (OCT) images to analyze structures represented by the OCT images, the system comprising:one or more processors configured to:obtain volumetric data comprising a plurality of cross-sectional images received from an OCT imaging device, the volumetric data representing three-dimensional image of an eyeball;extract a plurality of feature maps based on the volumetric data to determine a plurality of segmentations across the plurality of cross-sectional images;determine a first boundary of a structure and a second boundary of the structure based on at least one feature map of the plurality of feature maps;segment a plurality of substructures within a region established by the first boundary and the second boundary; andgenerate an image of at least a portion of the structure of the eyeball based on the volumetric data and the plurality of substructures to indicate locations of the plurality of substructures within the structure.

2. The system of claim 1, wherein the one or more processors configured to determine the first boundary and the second boundary are configured to:update the at least one feature map by executing one or more morphological operations to adjust a representation of the first boundary and the second boundary.

3. The system of claim 1, wherein the one or more processors configured to determine the first boundary and the second boundary are configured to:compare a relative position of the first boundary and the second boundary within the structure to determine a first type corresponding to the first boundary and a second type corresponding to the second boundary.

4. The system of claim 1, wherein the one or more processors are further configured to: determine one or more aspects of the plurality of substructures of the eyeball to augment in the image of the at least a portion of the structure,076333-1065 / 07025 PATENTwherein the one or more processors configured to determine the first boundary of the structure and the second boundary of the structure are configured to:determine the first boundary of a structure and the second boundary of the structure based on at least one feature map corresponding to the one or more aspects.

5. The system of claim 1, wherein the one or more processors are further configured to: determine one or more aspects of the plurality of substructures of the eyeball to augment in the image of the at least a portion of the structure,wherein the one or more processors configured to segment the plurality of substructures within the region are configured to:segment the plurality of substructures within the region based on the one or more aspects of the plurality of substructures of the eyeball to augment in the image.

6. The system of claim 1, wherein the one or more processors configured to segment the plurality of substructures within the region are configured to:generate a first updated feature map by executing one or more first morphological operations to adjust a representation of the first boundary and the second boundary;generate a second updated feature map by executing one or more second morphological operations to adjust the representation of the first boundary and the second boundary; and segment the plurality of substructures within the region based on the first updated feature map and the second updated feature map.

7. The system of claim 6, wherein the at least one feature map comprises a first feature map, the plurality of substructures comprise a first plurality of substructures, andwherein the one or more processors are further configured to:determine a first boundary of a second structure and a second boundary of the second structure based on a second feature map of the plurality of feature maps; and segment a second plurality of substructures within the region established by the first boundary and the second boundary.076333-1065 / 07025 PATENT8. The system of claim 7, wherein the one or more processors configured to generating the image of the at least a portion of the structure of the eyeball are configured to:generate the image of the at least a portion of the structure of the eyeball based on segmenting the plurality of substructures within the region.

9. A computer-implemented method, comprising:obtaining volumetric data comprising a plurality of cross-sectional images received from an optical coherence tomography (OCT) imaging device, the volumetric data representing three-dimensional image of an eyeball;extracting a plurality of feature maps based on the volumetric data to determine a plurality of segmentations across the plurality of cross-sectional images;determining a first boundary of a structure and a second boundary of the structure based on at least one feature map of the plurality of feature maps;segmenting a plurality of substructures within a region established by the first boundary and the second boundary; andgenerating an image of at least a portion of the structure of the eyeball based on the volumetric data and the plurality of substructures to indicate locations of the plurality of substructures within the structure.

10. The computer-implemented method of claim 9, wherein determining the first boundary and the second boundary comprises:updating the at least one feature map by executing one or more morphological operations to adjust a representation of the first boundary and the second boundary.

11. The computer-implemented method of claim 9, wherein determining the first boundary and the second boundary comprises:comparing a relative position of the first boundary and the second boundary within the structure to determine a first type corresponding to the first boundary and a second type corresponding to the second boundary.

12. The computer-implemented method of claim 9, further comprising:determining one or more aspects of the plurality of substructures of the eyeball to augment in the image of the at least a portion of the structure,076333-1065 / 07025 PATENTwherein determining the first boundary of the structure and the second boundary of the structure comprises:determining the first boundary of a structure and the second boundary of the structure based on at least one feature map corresponding to the one or more aspects.

13. The computer-implemented method of claim 9, further comprising:determining one or more aspects of the plurality of substructures of the eyeball to augment in the image of the at least a portion of the structure,wherein segmenting the plurality of substructures within the region comprises:segmenting the plurality of substructures within the region based on the one or more aspects of the plurality of substructures of the eyeball to augment in the image.

14. The computer-implemented method of claim 9, wherein segmenting the plurality of substructures within the region comprises:generating a first updated feature map by executing one or more first morphological operations to adjust a representation of the first boundary and the second boundary;generating a second updated feature map by executing one or more second morphological operations to adjust the representation of the first boundary and the second boundary; and segmenting the plurality of substructures within the region based on the first updated feature map and the second updated feature map.

15. The computer-implemented method of claim 14, wherein the at least one feature map comprises a first feature map, the plurality of substructures comprise a first plurality of substructures, the computer-implemented method further comprising:determining a first boundary of a second structure and a second boundary of the second structure based on a second feature map of the plurality of feature maps; andsegmenting a second plurality of substructures within the region established by the first boundary and the second boundary.

16. The computer-implemented method of claim 15, wherein generating the image of the at least a portion of the structure of the eyeball comprises:generating the image of the at least a portion of the structure of the eyeball based on segmenting the plurality of substructures within the region.076333-1065 / 07025 PATENT17. A non-transitoiy computer-readable medium for storing machine-readable processorexecutable instructions that, when executed by one or more processors, cause the one or more processors to:obtain volumetric data comprising a plurality of cross-sectional images received from an optical coherence tomography (OCT) imaging device, the volumetric data representing three-dimensional image of an eyeball;extract a plurality of feature maps based on the volumetric data to determine a plurality of segmentations across the plurality of cross-sectional images;determine a first boundary of a structure and a second boundary of the structure based on at least one feature map of the plurality of feature maps;segment a plurality of substructures within a region established by the first boundary and the second boundary; andgenerate an image of at least a portion of the structure of the eyeball based on the volumetric data and the plurality of substructures to indicate locations of the plurality of substructures within the structure.

18. The non-transitory computer-readable medium of claim 17, wherein the instructions that cause the one or more processors to determine the first boundary and the second boundary cause the one or more processors to:update the at least one feature map by executing one or more morphological operations to adjust a representation of the first boundary and the second boundary.

19. The non-transitory computer-readable medium of claim 17, wherein the instructions that cause the one or more processors to determine the first boundary and the second boundary cause the one or more processors to:compare a relative position of the first boundary and the second boundary within the structure to determine a first type corresponding to the first boundary and a second type corresponding to the second boundary.

20. The non-transitory computer-readable medium of claim 17, wherein the instructions further cause the one or more processors to:076333-1065 / 07025 PATENTdetermine one or more aspects of the plurality of substructures of the eyeball to augment in the image of the at least a portion of the structure,wherein the instructions that cause the one or more processors to determine the first boundary of the structure and the second boundary of the structure cause the one or more processors to:determine the first boundary of a structure and the second boundary of the structure based on at least one feature map corresponding to the one or more aspects.

21. A system for segmenting and analyzing optical coherence tomography (OCT) images to identify and quantify specific features within the OCT images, the system comprising:one or more processors configured to:obtain volumetric data comprising a plurality of cross-sectional images received from an OCT imaging device, the volumetric data representing three-dimensional image of an eyeball;determine a plurality of points associated with a boundary of a structure of the eyeball based on the cross-sectional images;extract a plurality of patches from the volumetric data that correspond to the boundary of the structure;determine a first set of patches having a first structural characteristic and a second set of patches having a second structural characteristic based on the plurality of patches; and generate a graphical user interface based on the plurality of patches indicating a first region within the eyeball where the first structural characteristic is present and a second region within the eyeball where the second structural characteristic is present.

22. The system of claim 21, wherein the one or more processors configured to extract the plurality of patches from the volumetric data are configured to:determine a patch size to be used when extracting the plurality of patches; and extract the plurality of patches from the volumetric data that correspond to the boundary of the structure based on the boundary of the structure and the patch size.

23. The system of claim 22, wherein the one or more processors configured to determine the patch size are configured to:076333-1065 / 07025 PATENTdetermine the patch size based on a location of one or more adjacent layers to a target layer within the eyeball and a location of a choroid within the eyeball relative to the target layer.

24. The system of claim 21, wherein the one or more processors configured to determine the first set of patches and the second set of patches are configured to:execute an artificial neural network based on the plurality of patches to generate a plurality of annotations corresponding to the plurality of patches; andsegment the plurality of patches to generate the first set of patches and the second set of patches based on the plurality of annotations.

25. The system of claim 21, wherein the plurality of patches comprises a first plurality of patches, andwherein the one or more processors configured to determine the first set of patches having the first structural characteristic and the second set of patches having the second structural characteristic are configured to:execute an artificial neural network by providing at least one patch of patches as an input to the artificial neural network to cause the artificial neural network to execute one or more operations to extract features from the at least one patch at a first resolution and a second resolution to generate at least one feature map, anddetermine the first set of patches having the first structural characteristic and the second set of patches having the second structural characteristic based on the at least one feature map.

26. The system of claim 25, wherein the one or more processors configured to determine the first set of patches having the first structural characteristic and the second set of patches having the second structural characteristic are configured to:generate a set of annotations based on determining the first set of patches and the second set of patches, andwherein the one or more processors configured to generate the graphical user interface are configured to:076333-1065 / 07025 PATENTdetermine the first region within the eyeball where the first structural characteristic is present and the second region within the eyeball where the second structural characteristic is present based on the at least one feature map; andgenerate the graphical user interface based on determining the first region within the eyeball and the second region within the eyeball.

27. The system of claim 26, wherein the artificial neural network comprises a first artificial neural network, andwherein the one or more processors configured to generate the set of annotations are configured to:provide the first set of patches and the second set of patches as inputs to a second artificial neural network in accordance with a sequence established by the first set of patches and the second set of patches to cause the second artificial neural network to generate the set of annotations.

28. The system of claim 21, wherein the one or more processors are further configured to: generate an en face image of at least a portion of the structure of the eyeball based on the volumetric data;execute one or more operations to determine a segmentation mask based on the en face image; andcompare the segmentation mask to the first region within the eyeball to determine a third region, the second region comprising a subset of the second region,wherein the one or more processors configured to generate the graphical user interface are configured to:generate the graphical user interface to indicate the third region based on determining the third region.

29. The system of claim 28, wherein the one or more processors configured to execute the one or more operations to determine the segmentation mask are configured to:execute an artificial neural network that is configured to receive the en face image as an input and generate the segmentation mask as an output.076333-1065 / 07025 PATENT30. The system of claim 29, wherein the one or more processors configured to execute the artificial neural network are configured to:provide the en face image as an input to the artificial neural network to cause the artificial neural network to execute one or more operations to:iteratively introducing noise to the en face image and, in response to introducing the noise, removing at least a portion of the noise from the en face image to generate the segmentation mask.

31. The system of claim 28, wherein the one or more processors configured to compare the segmentation mask to the first region within the eyeball are configured to:update the en face image to indicate the first region based on the plurality of patches indicating a first region within the eyeball where the first structural characteristic is present to generate a first annotated image;overlay the segmentation mask onto the first annotated image to generate a second annotated image; anddetermine the third region based on the second annotated image.

32. The system of claim 28, wherein the one or more processors are further configured to: determine an expected direction for disease progression based on one or more attributes of the third region; andgenerate the graphical user interface to the expected direction.

33. A computer-implemented method, comprising:obtaining volumetric data comprising a plurality of cross-sectional images received from an optical coherence tomography (OCT) imaging device, the volumetric data representing three-dimensional image of an eyeball;determining a plurality of points associated with a boundary of a structure of the eyeball; extract a plurality of patches from the volumetric data that correspond to the boundary of the structure;determining a first set of patches having a first structural characteristic and a second set of patches having a second structural characteristic based on the plurality of patches; and076333-1065 / 07025 PATENTgenerating a graphical user interface based on the plurality of patches indicating a first region within the eyeball where the first structural characteristic is present and a second region within the eyeball where the second structural characteristic is present.

34. The computer-implemented method of claim 33, wherein extracting the plurality of patches from the volumetric data comprises:determining a patch size to be used when extracting the plurality of patches; and extracting the plurality of patches from the volumetric data that correspond to the boundary of the structure based on the boundary of the structure and the patch size.

35. The computer-implemented method of claim 34, wherein determining the patch size comprises:determining the patch size based on a location of one or more adjacent layers to a target layer within the eyeball and a location of a choroid within the eyeball relative to the target layer.

36. The computer-implemented method of claim 33, wherein determining the first set of patches and the second set of patches comprises:executing an artificial neural network based on the plurality of patches to generate a plurality of annotations corresponding to the plurality of patches; andsegmenting the plurality of patches to generate the first set of patches and the second set of patches based on the plurality of annotations.

37. The computer-implemented method of claim 33, wherein the plurality of patches comprises a first plurality of patches, andwherein determining the first set of patches having the first structural characteristic and the second set of patches having the second structural characteristic comprises:executing an artificial neural network by providing at least one patch of patches as an input to the artificial neural network to cause the artificial neural network to execute one or more operations to extract features from the at least one patch at a first resolution and a second resolution to generate at least one feature map, anddetermining the first set of patches having the first structural characteristic and the second set of patches having the second structural characteristic based on the at least one feature map.076333-1065 / 07025 PATENT38. The computer-implemented method of claim 37, wherein determining the first set of patches having the first structural characteristic and the second set of patches having the second structural characteristic comprises:generating a set of annotations based on determining the first set of patches and the second set of patches, andwherein generating the graphical user interface comprises:determining the first region within the eyeball where the first structural characteristic is present and the second region within the eyeball where the second structural characteristic is present based on the at least one feature map; andgenerating the graphical user interface based on determining the first region within the eyeball and the second region within the eyeball.

39. The computer-implemented method of claim 38, wherein the artificial neural network comprises a first artificial neural network, andwherein generating the set of annotations comprises:providing the first set of patches and the second set of patches as inputs to a second artificial neural network in accordance with a sequence established by the first set of patches and the second set of patches to cause the second artificial neural network to generate the set of annotations.

40. A non-transitoiy computer-readable medium for storing machine-readable processorexecutable instructions that, when executed by one or more processors, cause the one or more processors to:obtain volumetric data comprising a plurality of cross-sectional images received from an optical coherence tomography (OCT) imaging device, the volumetric data representing three-dimensional image of an eyeball;determine a plurality of points associated with a boundary of a structure of the eyeball; extract a plurality of patches from the volumetric data that correspond to the boundary of the structure;determine a first set of patches having a first structural characteristic and a second set of patches having a second structural characteristic based on the plurality of patches; and076333-1065 / 07025 PATENTgenerate a graphical user interface based on the plurality of patches indicating a first region within the eyeball where the first structural characteristic is present and a second region within the eyeball where the second structural characteristic is present.