Machine learning based optical retinal imaging using phase sensitive optical coherence tomography
By combining phase-sensitive OCT and machine learning models, the problem of insufficient resolution of OCT technology in rodent optical retinal imaging was solved, and accurate resolution and nanoscale dynamic detection of outer retinal band signals were achieved.
Patent Information
- Application Number
- CN202480009645.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-01-26
- Filing Date
- 2024-01-26
- Publication Date
- 2025-09-16
AI Technical Summary
Existing OCT technology cannot effectively distinguish signals from different outer retinal bands in rodent optical retinal imaging, resulting in insufficient axial resolution and affecting the accuracy of dynamic imaging in response to visual stimuli.
Phase-sensitive OCT is combined with a machine learning model. By acquiring and marking the temporal phase signals of the OCT images, the machine learning model is trained to distinguish the signals of different outer retinal bands and improve imaging resolution.
The accuracy of dynamic imaging or detection of outer retinal tissue in response to visual stimuli is improved, enabling the detection of nanoscale tissue dynamics.
Smart Images

Figure CN120660122A_ABST
Abstract
Description
[0001] Cross-references
[0002] This application claims priority to Singapore Patent Application No. 10202300210V filed on January 26, 2023, the contents of which are incorporated herein by reference in their entirety for all purposes. Technical Field
[0003] The present invention generally relates to machine learning-based optoretinography (ORG) using phase-sensitive optical coherence tomography (OCT), and more particularly, to a method and system for training a machine learning model for classifying temporal phase signals obtained from OCT images generated for optoretinography, and a method and system for classifying temporal phase signals obtained from OCT images generated for optoretinography using the trained machine learning model. Background Art
[0004] Phototransduction involves changes in the concentration of ions and other solutes in the photoreceptor cells and the subretinal space, which affects osmotic pressure and associated water flow. The ion balance of the subretinal space is regulated by the retinal pigment epithelium (RPE), a single layer of cells between the outer segments of the photoreceptor cells and the choroid. And a variety of eye diseases (such as age-related macular degeneration, retinitis pigmentosa and Best yolk-like macular dystrophy) are associated with epithelial transport dysfunction of RPE cells. The in vivo assessment of RPE ion transport mainly relies on the c wave or electrooculogram (EOG) in the electroretinogram (ERG), which requires contact with electrodes and can provide low spatial resolution.
[0005] Optoretinography (ORG) is an emerging imaging technique for the noninvasive optical probing of retinal physiological activity in vivo. Specifically, optoretinography is an optical imaging modality used for the noninvasive diagnosis of photoreceptor cell function in vivo. It typically utilizes phase-sensitive optical coherence tomography (OCT) to detect mechanical deformations of retinal cells in response to visual stimuli, such as deformations of the outer segments (OS) of photoreceptor cells associated with photoisomerization and phototransduction. Therefore, optoretinography uses OCT to measure functionally relevant mechanical deformations of specialized neurons in the retina. Unlike conventional OCT, whose axial motion sensitivity is limited by bandwidth-constrained axial resolution (typically >1 μm), phase-sensitive OCT's axial sensitivity is determined by the signal-to-noise ratio, potentially enabling the detection of nanoscale tissue dynamics. For example, optoretinography has recently been applied to classify human color cone subtypes based on their responses to different colored flashes, and its clinical potential has been demonstrated in assessing the progression of retinitis pigmentosa. In addition, given the similarities in the structure and function of photoreceptor cells in mammals, small animals such as rodents can serve as convenient and efficient models for studying the mechanisms of photoreceptor cell degeneration using retinal imaging techniques. In genetically modified animals (such as Gnat1 and Gnat2 mutants and knockout models), light retinal imaging can reveal new insights into various stages of the phototransduction cascade. In addition, unlike human patients, animals usually do not have other ocular lesions, and there is less interference from motion artifacts during imaging under anesthesia. Compared with traditional methods such as electroretinography and visual field detection, light retinal imaging technology has become an urgent need in clinical applications due to its advantages of non-invasiveness, objectivity and ultra-high spatial resolution. For example, when no obvious signs of retinal degeneration are visible in structural images, this technology can serve as an ideal tool for the diagnosis / prognosis of early retinal degenerative diseases.
[0006] Despite these potential advantages of performing optical retinal imaging studies in rodents, the axial resolution of near-infrared OCT (NIR-OCT) systems, typically limited to a few micrometers in tissue, limits the ability to clearly distinguish the outer segment (OS) terminals of rodent photoreceptor cells from the RPE, particularly given the interdigitating presence of microvilli with the OS. For example, in typical structural images acquired by NIR-OCT, the OS and RPE are composited in a single speckle layer. Furthermore, due to the coherent convolution of backscattered light with the point spread function of the optical system, a single pixel (individual pixel) in an OCT image can contain mixed phase signals from multiple cell types. Consequently, separating optical retinal imaging signals from distinct outer retinal zones (e.g., OS and RPE) presents a significant technical challenge.
[0007] Therefore, there is a need to provide a method and system for performing optical retinal imaging using OCT, which seeks to overcome or at least improve one or more deficiencies in conventional methods of performing optical retinal imaging using OCT, and more specifically, to improve / improve the accuracy of imaging in response to visual stimuli or the accuracy (e.g., resolution) of detecting outer retinal tissue dynamics (functional responses, such as movement or deformation) in a manner that enables resolution of signals from different outer retinal zones, such as detecting tissue dynamics at the nanometer level. The present invention was developed in this context. Summary of the Invention
[0008] According to a first aspect of the present invention, there is provided a method for training a machine learning model for classifying a temporal phase signal obtained from an optical coherence tomography (OCT) image using at least one processor, the optical coherence tomography (OCT) image being generated for optical retinal imaging, the method comprising:
[0009] Acquiring a sequence of optical coherence tomography (OCT) images of the outer retina generated by a phase-sensitive optical coherence tomography (OCT) system for optical imaging of the retina, including a stimulus subsequence of OCT images generated in response to a visual stimulus applied to the outer retina;
[0010] determining, for each OCT image in the OCT image sequence and for each individual pixel in a target layer of the OCT image, a phase signal of the individual pixel in the target layer of the OCT image relative to a reference layer of the OCT image, so as to acquire, for each of a plurality of corresponding individual pixel sequences in the target layer across the OCT image sequence, a temporal phase signal of the corresponding individual pixel sequence, thereby generating a set of temporal phase signals of the plurality of corresponding individual pixel sequences, the target layer comprising a plurality of outer retinal zones;
[0011] Projecting multiple time phase signals in the time phase signal set to multiple feature points in a feature space;
[0012] marking each feature point in the plurality of feature points in the feature space; and
[0013] The machine learning model is trained based on the multiple labeled feature points in the feature space.
[0014] According to a second aspect of the present invention, there is provided a method for classifying a temporal phase signal obtained from an optical coherence tomography (OCT) image generated for optical retinal imaging, the method using a machine learning model trained according to the first aspect of the present invention, the method comprising:
[0015] Acquiring a sequence of optical coherence tomography (OCT) images of the outer retina generated by a phase-sensitive optical coherence tomography (OCT) system for optical imaging of the retina, including a stimulus subsequence of OCT images generated in response to a visual stimulus applied to the outer retina;
[0016] determining, for each individual pixel in a corresponding sequence of individual pixels in a target layer across the sequence of OCT images, a phase signal of the individual pixel in the target layer of the OCT image corresponding to a reference layer of the OCT image to obtain a temporal phase signal of the corresponding sequence of individual pixels, the target layer comprising a plurality of outer retinal bands;
[0017] Projecting the temporal phase signal of the corresponding individual pixel sequence to a feature point in a feature space; and
[0018] The time phase signal is classified using the trained machine learning model based on the feature points to which the time phase signal is projected.
[0019] According to a third aspect of the present invention, there is provided a system for training a machine learning model for classifying temporal phase signals obtained from OCT images generated for optical retinal imaging, the system comprising:
[0020] at least one memory, and
[0021] At least one processor is communicatively coupled to the at least one memory and is configured to execute the method of training a machine learning model according to the first aspect of the present invention.
[0022] According to a fourth aspect of the present invention, there is provided a system for classifying a temporal phase signal obtained from an optical coherence tomography (OCT) image, the system comprising:
[0023] at least one memory; and
[0024] At least one processor is communicatively coupled to the at least one memory and is configured to perform the method for classifying a time phase signal according to the second aspect of the present invention.
[0025] According to a fifth aspect of the present invention, there is provided a computer program product embodied in one or more non-temporary computer-readable storage media, comprising instructions executable by at least one processor, the instructions being used to execute the method for training a machine learning model according to the first aspect of the present invention.
[0026] According to a sixth aspect of the present invention, a computer program product is embodied in one or more non-temporary computer-readable storage media, comprising instructions executable by at least one processor, the instructions being used to perform the method for classifying a time phase signal according to the second aspect of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] Embodiments of the present invention will be better understood and apparent to those skilled in the art through the following description which is given by way of example only in conjunction with the accompanying drawings, in which:
[0028] Figure 1 A schematic flow chart of a method for training a machine learning model according to an embodiment of the present invention is shown, wherein the machine learning model is used to classify a temporal phase signal obtained from an OCT image generated for optical retinal imaging;
[0029] Figure 2 A schematic flow chart of a method for classifying a temporal phase signal obtained from an OCT image generated for optical retinal imaging according to an embodiment of the present invention is shown;
[0030] Figure 3 A schematic block diagram of a system for training a machine learning model for classifying temporal phase signals obtained from OCT images generated for optical retinal imaging according to an embodiment of the present invention is shown;
[0031] Figure 4 A schematic block diagram of a system for classifying a temporal phase signal obtained from an OCT image generated for optical retinal imaging using a trained machine learning model according to an embodiment of the present invention is shown;
[0032] Figure 5 A schematic block diagram of an exemplary computer system that can be used to implement or embody a system for training a machine learning model and / or a system for classifying a time phase signal according to an embodiment of the present invention is shown;
[0033] Figures 6A to 6I An example image and signal processing pipeline / framework for a method for training a machine learning model according to an embodiment of the present invention is shown;
[0034] Figure 7A shows unsupervised clustering using hierarchical clustering and classification of subsequent new OCT image datasets according to an embodiment of the present invention;
[0035] Figure 7B shows the trained SVM decision boundaries of Type-I and Type-II signals in principal component (PC) space according to an embodiment of the present invention;
[0036] Figure 7C shows average phase traces of classified Type-I and Type-II signals according to an embodiment of the present invention, wherein the band shows their standard deviations;
[0037] Figure 7D The embodiment of the present invention is shown Figure 7C The distribution of peak ΔOPL in and the corresponding delays of individual signals;
[0038] Figure 8 A schematic flow chart outlining the following is shown according to an embodiment of the present invention: processing a complex-valued OCT signal of an OCT image to obtain a temporal phase signal for training a machine learning model, training of the machine learning model, and classification of a new temporal phase signal obtained with respect to a target layer using the trained machine learning model;
[0039] Figures 9A to 9D Shown Figure 8 The various stages of the flow chart shown (labelled 'A', 'B', 'C' and 'D');
[0040] Figures 10A to 10D The retinal layers and their dynamics in response to visual stimulation in an experiment according to an embodiment of the present invention are shown;
[0041] Figures 11A to 11D shows unsupervised clustering in spatiotemporal feature space for signal classification according to an embodiment of the present invention;
[0042] Figures 12A to 12C shows the classification of new phase trajectories and verification of their origin according to an embodiment of the present invention;
[0043] Figures 13A to 13C shows the responses of the outer segment (OS) and subretinal space under different conditions according to an embodiment of the present invention;
[0044] Figures 14A to 14C shows a representative en-face functional map of OS and SRS signals according to an embodiment of the present invention;
[0045] Figures 15A to 15C shows an example system setup and example stimulation and acquisition protocols according to an embodiment of the present invention;
[0046] Figure 16A A table (Table 1) showing various details of animal experiments conducted according to an embodiment of the present invention is shown;
[0047] Figure 16B shows a table (Table 2) showing various details of an acquisition and stimulation protocol according to an embodiment of the present invention; and
[0048] Figures 17A to 17E Deformation of the cell layer over time and comparison of light retinal imaging with ERG are shown. DETAILED DESCRIPTION
[0049] Various embodiments of the present invention relate to machine learning-based optical retinal imaging (ORG) using phase-sensitive optical coherence tomography (OCT). To this end, various embodiments of the present invention provide a method and system for training a machine learning model for classifying temporal phase signals obtained from OCT images generated for ORG, as well as a method and system for classifying temporal phase signals obtained from OCT images generated for ORG using the trained machine learning model.
[0050] As described in the background, conventional methods for optical retinal imaging using optical coherence tomography (OCT) suffer from various deficiencies or shortcomings. In particular, due to limited spatial resolution, conventional methods are unable to distinguish signals scattered by individual cells in different outer retinal tissue layers in response to visual stimulation, resulting in a blurred speckle pattern in the OCT images used for optical retinal imaging. Consequently, this blurred speckle pattern (which may be referred to as a speckle layer) in the OCT image negatively impacts the accuracy (e.g., resolution) of imaging or detecting the dynamics (functional responses, such as movement or deformation) of the outer retinal tissue in response to visual stimulation. Therefore, various embodiments advantageously address technical issues associated with these speckle patterns in OCT images. More specifically, a method for optical retinal imaging using OCT is provided that is capable of distinguishing signals from different outer retinal wavelength bands, thereby enhancing / improving the accuracy (e.g., resolution) of imaging or detecting the dynamics of the outer retinal tissue in response to visual stimulation, enabling, for example, detection of tissue dynamics at the nanometer level. To this end, multiple embodiments of the present invention provide a method and system for training a machine learning model for classifying temporal phase signals obtained from OCT images generated for optical retinal imaging, as well as a method and system for classifying temporal phase signals obtained from OCT images generated for optical retinal imaging using a trained machine learning model.
[0051] Figure 1A flow chart of a method 100 for training a machine learning model by at least one processor according to various embodiments of the present invention is shown, wherein the machine learning model is used to classify a temporal phase signal obtained in an OCT image generated by optical retinal imaging. The method 100 includes: acquiring (at 106) an OCT image sequence (i.e., a series of OCT images) generated by a phase-sensitive OCT system for optical retinal imaging of the outer retina, the OCT image sequence including an OCT image stimulus subsequence generated in response to a visual stimulus applied to the outer retina; for each OCT image in the OCT image sequence and for each individual pixel in a target layer of the OCT image, determining (at 108): a phase signal of the individual pixel in the target layer of the OCT image relative to a reference layer of the OCT image. , to obtain a time phase signal of a plurality of corresponding individual pixel sequences in the target layer across the OCT image sequence, and generate a time phase signal set of the plurality of corresponding individual pixel sequences, wherein the target layer includes a plurality of outer retinal bands; (at 110) projecting the plurality of time phase signals in the time phase signal set to a plurality of feature points in a feature space; (at 112) labeling each of the plurality of feature points in the feature space; and (at 114) training the machine learning model based on the plurality of labeled feature points in the feature space.
[0052] Therefore, the method 100 for training a machine learning model for classifying temporal phase signals obtained from OCT images can advantageously distinguish signals from various outer retinal bands, thereby improving the accuracy (e.g., resolution) of imaging or detecting the dynamics of outer retinal tissue in response to visual stimuli. Specifically, for each corresponding individual pixel in a plurality of corresponding individual pixel sequences in a target layer across an OCT image sequence (one corresponding individual pixel in each OCT image in the OCT image sequence), a temporal phase signal for the corresponding individual pixel sequence is determined, thereby generating a set of temporal phase signals for a plurality of corresponding individual pixel sequences (e.g., for each individual pixel in the target layer). The plurality of temporal phase signals in the set of temporal phase signals are then projected onto a plurality of feature points in a feature space (e.g., a temporal feature space or a spatiotemporal feature space), such that the plurality of feature points can be grouped into a plurality of clusters (each cluster having a corresponding label assigned thereto, which can correspond to one of a plurality of outer retinal tissue types or an outlier type). The machine learning model can then be trained based on the plurality of labeled feature points in the feature space to classify new temporal phase signals. This method has been found to advantageously address the aforementioned technical issues associated with speckle patterns (e.g., a speckle layer, e.g., corresponding to the aforementioned target layer) in OCT images, which negatively impact the accuracy (e.g., resolution) of imaging or detecting the dynamics of outer retinal tissue in response to visual stimuli. Thus, training a machine learning model for classifying temporal phase signals according to method 100 enables the use of machine learning-based retinal light imaging using phase-sensitive OCT to resolve signals (temporal phase signals) from various outer retinal zones in the speckle layer (e.g., corresponding to the aforementioned target layer), thereby enhancing / improving the accuracy (e.g., resolution) of imaging or detecting the dynamics of outer retinal tissue in response to visual stimuli, for example, enabling the detection of nanoscale tissue dynamics. These advantages or technical effects and / or other advantages or technical effects will become more apparent to those skilled in the art as various embodiments and exemplary implementations of the present invention describe in more detail the method 100 for training a machine learning model and the method for classifying temporal phase signals using the trained machine learning model (described below), as well as corresponding systems.
[0053] In multiple embodiments, the above-mentioned determination (at 108) of the phase signal of an individual pixel in a target layer of the OCT image includes: correcting the systematic phase drift of the complex-valued OCT signal of the individual pixel in the target layer based on self-reference of the individual pixel to the reference layer of the OCT image to obtain a first complex-valued OCT signal of the individual pixel in the target layer.
[0054] In multiple embodiments, the above-mentioned correction of the systematic phase drift of the complex-valued OCT signal of the individual pixel includes: multiplying the complex-valued OCT signal of the individual pixel with the complex conjugate of the complex-valued OCT signal of the individual pixel of the reference layer of the OCT image to obtain a first complex-valued OCT signal of the individual pixel.
[0055] In multiple embodiments, for each OCT image in the OCT image stimulation subsequence, the above-mentioned determination (at 108) of the phase signal of the individual pixel in the target layer of the OCT image also includes: correcting the phase offset of the first complex-valued OCT signal of the individual pixel based on the pre-stimulation subsequence in the OCT image sequence acquired before the visual stimulation is applied to the outer retina to obtain the second complex-valued OCT signal of the individual pixel.
[0056] In multiple embodiments, the above-mentioned correction of the phase offset of the first complex-valued OCT signal of the individual pixel includes: determining the average value of the first complex-valued OCT signals of each corresponding individual pixel in the pre-stimulus subsequence of the OCT image to obtain an average pre-stimulus complex-valued OCT signal; and multiplying the first complex-valued OCT signal of the individual pixel with the average pre-stimulus complex-valued OCT signal to obtain a second complex-valued OCT signal of the individual pixel.
[0057] In various embodiments, for each OCT image in the OCT image stimulation subsequence, a phase signal of the individual pixel in the target layer of the OCT image is determined based on the second complex-valued OCT signal of the individual pixel. In various embodiments, similarly, for each OCT image in the OCT image pre-stimulation subsequence, determining (at 108) the phase signal of the individual pixel in the target layer of the OCT image may further include: correcting a phase offset of the first complex-valued OCT signal of the individual pixel based on the pre-stimulation subsequence in the OCT image sequence to obtain the second complex-valued OCT signal of the individual pixel; and determining the phase signal of the individual pixel in the target layer of the OCT image in the same or similar manner as described above for each OCT image in the OCT image stimulation subsequence.
[0058] In various embodiments, the phase signal of an individual pixel in a target layer of the OCT image is determined relative to a reference pixel region of a reference layer of the OCT image, wherein the reference pixel region includes a plurality of individual pixels. To this end, in order to correct for systematic phase drift in the complex-valued OCT signal of the individual pixel, the multiplication of the complex-valued OCT signal of the individual pixel includes multiplying the complex-valued OCT signal of the individual pixel by the complex conjugate of the complex-valued OCT signals of the plurality of individual pixels in the reference pixel region of the reference layer of the OCT image, to obtain a set of first complex-valued OCT signals for the individual pixels. In addition, for the above-mentioned correction of the phase offset of the complex-valued OCT signal of the individual pixel, the above-mentioned determination of the average value of the first complex-valued OCT signal corresponding to the individual pixel includes: determining the average value of the set of first complex-valued OCT signals of each corresponding individual pixel in the pre-stimulus subsequence of the OCT image to obtain a set of averaged pre-stimulus complex-valued OCT signals; and the above-mentioned multiplication of the first complex-valued OCT signal of the individual pixel includes: multiplying the set of first complex-valued OCT signals of the individual pixel with the set of averaged pre-stimulus complex-valued OCT signals to obtain a set of second complex-valued OCT signals of the individual pixel.
[0059] In various embodiments, when determining the phase signals of individual pixels in the target layer of the OCT image relative to a reference pixel region of the reference layer of the OCT image, for each OCT image in the OCT image stimulation subsequence, determining (at 108) the phase signals of the individual pixels in the target layer of the OCT image further includes: determining an average value of a set of second complex-valued OCT signals of the individual pixels to obtain an average complex-valued OCT signal of the individual pixels; and determining the phase signals of the individual pixels in the target layer of the OCT image based on the average complex-valued OCT signal of the individual pixels. In various embodiments, similarly, for each OCT image in the OCT image pre-stimulation subsequence, determining (at 108) the phase signals of the individual pixels in the target layer of the OCT image further includes: determining an average value of a set of second complex-valued OCT signals of the individual pixels to obtain an average complex-valued OCT signal of the individual pixels; and determining the phase signals of the individual pixels in the target layer of the OCT image based on the average complex-valued OCT signal of the individual pixels.
[0060] In various embodiments, the temporal phase signal of the sequence of corresponding individual pixels in the target layer across the OCT image sequence is obtained based on the phase signal of the sequence of corresponding individual pixels in the target layer across the OCT image sequence.
[0061] In multiple embodiments, the above-mentioned marking (at 112) of each feature point in the multiple feature points in the feature space includes: grouping the multiple feature points into multiple clusters in the feature space; and for each cluster, marking the feature points belonging to the cluster with a label assigned to the cluster.
[0062] In various embodiments, the plurality of feature points are grouped into a plurality of clusters based on an unsupervised clustering technique.
[0063] In various embodiments, the plurality of clusters are assigned different labels, and the plurality of different labels include a plurality of different outer retinal zone labels corresponding to a plurality of outer retinal zones of the target layer.
[0064] In various embodiments, the reference layer corresponds to the inner segment / outer segment (IS / OS) junction, and the plurality of outer retinal bands include two or more outer retinal bands corresponding to two or more of the photoreceptor cell outer segments, microvilli, retinal pigment epithelium (RPE), and Bruch's membrane (BrM), respectively.
[0065] In various embodiments, the method 100 further includes: performing band-stop filtering and low-pass filtering on the multiple time phase signals before projecting the multiple time phase signals to multiple feature points in the feature space; and normalizing the multiple time phase signals.
[0066] In various embodiments, the feature space is a temporal feature space or a spatiotemporal feature space.
[0067] Figure 2 A flow chart of a method 200 according to an embodiment of the present invention is shown, wherein the method 200 uses a machine learning model trained according to the method 100 according to an embodiment of the present invention to classify a temporal phase signal obtained from an OCT image generated for optical retinal imaging. The method 200 includes: acquiring (at 206) an OCT image sequence generated by a phase-sensitive OCT system performing optical retinal imaging of the outer retina, the OCT image sequence including an OCT image stimulus subsequence generated in response to a visual stimulus applied to the outer retina; determining (at 208) a phase signal of the individual pixel in a target layer of the OCT image relative to a reference layer of the OCT image for each individual pixel in a corresponding individual pixel sequence in a target layer across the OCT image sequence to obtain a temporal phase signal of the corresponding individual pixel sequence, the target layer including a plurality of outer retinal bands; projecting (at 210) the temporal phase signal of the corresponding individual pixel sequence to a feature point in a feature space; and classifying (at 212) the temporal phase signal based on the feature point to which the temporal phase signal is projected using the trained machine learning model.
[0068] Therefore, for the temporal phase signal classification method 200, the above-mentioned acquisition (at 206) of the OCT image sequence, the above-mentioned determination (at 208) of the phase signals of the individual pixels in the target layer of the OCT image, and the above-mentioned projection (at 210) of the temporal phase signals of the corresponding individual pixel sequence to the feature points in the feature space can be respectively performed in the same or similar manner as the above-described method 100 for training a machine learning model according to an embodiment of the present disclosure, wherein the acquisition (at 106) of the OCT image sequence, the determination (at 108) of the phase signals of the individual pixels in the target layer of the OCT image, and the projection (at 110) of the multiple temporal phase signals of the corresponding individual pixel sequence to the multiple feature points in the feature space are performed. In addition, in many embodiments, the above-mentioned projection (at 210) of the temporal phase signals of the corresponding individual pixel sequence to the feature points in the feature space uses the same feature space as the feature space used to train the machine model. In many embodiments, the feature space is a temporal feature space or a spatiotemporal feature space.
[0069] In multiple embodiments, the above-mentioned determination (at 208) of the phase signal of an individual pixel in the target layer of the OCT image includes: correcting the systematic phase drift of the complex-valued OCT signal of the individual pixel based on self-reference of the individual pixel to the reference layer of the OCT image to obtain a first complex-valued OCT signal of the individual pixel.
[0070] In multiple embodiments, the above-mentioned correction of the systematic phase drift of the complex-valued OCT signal of the individual pixel includes: multiplying the complex-valued OCT signal of the individual pixel with the complex conjugate of the complex-valued OCT signal of the individual pixel of the reference layer of the OCT image to obtain a first complex-valued OCT signal of the individual pixel.
[0071] In multiple embodiments, for each OCT image in the OCT image stimulation subsequence, the above-mentioned determination (at 208) of the phase signal of the individual pixel in the target layer of the OCT image also includes: correcting the phase offset of the first complex-valued OCT signal of the individual pixel based on the pre-stimulation subsequence collected before the visual stimulation is applied to the outer retina in the OCT image sequence to obtain a second complex-valued OCT signal of the individual pixel.
[0072] In multiple embodiments, the above-mentioned correction of the phase offset of the first complex-valued OCT signal of the individual pixel includes: determining the average value of the first complex-valued OCT signals of each corresponding individual pixel in the pre-stimulus subsequence of the OCT image to obtain an average pre-stimulus complex-valued OCT signal; and multiplying the first complex-valued OCT signal of the individual pixel with the average pre-stimulus complex-valued OCT signal to obtain a second complex-valued OCT signal of the individual pixel.
[0073] In various embodiments, for each OCT image in the OCT image stimulation subsequence, a phase signal of the individual pixel in the target layer of the OCT image is determined based on the second complex-valued OCT signal of the individual pixel. In various embodiments, similarly, for each OCT image in the OCT image pre-stimulation subsequence, determining (at 208) the phase signal of the individual pixel in the target layer of the OCT image may further include: correcting a phase offset of the first complex-valued OCT signal of the individual pixel based on the pre-stimulation subsequence in the OCT image sequence to obtain the second complex-valued OCT signal of the individual pixel; and determining the phase signal of the individual pixel in the target layer of the OCT image in the same or similar manner as described above for each OCT image in the OCT image stimulation subsequence.
[0074] In various embodiments, the phase signal of an individual pixel in a target layer of the OCT image is determined relative to a reference pixel region of a reference layer of the OCT image, wherein the reference pixel region includes a plurality of individual pixels. To this end, in order to correct for systematic phase drift in the complex-valued OCT signal of the individual pixel, the multiplication of the complex-valued OCT signal of the individual pixel includes multiplying the complex-valued OCT signal of the individual pixel by the complex conjugate of the complex-valued OCT signals of the plurality of individual pixels in the reference pixel region of the reference layer of the OCT image, to obtain a set of first complex-valued OCT signals for the individual pixels. In addition, for the above-mentioned correction of the phase offset of the complex-valued OCT signal of the individual pixel, the above-mentioned determination of the average value of the first complex-valued OCT signal corresponding to the individual pixel includes: determining the average value of the set of first complex-valued OCT signals corresponding to the individual pixels of the OCT image pre-stimulation subsequence corresponding to the individual pixels of the OCT image to obtain a set of averaged pre-stimulation complex-valued OCT signals; and the above-mentioned multiplication of the first complex-valued OCT signal of the individual pixel includes: multiplying the set of first complex-valued OCT signals of the individual pixel with the set of averaged pre-stimulation complex-valued OCT signals to obtain a set of second complex-valued OCT signals of the individual pixel.
[0075] In various embodiments, when determining the phase signals of individual pixels in the target layer of the OCT image relative to a reference pixel region of the reference layer of the OCT image, for each OCT image in the stimulation OCT image subsequence, determining (at 208) the phase signals of the individual pixels in the target layer of the OCT image further includes: determining an average value of a set of second complex-valued OCT signals of the individual pixels to obtain an average complex-valued OCT signal of the individual pixels; and determining the phase signals of the individual pixels in the target layer of the OCT image based on the average complex-valued OCT signal of the individual pixels. In various embodiments, similarly, for each OCT image in the OCT image and stimulation subsequence, determining (at 208) the phase signals of the individual pixels in the target layer of the OCT image further includes: determining an average value of a set of second complex-valued OCT signals of the individual pixels to obtain an average complex-valued OCT signal of the individual pixels; and determining the phase signals of the individual pixels in the target layer of the OCT image based on the average complex-valued OCT signal of the individual pixels.
[0076] In various embodiments, the temporal phase signal of the sequence of corresponding individual pixels in the target layer across the OCT image sequence is obtained based on the phase signal of the sequence of corresponding individual pixels in the target layer across the OCT image sequence.
[0077] In various embodiments, a trained machine learning model is used to classify the temporal phase signal into one of a plurality of outer retinal tissue types or an abnormality type corresponding to a plurality of outer retinal bands of a target layer.
[0078] In various embodiments, the reference layer corresponds to the inner segment / outer segment (IS / OS) junction, and the plurality of outer retinal bands include two or more outer retinal bands corresponding to two or more of the photoreceptor cell outer segments, microvilli, retinal pigment epithelium (RPE), and Bruch's membrane (BrM), respectively.
[0079] In various embodiments, the method 200 further includes: performing band-stop filtering and low-pass filtering on the time phase signal before projecting the time phase signal to the feature point in the feature space; and normalizing the time phase signal.
[0080] Figure 3 FIG. 3 shows an example block diagram of a system 300 for training a machine learning model according to an embodiment of the present invention, wherein the machine learning model is used to classify a temporal phase signal obtained from an OCT image generated for optical retinal imaging. Figure 1The machine learning model training method 100 described above according to an embodiment of the present invention. The system 300 includes: at least one memory 302; and at least one processor 304, which is communicatively coupled to the at least one memory 302 and configured to execute the machine learning model training method 100 described above according to an embodiment of the present invention. Therefore, the at least one processor 304 is configured to: acquire an OCT image sequence generated by a phase-sensitive OCT system performing optical retinal imaging on the outer retina, the OCT image sequence including an OCT image stimulus subsequence generated in response to a visual stimulus applied to the outer retina; determine, for each OCT image in the OCT image sequence and for each of a plurality of individual pixels in a target layer of the OCT image: a phase signal of the individual pixel in the target layer of the OCT image relative to a reference layer of the OCT image, so as to obtain a temporal phase signal of the individual pixel sequence for each of a plurality of corresponding individual pixel sequences in the target layer across the OCT image sequence, and generate a temporal phase signal set of the plurality of corresponding individual pixel sequences, the target layer including a plurality of outer retinal bands; project a plurality of temporal phase signals in the temporal phase signal set to a plurality of feature points in a feature space; label the temporal phase signals in the feature space; and train the machine learning model based on the plurality of labeled temporal phase signals in the feature space.
[0081] Those skilled in the art will appreciate that the at least one processor 304 may be configured to perform various functions or operations through an instruction set (e.g., a software module) executable by the at least one processor 304 to perform various functions or operations. Figure 3As shown, the system 300 may include: an OCT image module (or an OCT image circuit) 306, which is configured to acquire an OCT image sequence generated by a phase-sensitive OCT system for optical retinal imaging of the outer retina, wherein the OCT image sequence includes an OCT image stimulus subsequence generated in response to a visual stimulus applied to the outer retina; a phase signal determination (or extraction) module (or a phase signal determination (extraction) circuit 308, which is configured to: for each OCT image in the OCT image sequence and for each individual pixel in a target layer of the OCT image, determine: the phase of the individual pixel in the target layer of the OCT image relative to the reference layer of the OCT image phase signal to obtain a time phase signal of the individual pixel sequence for each corresponding individual pixel sequence in the target layer across the OCT image sequence, and generate a time phase signal set of the multiple corresponding individual pixel sequences, wherein the target layer includes multiple outer retinal bands; a phase signal projection module (or phase signal projection circuit) 310 is configured to project multiple time phase signals in the time phase signal set to multiple feature points in the feature space; a marking module (or marking circuit) 312 is configured to mark each of the multiple feature points in the feature space; and a training module (or training circuit) 314 is configured to train the machine learning model based on the multiple marked feature points in the feature space.
[0082] Those skilled in the art will understand that the above modules of the system 300 are not necessarily independent modules, and two or more modules may be implemented or implemented as one functional module (e.g., circuit or software program) as needed or appropriate without departing from the scope of the present invention. For example, two or more of the OCT image module 306, the phase signal determination module 308, the phase signal projection module 310, the labeling module 312, and the training module 314 may be implemented (e.g., compiled together) as one executable software program (e.g., a software application or simply "app"), for example, the executable software program may be stored in the at least one memory 302 and may be executed by the at least one processor 304 to perform the corresponding functions or operations described herein according to the embodiments of the present invention.
[0083] In various embodiments, the system 300 for training a machine learning model corresponds to the system described above with reference to Figure 1The machine learning model training method 100 described herein, therefore, the various operations, functions, or steps configured to be performed by at least one processor 304 may correspond to the various operations, functions, or steps of the method 100 described above according to multiple embodiments, and for the sake of clarity and brevity, no further description is required regarding the system 300. In other words, the multiple embodiments described herein in the context of a method (e.g., method 100) are similarly applicable to the corresponding system or device (e.g., system 300), and vice versa. For example, in multiple embodiments, at least one memory 302 may store an OCT image module 306; a phase signal determination module 308; a phase signal projection module 310; a labeling module 312 and / or a training module 314, which respectively correspond to the various operations, functions, or steps of the method 100 described above according to multiple embodiments, and may be executed by the at least one processor 304 to perform the corresponding operations, functions, or steps as described herein.
[0084] Figure 4 An example block diagram of a system 400 for classifying temporal phase signals obtained from OCT images generated for optical retinal imaging using a machine learning model trained according to method 100 or trained by system 300 according to an embodiment of the present invention is shown. The system 400 corresponds to the system 400 described above with reference to FIG. Figure 2 The method 200 for classifying a temporal phase signal according to an embodiment of the present invention is described. The system 400 includes: at least one memory 402; and at least one processor 404, the at least one processor being communicatively coupled to the at least one memory 402 and configured to execute the method 200 for classifying a temporal phase signal according to an embodiment of the present invention. Therefore, the at least one processor 404 is configured to: acquire an OCT image sequence generated by a phase-sensitive OCT system performing optical retinal imaging of the outer retina, the OCT image sequence comprising an OCT image stimulus subsequence generated in response to a visual stimulus applied to the outer retina; determine, for each individual pixel in each corresponding individual pixel sequence in a target layer across the OCT image sequence, a phase signal of the individual pixel in the target layer of the OCT image relative to a reference layer of the OCT image to obtain a temporal phase signal of the corresponding individual pixel sequence, the target layer comprising a plurality of outer retinal bands; project the temporal phase signal of the corresponding individual pixel sequence onto a feature point in a feature space; and classify the temporal phase signal using a trained machine learning model based on the feature point to which the temporal phase signal is projected.
[0085] Similarly, those skilled in the art will appreciate that the at least one processor 404 may be configured to perform various functions or operations through a set of instructions (e.g., software modules) executable by the at least one processor 404 to perform various functions or operations. Figure 4 As shown, the system 400 may include: an OCT image module (or OCT image circuit) 406, configured to acquire an OCT image sequence generated by a phase-sensitive OCT system for optical retinal imaging for the outer retina, including an OCT image stimulus subsequence generated in response to a visual stimulus applied to the outer retina; a phase signal determination (or extraction) module (or phase signal determination (extraction) circuit) 408, configured to: for each individual pixel of each corresponding individual pixel sequence in the target layer across the OCT image sequence, determine: a phase signal of the individual pixel in the target layer of the OCT image relative to a reference layer of the OCT image to obtain a temporal phase signal of the corresponding individual pixel sequence, wherein the target layer includes a plurality of outer retinal bands; a phase signal projection module (or phase signal projection circuit) 410, configured to project the temporal phase signal of the corresponding individual pixel sequence to a feature point in a feature space; a classification module (or classification circuit) 412, configured to classify the temporal phase signal using a trained machine learning model based on the feature point to which the temporal phase signal is projected.
[0086] Similarly, those skilled in the art will understand that the above modules of the system 400 are not necessarily separate modules, and that two or more modules may be implemented or implemented as one functional module (e.g., a circuit or software program) as needed or appropriate without departing from the scope of the present invention. For example, two or more of the OCT image module 406, the phase signal determination module 408, the phase signal projection module 410, and the classification module 412 may be implemented (e.g., compiled together) as one executable software program (e.g., a software application or simply "app"), for example, which may be stored in the at least one memory 402 and executed by the at least one processor 404 to perform the corresponding functions or operations described herein according to embodiments of the present invention.
[0087] In various embodiments, the system 400 for classifying a time phase signal corresponds to the system described above with reference to Figure 2The method 400 for classifying a time phase signal, as described above, may include various operations, functions, or steps configured to be performed by the at least one processor 304, which may correspond to various operations, functions, or steps of the method 200 described above according to various embodiments. For the sake of clarity and brevity, no further description is required regarding the system 400. In other words, various embodiments described herein in the context of a method (e.g., method 200) may also be similarly applicable to corresponding systems or devices (e.g., system 300), and vice versa. For example, in various embodiments, the at least one memory 402 may store an OCT image module 406; a phase signal determination module 408; a phase signal projection module 410, and / or a classification module 412, each of which may correspond to various operations, functions, or steps of the method 200 described above according to various embodiments, and may be executed by the at least one processor 404 to perform the corresponding operations, functions, or steps as described herein.
[0088] In various embodiments, the systems 300 and 400 may be implemented or integrated into one system (e.g., the OCT image module 406, the phase signal determination module 408, and the phase signal projection module 410 may be the same as the OCT image module 306, the phase signal determination module 308, and the phase signal projection module 310, respectively). Thus, the integrated / combined system may include the OCT image modules 306 / 406, the phase signal determination modules 308 / 408, the phase signal projection modules 310 / 410, the labeling module 312, the training module 314, and the classification module 412.
[0089] According to various embodiments of the present invention, a computing system, a controller, a microcontroller, or any other system providing processing capabilities may be provided. Such a system may be considered to include one or more processors and one or more computer-readable storage media. For example, the systems 300 and 400 described above may include at least one processor (or controller) and at least one computer-readable storage medium (or memory), such as for executing the various processes performed in the system. The memory or computer-readable storage medium used in various embodiments may be a volatile memory, such as a dynamic random access memory (DRAM) or a non-volatile memory, such as a programmable read-only memory (PROM), an erasable PROM (EPROM), an electrically erasable PROM (EEPROM), or a flash memory, such as a floating gate memory, a charge trap memory, a magnetoresistive random access memory (MRAM), or a phase change random access memory (PCRAM).
[0090] In a number of embodiments, a "circuit" may be understood as any type of logic implementation entity, which may be a dedicated circuit or a processor that executes software stored in a memory, firmware, or any combination thereof. Thus, in an embodiment, a "circuit" may be a hardwired logic circuit or a programmable logic circuit, such as a programmable processor, for example, a microprocessor (e.g., a complex instruction set computer (CISC) processor or a reduced instruction set computer (RISC) processor). A "circuit" may also be a processor that executes software (e.g., any type of computer program, for example, a computer program using virtual machine code (e.g., Java)). Any other type of implementation of various functions or operations according to various other embodiments may also be understood as a "circuit". Similarly, a "module" may be part of a system according to an embodiment of the present invention, and may include a "circuit" as described above, or may be understood as any type of logic implementation entity therein.
[0091] Some portions of this disclosure are presented, explicitly or implicitly, in terms of algorithms and functional or symbolic representations of operations on data within a computer memory. These algorithmic descriptions and functional or symbolic representations are the means used by those skilled in the art of data processing to most effectively convey the substance of their work to others skilled in the art. An algorithm is generally considered here to be a self-consistent sequence of steps leading to a desired result. These steps require physical manipulations of physical quantities, such as electrical, magnetic, or optical signals capable of being stored, transferred, combined, compared, and otherwise manipulated.
[0092] This specification also discloses systems (e.g., which may also be embodied as one or more devices or apparatuses) for performing the various operations, functions, or steps of the various methods described herein, such as system 300 and system 400. Such a system may be specially designed for the desired purpose, or it may comprise a general-purpose computer system that is selectively activated or reconfigured by a computer program stored in the computer system. In general, the various algorithms presented herein are not limited to being implemented or executed by any particular computer system. In addition, more specialized computer systems may be provided as needed or appropriate to perform the various operations, functions, or steps of the various methods described herein, and these embodiments do not depart from the scope of the present invention.
[0093] In addition, this specification also discloses computer programs or software / functional modules at least implicitly, because it is obvious to those skilled in the art that the various operations, functions or steps of the various methods described herein can be implemented by computer code. The computer program is not limited to any particular programming language and its implementation, and it will be appreciated by those skilled in the art that various programming languages and their encodings can be used to implement the computer program. In addition, the computer program is not intended to be limited to any particular control flow, because there are a variety of programming languages that can use different control flows. It will be appreciated by those skilled in the art that the computer program can be stored on any computer-readable storage medium (non-transitory computer-readable storage medium), such as, but not limited to, a disk, an optical disk, or a memory chip. For example, a computer program stored on a computer-readable storage medium can be loaded and executed on a computer system to implement the various operations, functions or steps of the various methods described herein according to embodiments of the present invention.
[0094] Thus, in various embodiments, a computer program product is provided, which is embodied in one or more computer-readable storage media (non-transitory computer-readable storage media), including instructions (e.g., OCT image module 306; phase signal determination module 308; phase signal projection module 310; labeling module 312 and / or training module 314), which are executable by one or more computer processors to implement the above-referenced methods according to embodiments of the present invention. Figure 1The machine learning model training method 100. Therefore, the various computer programs or software modules described herein can be stored in a system (such as Figure 3 300) for execution by at least one processor 304 of the system 300 to implement various operations, functions, or steps of various methods described herein according to embodiments of the present invention. Similarly, in various embodiments, a computer program product is provided, which is embodied in one or more computer-readable storage media (non-transitory computer-readable storage media), including instructions (e.g., OCT image module 406; phase signal determination module 408; phase signal projection module 410 and / or classification module 412), which can be executed by one or more computer processors to implement the methods described above with reference to the embodiments of the present invention. Figure 2 Thus, the various computer programs or software modules described herein may be stored in a system (e.g., Figure 4 In a computer program product received by the system 400 shown in FIG. 4 , the computer program product is provided for execution by at least one processor 404 of the system 400 to perform various operations, functions, or steps of various methods described herein according to embodiments of the present invention. In addition, in various embodiments, a computer program product may be provided, which is embodied in one or more computer-readable storage media (non-transitory computer-readable storage media), including instructions (e.g., OCT image module 306 / 406; phase signal determination module 308 / 408; phase signal projection module 310 / 410; labeling module 312, training module 314, and classification module 412), and the instructions may be executed by one or more computer processors to perform the machine learning model training method 100 and the method 200 for classifying time phase signals according to embodiments of the present invention.
[0095] Those skilled in the art will appreciate that the various modules of the system 300 described herein (e.g., OCT image module 306; phase signal determination module 308; phase signal projection module 310; labeling module 312, and / or training module 314) may be software modules implemented by computer programs or instruction sets that can be implemented or executed by a computer processor to perform various functions or operations. The various modules described herein (e.g., OCT image module 306; phase signal determination module 308; phase signal projection module 310; labeling module 312, and / or training module 314) may also be implemented as hardware modules that are designed to be functional hardware units that perform various functions or operations. Similarly, those skilled in the art will appreciate that the various modules of the system 400 described herein (e.g., OCT image module 406; phase signal determination module 408; phase signal projection module 410, and / or classification module 412) may be software modules implemented by computer programs or instruction sets that can be implemented or executed by a computer processor to perform various functions or operations. The various modules described herein (e.g., OCT image module 406; phase signal determination module 408; phase signal projection module 410 and / or classification module 412) may also be implemented as hardware modules, which are functional hardware units designed to perform various functions or operations. More specifically, in a hardware sense, a module refers to a functional hardware unit designed to be used in conjunction with other components or modules. For example, a module may be implemented using discrete electronic components, or a module may form part of an entire electronic circuit such as an application-specific integrated circuit (ASIC). Many other possibilities exist. Those skilled in the art will also appreciate that a combination of hardware and software modules may also be implemented. In addition, the various operations, functions, or steps of the various methods described herein may be performed in parallel as needed or appropriate (e.g., as long as this does not cause the method to fail to operate or to meet its intended purpose), rather than sequentially.
[0096] In various embodiments, the systems 300 and 400 may be implemented by any computer system (e.g., a desktop or portable computer system) including at least one processor and at least one memory, such as a computer system. Figure 5The example computer system 500 schematically shown in the figure is for example only and not for limitation. In multiple embodiments, as described above, system 300 and system 400 can be implemented or integrated into a computer system. Various methods / steps or functional modules can be implemented as software, such as a computer program executed within computer system 500, and instruct computer system 500 (particularly one or more processors therein) to perform various functions or operations as described herein according to various embodiments. For example, computer system 500 may include a system unit 502, one or more input devices 504 (such as a keyboard, touch screen and / or mouse), and multiple output devices (e.g., display 508). System unit 502 can be connected to a computer network 512 via a suitable transceiver device 514 to enable access to, for example, the Internet or other network systems, such as a local area network (LAN) or a wide area network (WAN). System unit 502 may include a processor 518 for executing various instructions, a random access memory (RAM) 520, and a read-only memory (ROM) 522. The system unit 502 may also include a number of input / output (I / O) interfaces, such as an I / O interface 524 for connecting to the display device 508 and an I / O interface 526 for connecting to one or more input devices 504. The components of the system unit 502 typically communicate via an interconnecting bus 528 in a manner known to those skilled in the art.
[0097] Those skilled in the art will appreciate that the terms used herein are only used to describe various embodiments and are not intended to limit the present invention. As used herein, the singular forms "a," "an," and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. It should be further understood that the terms "comprise" and / or "comprising" when used in this specification require the presence of stated features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or combinations thereof.
[0098] Unless otherwise specified or the context requires otherwise, any reference herein to an element or feature using designations such as "first," "second," etc. does not limit the quantity or order of those elements or features. For example, such designations may be used herein as a convenient way to distinguish between two or more elements or different instances of the same element. Thus, a reference to a first element and a second component does not necessarily mean that only two components may be employed, or that the first component must precede the second component, unless otherwise specified or the context requires otherwise. Furthermore, a phrase referring to "at least one" of a list of items refers to any single item in the list of items or any combination of two or more items in the list of items.
[0099] In order to make the present invention easy to understand and put into practice, the following description of the embodiments of the present invention is provided by way of example only and not by way of limitation. However, it will be appreciated by those skilled in the art that the present invention may be implemented in a variety of different forms or configurations and should not be construed as being limited to the exemplary embodiments described below. Rather, these exemplary embodiments are provided to make this disclosure comprehensive and complete and to fully convey the scope of the present invention to those skilled in the art.
[0100] Several exemplary embodiments provide a method for automated, unbiased optical retinal imaging with spatial heterogeneity suitable for preclinical ocular imaging studies. To this end, according to exemplary embodiments, extended and multi-layered optical retinal imaging (specifically, imaging the functional activity of multiple outer retinal strips) is used to reveal light-induced deformations in rod photoreceptors, the pigment epithelium, and the subretinal space. As an illustrative example, wide-field automated, unbiased optical retinal imaging techniques according to various embodiments reveal the comprehensive nanoscale dynamics (nanoscale tissue dynamics) of the outer retina in rodents.
[0101] As explained in the background, conventional methods for optical retinal imaging using optical coherence tomography (OCT) suffer from various deficiencies or shortcomings. In particular, due to limited spatial resolution, conventional methods are unable to resolve signals scattered by individual cells in different outer retinal tissue layers in response to visual stimulation, resulting in a blurred speckle pattern in the optical retinal imaging OCT images. Consequently, this blurred speckle pattern (which may be referred to as a speckle layer) in the OCT image negatively impacts the accuracy (e.g., resolution) of imaging or detecting the dynamics (functional responses, such as movement or deformation) of the outer retinal tissue layers in response to visual stimulation. Therefore, various embodiments advantageously address technical issues associated with these speckle patterns in OCT images. More specifically, a method for optical retinal imaging using OCT is provided that is capable of resolving signals from different outer retinal wavelength bands, thereby enhancing / improving the accuracy (e.g., resolution) of imaging or detecting the dynamics of the outer retinal tissue layers in response to visual stimulation, enabling, for example, detection of tissue dynamics at the nanometer level. To this end, multiple embodiments of the present invention provide a method and system for training a machine learning model for classifying a temporal phase signal obtained from an OCT image generated for optical retinal imaging, and a method and system for classifying a temporal phase signal obtained from an OCT image generated for optical retinal imaging using a trained machine learning model.
[0102] Phototransduction involves changes in the concentration of ions and other solutes within photoreceptor cells and in the subretinal space, which affects osmotic pressure and associated water flow. Based on phase-sensitive OCT, optical retinal imaging techniques can be used to capture the corresponding expansion and contraction of cell layers. To date, optical retinal imaging techniques have only been able to reliably detect photoisomerism and phototransduction in photoreceptor cells, primarily in cones under high-intensity light stimulation. By employing sub-pixel body motion correction methods or algorithms that can image nanoscale tissue dynamics during minute-long recordings, as well as unsupervised learning of spatiotemporal patterns, various exemplary embodiments have discovered optical signatures of the responses of other retinal structures to visual stimuli. These outer retinal structures typically include the inner and outer segments of rod photoreceptors, the retinal pigment epithelium (RPE), and the subretinal space. The high sensitivity of the methods for performing optical retinal imaging according to embodiments of the present invention enables detection of retinal responses to very dim stimuli, for example, down to 0.01% bleaching levels, equivalent to natural scotopic light levels. Various embodiments also demonstrate that, using a single flash, a photoretinogram can map retinal responses (e.g., across a 12° field of view), potentially replacing multifocal electroretinograms that have long acquisition times and low spatial resolution. Therefore, the method of photoretinography performed according to embodiments of the present invention expands the diagnostic capabilities and practical applicability of photoretinography, while combining retinal structural imaging with functional imaging in the same OCT machine or system, providing a more comprehensive alternative to electroretinograms. In particular, various embodiments provide a robust unsupervised learning method for discovering hidden spatiotemporal patterns in phase signals measured from speckle in OCT images, revealing, for example, different features of the subretinal space (SRS), photoreceptor cell inner segments (IS) and outer segments (OS), and RPE responses to light (visual stimuli). For example, image registration using a phase recovery subpixel motion correction method or algorithm according to an embodiment can achieve nanometer-scale photoretinography recordings within tens of seconds, thereby enabling the detection of slower and more subtle phenomena in the retinal response to visual stimuli. Various example embodiments validate such light-evoked responses under various scotopic and photopic conditions with a single flash, and achieve wide-field mapping of OS and SRS dynamics.
[0103] For example, small animals such as rodents have become advantageous models for studying the characteristics of photoreceptor cell degeneration due to their wide availability and versatility in disease models and genetic manipulation. This advantage of small animal models can also be applied to optical retinal imaging technology, an emerging optical imaging modality that can measure photoreceptor cell function in vivo. However, due to the extreme difficulty in distinguishing individual photoreceptor cells in small animals such as rodents, there is a lack of objective quantitative methods that can measure their outer retinal dynamics while maintaining spatial heterogeneity. Such challenges have hindered in-depth research on rod photoreceptor cell optical retinal imaging signals and their correlation with early photoreceptor cell degeneration. To this end, as illustrative examples, various exemplary embodiments demonstrate automatic unbiased optical retinal imaging technology for rodent imaging based on unsupervised machine learning. The method according to the embodiment can automatically search and classify optical retinal imaging (ORG) signals (or more specifically, temporal phase signals, which can also be referred to as phase trajectory signals or simply phase trajectory) from the blurred speckle pattern of the outer retina captured by a low-cost phase-sensitive OCT system. Using this method, various embodiments observed reproducible, comprehensive nanoscale dynamics of the outer retina in response to visual stimuli in a rat model. Various exemplary embodiments further classified or categorized these light retinal imaging signals into Type-I and Type-II signals, which were found to be associated with different parts / layers of the outer retina. Various embodiments also characterized the light-induced light retinal imaging response of the outer retina under dark and bright conditions, providing a new perspective for understanding the physiological origin of light retinal imaging signals. In addition, various exemplary embodiments demonstrated frontal mapping imaging of light retinal imaging signals in a wide field of view of 12°, similar to multifocal electroretinography, but with higher resolution, which can reveal the spatial distribution of outer retinal function. The method according to the exemplary embodiment can be widely used to study tissue-specific dynamics in various animal models, and also has great potential in studying visual sensory pathways.
[0104] Thus, as an illustrative example, various embodiments provide an automated, unbiased method for studying the dynamics of the outer retina of rodents in response to visual stimulation. In various embodiments, the method uses a phase-sensitive spectral domain OCT system to analyze the integrated phase signals acquired from all pixels in a speckle pattern in an OCT image (e.g., the speckle layer, which may also be referred to herein as the target layer). Various exemplary embodiments further employ a phase-recovery sub-pixel motion correction method to correct for bulk tissue motion that reduces the accuracy of phase measurements. The temporal phase signals at individual pixels can be projected into a feature space (e.g., principal component (PC) space) and then automatically classified using hierarchical clustering. For example, to ensure consistency between data from different animals, a support vector machine (SVM) is trained using labels obtained from the hierarchical clustering analysis in the established principal component (PC) space. All subsequent analyses across different animals can be performed using the trained SVM. The reproducibility of the method has been verified, and the nanoscale dynamics of the outer retina under scotopic and photopic visual stimulation have been studied.
[0105] Therefore, it is necessary to automatically extract and classify the temporal phase signals extracted from each pixel of the OCT image to detect micrometer-scale or nanometer-scale movements of various different outer retinal tissue types in response to visual stimuli. To this end, multiple embodiments provide a method for performing optical retinal imaging using OCT, which automatically extracts the dynamics of the outer retinal tissue in response to visual stimuli. In multiple embodiments, the method uses unsupervised machine learning techniques to cluster temporal optical retinal imaging signals (or more specifically, temporal phase signals) in a feature space (e.g., time or spatiotemporal feature space) to distinguish responses from photoreceptor cells from responses from other outer retinal tissues (such as microvilli, RPE, and BrM). The method does not require manual intervention or prior information on tissue distribution, and the method according to the embodiment can be extended to capture heterogeneous tissue responses triggered by other external stimuli (such as electrical stimulation and photothermal stimulation). In addition, in multiple embodiments, the obtained labels (labeled feature points) are used to train a machine learning model (e.g., a classifier) so that new temporal phase signals can be classified based on the same criteria.
[0106] As previously mentioned, due to the limited optical resolution provided by OCT and the influence of eye aberrations, the outer retina typically exhibits a speckle pattern in OCT images. To reveal the dynamics of photoreceptor cells, some previous studies have used adaptive optics technology to resolve individual photoreceptor cells. However, multiple embodiments point out that the field of view of adaptive optics technology is small, limiting its potential applicability in clinical practice. Another conventional strategy is to perform an overall average of the signals from a selected region of interest, which may mix heterogeneous responses from different retinal tissue types. In contrast, the method of performing optical retinal imaging using OCT according to the embodiment is able to reveal comprehensive nanoscale dynamics from the speckle pattern using a machine learning-assisted process.
[0107] A method for training a machine learning model for classifying temporal phase signals obtained from OCT images generated for optical retinal imaging (e.g., corresponding to the above-referenced method) will now be described according to an embodiment of the present invention. Figure 1 In particular, according to various embodiments, a phase response (a temporal phase signal, also referred to as a phase trajectory) is extracted from the outer retina. As an illustrative example, 6A to 6I An example image and signal processing flow / framework for a method for training a machine learning model according to an embodiment of the present invention is shown. Thus, various embodiments provide a method for automatically extracting temporal optical retinal imaging signals (or more specifically, temporal phase signals) from phase-sensitive OCT, preprocessing the extracted temporal phase signals, feature extraction, unsupervised clustering, and training a machine learning model (e.g., a classifier).
[0108] In various embodiments, the raw interference fringes in the OCT image can first undergo standard OCT post-processing steps, including spectral calibration, k-linearity, dispersion compensation, and discrete Fourier transform (DFT), to obtain a complex-valued OCT image. During this process, the DFT of the processed spectral signal is used to reconstruct a depth-resolved complex-valued OCT image.
[0109] The process of processing the phase signals of an OCT image and its pixels according to an embodiment of the present invention will now be described.
[0110] To suppress the effects of tissue motion on phase-sensitive measurements, a single-step DFT algorithm can be used to estimate the subpixel translational displacement between repeated cross-sectional scans (B-scans) at the same location. Complex-valued OCT images can then be registered using a phase-recovery subpixel motion correction method that corrects for lateral and axial displacements by multiplying with an exponential term in the spatial frequency domain and spectral domain, respectively.
[0111] In various embodiments, in phase-sensitive light retinal imaging measurements, various embodiments determine or calculate temporal phase changes (also referred to as temporal phase differences) of a target layer relative to a reference layer. In phase-sensitive light retinal imaging measurements, light-induced dynamics of the outer retina can be assessed by calculating the temporal phase difference between two highly reflective outer retinal bands, i.e., one from a reference layer (e.g., the IS / OS junction) and the other from a target layer (which may be referred to as a composite layer or mixed layer), the target layer comprising multiple outer retinal bands, such as outer retinal bands corresponding to photoreceptor cell outer segments, microvilli, retinal pigment epithelium (RPE), and / or Bruch's membrane (BrM). Thus, the dynamics of the outer retina can be assessed by calculating / extracting temporal optical length (OPL) changes of the outer retina using the IS / OS as a reference.
[0112] In various embodiments, the reference layer and the target layer in the obtained OCT image can be automatically segmented using an automatic image segmentation algorithm, such as an automatic image segmentation algorithm based on graph theory and dynamic programming, or manual segmentation can also be used. In various embodiments, in order to improve the signal-to-noise ratio of motion detection, for each pixel in the target layer, a reference pixel or a reference pixel area can be selected from the reference layer (e.g., the IS / OS junction), and the reference pixel area is centered on the same A-scan line as the pixel in the target layer. For example, each reference pixel area may include 5 adjacent A-lines (e.g., approximately 7.1 μm in the horizontal direction (the total width of 5 adjacent A-lines)).
[0113] In the illustrative example, Figure 6A An example of time-sequential cross-sectional OCT images (OCT image sequence) acquired from a 12° field of view of a rat retina is shown (more specifically, a sequence of repeated cross-sectional scans (or B-scans) acquired from the same location). In the outer retina, one highly reflective layer band corresponds to the IS / OS junction, while another highly reflective layer band is a composite layer (also referred to herein as a speckle layer or target layer) including, for example, the outer segment tips of photoreceptor cells, microvilli, RPE, and BrM. Figure 6A As shown, the OCT structural images are in good agreement with the histological sections (inset) ( Figure 6A , the horizontal and vertical scale bars shown correspond to 100 μm). Therefore, an OCT image sequence generated by a phase-sensitive OCT system performing optical retinal imaging of the outer retina can be obtained from the same position in the outer retina, the sequence including the images in response to a visual stimulus applied to the outer retina (e.g., Figure 6AEach OCT image in the OCT image sequence has a reference layer and a target layer (including a plurality of outer retinal bands). In this regard, as described above, the same reference layer and the same target layer in each OCT image can be automatically segmented using an automatic image segmentation algorithm or can be manually segmented.
[0114] For each OCT image in the OCT image sequence and for each of a plurality of individual pixels in the target layer of the OCT image, a phase signal of the individual pixel in the target layer of the OCT image relative to a reference layer of the OCT image can be determined to obtain a temporal phase signal for each of a plurality of corresponding individual pixel sequences in the target layer across the OCT image sequence, thereby generating a set of temporal phase signals for the plurality of corresponding individual pixel sequences. In various embodiments, the plurality of corresponding individual pixel sequences cover all (or substantially all) pixels in the target layer across the OCT image sequence (i.e., each pixel of the target layer in each OCT image). In various embodiments, the corresponding individual pixel sequence in the target layer across the OCT image sequence refers to a sequence of corresponding pixels in the target layer of each OCT image in the OCT image sequence (e.g., corresponding to the same portion of the target layer across the OCT image sequence). Various embodiments correct or eliminate systematic phase drift by self-referencing and eliminate any phase offset of independent phase traces (individual temporal phase signals) by referencing their pre-stimulus frames.
[0115] In various embodiments, determining the phase signal of an individual pixel in the target layer of the OCT image includes correcting a systematic phase drift of the complex-valued OCT signal of the individual pixel based on self-referencing of the individual pixel with respect to a reference layer of the OCT image to obtain a first complex-valued OCT signal of the individual pixel (which may be referred to as a paired self-referenced complex-valued OCT signal). In this regard, in various embodiments, the systematic phase drift may be corrected (e.g., eliminated) by calculating the product of the complex-valued OCT signal of the individual pixel (pixel of interest) in the target layer of the OCT image and the complex conjugate of the complex-valued OCT signal of the individual pixel in the reference layer of the OCT image, which may be defined by the following example expression:
[0116]
[0117] in represents the complex-valued OCT signal of the pixel of interest in the target layer, i represents the frame number index, and Represents the complex-valued OCT signal of an individual pixel in the reference layer. represents a pair of self-referenced complex-valued OCT signals of a pixel of interest in the target layer (corresponding to the first complex-valued OCT signal described above). In equation (1), '*' represents a complex conjugate.
[0118] In various embodiments, to eliminate any phase offset, each complex-valued phase trace (each independent phase signal) can be referenced to its pre-stimulus frame. To this end, in various embodiments, for each OCT image in a stimulation subsequence of OCT images, determining the phase signal of an individual pixel in the target layer of the OCT image further includes correcting the phase offset of a first (pair-wise self-referenced) complex-valued OCT signal of the individual pixel based on a pre-stimulus subsequence generated before applying a visual stimulus to the outer retina in the OCT image sequence to obtain a second complex-valued OCT signal of the individual pixel (which may be referred to as a time-referenced complex-valued OCT signal). In this regard, in various embodiments, the phase offset of the first complex-valued OCT signal of the individual pixel can be corrected (e.g., eliminated) by determining the average of the first complex-valued OCT signals of each corresponding individual pixel in the pre-stimulus subsequence of the OCT images to obtain an average pre-stimulus complex-valued OCT signal; and multiplying the first complex-valued OCT signal of the individual pixel by the average pre-stimulus complex-valued OCT signal to obtain the second complex-valued OCT signal of the individual pixel. Similarly, for each OCT image in the pre-stimulation subsequence of OCT images, determining the phase signal of the individual pixel in the target layer of the OCT image may further include: correcting the phase offset of the first complex-valued OCT signal of the individual pixel based on the pre-stimulation subsequence in the OCT image sequence to obtain a second complex-valued OCT signal of the individual pixel; and determining the phase signal of the individual pixel in the target layer of the OCT image in a manner identical or similar to that described above for each OCT image in the stimulation subsequence of OCT images. By way of example only and not limitation, the correction process for the phase offset of the first (paired self-referenced) complex-valued OCT signal of the individual pixel may be defined by the following example expression:
[0119]
[0120] in represents the time (pre-stimulation time) reference signal (corresponding to the above-mentioned second complex-valued OCT signal), and N represents the number of frames acquired before light stimulation.
[0121] In various embodiments, for each OCT image in the OCT image sequence, a phase signal of the individual pixel in the target layer of the OCT image is determined based on the second complex-valued OCT signal of the individual pixel. As an example, the phase signal (corresponding to the phase change / difference) of the individual pixel in the target layer of the OCT image can be extracted from the time reference signal according to the following example expression:
[0122]
[0123] Where Δφ(i) represents the phase signal extracted from an individual pixel in the target layer (the pixel of interest), and '∠' represents the phase angle calculation. Therefore, a temporal phase signal (also referred to as a phase trajectory) for a corresponding sequence of individual pixels in the target layer across the OCT image sequence can be obtained based on the phase signal of the corresponding sequence of individual pixels in the target layer across the OCT image sequence.
[0124] In multiple embodiments, as described above, the above-mentioned process for determining the phase signal of a pixel of interest is applied separately to each pixel in the target layer of each OCT image to extract the corresponding phase signal therefrom, thereby obtaining a set of time phase signals of multiple corresponding individual pixel sequences in the target layer across the OCT image sequence.
[0125] In various embodiments, the complex-valued OCT signals of individual pixels in the target layer can be spatially averaged over multiple pixels of the reference layer (IS / OS interface) or over multiple pixels of the target layer, or both, to improve the signal-to-noise ratio (SNR). In various embodiments, the complex-valued OCT signals are spatially averaged over multiple pixels of a reference pixel region of the reference layer, and correction (or elimination) of systematic phase drift and correction (or removal) of any phase offset can be performed as described below. That is, the phase signals of individual pixels in the target layer of the OCT image can be determined relative to the reference pixel region (including multiple individual pixels) of the reference layer of the OCT image.
[0126] In various embodiments, when phase signals of individual pixels in a target layer of an OCT image are determined relative to a reference pixel region of a reference layer of the OCT image, systematic phase drift of the complex-valued OCT signals of the individual pixels can be corrected (e.g., eliminated) by calculating the product of the complex-valued OCT signals of the individual pixels of interest in the target layer and the complex conjugate of the complex-valued OCT signals of the individual pixels in a selected reference pixel region in the reference layer, to obtain a set of first (paired self-referenced) complex-valued OCT signals of the individual pixels (each first complex-valued OCT signal is paired with a corresponding reference pixel in the reference pixel region), which can be defined by the following example expression:
[0127]
[0128] in represents the complex-valued OCT signal of the pixel of interest in the target layer, i represents the frame number index, represents the complex-valued OCT signal of an individual pixel in the corresponding reference pixel region, and s represents the pixel index in the reference region. represents a pairwise self-referenced complex-valued OCT signal of a single pixel, where the pixel of interest in the target layer ergodicly references all pixels (every pixel) in the reference pixel region. In equation (4) above and equation (5) below, '*' represents complex conjugation.
[0129] In various embodiments, to eliminate any phase offset, each independent phase signal can be referenced to its pre-stimulus frame. In this regard, in various embodiments, the average of a first (pair-wise self-referenced) set of complex-valued OCT signals corresponding to individual pixels of a pre-stimulus subsequence of OCT images corresponding to the individual pixels of the OCT image can be determined to obtain a set of averaged pre-stimulus complex-valued OCT signals; and then the first set of complex-valued OCT signals for the individual pixels can be multiplied by the set of averaged pre-stimulus complex-valued OCT signals to obtain a second (time (pre-stimulus time) referenced) set of complex-valued OCT signals for the individual pixels, which can be defined by the following example expression:
[0130]
[0131] in represents the time (pre-stimulus time) of a single pixel referenced to the complex-valued OCT signal set, and N represents the number of frames acquired before light stimulation.
[0132] The pre-stimulus time-referenced complex-valued OCT signal set can then be averaged across the pixels in the reference region, and phase information (phase signal) can be extracted from the averaged complex-valued OCT signal. In this regard, in various embodiments, an average value of the second complex-valued OCT signal set of individual pixels can be determined to obtain an average complex-valued OCT signal for the individual pixels; and a phase signal for an individual pixel in the target layer of the OCT image can be determined based on the average complex-valued OCT signal for the individual pixels, which can be defined by the following example expression:
[0133]
[0134] Where Δφ(i) represents the phase signal extracted from the individual pixel of interest in the target layer, 'S' represents the total number of pixels in the reference pixel region, and '∠' represents the phase angle calculation. Therefore, the phase signal (corresponding to the phase change / difference) of the individual pixel in the target layer of the OCT image can be extracted from the averaged complex-valued OCT signal according to Equation (7). Therefore, the temporal phase signal (also referred to as the phase trajectory) of the corresponding individual pixel sequence in the target layer across the OCT image sequence can be obtained based on the phase signal of the corresponding individual pixel sequence in the target layer across the OCT image sequence.
[0135] In various embodiments, the time phase signal may then be converted to an OPL change (ΔOPL) by:
[0136]
[0137] or
[0138]
[0139] Where λ0 is the center wavelength of the OCT system. In various embodiments, if the reference layer precedes the target layer, the OPL change can be calculated using equation (8). Otherwise, it can be calculated using equation (9). This ensures that an increase in the OPL change (ΔOPL) always indicates an expansion between the target and reference layers.
[0140] In various embodiments, as described above, the aforementioned process for determining the phase signal of a pixel of interest is applied to each pixel in the target layer of each OCT image to extract the corresponding phase signal, thereby obtaining a set of temporal phase signals corresponding to multiple individual pixel sequences in the target layer across the OCT image sequence. In various embodiments, when averaging the phase signals of different pixels, the calculation can be performed in the complex plane to avoid bias introduced by phase wrapping.
[0141] In the illustrative example, Figure 6A after, Figure 6B The upsampled cross-correlogram calculated between repeated cross-sectional scans (B-scans) is shown, where the position of its peak is used as the estimated displacement for phase recovery sub-pixel motion correction. Figure 6C Shown are the temporal intensity (upper panel) and phase (lower panel) M-scans without phase-recovery subpixel correction (scale bar: 500 ms (horizontal), 100 μm (vertical)), the estimated subpixel global motion (estimated temporal global tissue motion) by locating the peak of the upsampled cross-correlogram between repeated B-scans (scale bar: 500 ms (horizontal), 5 μm (vertical)), and the corresponding temporal intensity (upper panel) and phase (lower panel) M-scans after phase-recovery subpixel motion correction (scale bar: 500 ms (horizontal), 100 μm (vertical)) when no light stimulus is delivered to the retina. Figure 6C In FIG, white and gray arrows mark the locations of the outer retina and choroid, respectively. 6D shows an OCT structural image showing the segmentation of two hyperreflective bands in the outer retina, one hyperreflective layer corresponding to the IS / OS junction, and the other hyperreflective layer corresponding to a composite layer including the outer segment tips of photoreceptor cells, RPE, and BrM (e.g., corresponding to the target layer). Figure 6D As shown, the OCT structural images are in good agreement with the histological sections (inset) ( Figure 6D , the horizontal and vertical scale bars shown correspond to 100 μm).
[0142] In a number of embodiments, the extracted optical retinal imaging signal (more specifically, the time phase signal, also referred to as the phase trajectory) can then be bandpass filtered to minimize residual artifacts caused by heartbeat and respiration. Subsequently, a low-pass filter (e.g., with a cutoff frequency of 10 Hz) can be used to filter out high-frequency oscillations. Alternatively, the extracted time phase signal can be low-pass filtered and the remaining periodic oscillations caused by heartbeat and respiration can be suppressed by a bandpass filter. After filtering, time phase signals with high variations in the baseline can be regarded as noise and excluded from further analysis. For example, due to edge effects, the two ends of the time phase signal exhibit high variance, so only the middle segment can be selected for further analysis. For example, time phase signals with a baseline standard deviation greater than 60 mrad can be excluded. The independent time phase signals (independent signal trajectories) can then each be normalized by subtracting their mean and then dividing by their standard deviation SD to eliminate amplitude differences.
[0143] In the illustrative example, Figure 6D after, Figure 6E Example representative time phase signals before and after band-stop and low-pass filtering are shown.
[0144] According to multiple embodiments, in order to facilitate the analysis of high-dimensional time phase signals, the extracted time phase signals can be projected into a feature space (e.g., a time feature space or a spatiotemporal feature space) established by an automatic feature extraction method or manually selected features. For example, principal component analysis (PCA), self-organizing maps, or autoencoders can be used to compress the data set into a low-dimensional spatiotemporal feature space. Alternatively, representative features such as amplitude, peak time, etc. can be manually selected to construct a spatiotemporal feature space. As an illustrative example, feature extraction using PCA is now described according to multiple embodiments of the present invention. Specifically, PCA is used to compress the data set into a low-dimensional feature space to facilitate the analysis of high-dimensional time phase signals.
[0145] In various embodiments, an outlier detection method may be employed to avoid the appearance of undesirable elongated clusters in subsequent unsupervised cluster analysis. For example, a distance-based outlier detection method may be used to exclude outliers that are distributed in low-density areas in the common PC space. An example approach is that for each data point (feature point) in the feature space (common PC space), for example, a minimum radius that includes 2% of the group size may be calculated. Then, any data point with a radius greater than Q3+(Q3-Q1) / 5 may be considered (marked as) an outlier and excluded from the cluster analysis, where Q1 and Q3 are the first and third quartiles of the calculated radius.
[0146] In the illustrative example, to account for inter-subject variability, for example, the normalized temporal phase signals extracted from five rats are integrated to calculate the PC coefficients. For example, the first three PCs that capture 83.1% of the total variance are selected to form the common PC space. In contrast, assuming that each dataset is projected into its own unique PC space, the first three PCs account for an average of 83.4% (1.5%) of the variance, which is only 0.3% higher than the variance in the common PC space, indicating a high degree of consistency between the datasets. In the illustrative example, Figure 6F Filtered temporal phase signals (phase traces) extracted from five rats are shown. Temporal phase signals outside the dashed rectangle suffer from edge effects and are excluded from further analysis. Figure 6F after, Figure 6G The distribution density of stable phase responses under conditions of applied light stimulation (upper) and no applied light stimulation (lower) is shown. Figure 6G shows the distribution density of the normalized time phase signal, where Figure 6G The independent time phase signals in the upper part of the figure have been processed by subtracting their mean values and dividing them by their respective standard deviations. Thus, the processing of the time phase signals and the extraction of time series features using PCA have been described. In particular, Figure 6G The distribution density of the phase traces before the light stimulus is applied (before 0 seconds) is shown. When the flash is delivered to the retina (at t = 0 seconds), the functionally relevant phase response ( Figure 6G ), while in the absence of applied light stimulation, the phase traces gradually decorrelated without any obvious response pattern (lower panel). Figure 6H Shows the Figure 6G The normalized results of each phase trajectory in the upper part of the figure are processed by subtracting the mean value of the phase trajectory and dividing it by its respective standard deviation SD.
[0147] As an illustrative example, Figure 6I shows the high-dimensional phase trajectory projected into the temporal feature space (or more specifically, the PC space consisting of the first three PCs), where outliers have been removed using a density-based approach. Specifically, Figure 6I The distribution of normalized phase trajectories in the common PC space is shown, which includes the first three PCs and has been stripped of outliers using a density-based outlier detection method. Figure 6I In , the heat map shows the distribution density of the remaining data points (gray points) when projected onto different feature planes.
[0148] Various unsupervised clustering techniques (such as k-means clustering, hierarchical clustering, density-based spatial clustering with noisy applications, etc.) can be used to explore common patterns in feature space. As an illustrative example, unsupervised clustering using hierarchical clustering methods or algorithms is now described according to multiple embodiments of the present invention. In multiple embodiments, an agglomerative hierarchical clustering algorithm can be used in conjunction with the Ward criterion to group the points in the common PC space. This technique calculates the Euclidean distance for each pair of points and iteratively merges similar subclusters (e.g., two similar subclusters) into larger clusters. Under the guidance of the Ward method / criterion or the minimum variance method, each merge ensures that the increment of the variance within the total cluster cluster is minimized.
[0149] After generating the agglomerative hierarchical clustering tree, various clustering results can be constructed by truncating the dendrogram at different levels. There are many methods and criteria for determining the number of clusters. In many embodiments, the optimal number of clusters can be set manually or determined by metrics such as the silhouette method and the gap statistic method. As an illustrative example, various embodiments seek to find the optimal number of clusters that can maximize the overall inter-cluster dissimilarity (average dissimilarity) between the reconstructed phase trajectories. In particular, when the phase trajectories are grouped into k clusters, the filtered phase trajectories can be retrieved and averaged within each cluster to reconstruct the phase trajectories. In other words, for each clustering result, the filtered phase trajectories can be retrieved and averaged in each cluster. Then, for all n clusters, the number of filtered phase trajectories can be retrieved and averaged. c =(k-1)k / 2 reconstructed phase traces (between any two reconstructed phase traces), calculate the pairwise Pearson cross-correlation coefficient ρ, where k is the number of clusters. For example, the root mean square error of these cross-correlation coefficients relative to 1 (perfect correlation) can be calculated as follows:
[0150]
[0151] With the largest C RMS The number of clusters can be determined as the optimal number of clusters. Accordingly, in many embodiments, the maximum inter-cluster difference is sought.
[0152] Therefore, as previously described, according to various embodiments, the obtained time phase signal (phase trajectory) can be projected onto feature points in a feature space (e.g., a temporal feature space or a spatiotemporal feature space). In addition, these feature points in the feature space can be labeled. In various embodiments, labeling the feature points in the feature space can include: grouping the plurality of feature points into a plurality of clusters in the feature space (e.g., using an unsupervised clustering technique); and for each of the plurality of clusters, labeling the feature points belonging to the cluster with a label assigned to the cluster. In various embodiments, the plurality of clusters are assigned different labels. In this case, the plurality of different labels include a plurality of different outer retinal band labels corresponding to a plurality of outer retinal bands of the target layer.
[0153] According to embodiments of the present invention, training a machine learning model (e.g., a support vector machine (SVM) model) in an established / identical feature space (e.g., an established / identical PC space) is now described. To classify new time phase signals based on the same criteria, a machine learning model, such as a classifier (e.g., an SVM, a Bayesian classifier, a neural network, etc.), can be trained using labeled feature points in the same feature space. In various embodiments, each phase trajectory is projected into a PC space comprising the first three principal components. For each phase trajectory, its value in the directions of these three principal components or its coordinates in the principal component space can be referred to as a score in the PC space. Therefore, to train the machine learning model, the scores in the PC space serve as features, and the labels of the phase trajectories are obtained through the unsupervised clustering described above. The goal is to assign each feature point (data point) in the feature space (in this example, the PC space) to a cluster. Therefore, unsupervised learning methods are first used to label the feature points in the feature space, and then the labeled feature points are used to train a machine learning model (e.g., a classifier) in the same feature space.
[0154] For example, training an SVM can be achieved by using the "templateSVM" and "fitcecoc" functions in MATLAB. The input features can be standardized before training, a Gaussian kernel can be selected, and automatic hyperparameter optimization can be enabled. Subsequently, the trained classifier in the feature space can be used to classify new time phase signals (extracted from new OCT images in the same or similar manner as described above according to the embodiments). As an illustrative example, the decision boundary of each cluster cluster can be extracted by training an SVM in a common PC space using previously obtained labels (including abnormalities, Type-I signals, and Type-II signals). The trained SVM achieved a classification accuracy of 99.5% as evaluated by a 10-fold cross-validation strategy. Therefore, the trained SVM is useful for processing new data sets (new OCT image sequences) using the same criteria. In this regard, the time phase signals (phase trajectories) extracted from the new data set can be preprocessed and projected into the common PC space in the same or similar manner as described above according to multiple embodiments. For example, each new phase trajectory can be projected onto a feature space such as the PC space based on the principal component coefficients. In PC space, each feature point (data point) corresponding to a pixel in the target layer of the OCT image can be assigned to a cluster using a trained support vector machine (SVM). Type-I and Type-II signals can then be reconstructed based on the classification results from the trained SVM.
[0155] Thus, in various embodiments, a method for classifying temporal phase signals obtained from optical coherence tomography (OCT) images generated for optical retinal imaging using a trained machine learning model is provided. The method comprises: acquiring a sequence of optical coherence tomography (OCT) images generated by a phase-sensitive optical coherence tomography (OCT) system for optical imaging of the outer retina, including a stimulus subsequence of OCT images generated in response to a visual stimulus applied to the outer retina; determining, for each individual pixel in a corresponding sequence of individual pixels in a target layer across the sequence of OCT images, a phase signal of the individual pixel in the target layer of the OCT image corresponding to a reference layer of the OCT image to obtain a temporal phase signal for the corresponding sequence of individual pixels, the target layer comprising a plurality of outer retinal bands; projecting the temporal phase signals for the corresponding sequence of individual pixels onto feature points in a feature space; and classifying the temporal phase signals using the trained machine learning model based on the feature points to which the temporal phase signals are projected. In various embodiments, a sequence of OCT images can be acquired, phase signals of individual pixels in a target layer of the OCT images can be determined, and the temporal phase signals of the corresponding sequence of individual pixels can be projected onto feature points in a feature space in the same or similar manner as described above for training a machine learning model according to various embodiments. For example, a new temporal phase signal extracted from a new OCT image of the outer retina can be classified using the trained machine learning model as one of a plurality of outer retinal tissue types corresponding to a plurality of outer retinal bands of the target layer or as an outlier type.
[0156] For example, for repeated volume scans, in multiple embodiments, the same method as the repeated cross-sectional scans (B-scans) described above according to multiple embodiments can be used to extract temporal phase signals (phase trajectories) from OCT images. The phase trajectories can be thresholded, normalized, and classified via a pre-trained machine learning model (such as a support vector machine SVM) in a common downsampled principal component space (d-PC space). A new machine learning model (such as an SVM) can then be trained to classify the phase signals extracted from the volume scans. The first three PCs previously extracted can be downsampled at the same time point as the repeated volume scans to form a common d-PC space. As an illustrative example, the original phase signals extracted from five healthy rats using a repeated B-scan protocol are combined and downsampled at the same time point as the repeated volume scans. Signals with a baseline standard deviation greater than 0.3 rad can be considered decorrelated noise and can be removed from further analysis. Each phase trajectory can then be normalized by subtracting its mean and then dividing by its standard deviation. The normalized phase trajectory can be projected into the common d-PC space (to feature points in the common d-PC space) and used as features for SVM training. Labels can be obtained by extracting the corresponding complete phase trajectory and, after preprocessing, classifying it with a previously trained support vector machine (SVM) in the common principal component space (PC space) (without additional thresholding). In experiments, a new SVM classifier was trained in the common dynamic principal component space (d-PC space) using the same process described above, achieving a classification accuracy of 83.4% using 10-fold cross-validation.
[0157] Specifically, as an illustrative example, 6A to 6I The following is a diagram illustrating an example image and signal processing pipeline / framework for a method for training a machine learning model based on optical retinal imaging of wild-type rats, wherein the machine learning model is trained in a temporal feature space. Specifically, the method for training a machine learning model for classifying temporal phase signals obtained from OCT images has been validated in optical retinal imaging experiments of wild-type rats, and the corresponding example image and signal processing pipeline / framework are shown in FIG. 6A to 6I As shown in .
[0158] Figure 6A Serial cross-sectional scans (or B-scans, corresponding to an OCT image sequence) acquired from a rat retina at the same location are shown. Figure 6B and Figure 6C As shown, a phase-recovery sub-pixel motion correction method is used to correct for global motion with sub-pixel accuracy. To this end, a single-step DFT method is used to estimate the volumetric tissue motion between repeated cross-sectional scans, where the peak of the upsampled cross-correlogram is found as the estimated displacement. Figure 6CAlso shown are time-series intensity and phase maps extracted from a dataset of repeated cross-sectional scans (B-scans) at the same location without light stimulation, which demonstrate the high stability of the photoreceptor cell layer. The IS / OS and outer retina were then automatically segmented based on the average B-scans, as shown in Figure 5. Figure 6D shown.
[0159] By taking IS / OS as reference, the phase signal is extracted from each pixel in the outer retina (or more specifically, in the target layer) and processed through a band-stop filter and a low-pass filter. Figure 6E Compared with the light grey line in the figure, the periodic oscillations and high frequency fluctuations ( Figure 6E The dark grey line in the figure was significantly suppressed. Figure 6F Filtered phase traces extracted from five rats are shown, where the ends of the signal with high variations due to edge effects were excluded from further analysis ( Figure 6F The signal outside the black dashed rectangle is removed). Select independent phase traces with low phase fluctuation in the baseline (see Figure 6G ) and normalized by subtracting the mean of the independent phase traces and dividing by their respective standard deviations ( Figure 6H ). For example, the PCA method is then used to extract the principal components PC.
[0160] Therefore, in various embodiments, agglomerative hierarchical clustering and machine learning (e.g., SVM) are applied in PC space to achieve automatic, unbiased light retinal maps. In order to reveal the comprehensive dynamics of the outer retina in response to light stimulation in an unbiased manner, unsupervised learning is used to group independent temporal phase signals (phase trajectories) into different types through agglomerative hierarchical clustering. The labels obtained in the cluster analysis and the features of the independent phase signals in the common PC space are then used to train the SVM model, so that the temporal phase signals extracted from new animals can be classified through the same decision boundary. Specifically, Figure 6A The cross-sectional scans or volume scans shown in the figure were acquired in time series from a 12° field of view (FOV) of the wild-type rat retina. The phase difference between two hyperreflective bands in the outer retina was calculated to reveal their dynamics. Figure 6A As shown, the upper hyperreflective band corresponds to the IS / OS junction of the photoreceptor cells, while the lower hyperreflective band corresponds to the target layer (also called the composite layer or mixed layer), which includes multiple types of tissue that cannot be resolved by existing imaging modalities, as described in the background. In various embodiments, the target layer includes the outer segment tips, microvilli, RPE, and BrM of the photoreceptors. Although phase-sensitive OCT has high motion detection sensitivity, its stability is easily affected by body tissue motion, which will cause significant image distortion in the time-series intensity map and destroy the phase stability (see Figure 6CIn order to eliminate the negative impact of tissue motion, the phase recovery sub-pixel motion correction method is used to align the complex-valued OCT images to sub-pixel accuracy. Figure 6C As shown in , it can be observed that the estimated lateral and axial displacements have periodic oscillations of several micrometers that match the heart beat frequency (about 4 Hz). Figure 6C As shown, excellent motion stability is achieved in the photoreceptor cell layer (white arrows), while slight periodic oscillations can be observed in the choroid (grey arrows).
[0161] For each pixel in the target layer, the temporal phase change of that pixel was calculated relative to the IS / OS layer. Taking into account inter-subject variability, the temporal phase signals from five healthy rats were combined for subsequent feature extraction and unsupervised clustering. The extracted temporal phase signals were filtered and thresholded in the baseline. When the flash was delivered to the retina, a functionally relevant phase response was clearly observed (see Figure 6G As a control, in the absence of light stimulation, the phase traces gradually decorrelated without any obvious response pattern (see Figure 6G (lower part of the figure). Figure 6G The individual phase traces in the figure are normalized by subtracting the mean and then dividing by their own standard deviation. The normalization process compares the two signal modes in the distribution density plot (see Figure 6H Then, the complex time phase signal (phase trajectory) features are extracted into the first three components to form a common PC space, such as Figure 6I As shown. Figure 6I In the dataset, outliers have been removed, and the remaining data points (gray points) are distinguished using a distance-based outlier detection algorithm to avoid undesirable elongated clusters in the subsequent unsupervised clustering.
[0162] like Figure 7A As shown in Figure 2, different levels of threshold processing on the dendrogram can lead to different clustering results. For example, Figure 7A The clustering results are shown in Figure 1. The signal is divided into 7 clusters after the hierarchical tree is truncated along the dotted line. The independent time phase signals (phase traces) in each cluster are retrieved and averaged. Specifically, Figure 7AFigure 2 shows unsupervised clustering using hierarchical clustering and subsequent classification of a new dataset using a trained SVM, including: a dendrogram showing the cluster structure of the signals in the common PC space, with only the first 100 subclusters shown; clustering the signals in the PC space into two clusters by cutting the dendrogram along the black solid line, and showing the corresponding average reconstructed Type-I and Type-II signals (Type-I and Type-II signals are obtained by retrieving the time phase signals and averaging them); C with respect to different numbers of clusters RMS Value (the optimal number of clusters is the value corresponding to the maximum C RMS By cutting the dendrogram along the black dashed line, the signals in the PC space are clustered into 7 clusters, and the average reconstructed signal corresponding to each cluster is displayed.
[0163] In many embodiments, in order to obtain the optimal number of clusters, C is calculated. RMS To quantify the average difference between the reconstructed signals, the larger the C RMS A value of indicates a larger inter-cluster dissimilarity. Figure 7A As shown, when the number of clusters is equal to 2, C RMS reaches its maximum value. Therefore, by following Figure 7A The black solid line in Figure 7 cuts through the cluster tree, grouping the signals into two clusters. Figure 7A also shows the corresponding clustering results and reconstructed signals. The first type of signal (Type-I signal) exhibits rapid elongation followed by gradual recovery, while the second type of signal (Type-II signal) exhibits slow elongation.
[0164] Therefore, in order to automatically classify temporal phase signals (phase traces) in a common PC space (e.g. Figure 6I ), multiple embodiments may use an agglomerative hierarchical clustering algorithm based on the Ward criterion. Thresholding different levels of the dendrogram may result in different clusters (see Figure 7A ). As an example, Figure 7A The figure shows the clustering result of the signal being divided into 7 groups in the PC space after the dendrogram is cut along the black dashed line. The corresponding reconstructed phase signal can be averaged in the corresponding cluster and converted into the optical length change ΔOPL. For example, Figure 7A As shown, two significant types (Type-I and Type-II signals) and some intermediate types are detected. In order to determine the optimal number of clusters for further quantitative analysis, various embodiments can calculate the inter-cluster correlation coefficient between the average reconstructed phase signals and estimate their root mean square RMS relative to 1 (perfect correlation), referred to as C RMS In this regard, the higher the C RMS Indicates greater inter-cluster dissimilarity. In the illustrative example, Figure 7A As shown, when the number of clusters is equal to 2, C RMS reaches its maximum value, at which point the reconstructed phase trajectory has the best effect in distinguishing between clusters. Figure 7A The black solid line in the figure performs thresholding on the dendrogram, grouping the signals in PC space into two clusters. After averaging the reconstructed signals within the clusters, the first type of signal (Type-I signal) exhibits a biphasic trend of rapid elongation followed by gradual recovery, while the second type of signal (Type-II signal) exhibits a monophasic characteristic of slow elongation.
[0165] In various embodiments, in order to classify new signals using the same criteria, the obtained labels can be used to train an SVM model to set the boundaries between Type-I and Type-II signals in a common PC space, such as Figure 7B As shown in Figure 3, using this trained SVM, temporal phase signals extracted from new OCT image datasets can be preprocessed and projected into the same established PC space. Based on the trained SVM classification boundaries, these new temporal phase signals (phase trajectories) are successfully classified as Type-I signals, Type-II signals, and abnormalities. Figure 7C The average phase traces of the classified Type-I and Type-II signals are shown, with their standard deviations shown in the band. Figure 7D As shown in , the ΔOPL peaks and their corresponding delays of the Type-I and Type-II signals exhibit different distributions. Specifically, Figure 7C The corresponding reconstructed time phase signal is shown. The ΔOPL peak and delay of the individual time phase signals (phase traces) are plotted on Figure 7D In the figure, Type-I and Type-II signals show different distributions.
[0166] therefore, Figure 7B The trained SVM decision boundaries of Type-I and Type-II signals in PC space are shown. Figure 7B In FIG7 , the points represent phase trajectories extracted from a new OCT image dataset and preprocessed and projected into the same PC space. The trained SVM is used to classify these phase trajectories into Type-I signals, Type-II signals, and abnormalities. FIG7C shows the phase trajectories according to the Figure 7B The classification results in the reconstruction are Figure 7B The points in correspond to the reconstructed Type-I signal and the reconstructed Type-II signal. Figure 7C In the figure, the solid line and the band represent the mean and standard deviation range (± standard deviation), respectively. Figure 7C Distribution of ΔOPL peaks and corresponding delays of independent signals in .
[0167] According to various embodiments of the present invention, Figure 8 An overview flow chart is shown illustrating processing complex-valued OCT signals of an OCT image to obtain (extract) a time phase signal for training a machine learning model (e.g., an SVM model), training the machine learning model, and classifying new time phase signals obtained relative to a target layer (mixed layer) using the trained machine learning model. In various embodiments, the machine learning model may be trained in a temporal feature space or a spatiotemporal feature space. In this regard, as an illustrative example, Figure 8 shows a machine learning model trained in spatiotemporal feature space. Figure 8 As shown, the phase traces extracted from the mixed layer are preprocessed and projected into the spatiotemporal feature space. Hierarchical clustering is used to identify different phase responses (phase traces). SVMs can then be trained in the same spatiotemporal feature space to facilitate classification of new phase traces using the same criteria. For a better understanding, 9A to 9D Shown Figure 8 The various stages of the flow chart shown are marked 'A', 'B', 'C' and 'D'. Further experimental results according to embodiments of the present invention will now be described or discussed.
[0168] 10A to 10D illustrate the retinal layers and their dynamics in response to visual stimulation during the experiments performed. Figure 10A Shown are average retinal B-scans acquired using near-infrared optical coherence tomography (NIR-OCT). Figure 10B Ultra-high-resolution visible-OCT (vis-OCT) images are shown, which help distinguish the OS, RPE, and BrM layers of photoreceptor cells. The enlarged view of the dotted box is in Figure 10B Shown in (scale bar = 50 μm). NIR-OCT was used for optical retinal imaging experiments. Figure 10C Shown are the optical retinal imaging signals obtained from various hyperreflective bands in the outer retina using BrM as a reference. Figure 10D The optical retinal imaging signals obtained at different depths relative to IS / OS in the mixed layer (target layer) are shown, corresponding to Figure 10A In this article, RNFL denotes retinal nerve fiber layer; GCL denotes ganglion cell layer; IPL denotes inner plexiform layer; INL denotes inner nuclear layer; OPL denotes outer plexiform layer; ONL denotes outer nuclear layer; ELM denotes outer limiting membrane; IS / OS denotes inner segment / outer junction; OS denotes outer segment; RPE denotes retinal pigment epithelium; and BrM denotes Bruch's membrane.
[0169] Specifically, regarding 10A to 10DAll optical retinal imaging experiments were performed in vivo using a custom-built NIR-OCT with an axial resolution of 2.0 μm within the tissue. Sectional scans were acquired from a time series of a 12° field of view (FOV) of the wild-type rat retina (see Figure 10A Custom ultra-high-resolution visible-light OCT (vis-OCT) uses the same optical design as NIR-OCT but provides 1.1 μm axial resolution in tissue for verification of retinal layer delineation (see Figure 10B ).like Figure 10A and 10B As shown, several super-reflective layers are observed in the outer retina, including the outer limiting membrane ELM, the IS / OS junction, the BrM, and a thick speckle layer between the IS / OS and the BrM, which may be referred to herein as the target layer or mixed layer. As verified by vis-OCT, the mixed layer includes the outer ends of photoreceptor cells and RPE cells that cannot be resolved by conventional NIR-OCT systems. In order to reveal the light-induced dynamics in the outer retina, according to multiple embodiments, the changes in the optical length OPL (phase difference) between the BrM and the three super-reflective bands in the outer retina are first calculated, and the three super-reflective bands include the ELM, IS / OS, and the top band of the mixed layer ( Figure 10A To achieve high phase stability / sensitivity in vivo, according to various embodiments, bulk tissue motion is corrected. For example, a phase-recovery motion correction method for registering complex-valued OCT images with sub-pixel accuracy can be employed (e.g., as described in “Shot-noise-limited phase-sensitive imaging of moving samples by phase-recovery sub-pixel motion correction of Fourier-domain optical coherence tomography” by H. Li et al., June 2022, bioRxiv). Phase traces extracted from an extended light retinal imaging experiment (55 seconds in duration) are spatially averaged across pixels within each band. In various embodiments, when converting phase changes to OPL changes, a negative sign is added if the target layer is ahead of the reference layer. This ensures that an increase in OPL always represents an expansion between the two layers.
[0170] like Figure 10C As shown, the distance between the ELM and BrM increases at a rate of approximately 26 nm / s after stimulation (1 ms, 500 nm, 0.18% bleaching level), reaching 220 nm around 20 seconds before slowly recovering. The IS / OS layer rapidly (approximately 1 second) moves away from the BrM, rebounds within 2 seconds, and then continues to move slowly away from the BrM. The top band (first layer) of the mixed layer (probably the OS tip) rapidly (approximately 1 second) moves toward the BrM by approximately 35 nm, then moves away from the BrM over the next 20 seconds before recovering more slowly.
[0171] Traditionally, optical retinal imaging monitors the movement of the OS tip relative to the IS / OS layer. Due to the ambiguity of multiple tissue layers in NIR-OCT, several embodiments have documented OPL changes in several outer retinal bands in the mixed layer relative to the IS / OS, as determined by Figure 10A Various pattern bars are indicated on the lower right corner. Figure 10D As shown, the OPL variations in different bands include a mixture of two different characteristics: (a) rapid (about 1 second) expansion followed by a slower (about 10 seconds) decline, and (b) slow (about 20 seconds) expansion followed by an even slower recovery. Therefore, various embodiments address the technical issue of superposition of different signals by performing signal decomposition and classification to determine the tissue origin associated with each individual time phase signal.
[0172] Unsupervised learning of spatiotemporal patterns for signal classification is now described according to multiple embodiments of the present invention. The limitation of the axial resolution of NIR-OCT causes the light-evoked responses from OS and RPE to be mixed together, thereby complicating the interpretation of the results. In order to identify different signal patterns in an unbiased manner and group them into different types, individual phase trajectories are projected onto the spatiotemporal feature space, and unsupervised learning is performed using agglomerative hierarchical clustering based on the Ward criterion. The SVM model is then trained on these type labels within the established feature space, making it possible to extract the decision boundary of each cluster and classify the new phase trajectories using the same decision boundary.
[0173] Figures 11A to 11D The unsupervised clustering in the spatiotemporal feature space for signal classification according to an embodiment of the present invention is shown. To this end, a machine learning model is trained in the spatiotemporal feature space. Figure 11A The distribution of the signal in the 3D spatiotemporal feature space is shown. Outliers have been removed using a distance-based detection method, and the remaining data points (feature points) are shown. The heat map shows the distribution density of the remaining data points projected onto the temporal feature plane. Figure 11B The residual phase trajectory in the spatiotemporal feature space (corresponding to Figure 11A ), only the first 100 sub-clusters are shown in the spatiotemporal feature space. Figure 11C Shown by Figure 11B The dendrogram is thresholded along the black solid line shown in , and the remaining phase trajectories in the spatiotemporal feature space are clustered into three clusters. Points with different grayscales correspond to different groups. Figure 11D Corresponding representative Type-I and Type-II signals obtained by averaging the individual phase traces within each cluster are shown.
[0174] To account for inter-subject variability, phase trajectories from five rats were combined for feature extraction and subsequent unsupervised clustering. The phase trajectories were preprocessed and subjected to principal component analysis (PCA) to extract temporal features of the phase trajectories. A three-dimensional spatiotemporal feature space was then constructed using the distance to the BrM (depth) as the spatial feature and the first two principal components (PCs) as the temporal features (see
[15] ). Figure 14A ). Use a distance-based algorithm to remove outliers in low-density areas ( Figure 14A gray dots in the image).
[0175] In order to distinguish between two different signatures (see Figure 10D ), using the agglomerative hierarchical clustering algorithm based on Ward's criterion. Figure 11B The black solid line in the dendrogram is thresholded and the remaining phase traces (corresponding to Figure 11A The remaining points in the spacetime feature space ( Figure 11C ) in the three clusters, among which the transition zone ( Figure 11C The points with intermediate gray levels shown in help to better distinguish the two different dynamic features. Representative signals (e.g., Figure 11D The first type of signal (Type-I) shows a rapid rise, reaches a peak around 0.5 seconds, then gradually decreases, and has a negative overshoot after 2.5 seconds ( Figure 11D The second type of signal (Type-II) is characterized by a significantly slower rise and a Figure 11D Its peak value is not reached within the 3.5 second time frame shown in FIG.
[0176] Then, use Figure 11C The labels obtained by unsupervised clustering in are used to train SVMs to set boundaries for Type-I and Type-II signals in the spatiotemporal feature space (see Figure 12A ). Figures 12A to 12C The classification of the new phase trajectory and the verification of its origin according to an embodiment of the present invention are shown. Specifically, Figure 12A The decision boundaries of the trained SVM for Type-I and Type-II signals are shown. The points represent phase trajectories extracted from the new dataset, preprocessed, and projected onto the same 3D spatiotemporal feature space. These phase trajectories are classified as Type-I signals, Type-II signals, intermediate phase trajectories, and outliers (not shown). Figure 12B Shown Figure 10BA magnified view of the dashed box in Figure 3, with contrast adjusted to enhance visibility of the RPE. Histograms of Type-I and Type-II signals fitted by a Gaussian function (solid line) show their depth distribution on the average structural image acquired by vis-OCT, with the white line marking their intensity distribution. Figure 12B The bar chart on the left is extracted Figure 10D 12C shows a plot of Type-I, Type-II, and SRS signals from an extended (55 s) recording.
[0177] Therefore, using the pre-trained SVM, multiple embodiments successfully extracted Type-I and Type-II signals from the new OCT image dataset. Type-I signals are located further in front than Type-II signals, and their distribution is close to the normal distribution ( Figure 12B The gray curve in Figure 2). The Type-II signal is located before the BrM and is fitted with a Gaussian function ( Figure 12B The positions of the peaks were significantly different (*P<0.05, t-test). The positions of the Type-I signals corresponded to the intensity distribution of the OS, while the Type-II signals corresponded to the positions of the RPE, as shown in vis-OCT ( Figure 15B The observations indicate that the Type-I and Type-II signals correspond to the dynamics of the OS and RPE relative to the IS / OS, respectively. Phase traces from the extended (55 s) recordings were interpolated and clustered using SVM. The Type-I signal peaked within one second, reached a trough with a negative undershoot at 10 s, and returned to baseline within 30 s (see Figure 12C By combining the Type-II signal with the dynamics between IS / OS and ELM, the dynamics of SRS (from ELM to RPE) are obtained. Figure 12C The SRS trace shown in Figure 3 increases more slowly, peaking at approximately 250 nm around 20 seconds, and then recovers even more slowly. In subsequent quantitative analysis, we will focus on the dynamics of OS and SRS.
[0178] The dependence of the signal on the stimulation parameters will now be described according to an embodiment of the present invention. Figures 13A to 13C The responses of the outer segment OS and subretinal space SRS under different conditions are shown. Specifically, Figure 13A Representative traces of light-evoked responses of photoreceptor cells OS and SRS are shown. Figure 13B Shown are the amplitude and latency of the OS response, and the slope of the SRS dilation as a function of stimulus intensity on a scotopic background. Figure 13C Photopic background intensity plots are shown. For boxplots, the horizontal bar represents the mean, the box edges represent the 25th and 75th percentiles, and the whiskers represent 1.5× standard deviation (SD).
[0179] Specifically, the signal under scotopic and photopic conditions was quantified by plotting the amplitude and latency of the OS dilation peak and the slope (dilation rate) of the SRS response (see Figure 13A Under dark conditions (see Figure 13B ), the peak amplitude of the OS signal increased logarithmically with increasing stimulus intensity, from a ΔOPL of approximately 10 (1.72) nm [mean (standard deviation)] at a bleaching level of 0.002% to approximately 53 (8.68) nm when the flash intensity was increased 140-fold to 0.28%. At bleaching levels <0.1%, the peak latency remained within 440 ms to 500 ms, but increased to 877 (130) ms and 1063 (108) ms in response to bleaching levels of 0.19% and 0.28%, respectively. In contrast, the slope of the SRS signal increased rapidly with increasing stimulus intensity until the bleaching level reached 0.06%, after which it stabilized at a level of approximately 25 nm / s. Under photopic conditions ( Figure 13C ), the retina was pre-irradiated with a 500 nm flash for 5 minutes, resulting in a bleaching level of 0.28%. Background illumination below 0.078 × 106 photons / (μm2·s) did not reduce the OS response, but at background illumination levels of 7.8 × 106 photons / (μm2·s), the OS response decreased fivefold. The OS peak latency remained unchanged at a stronger background illumination of 0.78 × 106 photons / (μm2·s), but decreased by approximately 30% at 7.8 × 106 photons / (μm2·s). Similarly, the SRS expansion rate was affected by background illumination levels just above 0.078 × 106 photons / (μm2·s), which decreased the slope by approximately 30% to 17 nm / s.
[0180] Figures 14A to 14C Representative en face functional maps of the outer segment OS and subretinal space SRS signals are shown. Specifically, Figure 14A A volume scan covering a 12° field of view with structural contrast is shown. 14B shows the OS and SRS signals at selected time points. Each grid represents the average response from a 0.48° × 0.48° (x × y) region. Locations obscured by large blood vessels are indicated by dashed lines. Figure 14C The spatiotemporal evolution of OS and SRS signals over the entire field of view (FOV) is shown. Each curve represents the average response from a 1.2° × 0.48° (x × y) region (scale bar: 200 μm).
[0181] Specifically, repeated volume scans were used to map the dynamics of SRS and OS, an approach similar to multifocal ERG but utilizing a single flash. Figure 14AAn OCT volume with a 12° field of view (FOV) and structural contrast is shown. The spatiotemporal distribution of SRS and OS signals was acquired at a bleaching level of 0.10%. The temporal resolution was 8 Hz, and the total recording time was 5 seconds, with a baseline of 1 second. Figure 14B OS and SRS signals at specific time points are shown. Figure 14C The spatial distribution of the phase trace is shown, where each grid represents the average response from a 1.2° × 0.48° region. Overall, the OS and SRS signal plots show high-fidelity detection across the entire FOV, and Figure 14B and Figure 14C High signal variance is observed beneath the two large blood vessels outlined by dashed lines in the middle image.
[0182] Compared to the slow, low-resolution multifocal ERG, recording signals from SRS and OS in response to a single flash over a wide field of view provides a convenient method for diagnostic mapping. This technique expands the practical applicability of photoretinography to study not only photoreceptor cells but also the RPE's control of subretinal water dynamics in health and disease, thus enabling a more complete replacement for ERG.
[0183] Now refer to Figures 15A to 15C An example system setup as well as example stimulation and acquisition protocols according to an embodiment of the present invention are described. Figure 15A Shown is imaging of the posterior segment of a rat eye using spectral domain OCT (ultra-high axial resolution). Interference fringes were acquired using a line scan camera interfaced with a spectrometer. A function generator synchronized the line scan camera acquisition, galvanometer scanner rotation, and flash timing. Figure 15A In the figure, L1-L6 represent doublet lenses, CL represents a condenser lens, GS represents a galvanometer scanner, and the spectral window of the filter is 500±5 nm. Figure 15B The three acquisition protocols used are shown, namely, repeated B-scan (Protocol 1: 1000 A-scans per B-scan, 200 B-scans per second, Protocol 2: 1000 A-scans per B-scan, 25 B-scans per second) and repeated volume (Protocol 3: 1000 A-scans per B-scan, 25 B-scans per volume, 8 B-scans per second). Details of the acquisition and stimulation protocols are presented in Figure 16B As an illustrative example described above, Figure 6CShown are the uncorrected temporal intensity (upper panel) and phase (lower panel) M-scans (scale bar: 500 ms (horizontal), 100 μm (vertical)) when no light stimulus was delivered to the retina, the subpixel overall motion estimated by locating the peak of the upsampled cross-correlogram between repeated B-scans (scale bar: 500 ms (horizontal), 5 μm (vertical)), and the corresponding temporal intensity (upper panel) and phase (lower panel) M-scans after phase-recovery subpixel motion correction (scale bar: 500 ms (horizontal), 100 μm (vertical)).
[0184] Optical retinal imaging experiments were performed using Figure 15A The custom spectral domain OCT system shown here is completed. The NIR-OCT system uses a broadband superluminescent diode (c BLMD-T-850-HP-I, λ c =840 nm, Δλ=146 nm, Superlum, Ireland), providing an axial resolution of 2.0 μm in tissue. A spectrometer interfaced with a line scan camera (2048 pixels, 250,000 Hz Cobra-800, Octoplus, E2V) collected spectral interference fringes at an A-scan rate of 250 kHz, corresponding to an imaging depth of 1.07 mm in air. For imaging the posterior segment of the rodent eye, a lens-based afocal telescope conjugates the axis of the galvanometer scanner to the pupil. To reduce the incident beam size and increase the scanning angle, the telescope has a magnification of 0.17 (scan lens: focal length 80 mm; eyepiece: focal length 30 mm + 25 mm). For a standard rat eye model, the theoretical diffraction-limited lateral resolution is 7.2 μm. A function generator (PCIe-6363, National Instruments) synchronizes camera acquisition (both A-scan and B-scan acquisition), galvanometer scanner scanning, and visual stimulation.
[0185] For visual stimulation, a light-emitting diode (LED) (MBB1L3, Thorlabs, USA) was collimated by an aspheric condenser lens and its spectrum was reshaped by a narrow bandpass filter (500 ± 5 nm, #65-694, Edmund Optics, Singapore) to optimize the sensitivity to rhodopsin (32). The response time of the LED was about 300 μs, which is equivalent to 6% of the B-scan frame acquisition time (5 ms). The trigger response waveform of the LED is shown in Figure 2. Figure 15C A 43.4° Maxwell illumination was projected onto the posterior segment of the eye to cover an area of approximately 6.75 mm2. The power and duration of the light stimulus were converted into a percentage of rhodopsin bleaching level.
[0186] To resolve the ultra-fine laminar structure of the outer retina, a visible light optical coherence tomography (vis-OCT) system with high axial resolution (1.1 μm) within the tissue was constructed. Briefly, the system uses a supercontinuum laser (SuperKExtreme, NKT Photonics, Denmark) that cuts off the spectrum between 435 and 650 nm. Vis-OCT uses an optical design similar to that of the NIR system, although all components operate in the visible spectrum. A spectrometer (Cobra VIS, Wasatch, USA) interfaced with a line scan camera acquires the spectrum of the interference fringes at an A-scan rate of 50 kHz. Volume raster scans (500 A-scans × 500 B-scans) are acquired within a 12° × 1.2° (1.14 mm × 0.11 mm) rectangular field of view. Continuous B-scans are aligned and then averaged into a single frame.
[0187] The animal experimental protocol will now be described. These experiments were conducted in accordance with the guidelines and approval of the Institutional Animal Care and Use Committee (IACUC) of SingHealth Group (2020 / SHS / 1574). Brown Norway rats (N=39) were used for the experiments, and the details are shown in Figure 16A Table 1. Animals were anesthetized with a ketamine / xylazine combination, which better maintains retinal functional responses while minimizing eye movements compared to other commonly used anesthetics (such as isoflurane and urethane). Vital signs, including heart rate and respiratory rate, were monitored throughout the imaging process. The animals were placed in a prone position with their heads fixed stereotaxically. Prior to imaging, two drops of mydriatic drops, 1% tropicamide (Alcon, Geneva, Switzerland) and 2.5% phenylephrine (Alcon, Geneva, Switzerland) were applied to the cornea. The cornea was frequently moistened with a balanced salt solution throughout the imaging process.
[0188] M-opsin and rhodopsin have similar sensitivity spectra (peaking at approximately 500 nm), while S-opsin's sensitivity peaks at 350 nm, with minimal overlap with the rhodopsin sensitivity spectrum. Because the co-expression rate of M-opsin decreases from ventral to dorsal, the scanning area was limited to the dorsal region to minimize the effect of M-opsin on cones.
[0189] Details of the acquisition and stimulation protocols are presented in Figure 16BTable 2 in the accompanying figures. Three scanning protocols were used in the experiments performed. In the first protocol, each B-scan was acquired with 1000 A-scans and 200 B-scans per second, for a total acquisition time of 5 seconds. In the second protocol, the B-scan interval was increased to 40 ms, and 1375 B-scans were acquired over a 55-second period. In the third protocol, 40 replicate volumes (25 B-scans per volume) were acquired, corresponding to 8 volume scans per second, for a duration of 5 seconds. Flash intensity, dark / bright adaptation, and inter-flash interval were varied between experiments.
[0190] Automatic extraction of outer retinal dynamics from speckle patterns will now be described according to an embodiment of the present invention.
[0191] First, the OCT images and phase traces are preprocessed. In this regard, the raw interference fringes first undergo standard OCT post-processing steps, including spectral calibration, k-linearity, dispersion compensation, and discrete Fourier transform (DFT), to obtain a complex-valued OCT image.
[0192] Phase-sensitive OCT is highly susceptible to global tissue motion, which can lead to significant image distortion and reduced phase stability in sequential intensity M scans (see Figure 6C To correct for motion-induced phase errors, a single-step DFT algorithm is used to estimate subpixel translational displacements between repeated B-scans (see Figure 6C The complex-valued OCT images are registered using a phase-recovery sub-pixel motion correction algorithm, where lateral and axial displacements are corrected by multiplying the corresponding exponential terms in the spatial frequency domain and the spectral domain, respectively. After image registration (see Figure 6C (right part of the image), excellent motion stability is achieved in the photoreceptor cell layer (white arrows), while periodic oscillations from vascular pulsation can be observed in the choroid (grey arrows).
[0193] The light-evoked dynamics of the outer retina are then extracted by calculating the temporal phase difference between pairs of pixels from these outer retinal zones. Graph theory and dynamic programming are used to automatically segment several hyperreflective zones, including the ELM, IS / OS, BrM, and a thick speckle layer containing OS photoreceptors and RPE cells. In a self-referenced measurement, for each pixel in the target layer, its temporal phase change is calculated relative to a reference layer.
[0194] Temporal filtering was used to remove unwanted signal frequencies. The extracted signal was first processed with a band-stop filter to minimize residual artifacts caused by heartbeat and respiration. A low-pass filter with a cutoff frequency of 10 Hz was then used to remove high-frequency oscillations. After filtering, the two ends of the phase trace exhibited high variance due to edge effects and were therefore excluded from subsequent analysis.
[0195] Regarding the construction of the spatiotemporal feature space according to multiple embodiments, PCA is used to compress high-dimensional phase trajectories into a low-dimensional feature space for ease of analysis. Phase trajectories with a standard deviation SD greater than 60 mrad before stimulation were removed, which mainly came from pixels with low signal-to-noise ratio or under blood vessels. Subsequently, each trajectory was normalized by subtracting its mean and dividing by its standard deviation SD. In order to take into account inter-subject differences, the normalized trajectories extracted from five rats were combined to calculate the PC coefficients. A 3D spatiotemporal feature space was constructed, in which the first two PCs that captured 73.3% of the total variance were selected as temporal features, and their axial distances to the BrM were selected as spatial features.
[0196] In order to avoid the appearance of undesirable elongated clusters in the subsequent unsupervised cluster analysis, a distance-based outlier detection method is used to exclude outliers distributed in low-density areas in the spatiotemporal feature space. For each data point in the spatiotemporal feature space, the minimum radius of a sphere centered at that point and covering 2% of the remaining data points is determined. This radius reflects the local distribution density around each data point in the spatiotemporal feature space. If the corresponding radius of a given point is greater than Q3+(Q3-Q1) / 5, the given point will be marked as an outlier, where Q1 and Q3 are the first and third quartiles of the calculated radius.
[0197] In various embodiments, unsupervised clustering using a hierarchical clustering algorithm is performed using an agglomerative hierarchical clustering algorithm based on the Ward criterion to group points in the established spatiotemporal feature space. The Euclidean distance between each pair of points is calculated, and similar subclusters are then iteratively merged into larger clusters. Each merge is performed to minimize the increase in the total variance within the cluster, using either the Ward method or the minimum variance method.
[0198] Regarding training a support vector machine model in an established spatiotemporal feature space according to multiple embodiments, the SVM is trained in the spatiotemporal feature space using previously obtained labels (labeled feature points), including outliers, intermediate phase trajectories, Type-I signals, and Type-II signals, to obtain their decision boundaries. The input features can be standardized before training, a Gaussian kernel can be selected, and automatic optimization of hyperparameters can be enabled. The trained SVM achieved a classification accuracy of 99.2% after evaluation using a 10-fold cross-validation strategy. Therefore, the trained SVM is useful for processing new data sets using the same standard. In this regard, the time phase signal (phase trajectory) extracted from the new data set can be preprocessed and projected into the same spatiotemporal feature space. Then, for example, Type-I and Type-II signals can be automatically extracted based on the classification results of the trained SVM.
[0199] Regarding processing datasets with lower temporal resolution according to various embodiments, extended recordings (Scheme 2) and volume scans (Scheme 3) have temporal sampling rates 8 and 25 times lower, respectively, than the repeated B-scans (Scheme 1). To cluster these phase traces in the same spatiotemporal feature space, they are first linearly interpolated to 5 ms intervals and smoothed using a Gaussian filter. The remaining processing flow is similar to that used for processing the repeated B-scan datasets.
[0200] According to various embodiments, a supercontinuum laser (SuperK Extreme, NKT Photonics, Denmark) covering a spectrum of 400-2300 nm was used as a light source to construct visible light spectral domain optical coherence tomography (vis-OCT). A spectral range of 435 nm to 650 nm (full width at half maximum bandwidth of 120 nm) was selected by a spectral beam splitter (SuperK SPLIT, NKT Photonics, Denmark) and a short-pass filter (#47-290, Edmund Optics, USA), and further shaped by a bandpass filter (#16-362, Edmund Optics, USA). In the sample arm, a custom achromatic triplet lens was placed between a reflective collimator and a galvanometer scanner (Saturn5, ScannerMax, USA) to reduce chromatic aberration in the rat eye. A lens-based afocal telescope conjugates the center of the galvanometer scanning pair to the pupil with a magnification of 0.18 (scan lens: 75 mm focal length; eyepiece: 30 mm + 25 mm focal length) to reduce the incident beam diameter at the pupil to 310 μm.
[0201] Thus, various embodiments demonstrate that non-adaptive optics (non-AO)-based point-scanning frequency-domain optical coherence tomography (SD-OCT) can enable automated, unbiased optical retinal imaging of rodents. Thus, an SVM model is trained to explore the decision boundaries of each type of signal in the established PC space, thereby facilitating the classification of new phase trajectories extracted from, for example, new animals. For example, various embodiments can simultaneously investigate outer retinal dynamics from various tissue types and generate a geographical map of tissue-specific dynamics using only a single, brief flash, similar to a multifocal electroretinogram. Furthermore, various embodiments develop a general protocol applicable to conventional OCT systems in animal laboratories and clinics, requiring no hardware or human intervention, and possessing significant potential for exploring tissue-specific dynamics without prior knowledge.
[0202] According to various embodiments, after disambiguating various signals in the outer retina, the dynamics of the outer retinal cell layers can be independently described. The time derivative of tissue deformation (i.e., its dilation rate) is very informative because it can reveal the water influx rate. Figures 17A to 17E The temporal deformation of the cell layer and the comparison of light retinal imaging and ERG are shown. Figure 17A Shown are the dynamics of deformation of the rod outer segments OS, inner segments IS, retinal pigment epithelium RPE, and subretinal space (SRS, from BrM to ELM) after 1 ms green stimulation at a 0.26% bleaching level. Figure 17B Shown are the mean expansion rates of OS and RPE, respectively, over 5 measurements. Figure 17C and Figure 17D Shown respectively Figure 17A and 17B A magnified view of the dashed box data in . Figure 17E An example electroretinogram (ERG) trace in response to a white flash is shown, with the a, b, and c waves labeled. The OPL change translates into a physical deformation with a refractive index of 1.41.
[0203] like 17A to 17E As shown, the expansion rate of OS reaches a maximum of about 165nm / s within 0.1 seconds after stimulation and falls back to zero within the next 1 second. On the other hand, RPE reaches a maximum contraction rate of about -9nm / s 2 seconds after stimulation, resulting in a maximum contraction of about 20nm at 3 seconds, and then slowly recovers. The expansion of SRS begins about 1.5 seconds after stimulation, reaches an expansion rate of 15nm / s within the following 10 seconds, after which the expansion rate gradually decreases. The expansion of SRS may be derived from a response to a decrease in K+ and an increase in Na+ concentration due to phototransduction in photoreceptor cells, thereby causing water transport across RPE, as previously observed using other methods. The inner segment is compressed by about 10nm during 1.5 seconds, after which it follows the dynamic changes of SRS, although with a smaller amplitude (maximum expansion of 100nm).
[0204] Notably, it was possible to correlate the retinal optical imaging signals with the presence of white flash-evoked ERGs previously recorded in pigmented rats of the same species, e.g. Figure 17EAs shown. The a-wave in the ERG (i.e., the response of the photoreceptor cells to light) begins immediately after the flash (within 10 ms) and corresponds to the onset of OS dilation in photoretinal imaging. The continuation of the a-wave is obscured by the b-wave, which begins a few tens of milliseconds later and corresponds to the current generated by the secondary retinal neurons. The C-wave in the ERG appears after about 0.5 seconds and reaches a maximum at about 1.5 seconds. It corresponds to the potassium current through the RPE caused by the response of the photoreceptor cells to light. The RPE contraction in photoretinal imaging reaches its maximum rate between 1 and 2 seconds, overlapping with the peak of the c-wave (see Figure 17D and 17E This new optical feature opens a window into the physiological activities of photoreceptor cell-RPE interactions and the photoreceptor intercellular matrix. Water transport is limited by membrane permeability and is therefore much slower than electrical current. Using a model of OS elongation due to osmotic imbalance during phototransduction, the membrane permeability coefficient can be estimated based on the OS expansion rate. The water permeability coefficient of the rod OS membrane was measured to be 5.9×10 -3 cm·s -1 , which is consistent with the previously measured in vitro value of 2.6×10 -3 cm·s -1 Not too far off.
[0205] Interestingly, the hyperpolarized RPE layer is compressed during this process, which differs from the expansion of the hyperpolarized outer segments during phototransduction that is driven by osmotic changes in response to released osmolytes.
[0206] Rods and cones differ in their responses to light in several ways, including sensitivity, time constants, and light adaptation. These differences are attributed to differential expression of isoenzymes involved in the phototransduction cascade and differences in the structural organization of the rod and cone plasma membranes. Until now, light retinal imaging has primarily been used to visualize cones, as detecting rod signals in humans is more challenging because rods are typically smaller and densely packed around cones. However, in rats, where 97% of photoreceptors are rods, OS signals may be dominated by rod outer segment ROS. Adaptive optics (AO) has enabled imaging of single rods in the human peripheral retina, and one AO-OCT study reported that 0.05% rhodopsin bleaching resulted in an approximately 60 nm elongation of the rod outer segment, whereas 0.2% opsin bleaching under the same flash conditions did not result in detectable cone outer segment OS elongation. These results are consistent with our observation that 0.06% bleaching level caused rat rod OS to elongate by 20–30 nm.
[0207] In a previous study using conventional intensity-based OCT, Zhang et al., in their April 2017 publication of the Proceedings of the National Academy of Sciences, published "In vivo photophysiology reveals osmotic swelling and enhanced light scattering of rod photoreceptors triggered by G protein activation," observed a slow increase in the distance between the IS / OS and the BrM in response to light (peak delay of 10–100 seconds, depending on the level of bleaching) and attributed this to OS elongation. Notably, this slow signal resembled the SRS expansion reported here, while the actual rise and fall of the OS signal in this study was much faster than in previous reports.
[0208] Another interesting finding is the undershoot of the OS signal, which has a recovery process of tens of seconds ( Figure 17A ), which had not been reported in previous studies due to the short observation time. This phenomenon may be related to water transport from the OS to the SRS during phototransduction inactivation, when osmolytes (Gαt, Gβ1γ1) re-bind to the cell membrane. During SRS expansion, the recovery of OS osmotic pressure may be accompanied by a slight 20 nm compression, which is later restored along with other cellular structures in the outer retina.
[0209] The expansion of the SRS, which is closely related to phototransduction and water transport through the RPE, is greater than that of the OS. This expansion can be used as a more sensitive measure of retinal physiology, providing additional diagnostic insights into various diseases involving the outer retina. Optical retinal imaging conveniently integrates structural and functional retinal imaging in the same instrument.
[0210] Thus, according to various embodiments, an automated, unbiased method for extracting wide-field, depth-resolved optical retinal imaging signals is disclosed, revealing comprehensive nanoscale dynamics of the outer retina in preclinical models. This method can be broadly applied to studying tissue-specific dynamics in various animal models. For example, the detected outer retinal dynamics can serve as novel evidence to improve our understanding of the phototransduction cascade. Furthermore, it has significant potential for studying disease models and promoting clinical translation.
[0211] For example, the method of optical retinal imaging using OCT according to multiple embodiments can significantly facilitate non-invasive imaging of nanoscale motion or cell deformation in clinical applications on various FD-OCT systems (including point scanning, line scanning, and full-field systems). Although OCT has been widely used for structural imaging in ophthalmology, the method according to multiple embodiments improves the accuracy of cell dynamic imaging, which may bring new market space for functional diagnosis, especially for emerging optical retinal maps that measure the physiological response of the retina in a non-invasive and all-optical manner. For example, because the method according to multiple embodiments allows for accurate correction of the phase component that is disturbed by the inevitable overall motion in living tissue without additional hardware, it can be easily applied to any clinical OCT system that uses the phase component for imaging, thereby achieving large-scale rapid commercialization.
[0212] Although the embodiments of the present invention have been particularly shown and described with reference to specific embodiments, it will be understood by those skilled in the art that various changes in form and details may be made without departing from the scope of the invention as defined by the appended claims. The scope of the present invention is therefore indicated by the appended claims, and all changes that come within the meaning and range of equivalence of the claims are intended to be embraced therein.
Claims
1. A method for training a machine learning model for classifying temporal phase signals obtained from optical coherence tomography (OCT) images generated for optical retinal imaging, the method being performed using at least one processor, the method comprising: Acquiring an OCT image sequence generated by a phase-sensitive OCT system performing retinal optical imaging of the outer retina, the OCT image sequence including an OCT image stimulus subsequence generated in response to a visual stimulus of the outer retina; determining, for each OCT image in the OCT image sequence and for each individual pixel in a target layer of the OCT image, a phase signal of the individual pixel in the target layer of the OCT image relative to a reference layer of the OCT image, so as to acquire, for each of a plurality of corresponding individual pixel sequences in the target layer across the OCT image sequence, a temporal phase signal of the corresponding individual pixel sequence, thereby generating a set of temporal phase signals of the plurality of corresponding individual pixel sequences, the target layer comprising a plurality of outer retinal zones; Projecting multiple time phase signals in the time phase signal set to multiple feature points in a feature space; marking each feature point of the plurality of feature points in the feature space; as well as The machine learning model is trained based on a plurality of labeled feature points in the feature space.
2. The method according to claim 1, wherein Determining the phase signal of the individual pixel in the target layer of the OCT image includes: correcting the systematic phase drift of the complex-valued OCT signal of the individual pixel based on the individual pixel of the reference layer of the self-referenced OCT image to obtain a first complex-valued OCT signal of the individual pixel.
3. The method according to claim 2, wherein: Correcting the systematic phase drift of the complex-valued OCT signal of the individual pixel includes multiplying the complex-valued OCT signal of the individual pixel with the complex conjugate of the complex-valued OCT signal of the individual pixel of the reference layer of the OCT image to obtain the first complex-valued OCT signal of the individual pixel.
4. The method according to claim 3, wherein: For each OCT image in the OCT image stimulation subsequence, determining the phase signal of the individual pixel in the target layer of the OCT image further includes: correcting the phase offset of the first complex-valued OCT signal of the individual pixel based on a pre-stimulation subsequence in the OCT image sequence generated before the visual stimulation of the outer retina to obtain a second complex-valued OCT signal of the individual pixel.
5. The method according to claim 4, wherein The correcting the phase offset of the first complex-valued OCT signal of the individual pixel comprises: determining an average value of the first complex-valued OCT signals of corresponding individual pixels of a pre-stimulation subsequence of the OCT images corresponding to the individual pixels of the OCT image to obtain an average pre-stimulation complex-valued OCT signal; and The first complex-valued OCT signal of the individual pixel is multiplied by the average pre-stimulus complex-valued OCT signal to obtain the second complex-valued OCT signal of the individual pixel.
6. The method according to claim 4 or 5, wherein: For each OCT image in the OCT image stimulation subsequence, the phase signal of the individual pixel in the target layer of the OCT image is determined based on the second complex-valued OCT signal of the individual pixel.
7. The method according to claim 5 or 6, wherein: The phase signal of the individual pixel in the target layer of the OCT image is determined relative to a reference pixel region of the reference layer of the OCT image, the reference pixel region including a plurality of individual pixels, For correcting the systematic phase drift of the complex-valued OCT signal of the individual pixel, multiplying the complex-valued OCT signal of the individual pixel comprises: multiplying the complex-valued OCT signal of the individual pixel by the complex conjugate of the complex-valued OCT signals of the plurality of individual pixels of the reference pixel region of the reference layer of the OCT image, respectively, to obtain a set of first complex-valued OCT signals of the individual pixel; and for correcting the phase shift of the complex-valued OCT signal of an individual pixel, The determining the average value of the first complex-valued OCT signal corresponding to the individual pixel comprises: determining an average value of a set of the first complex-valued OCT signals corresponding to the individual pixels of the pre-stimulation subsequence of the OCT images corresponding to the individual pixels of the OCT image to obtain a set of average pre-stimulation complex-valued OCT signals; and The multiplying the first complex-valued OCT signals of the individual pixels includes multiplying the set of the first complex-valued OCT signals of the individual pixels with the set of the averaged pre-stimulus complex-valued OCT signals to obtain the set of second complex-valued OCT signals of the individual pixels.
8. The method of claim 7, wherein for each OCT image in the OCT image stimulation subsequence, determining the phase signal of the individual pixel in the target layer of the OCT image further comprises: determining an average value of a set of the second complex-valued OCT signals of the individual pixels to obtain an average complex-valued OCT signal of the individual pixels; as well as The phase signals of the individual pixels in the target layer of the OCT image are determined based on the average complex-valued OCT signals of the individual pixels.
9. The method according to claim 6 or 8, wherein the temporal phase signal of the corresponding individual pixel sequence in the target layer across the OCT image sequence is obtained based on the phase signal of the corresponding individual pixel sequence in the target layer across the OCT image sequence.
10. The method according to any one of claims 1 to 9, wherein marking each feature point in the plurality of feature points in the feature space comprises: Grouping the plurality of feature points into a plurality of clusters in the feature space; as well as For each of the plurality of clusters, feature points belonging to the cluster are marked using a label assigned to the cluster.
11. The method according to claim 10, wherein: The plurality of feature points are grouped into the plurality of clusters based on an unsupervised clustering technique.
12. The method according to claim 10 or 11, wherein: The plurality of clusters are assigned a plurality of different labels including a plurality of different outer retinal zone labels corresponding to the plurality of outer retinal zones of the target layer.
13. The method of any one of claims 1 to 12, wherein the reference layer corresponds to the inner segment / outer segment IS / OS junction, and the plurality of outer retinal bands include two or more outer retinal bands corresponding to two or more of the photoreceptor cell outer segments, microvilli, retinal pigment epithelium RPE, and Bruch's membrane, respectively.
14. The method according to any one of claims 1 to 13, further comprising: Before projecting the multiple time phase signals to the multiple feature points in the feature space, performing band-stop filtering and low-pass filtering on the multiple time phase signals; and normalizing the multiple time phase signals.
15. A method for classifying a temporal phase signal obtained from an optical coherence tomography (OCT) image generated for optical retinal imaging, the method using a machine learning model trained according to any one of claims 1 to 14, the method comprising: Acquiring an OCT image sequence generated by a phase-sensitive OCT system performing retinal optical imaging on an outer retina, the OCT image sequence including an OCT image stimulus subsequence generated in response to a visual stimulus of the outer retina; determining, for each individual pixel in a corresponding sequence of individual pixels in a target layer across the sequence of OCT images, a phase signal of the individual pixel in the target layer of the OCT image relative to a reference layer of the OCT image to obtain a temporal phase signal of the corresponding sequence of individual pixels, the target layer comprising a plurality of outer retinal bands; Projecting the time phase signal corresponding to the individual pixel sequence to a feature point in a feature space; as well as The time phase signal is classified using the trained machine learning model based on the feature points to which the time phase signal is projected.
16. The method according to claim 15, wherein Determining the phase signal of the individual pixel in the target layer of the OCT image includes: correcting the systematic phase drift of the complex-valued OCT signal of the individual pixel based on the individual pixel of the reference layer of the self-referenced OCT image to obtain a first complex-valued OCT signal of the individual pixel.
17. The method according to claim 16, wherein Correcting the systematic phase drift of the complex-valued OCT signal of the individual pixel includes multiplying the complex-valued OCT signal of the individual pixel with the complex conjugate of the complex-valued OCT signal of the individual pixel of the reference layer of the OCT image to obtain the first complex-valued OCT signal of the individual pixel.
18. The method according to claim 17, wherein: For each OCT image in the OCT image stimulation subsequence, determining the phase signal of the individual pixel in the target layer of the OCT image further includes: correcting the phase offset of the first complex-valued OCT signal of the individual pixel based on a pre-stimulation subsequence in the OCT image sequence generated before the visual stimulation of the outer retina to obtain a second complex-valued OCT signal of the individual pixel.
19. The method according to claim 18, wherein The correcting the phase offset of the first complex-valued OCT signal of the individual pixel comprises: determining an average value of the first complex-valued OCT signals of corresponding individual pixels of a pre-stimulation subsequence of OCT images corresponding to the individual pixels of the OCT image to obtain an average pre-stimulation complex-valued OCT signal; and The first complex-valued OCT signal of the individual pixel is multiplied by the average pre-stimulus complex-valued OCT signal to obtain the second complex-valued OCT signal of the individual pixel.
20. The method according to claim 18 or 19, wherein For each OCT image in the OCT image stimulation subsequence, the phase signal of the individual pixel in the target layer of the OCT image is determined based on the second complex-valued OCT signal of the individual pixel.
21. The method according to claim 19 or 20, wherein The phase signal of the individual pixel in the target layer of the OCT image is determined relative to a reference pixel region of the reference layer of the OCT image, the reference pixel region including a plurality of individual pixels, For correcting the systematic phase drift of the complex-valued OCT signal of the individual pixel, multiplying the complex-valued OCT signal of the individual pixel comprises: multiplying the complex-valued OCT signals of the individual pixels by the complex conjugates of the complex-valued OCT signals of the plurality of individual pixels of the reference pixel region of the reference layer of the OCT image to obtain a set of first complex-valued OCT signals of the individual pixels, and for correcting the phase shift of the complex-valued OCT signal of an individual pixel, The determining the average value of the first complex-valued OCT signal corresponding to the individual pixel comprises: determining an average value of a set of the first complex-valued OCT signals corresponding to the individual pixels of the pre-stimulation subsequence of the OCT images corresponding to the individual pixels of the OCT image to obtain a set of average pre-stimulation complex-valued OCT signals; and The multiplying the first complex-valued OCT signals of the individual pixels includes multiplying the set of the first complex-valued OCT signals of the individual pixels with the set of the averaged pre-stimulus complex-valued OCT signals to obtain the set of second complex-valued OCT signals of the individual pixels.
22. The method of claim 21 , wherein for each OCT image in the OCT image stimulation subsequence, determining the phase signal of the individual pixel in the target layer of the OCT image further comprises: determining an average value of a set of the second complex-valued OCT signals of the individual pixels to obtain an average complex-valued OCT signal of the individual pixels; as well as The phase signals of the individual pixels in the target layer of the OCT image are determined based on the average complex-valued OCT signals of the individual pixels.
23. A method according to claim 20 or 22, wherein the temporal phase signal of the corresponding individual pixel sequence in the target layer across the OCT image sequence is obtained based on the phase signal of the corresponding individual pixel sequence in the target layer across the OCT image sequence.
24. The method according to any one of claims 15 to 23, wherein The temporal phase signal is classified into one of a plurality of types corresponding to a plurality of outer retinal bands of the target layer or an abnormal type using the trained machine learning model.
25. The method of any one of claims 15 to 24, wherein the reference layer corresponds to the inner segment / outer segment IS / OS junction, and the plurality of outer retinal bands include two or more outer retinal bands corresponding to two or more of the photoreceptor cell outer segments, microvilli, retinal pigment epithelium RPE, and Bruch's membrane, respectively.
26. The method according to any one of claims 15 to 25, further comprising: Before projecting the time phase signal onto the feature point in the feature space, performing band-stop filtering and low-pass filtering on the time phase signal; and normalizing the time phase signal.
27. A system for training a machine learning model for classifying temporal phase signals obtained from optical coherence tomography (OCT) images generated for optical retinal imaging, the system comprising: at least one memory; as well as At least one processor, the at least one processor being communicatively coupled to the at least one memory and configured to execute the method for training a machine learning model according to any one of claims 1 to 14.
28. A system for classifying a temporal phase signal obtained from an optical coherence tomography (OCT) image generated for optical retinal imaging, the system comprising: at least one memory; as well as At least one processor is communicatively coupled to the at least one memory and is configured to perform the method of classifying a time phase signal according to any one of claims 15 to 26.