Acquisition of high-resolution eye measurement parameters
Patent Information
- Application Number
- JP2023568476
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2021-05-03
- Filing Date
- 2022-05-02
- Publication Date
- 2025-05-19
- Estimated Expiration
- 2042-05-02
AI Technical Summary
Existing methods for measuring micro-eye movements are costly, time-consuming, and require controlled laboratory environments, often failing to achieve the necessary resolution for neurological disorder assessments, and do not utilize standardized stimuli effectively.
Utilizing standard cameras to capture video without controlled settings, employing probabilistic methods and machine learning models to enhance resolution and extract high-resolution ocular parameters like eyelid and iris data, and generating digital biomarkers for neurological assessments.
Enables accurate, remote, and distributed measurement of high-resolution ocular parameters, providing objective digital biomarkers for neurological conditions, overcoming the limitations of existing technologies.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[Technical field]
[0001] CROSS-REFERENCE TO RELATED PATENT APPLICATIONS This patent application claims the benefit of U.S. Provisional Patent Application No. 63 / 183,388, entitled “MEASURING HIGH RESOLUTION EYE MOVEMENTS USING BLIND DECONVOLUTION TECHNIQUES,” filed May 3, 2021, which is incorporated by reference in its entirety. [Background technology]
[0002] background Several papers have been published that demonstrate the association between micro-eye movements and the progression of neurological disorders (see, for example, the References section below). Typically, these eye movements are measured in a well-controlled laboratory environment (e.g., no movement, controlled ambient light, or other such parameters) using dedicated equipment (e.g., infrared eye-tracking equipment, pupillometers, or other such equipment), which can be difficult to set up, prohibitively expensive, or involve significant time and effort to create or maintain a controlled environment. Some prior art uses neural networks to obtain these eye movements, but does not achieve the resolution required for the measurements (e.g., about 0.1 mm or higher resolution). Furthermore, the prior art focuses on measuring the eye's response to specific standardized stimuli (e.g., controlled light stimuli, stimuli provided to measure specific parameters, stimuli that the patient must be aware of, or other such stimuli) in controlled ambient conditions. These and other shortcomings exist. Summary of the Invention
[0003] overview The disclosed embodiments relate to systems and methods for facilitating measurement of micro-ocular parameters or eye movements using video captured by a standard camera (e.g., a smartphone camera, a webcam, or another video capture device) without the need for controlled settings. The embodiments acquire various high-resolution eye measurement parameters. For example, the embodiments may acquire eyelid data such as coordinates of the eyelid border. In another example, the embodiments acquire iris data such as iris translation or iris center, iris rotation, iris radius, iris visible ratio, or iris coverage asymmetry. In yet another example, the embodiments acquire pupil data such as pupil center, pupil radius, pupil visible ratio, or pupil coverage asymmetry. The eye measurement parameters are acquired at high resolution, e.g., at sub-pixel level, 0.1 mm, or other higher resolution. In some embodiments, the eye measurement parameters may be used in determining eye movement measurements such as pupil response parameters, gaze, saccades, fixations, etc., which may be used to generate digital biomarkers that may be further used to diagnose or measure the progression of a neurological condition or disorder. Digital biomarkers are objective, sensitive, accurate, correlate with disease progression, and can be done remotely and even in a distributed manner. Ocular measurement or eye movement parameters can be acquired using non-standard stimuli presented to a user on a user device (e.g., stimuli that the patient does not need to be aware of, stimuli that can be used to measure multiple parameters, stimuli such as a video being played on a display device being viewed by the user, or other such stimuli that do not require controlled ambient lighting conditions).
[0004] In some aspects, the disclosed aspects may use probabilistic methods (e.g., maximum likelihood estimation or other such methods) or signal processing methods (e.g., blind deconvolution or other such methods) to increase the resolution and ascertain eye measurement parameters, eye movement parameters at ultra-high resolution (e.g., at sub-pixel level, 0.1 mm, or other higher resolution) against data layers present in a video frame sequence of eye movements captured by a standard camera in an uncontrolled setting. The disclosed aspects may use predictive models (e.g., machine learning (ML) models such as neural networks) to obtain one or more eye measurement parameters at high resolution.
[0005] Various other aspects, features, and advantages of the present invention will become apparent through the detailed description of the present invention and the accompanying drawings. It should also be understood that both the foregoing general description and the following detailed description are exemplary and are not intended to limit the scope of the present invention. As used in this specification and the claims, the singular forms "a," "an," and "the" include plural references unless the context clearly dictates otherwise. Furthermore, as used in this specification and the claims, the term "or" means "and / or" unless the context clearly dictates otherwise. [Brief description of the drawings]
[0006] [Figure 1A] 1 illustrates a system for facilitating measurement of a user's eye measurement parameters and eye movements, consistent with various aspects. [Figure 1B] 1 illustrates a process for extracting high resolution (HR) eye measurement parameters associated with an eye, consistent with various aspects. [Diagram 2] 1 illustrates a machine learning model configured to facilitate prediction of an ocular measurement parameter, according to one or more embodiments. [Diagram 3] 3A and 3B illustrate a device coordinate system and a head coordinate system, respectively, consistent with various embodiments. [Figure 4] 1 is a flow diagram of a process for obtaining eye measurement parameters at high resolution, consistent with various aspects. [Diagram 5] 1 is a flow diagram of a process for deconvolving a video stream to obtain adjusted eye measurement parameters, consistent with various aspects. [Figure 6] FIG. 1 is a block diagram of a computer system that can be used to implement features of the disclosed aspects. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0007] Detailed Description In the following description, for purposes of explanation, numerous specific details are set forth in order to provide a thorough understanding of aspects of the present invention. However, it will be understood by those skilled in the art that aspects of the present invention may be practiced without these specific details or with equivalent configurations. In other instances, well-known structures and devices are shown in block diagram form in order to avoid unnecessarily obscuring aspects of the present invention.
[0008] FIG. 1A illustrates a system 100 for facilitating measurement of a user's eye measurement parameters and eye movements, consistent with various aspects. FIG. 1B illustrates a process 175 for extracting high-resolution (HR) eye measurement parameters associated with an eye, consistent with various aspects. As shown in FIG. 1A, the system 100 may include a computer system 102, a client device 106, or other components. By way of example, the computer system 102 may include any computing device, such as a server, a personal computer (PC), a laptop computer, a tablet computer, a handheld computer, or other computing equipment. The computer system 102 may include a video management subsystem 112, an eye measurement parameter subsystem (OPS) 114, a marker subsystem 116, or other components. The client device 106 may include any type of mobile terminal, fixed terminal, or other device. By way of example, the client device 106 may include a desktop computer, a notebook computer, a tablet computer, a smartphone, a wearable device, or other client device. A user, such as user 140, may utilize, for example, client device 106 to interact with one or more servers or other components of system 100. Camera 120 may capture still or video images of user 140, such as video of the face of user 140. Camera 120 may be integrated with client device 106 (e.g., a smartphone camera) or may be a stand-alone camera (e.g., a webcam or other such camera) connected to client device 106.
[0009] The components of the system 100 may communicate with one or more components of the system 100 via a communications network 150 (e.g., the Internet, a cellular network, a mobile voice or data network, a cable network, a public switched telephone network, or other type of communications network, or a combination of communications networks). The communications network 150 may be a wireless or wired network. As an example, the client device 106 and the computer system 102 may communicate wirelessly.
[0010] It should be noted that while one or more operations are described herein as being performed by particular components of computer system 102, those operations may in some aspects be performed by other components of computer system 102 or other components of system 100. As an example, while one or more operations are described herein as being performed by components of computer system 102, those operations may in some aspects be performed by components of client device 106.
[0011] It should be noted that while some aspects are described herein with respect to machine learning models, in other aspects, other predictive models (e.g., statistical models or other analytical models) may be used instead of or in addition to the machine learning models (e.g., in one or more aspects, a statistical model replaces the machine learning model, and a non-statistical model replaces the non-machine learning model).
[0012] In some embodiments, the system 100 uses a video stream 125 of eye movements (e.g., a sequence of frames of images of the eye) captured by the camera 120 to facilitate the determination of eye measurement parameters 170-178 of the user 140 at high resolution (e.g., sub-pixel level, 0.1 mm or higher resolution). The video management subsystem 112 obtains the video stream 125 from the client device 106 and stores the video stream 125 in the database 132. The OPS 114 processes the video stream 125 to extract eye measurement parameters such as (i) pupil center 176, (ii) pupil radius 178, (iii) iris radius 174, (iv) iris translation 172, and (v) iris rotation 170, among other parameters, for each of the eyes at high resolution. These eye measurement parameters may determine eye movement parameters (e.g., pupil response parameters and gaze parameters) indicative of eye movements in response to non-standard stimuli (e.g., videos or images displayed on the client device 106). The marker subsystem 116 may generate digital biomarkers (e.g., parameters) based on the eye movement parameters that may serve as indicators of one or more neurological disorders.
[0013] The video management subsystem 112 may store the video stream 125 in the database 132 in any of several formats (e.g., WebM, Windows Media Video, Flash Video, AVI, QuickTime, AAC, MPEG-4, or another file format). The video management subsystem 112 may also provide the video stream 125 as an input (e.g., in real time) to the OPS 114 for generating the eye measurement parameters. In some aspects, the video management subsystem 112 may perform pre-processing 191 of the video stream 125 to convert the video stream 125 into a format suitable for eye measurement parameter extraction by the OPS 114. For example, the video management subsystem 112 may perform an extract, transform, and load (ETL) process on the video stream 125 to generate a pre-processed video stream 151 for the OPS 114. The ETL process is typically a multi-phase process in which data is first extracted, then transformed (e.g., cleaned, sanitized, scrubbed), and finally loaded into the target system. Data may be collated from one or more sources and output to one or more destinations. In one example, the ETL process performed by the video management subsystem 112 may include adjusting color and brightness (e.g., white balance) in the video stream 125. In another example, the ETL process may include reducing noise from the video stream 125. In yet another example, the ETL process may include improving the resolution of the video stream 125 based on the multi-frame data, for example, by implementing one or more multi-frame super-resolution techniques that reconstruct a high-resolution image (or sequence) from a sequence of low-resolution images. In some aspects, super-resolution is a digital image processing technique for obtaining a single high-resolution image (or sequence) from multiple blurred low-resolution images.The basic idea of super-resolution is that low-resolution images of the same scene contain different information due to relative sub-pixel shifts, and thus a high-resolution image with higher spatial information can be reconstructed by image fusion. The video management subsystem 112 may perform the above pre-processing operations and other similar pre-processing operations to generate a pre-processed video stream 151 in a format suitable for the OPS 114 to perform ophthalmometry parameter extraction.
[0014] In some aspects, the video management subsystem 112 may obtain the preprocessed video stream 151 (e.g., based on the video stream 125 obtained from the client device 106) via a predictive model. As an example, the video management subsystem 112 may input the video stream 125 obtained from the client device 106 into the predictive model, which then outputs the preprocessed video stream 151. In some aspects, the system 100 may train or configure the predictive model to facilitate the generation of the preprocessed video stream. In some aspects, the system 100 may obtain a video stream from a client device associated with a user (e.g., a video stream 125 having a video of the user's 140 face) and provide such information as input to the predictive model to generate a prediction (e.g., a preprocessed video stream 151 with adjusted color or brightness, improved video resolution, etc.). The system 100 may provide reference feedback to the predictive model, which may update one or more portions of the predictive model based on the prediction and the reference feedback. As an example, if the predictive model generates predictions based on video streams obtained from a client device, one or more preprocessed video streams associated with such input video streams may be provided as reference feedback to the predictive model. As an example, a particular preprocessed video stream may be verified as a suitable video stream for eye measurement parameter extraction by OPS 114 based on one or more user responses (e.g., via user confirmation of the preprocessed video, via one or more subsequent actions that substantiate such a goal, etc.) or one or more other user actions via one or more services. The aforementioned user input information may be provided as input to the predictive model to cause the predictive model to generate predictions for the preprocessed video stream, and the verified preprocessed video stream may be provided as reference feedback to the predictive model to update the predictive model.In this manner, for example, a predictive model may be trained or configured to generate more accurate predictions. In some aspects, the training dataset may include several data items, where each data item includes at least a video acquired from a client device and a corresponding pre-processed video as ground truth.
[0015] In some aspects, the aforementioned operations for updating the predictive model may be performed using a training dataset associated with one or more users (e.g., a training dataset associated with a given user to specifically train or configure a predictive model for the given user, a training dataset associated with a given cluster, population, or other group to specifically train or configure a predictive model for a given group, a training dataset associated with any given set of users, or other training dataset). Thus, in some aspects, following updating the predictive model, system 100 may use the predictive model to facilitate generation of a pre-processed video stream to facilitate ocular measurement parameter extraction by OPS 114.
[0016] Video management subsystem 112 may generate preprocessed video stream 151 before inputting video stream 125 to OPS 114 or after storing video stream 125 in database 132. Additionally, video management subsystem 112 may store preprocessed video stream 151 along with video stream 125 in database 132.
[0017] The OPS 114 may process the video stream 125 or the pre-processed video stream 151 to extract, determine, derive, or generate eye measurement parameters (e.g., eye measurement parameters 170-178) that are characteristics of the eye 145 of the user 140. In some embodiments, the eye measurement parameters include parameters such as (i) pupil center 176, (ii) pupil radius 178, (iii) iris radius 174, (iv) iris translation 172, and (v) iris rotation 170. The pupil center 176 is the center of the pupil and may be represented using coordinates (e.g., a three-dimensional (3D) lateral displacement vector (x,y,z)) that identify the location of the pupil center in a coordinate system. The pupil radius 178 may be the radius of the pupil and may be represented as a scalar having distance units. The iris translation 172 (or iris translation 160) is the center of the iris and may be represented using coordinates (e.g., a 3D lateral displacement vector (x, y, z)) that identify the location of the iris center in a coordinate system. The iris radius 174 (or iris radius 162) may be the radius of the iris and may be represented as a scalar with distance units. In some embodiments, the iris radius 174 is assumed to have a small variance in the human population (e.g., 10%-15% considering the iris radius has a range of 11 mm-13 mm among the human population). The iris rotation 170 (or iris rotation 158) represents the rotation of the iris and is taken as an iris normal vector (e.g., perpendicular to the iris circle plane). In some embodiments, the iris rotation 170 may be projected to a 3D vector (e.g., x, y, z) in the device coordinate system or may be represented by a 2D vector (e.g., azimuth and elevation) in the head coordinate system. In some embodiments, OPS 114 may obtain coordinates of eye measurement parameters in one or more coordinate systems, such as the device coordinate system and the head coordinate system illustrated in Figures 3A and 3B, respectively.
[0018] 3A illustrates a device coordinate system consistent with various aspects. In some aspects, the device coordinate system is a Cartesian coordinate system, and the center of the display 302 of a client device (e.g., client device 106) is considered to be the origin 304 with (x,y,z) coordinates as (0,0,0). The device coordinate system defines three orthogonal axes, e.g., an x-axis in the direction of a horizontal edge of the display 302, a y-axis in the direction of a vertical edge of the display 302 and perpendicular to the x-axis, and a z-axis in the direction perpendicular to the plane of the display 302.
[0019] FIG. 3B illustrates a head coordinate system consistent with various embodiments. In some embodiments, the head coordinate system is defined by a plane that includes the following points in the face bounding polygon: the leftmost, rightmost, and topmost (points 1, 2, and 3 in FIG. 3B) detectable points in the face polygon. For example, the topmost point 3 may be defined as the topmost detectable point in the face bounding polygon and may be located in the cross section between the y-axis and the head. When acquiring a face image, the horizontal axis (x) may be defined as the line between the leftmost and rightmost points, and the perpendicular vector to that line on the defined plane may be defined as the vertical axis (y). The depth axis (z) may be orthogonal to both the x-axis and the y-axis. The origin of the head coordinate system (point 0 in FIG. 3B) is located centered between the eyes, projected onto the defined plane by a displacement in the z-direction. In some embodiments, iris rotation may be represented on the head coordinate system as eye rotation, given two angles (e.g., azimuth and elevation).
[0020] With respect to the OPS 114, the OPS 114 may be configured to represent coordinates of the ocular measurement parameters in any of a number of coordinate systems, including a head coordinate system or a device coordinate system, and further, the OPS 114 may be configured to transform coordinates from one coordinate system to another, such as from the head coordinate system to the device coordinate system or vice versa.
[0021] OPS 114 may also acquire eyelid data 164, such as coordinates of the upper and lower eyelid borders of both eyes. Eyelid data 164 may be used in determining or improving the accuracy of eye measurement parameters. In some aspects, OPS 114 may acquire the above data (e.g., eye measurement parameters 170-178, eyelid data 164, or other data) as time series data. For example, OPS 114 may extract a first set of values of eye measurement parameters 170-178 at a first time point in video stream 125, a second set of values of eye measurement parameters 170-178 at a second time point in video stream 125, etc. That is, OPS 114 may continuously extract eye measurement parameters 170-178 at a specified time interval (e.g., every millisecond, every few milliseconds, or other time resolution). Additionally, OPS 114 may acquire the above data for one or both of eyes 145 of user 140.
[0022] OPS 114 may perform several processes to extract the above eye measurement parameters (e.g., eye measurement parameters 170-178) from preprocessed video stream 151. For example, in extraction process 192, OPS 114 may obtain a first set of eye measurement parameters, such as iris data 153, eyelid data 155, and pupil center 157. In some aspects, OPS 114 may use computer vision techniques on preprocessed video stream 151 to model the translation and rotation of user 140's face relative to camera 120 and identify the eyes and mouth of the face. After identifying the eye locations, OPS 114 may obtain eyelid data 155 (e.g., coordinates of an array of points describing a polygon representing the eyelids) for each frame (e.g., using the fact that their edges can be approximated with quadratic functions). In some aspects, OPS 114 may determine iris data 153 (e.g., a shape such as an ellipsoid representing the iris) by determining which pixels are within or outside the iris radius at any time based on the position of the eyelid as a time series. In some aspects, OPS 114 may be configured to consider the iris and pupil to be ellipsoids having a relatively uniform and dark color compared to the sclera (e.g., the white portion of the eye) when determining which pixels are within or outside. Iris data 153 may include several coordinates, such as first and second coordinates corresponding to the left and right coordinates of a shape representing the iris (e.g., an ellipse), third and fourth coordinates corresponding to the top and bottom coordinates of the shape, and a fifth coordinate corresponding to the center of the iris. Pupil center 157 may also be determined based on the time series of eyelid data 155 and iris data 153. In some aspects, pupil center 157 may include coordinates representing the location of the center of the pupil.
[0023] In some aspects, OPS 114 may implement a predictive model in extraction process 192 and obtain a first eye measurement parameter set (e.g., based on video stream 125 or preprocessed video stream 151) through the predictive model. As an example, OPS 114 may input video stream 125 or preprocessed video stream 151 into the predictive model, which then outputs the first eye measurement parameter set. In some aspects, system 100 may train or configure a predictive model to facilitate generation of the first eye measurement parameter set. In some aspects, system 100 may obtain a video stream having a video of a user's face (e.g., video stream 125 or preprocessed video stream 151) and provide such information as input to the predictive model to generate a prediction (e.g., a first eye measurement parameter set, such as iris data, eyelid data, pupil center, etc.). System 100 may provide reference feedback to the predictive model, which may update one or more portions of the predictive model based on the prediction and the reference feedback. As an example, if the predictive model generates predictions based on the video stream 125 or the preprocessed video stream 151, a first of the eye measurement parameters associated with such input video stream may be provided as reference feedback to the predictive model. As an example, a particular eye measurement parameter set may be verified as a suitable eye measurement parameter set (e.g., via user confirmation of the eye measurement parameter set, via one or more subsequent actions that substantiate such a goal, etc.). Such user input information may be provided as input to the predictive model to cause the predictive model to generate predictions of the first eye measurement parameter set, and the verified eye measurement parameter set may be provided as reference feedback to the predictive model to update the predictive model. In this manner, for example, the predictive model may be trained or configured to generate more accurate predictions.
[0024] In some embodiments, the aforementioned operations for updating the predictive model may be performed using a training dataset associated with one or more users (e.g., a training dataset associated with a given user to specifically train or configure a predictive model for the given user, a training dataset associated with a given cluster, population, or other group to specifically train or configure a predictive model for a given group, a training dataset associated with any given set of users, or other training dataset). Thus, in some embodiments, following updating the predictive model, the system 100 may use the predictive model to facilitate generation of a first set of eye measurement parameters.
[0025] In some aspects, OPS 114 may also further adjust, modify, or correct the first set of parameters based on other factors, such as optometric data associated with the user. For example, user 140 may have a particular vision condition, such as myopia, hyperopia, astigmatism, or other vision condition, and the user may wear corrective lenses, such as glasses or contact lenses. OPS 114 may take such optometric conditions of user 140 into account and correct, adjust, or modify one or more values of the first set of eye measurement parameters. In some aspects, OPS 114 may obtain the optometric data associated with user 140 from user profile data (e.g., stored in database 132).
[0026] In some embodiments, OPS 114 may further process 193 the iris data 153 to extract iris-related parameters such as iris rotation 158, iris translation 160, and iris radius 162. In some embodiments, the above iris-related parameters are obtained using ML techniques. For example, OPS 114 may perform a maximum likelihood estimation (MLE)-based curve-fitting method to estimate the iris-related parameters. The iris data is provided as input to an MLE-based curve-fitting method (e.g., ellipse fitting, based on the assumption that the iris is circular and therefore oval / elliptical when projected onto a 2D plane) that generates the above iris-related parameters. In some embodiments, in statistics, MLE is a method of estimating parameters of a probability distribution by maximizing a likelihood function, such that the observed data is most likely under an assumed statistical model. The point in the parameter space that maximizes the likelihood function is referred to as the maximum likelihood estimate. In some embodiments, the curve-fitting operation is a process of constructing a curve or mathematical function that has an optimal solution to a set of data points, subject to constraints (e.g., ellipse fitting). Curve fitting may involve either interpolation, where an exact fit to the data is required, or smoothing, where a "smooth" function is constructed that approximately fits the data. Fitted curves may be used as an aid in data visualization, to infer values of functions for which data are not available, and to summarize relationships between two or more variables.
[0027] In some embodiments, the accuracy of at least some of the iris-related parameters may be further improved by performing an MLE-based curve fitting operation. For example, a shape describing the iris boundary may be further refined from the coordinates of the shape in the iris data 153, and thus an improved or more accurate iris radius 162 may be determined. In some embodiments, the iris translation 160, also referred to as the iris center, is similar to the coordinates of the iris center in the iris data 153. In some embodiments, in process 193, OPS 114 may consider different iris centers, construct an iris shape for each of the candidate iris centers, and assign a confidence score to each of the shapes. Such a method may be repeated for different iris centers, and the shape with a score (e.g., the best score) that meets the score criteria is selected. In some embodiments, the score of each iteration may reflect the log-likelihood of the examined parameter set (e.g., iris shape, radius, and center coordinates) based on known physical constraints and assumptions. In some embodiments, the physical constraints and assumptions may improve the efficiency of the MLE process by introducing better estimated initial conditions before iterating toward convergence. For example, some assumptions may include that the iris is circular (represented as an ellipse / oval when projected onto a plane in 2D, that the iris center is expected to be at the center of gravity of the circle, high contrast between the iris and the sclera, similarity of orientation between the left and right eye, stimulus-based assumptions regarding the expected point of gaze, and brightness of the display of the client device on which the video stimuli are shown to the user. In some aspects, examples of physical constraints may include pupil dilation dependency on overall brightness (light conditions), a limited range of facial orientations when the user views the display, blinking may cause uncertainty due to realignment time, or physically possible range of motion (both face and eyeball).
[0028] The OPS 114 may perform a deconvolution process 194 on the video stream 125 or the preprocessed video stream 151 to adjust (e.g., improve resolution or increase accuracy) some eye measurement parameters, such as eyelid data 155 and pupil center 166, and determine other eye measurement parameters, such as pupil radius. In some aspects, the OPS 114 may perform a blind deconvolution process to improve the accuracy of the eye measurement parameters. In image processing, blind deconvolution is a deconvolution technique that allows the reconstruction of a target scene from a single "blurred" image or a set of "blurred" images in the presence of a poorly determined or unknown point spread function (PSF). Conventional linear and nonlinear deconvolution techniques may utilize a known PSF. For blind deconvolution, the PSF is estimated from an image or set of images, allowing the deconvolution to be performed. The blind deconvolution may be performed iteratively, whereby each iteration improves the PSF and scene estimates, or non-iteratively, where a single application of the algorithm extracts the PSF based on external information. After determining the PSF, the PSF may be used in deconvolving the video stream 125 or the preprocessed video stream 151 to obtain adjusted ophthalmometry parameters (e.g., adjusted eyelid data 164, adjusted pupil center 166, or pupil radius 168).
[0029] In some aspects, applying a blind deconvolution algorithm may help reduce or remove blur from the image / video and estimate irradiance (e.g., based on parameters of the estimated convolution vector (e.g., PSF)). In some aspects, the estimated irradiance is the estimated irradiance reflected from the eye of the user 140, which is an indication of the irradiance to which the eye is exposed. After the blur is reduced or removed and the irradiance is obtained, the deconvolution process may be repeated one or more times to obtain adjusted eye measurement parameters (e.g., whose accuracy or resolution is improved compared to the eye measurement parameters before the deconvolution process 194).
[0030] In some embodiments, the blind deconvolution process considers various factors when determining the PSF. For example, the OPS 114 may input (a) a video stream 125 or a preprocessed video stream 151 (e.g., a time-series image sequence) with high temporal resolution (e.g., 30 ms or faster per frame), (b) stimulus data, such as spatiotemporal information about the stimuli presented on the display of the client device 106 and optical property information about the stimuli, including spectral characteristics (e.g., color) and intensity (e.g., brightness), (c) environmental data, such as lighting in the environment (e.g., room) in which the user is located, which can be measured using information obtained from the camera 120, and (d) device data, such as orientation information of the client device 106, information from one or more sensors associated with the client device 106, such as an acceleration sensor. Such factors aid in efficient calculation of the PSF, minimize noise uncertainty, and lead to better accuracy of the overall deconvolution.
[0031] In some embodiments, the deconvolution process 194 is integrated with an MLE process to further improve the accuracy of the eye measurement parameters. For example, the MLE process may be used to improve the accuracy of the eyelid data 155. As described above, the eyelid data 155 includes coordinates of an array of points that describe the shape of the eyelid (e.g., a polygon). The MLE process performs a parabolic curve fitting operation on the eyelid data 155 under the constraint that the shape of the eyelid is parabolic to obtain a more accurate representation of the eyelid as the adjusted eyelid data 164. In some embodiments, the adjusted eyelid data 164 may include coordinates of a set of points that describe the shape of the eyelid (e.g., a parabola). Such adjusted eyelid data 164 may be obtained for both eyelids of an eye and for both eyes. In some embodiments, OPS 114 may use adjusted eyelid data 164 information to predict precise eye measurement parameter values, such as pupil radius center, even when some are physically difficult (e.g., when the upper eyelid covers the pupil center or when only a portion of the iris is exposed to the camera).
[0032] In some aspects, the MLE process may also be used to obtain or refine pupil-related data, such as pupil radius 168. In some aspects, the accuracy of the pupil-related data may be further improved by performing an MLE-based curve-fitting operation. In some aspects, in the MLE process, OPS 114 may consider different pupil centers, construct a pupil shape for each of the candidate pupil centers, and assign a confidence score to each of the shapes. Such a method may be repeated for different pupil centers, and the shape with a score (e.g., the best score) that meets a score criterion is selected. Once a shape is selected, the corresponding center may be selected as the adjusted pupil center 166, and the pupil radius 168 may be determined based on the selected shape and the adjusted pupil center 166. In some aspects, the score of each iteration may reflect the log-likelihood of the examined parameter set (e.g., pupil radius and center) based on known physical constraints and assumptions. In some aspects, the physical constraints and assumptions may improve the efficiency of the MLE process by introducing better estimated initial conditions before iterating toward convergence. For example, some assumptions may include that the pupil is circular (represented as an ellipse / oval when projected onto a 2D plane, that the pupil center is expected at the center of gravity of the circle, the degree of orientation similarity between the left and right eye, stimulus-based assumptions regarding the expected point of gaze, and the brightness of the display of the client device on which the video stimuli are shown to the user. In some aspects, examples of physical constraints may include pupil dilation dependency on overall brightness (light conditions), a limited range of facial orientations when the user views the display, that blinking may cause uncertainty due to realignment time, or a physically possible range of motion (both face and eye). Thus, OPS 114 may obtain adjusted eye measurement parameters such as adjusted eyelid data 164, adjusted pupil center 166, and pupil radius 168 from the deconvolution process 194.In some embodiments, the above process 193 for processing iris data and deconvolution process 194 may output the following eye measurement parameters at a first resolution: adjusted eyelid data 164, adjusted pupil center 166 and pupil radius 168, iris rotation 158, iris translation 160 and iris radius 162 (also referred to as the "first set of eye measurement parameters"). For example, the first resolution may be pixel level or other lower resolution.
[0033] In some embodiments, OPS 114 may further improve the resolution of the first eye measurement parameter set. For example, OPS 114 may improve the resolution from a first resolution to a second resolution (e.g., 0.1 mm, sub-pixel level, or some other resolution higher than the first resolution). In some embodiments, OPS 114 may input the first eye measurement parameter set to a resolution improvement process 195 to obtain a second eye measurement parameter set (e.g., eye measurement parameters 170-178) at the second resolution. In some embodiments, the resolution improvement process 195 may obtain the second eye measurement parameter set via a predictive model (e.g., based on the first eye measurement parameter set). As an example, OPS 114 may input the first eye measurement parameter set obtained at the first resolution to a predictive model, which then outputs the second eye measurement parameter set at the second resolution. In some embodiments, system 100 may train or configure the predictive model to facilitate generation of the second eye measurement parameter set. In some aspects, the system 100 may obtain input data such as (a) a video stream having video of the user's face (e.g., video stream 125 or preprocessed video stream 151); (b) a first set of eye measurement parameters (obtained at a first resolution as described above); (c) environmental data such as lighting in the environment (e.g., a room) in which the user is located, which can be measured using information obtained from the camera 120; (d) device data such as display size, display resolution, display brightness, or display contrast associated with the display of the client device 106, model and manufacturer information of the camera 120, or (e) user information such as demographics, medical history, optometric data, etc.
[0034] The system 100 may provide such input data to the predictive model to generate predictions (e.g., a second set of eye measurement parameters, such as iris rotation 170, iris translation 172, iris radius 174, pupil center 176, and pupil radius 178 at the second resolution). The system 100 may provide reference feedback to the predictive model, and the predictive model may update one or more portions of the predictive model based on the predictions and the reference feedback. As an example, if the predictive model generates predictions based on the above input data, the second eye measurement parameter set associated with such input data may be provided as reference feedback to the predictive model. As an example, a particular eye measurement parameter set acquired at the second resolution may be verified as a suitable eye measurement parameter set (e.g., via user confirmation of the eye measurement parameter set, via one or more subsequent actions that substantiate such a goal, etc.). The aforementioned user input information may be provided as input to the predictive model to cause the predictive model to generate predictions of the second eye measurement parameter set, and the verified eye measurement parameter set may be provided as reference feedback to the predictive model to update the predictive model. In this manner, for example, the predictive model may be trained or configured to generate more accurate predictions. In some aspects, the reference feedback having eye measurement parameters at a second resolution may be obtained, determined, or derived from information acquired using any of several eye tracking devices that generate eye measurement parameters at a high resolution (e.g., the second resolution). For example, some tracking devices generate eye measurement parameters such as gaze origin, gaze point, and pupil diameter at a second resolution. OPS 114 may derive a second set of eye measurement parameters, such as iris rotation, iris translation, iris radius, pupil center, and pupil radius, from the eye measurement parameters generated using the eye tracking device, and provide the derived second set of eye measurement parameters as reference feedback to train the predictive model. Such reference feedback may be acquired for several videos and provided as a training data set to train the predictive model.
[0035] In some embodiments, the aforementioned operations for updating the predictive model may be performed using a training dataset associated with one or more users (e.g., a training dataset associated with a given user to specifically train or configure a predictive model for the given user, a training dataset associated with a given cluster, population, or other group to specifically train or configure a predictive model for a given group, a training dataset associated with any given set of users, or other training dataset). Thus, in some embodiments, following updating the predictive model, the system 100 may use the predictive model to facilitate generation of a first set of eye measurement parameters.
[0036] In some embodiments, OPS 114 may also extract additional eye measurement parameters such as pupil visible ratio, pupil coverage asymmetry, iris visible ratio, or iris coverage asymmetry. In some embodiments, pupil visible ratio is calculated as the ratio of pupil area not covered by eyelid to pupil iris area. Pupil coverage asymmetry may be defined as the average of pupil upper eyelid coverage ratio and pupil lower eyelid coverage ratio, normalized by total iris coverage area, with upper eyelid coverage ratio represented as a positive value and lower eyelid coverage ratio represented as a negative value. The value of this parameter may vary from -1 to 1 to project the asymmetry of eyelid coverage between the upper and lower eyelids (e.g., "-1" may represent that all covered area is covered by the lower eyelid, "1" may represent that all covered area is covered by the upper eyelid, and "0" may represent that the upper and lower eyelids cover equal areas).
[0037] In some embodiments, the iris visible ratio is calculated as the ratio of the iris area not covered by the eyelid to the total area of the iris. In some embodiments, the iris coverage asymmetry can be determined as the average of the iris upper eyelid coverage ratio and the iris lower eyelid coverage ratio, normalized by the total iris coverage area, with the upper eyelid coverage ratio represented as a positive value and the lower eyelid coverage ratio represented as a negative value. The value of this parameter varies from -1 to 1, projecting the asymmetry of eyelid coverage between the upper and lower eyelids (e.g., "-1" is all the coverage area covered by the lower eyelid, "1" is all the coverage area covered by the upper eyelid, and "0" is the upper and lower eyelids covering equal areas).
[0038] The OPS 114 may extract additional eye measurement parameters based on the second eye measurement parameter set 170-178. For example, the OPS 114 may perform geometric projections and calculations using the second eye measurement parameter set 170-178 to determine the additional eye measurement parameters.
[0039] In some aspects, OPS 114 may also extract, generate, or derive eye movement parameters (e.g., pupil response parameters and gaze parameters) indicative of eye movement in response to non-standard stimuli (e.g., videos or images displayed on client device 106). The eye movement parameters may be derived using a second set of eye measurement parameters (e.g., obtained as described above). In some aspects, the pupil response parameters indicate the eye's response to a particular stimulus, and the gaze parameters indicate where the eye is focusing or looking. OPS 114 may obtain different types of gaze parameters, such as fixation, saccade, pursuit, or another gaze parameter. In some aspects, fixation is defined as the distribution of eye movement while inspecting a particular region of a stimulus. In some aspects, saccade is defined as the distribution of eye movement between inspected regions. In some aspects, pursuit is defined as the distribution of eye movement during a pursuit movement. These eye movement parameters may be used in generating various digital biomarkers.
[0040] In some embodiments, the marker subsystem 116 may generate digital biomarkers (e.g., parameters) that may serve as indicators of one or more neuropathy. The marker subsystem 116 may generate digital biomarkers based on ocular measurement parameters or eye movement parameters (e.g., as described above). In some embodiments, these digital biomarkers are strongly correlated with disease progression or severity (and thus can serve as surrogates). The digital biomarkers generated by the marker subsystem 116 may be objective, sensitive, accurate (0.1 mm or less) or correlated with clinical progression of the disease, and may be acquired remotely, outside of a laboratory environment, and even in a distributed manner.
[0041] It should be noted that the eye measurement parameters, eye movement parameters, or digital biomarkers may be acquired in real time (e.g., based on a real-time video stream acquired from the camera 120) or offline (e.g., using a video stored in the database 132). Furthermore, the eye measurement parameters (e.g., the second set of eye measurement parameters 170-178 or other eye measurement parameters) are acquired for one or more eyes of the user as time series data. That is, the eye measurement parameters 170-178 may be extracted continuously at a specified time interval (e.g., every millisecond, every few milliseconds, or other time resolution).
[0042] In some embodiments, the predictive model described above may include one or more neural networks or other machine learning models. As an example, the neural network may be based on a large collection of neural units (or artificial neurons). The neural network may roughly mimic the way that a biological brain works (e.g., via a large cluster of biological neurons connected by axons). Each neural unit of the neural network may be connected with many other neural units of the neural network. Such connections may enhance or suppress their influence on the activation state of the connected neural units. In some embodiments, each individual neural unit may have a summing function that combines the values of all its inputs. In some embodiments, each connection (or the neural unit itself) may have a threshold function such that a signal must exceed a threshold before it propagates to other neural units. These neural network systems may be self-learning and trained rather than explicitly programmed, and can perform significantly better in certain areas of problem solving compared to traditional computer programs. In some embodiments, the neural network may include multiple layers (e.g., where a signal path traverses from a front layer to a back layer). In some aspects, backpropagation techniques may be utilized by the neural network, where forward stimuli are used to reset the weights of the "forward" neural units. In some aspects, the neural network's stimuli and inhibitions may flow more freely, with connections interacting in more chaotic and complex ways.
[0043] A neural network may be trained (i.e., its parameters are determined) using a training data set (e.g., ground truth). The training data may include a set of training samples. Each sample may be a pair including an input object (typically an image, measurement, tensor or vector, which may be called a feature tensor or vector) and a desired output value (also called a teacher signal). A training algorithm analyzes the training data and adjusts the behavior of the neural network by adjusting the parameters of the neural network (e.g., the weights of one or more layers) based on the training data. For example, x i is the feature tensor / vector of the i-th example, and y i is the teacher signal {(x 1 ,y 1 ),(x 2 ,y 2 ),…(x N ,y N Given a set of N training samples of the form {N, ...
[0044] As an example, with respect to FIG. 2, the machine learning model 202 may receive an input 204 and provide an output 206. In one use case, the output 206 may be fed back to the machine learning model 202 as an input to train the machine learning model 202 (e.g., alone or in combination with a user indication of accuracy of the output 206, a label associated with the input, or other reference feedback information). In another use case, the machine learning model 202 may update its configuration (e.g., weights, biases, or other parameters) based on an evaluation of its prediction (e.g., the output 206) and the reference feedback information (e.g., a user indication of accuracy, a reference label, or other information). In another use case, if the machine learning model 202 is a neural network, it may adjust connection weights to adjust for the difference between the neural network's prediction and the reference feedback. In a further use case, one or more neurons (or nodes) of the neural network may require that the neuron's respective errors be sent back to them through the neural network to facilitate the update process (e.g., backpropagation of errors). The connection weight updates may, for example, reflect the magnitude of the error that is propagated backwards after a forward pass is completed. In this manner, for example, the machine learning model 202 may be trained to generate better predictions.
[0045] 4 is a flow diagram of a process 400 for acquiring eye measurement parameters at high resolution, consistent with various embodiments. In some embodiments, the process 400 may be performed in the system 100 of FIG. 1A.
[0046] At operation 402, a video stream is obtained from a client device associated with a user. The video stream may include a video of the user's face. For example, video stream 125 is obtained from client device 106. In some aspects, the video stream may be obtained in real time from the client device, for example, captured by a camera associated with the client device, such as camera 120 of client device 106, or may be a recorded video obtained from a data store, such as database 132, a data store or memory on the client device, or other data store.
[0047] In operation 404, the video stream is input to the first predictive model to obtain a first set of eye measurement parameters. In some aspects, the video stream may be pre-processed before being input to the first predictive model. In some aspects, pre-processing the video stream may include performing an ETL process on the video stream to adjust the color and brightness (e.g., white balance) of the video stream, reduce noise from the video stream, and improve the resolution of the video stream by performing one or more multi-frame super-resolution techniques, or other such video processing. For example, the video stream input to the first predictive model may be the video stream 125 or the pre-processed video stream 151.
[0048] In operation 406, a first set of eye measurement parameters is obtained from the first predictive model. The eye measurement parameters correspond to various characteristics of the eye. For example, the first set of eye measurement parameters may include iris data, eyelid data, and a pupil center. The iris data may include coordinates of several points describing the center and shape (e.g., an ellipse) of the iris. The eyelid data may include coordinates of an array of points describing a shape (e.g., a polygon) representing the eyelid. The pupil center may include coordinates of a point representing the center of the pupil. In some aspects, the first set of eye measurement parameters is obtained at a first resolution. For example, the first resolution may be at the pixel level or other lower resolution. In some aspects, the iris data is further processed to obtain an iris rotation, an iris translation, and an iris radius from the iris data. For example, the iris data 153 is processed to obtain an iris rotation 158, an iris translation 160, and an iris radius 162. The iris rotation may represent the rotation of the iris and may be taken as an iris normal vector (e.g., perpendicular to the iris circle plane). The iris translation may be the center of the iris and may be expressed using coordinates that identify the location of the iris center. The iris radius may be the radius of the iris and may be expressed as a scalar with distance units. In some aspects, an MLE-based curve fitting method may be performed to estimate the iris-related parameters. The iris data is provided as an input to an MLE-based curve fitting method (e.g., an ellipse fit) that generates the above iris-related parameters.
[0049] In operation 408, the video stream is deconvolved to adjust (e.g., improve accuracy or resolution) the first set of eye measurement parameters. The video stream input to the deconvolution process may be the video stream 125 or the preprocessed video stream 151. In some aspects, the video stream is deconvolved using a blind deconvolution process integrated with an MLE process, further details of which are described above with reference to at least FIG. 1A and FIG. 1B and below with reference to at least FIG. 5. In some aspects, applying a blind deconvolution algorithm may help reduce or remove blurring in the image / video or any other noise from the video stream to improve the accuracy of the extracted eye measurement parameters. The deconvolution process may be repeated one or more times to obtain adjusted eye measurement parameters (e.g., whose accuracy or resolution is improved compared to the eye measurement parameters before the deconvolution process). The adjusted parameters may include adjusted eyelid data, adjusted pupil center and pupil radius.
[0050] In operation 410, the video stream and the first parameter set are input to a second predictive model to obtain eye measurement parameters at high resolution. For example, the first eye measurement parameter set including iris rotation 158, iris translation 160, iris radius 162, adjusted eyelid data 164, adjusted pupil center 166 and pupil radius 168, etc. are input to the second predictive model. In some aspects, the second predictive model is trained to predict eye measurement parameter sets at high resolution. For example, the second predictive model is trained with several training data sets, each data set including a video of a user's face, device data of a client device, environmental data of an environment in which the user is located, and user information of a user as input data, and a corresponding eye measurement parameter set at high resolution as ground truth. In some aspects, the high resolution eye measurement parameter set is obtained using one or more eye tracking devices configured to generate eye measurement parameters at high resolution.
[0051] In operation 412, a second set of eye measurement parameters is obtained from the second predictive model at a second resolution, for example, a second set of eye measurement parameters including iris rotation 170, iris translation 172, iris radius 174, pupil center 176, and pupil radius 178, at a second resolution (e.g., 0.1 mm, sub-pixel level, or some other resolution higher than the first resolution).
[0052] Optionally, in operation 414, additional eye measurement parameter sets may be obtained. For example, additional eye measurement parameters such as pupil visible ratio, pupil coverage asymmetry, iris visible ratio, or iris coverage asymmetry may be obtained based on the second eye measurement parameter set. OPS 114 may obtain the additional eye measurement parameters by performing geometric projection and calculation using the second eye measurement parameter set. In some embodiments, pupil visible ratio is calculated as the ratio of pupil area not covered by eyelid to pupil iris area. Pupil coverage asymmetry may be defined as the average of pupil upper eyelid coverage ratio and pupil lower eyelid coverage ratio, normalized by the total iris coverage area, with upper eyelid coverage ratio expressed as a positive value and lower eyelid coverage ratio expressed as a negative value. The value of this parameter varies from -1 to 1 and may project the asymmetry of eyelid coverage between the upper and lower eyelids (e.g., "-1" may represent that all covered area is covered by the lower eyelid, "1" may represent that all covered area is covered by the upper eyelid, and "0" may represent that the upper and lower eyelids cover equal areas).
[0053] Operations 402-414 may be performed by the same or similar subsystem as OPS 114, according to one or more aspects.
[0054] 5 is a flow diagram of a process 500 for deconvolving a video stream to obtain adjusted eye measurement parameters, consistent with various aspects. In some aspects, process 500 may be performed as part of operation 408 of process 400.
[0055] In operation 502, input data such as a video stream, stimulus data, environmental data, device data or other data is obtained to perform deconvolution of the video stream. The video stream input to the deconvolution process may be the video stream 125 or the preprocessed video stream 151. The stimulus data may include spatiotemporal information about the stimulus presented on the display of the client device 106, or optical property information of the stimulus including spectral characteristics (e.g., color) and intensity (e.g., brightness). The environmental data may include information such as lighting in the environment (e.g., room) in which the user is located, which can be measured using information obtained from the camera 120. The device data may include information such as orientation information of the client device 106, information from one or more sensors associated with the client device 106, such as an acceleration sensor. Additionally, the input data may include eyelid data (e.g., eyelid data 155) and pupil center (e.g., pupil center 157).
[0056] In operation 504, a point spread function of a blind deconvolution algorithm is determined based on the input data.
[0057] In operation 506, the video stream is deconvolved based on the point spread function. The deconvolution process removes or reduces any blurring or any other noise introduced into the video stream due to user environment related factors, client device related factors, camera related factors, video stimulus related factors, etc., to improve the video stream and thus improve the accuracy of the extracted eye measurement parameters.
[0058] In operation 508, adjusted eye measurement parameters are derived from the deconvolved video stream as described above. For example, the adjusted eye measurement parameters may include adjusted eyelid data 164, adjusted pupil center 166, or pupil radius 168. In some embodiments, the deconvolution process is integrated with an MLE process to further improve the accuracy of the eye measurement parameters. For example, the MLE process may be used to improve the accuracy of the eyelid data 155. The MLE process may be used to perform a parabolic curve fitting operation on the eyelid data 155 to obtain a more accurate representation of the eyelid as the adjusted eyelid data 164. In another example, the MLE process may be used to obtain or improve the accuracy of pupil-related data, such as the pupil radius 168. The MLE process may be used to perform a curve fitting operation (e.g., an ellipse fit) to improve the shape of the pupil. For example, the MLE process may consider different pupil centers and construct a pupil shape for each of the candidate pupil centers and assign a confidence score to each of the shapes. Such a method may be repeated for different pupil centers, and the shape having a score (e.g., the best score) that meets the score criteria is selected. Once a shape is selected, the corresponding center may be selected as the adjusted pupil center 166, and a pupil radius 168 may be determined based on the selected shape and the adjusted pupil center 166.
[0059] In some embodiments, process 500 may be repeated one or more times to improve the accuracy or resolution of the adjusted eye measurement parameter (e.g., until the accuracy or resolution of the eye measurement parameter meets a criterion (e.g., exceeds a threshold) or the PSF meets a criterion).
[0060] In some embodiments, the various computers and subsystems illustrated in FIG. 1A may include one or more computing devices programmed to perform the functions described herein. A computing device may include one or more electronic storage devices (e.g., training data database(s) 134 for storing training data, model database(s) 136 for storing predictive models, etc., or prediction database(s) 132, which may include other electronic storage devices), one or more physical processors programmed with one or more computer program instructions, and / or other components. A computing device may include communication lines or ports that enable exchange of other computing platform information in a network (e.g., network 150) or via wired or wireless technology (e.g., Ethernet, fiber optic, coaxial cable, Wi-Fi, Bluetooth, near field communication, or other technology). A computing device may include multiple hardware, software, and / or firmware components operating together. For example, a computing device may be implemented by a cloud of computing platforms operating together as a computing device.
[0061] The electronic storage device may include a non-transitory storage medium that electronically stores information. The storage medium of the electronic storage device may include one or both of (i) system storage that is integral with the server or client device (e.g., substantially non-removable), or (ii) removable storage that is removably connectable to the server or client device, for example, via a port (e.g., USB port, FireWire port, etc.) or drive (e.g., disk drive, etc.). The electronic storage device may include one or more of an optically readable storage medium (e.g., optical disk, etc.), a magnetically readable storage medium (e.g., magnetic tape, magnetic hard drive, floppy drive, etc.), a charge-based storage medium (e.g., EEPROM, RAM, etc.), a solid-state storage medium (e.g., flash drive, etc.), and / or other electronically readable storage medium. The electronic storage device may include one or more virtual storage resources (e.g., cloud storage, virtual private network, and / or other virtual storage resources). The electronic storage device may store software algorithms, information determined by a processor, information obtained from a server, information obtained from a client device, or other information that enables the functionality described herein.
[0062] A processor may be programmed to provide information processing capabilities in a computing device. Thus, a processor may include one or more of a digital processor, an analog processor, a digital circuit designed to process information, an analog circuit designed to process information, a state machine, and / or other mechanisms for electronically processing information. In some aspects, a processor may include multiple processing units. These processing units may be physically located in the same device, or a processor may represent the processing functions of multiple devices operating in concert. A processor may be programmed to execute computer program instructions to perform the functions described herein of subsystems 112-116 or other subsystems. A processor may be programmed to execute computer program instructions by software, hardware, firmware; some combination of software, hardware, or firmware; and / or other mechanisms for configuring processing capabilities on the processor.
[0063] The description of the functionality provided by the different subsystems 112-116 described herein is for purposes of illustration and not intended to be limiting, and it should be understood that any of the subsystems 112-116 may provide more or less functionality than described. For example, one or more of the subsystems 112-116 may be eliminated and some or all of its functionality may be provided by other of the subsystems 112-116. As another example, additional subsystems may be programmed to perform some or all of the functionality attributed herein to one of the subsystems 112-116.
[0064] 6 is a block diagram of a computer system that may be used to implement features of the disclosed aspects. Computer system 600 may be used to implement any of the entities, subsystems, components, or services illustrated in the example diagrams above (as well as any other components described herein). Computer system 600 may include one or more central processing units ("processors") 605, memory 610, input / output devices 625 (e.g., keyboard and pointing devices, display devices), storage devices 620 (e.g., disk drives), and network adapters 630 (e.g., network interfaces) connected to an interconnect 615. Interconnect 615 is illustrated as an abstraction representing any one or more separate physical buses, point-to-point connections, or both connected by appropriate bridges, adapters, or controllers. Thus, the interconnect 615 may include, for example, a system bus, a Peripheral Component Interconnect (PCI) bus or PCI-Express bus, a HyperTransport or Industry Standard Architecture (ISA) bus, a Small Computer System Interface (SCSI) bus, a Universal Serial Bus (USB), an IIC (I2C) bus, or the Institute of Electrical and Electronics Components (IEEE) standard 1394 bus, also known as "FireWire."
[0065] The memory 610 and the storage device 620 are computer-readable storage media that can store instructions that implement at least a portion of the described embodiments. In addition, the data structures and message structures can be stored or transmitted via a data transmission medium, such as a signal on a communication link. Various communication links can be used, such as the Internet, a local area network, a wide area network, or a point-to-point dial-up connection. Thus, the computer-readable medium can include a computer-readable storage medium (e.g., a "non-transitory" medium) and a computer-readable transmission medium. The storage medium of the electronic storage device can include one or both of (i) system storage that is integral with the server or client device (e.g., substantially non-removable), or (ii) removable storage that is removably connectable to the server or client device, for example, via a port (e.g., a USB port, a FireWire port, etc.) or a drive (e.g., a disk drive, etc.). The electronic storage may include one or more of an optically readable storage medium (e.g., optical disk, etc.), a magnetically readable storage medium (e.g., magnetic tape, magnetic hard drive, floppy drive, etc.), a charge-based storage medium (e.g., EEPROM, RAM, etc.), a solid-state storage medium (e.g., flash drive, etc.), and / or other electronically readable storage medium. The electronic storage may include one or more virtual storage resources (e.g., cloud storage, virtual private networks, and / or other virtual storage resources). The electronic storage may store software algorithms, information determined by a processor, information obtained from a server, information obtained from a client device, or other information enabling the functionality described herein.
[0066] The instructions stored in memory 610 may be implemented as software and / or firmware for programming processor(s) 605 to perform the actions described above. The processor may be programmed to execute computer program instructions by software, hardware, firmware; some combination of software, hardware, or firmware; and / or other mechanisms for configuring processing capabilities on the processor. In some aspects, such software or firmware may be provided to computer system 600 by initially downloading it by computer system 600 from a remote system (e.g., via network adapter 630).
[0067] The aspects introduced herein may be implemented, for example, by programmable circuitry (e.g., one or more microprocessors) programmed with software and / or firmware, or in entirely dedicated hardwired (non-programmable) circuitry, or in a combination of such forms. Dedicated hardwired circuitry may be in the form of, for example, one or more ASICs, PLDs, FPGAs, etc.
[0068] remarks The above description and drawings are illustrative and should not be construed as limiting. Numerous specific details are described to provide a thorough understanding of the present disclosure. However, in some cases, well-known details are not described to avoid obscuring the description. Furthermore, various modifications may be made without departing from the scope of the aspects. Thus, the aspects are not limited except as by the appended claims.
[0069] References herein to "one embodiment" or "an embodiment" mean that a particular feature, structure, or characteristic described in connection with that embodiment is included in at least one embodiment of the disclosure. The appearances of the phrase "in one embodiment" in various places in this specification do not necessarily all refer to the same embodiment, nor to separate or alternative embodiments that are mutually exclusive with other embodiments. Furthermore, various features are described that may be exhibited by some embodiments and not by other embodiments. Similarly, various features are described that may be requirements of some embodiments but not other embodiments.
[0070] The terms used in this specification generally have their ordinary meaning in the art, within the context of this disclosure and within the specific context in which each term is used. The terms used to describe this disclosure are explained below or elsewhere in this specification to provide additional guidance to practitioners regarding the description of this disclosure. For convenience, some terms may be highlighted, for example, using italics and / or quotation marks. The use of highlighting does not affect the scope and meaning of a term, which is the same in the same context whether or not it is highlighted. It will be understood that the same thing can be said in multiple ways. It will be understood that "memory" is a form of "storage device" and that these terms may be used interchangeably in some cases.
[0071] Thus, alternative wording and synonyms may be used for any one or more of the terms discussed herein, and no special meaning should be attached to whether a term is detailed or discussed herein. Synonyms of some terms are provided. The description of one or more synonyms does not exclude the use of other synonyms. The use of examples anywhere in this specification, including examples of any term discussed herein, is merely illustrative and is not intended to further limit the scope and meaning of the disclosure or any exemplified term. Similarly, the disclosure is not limited to the various aspects provided herein.
[0072] Those skilled in the art will appreciate that the logic illustrated in each of the above flow charts may be modified in various ways, such as by reordering the logic, performing sub-steps in parallel, omitting illustrated logic, or including other logic.
[0073] Without intending to further limit the scope of the present disclosure, examples of instruments, devices, methods and their related results according to the embodiments of the present disclosure are given below. Note that in the examples, titles or subtitles may be used for the convenience of the reader, and they should not limit the scope of the present disclosure in any way. Unless otherwise defined, all scientific and technical terms used herein have the same meaning as commonly understood by those skilled in the art to which the present disclosure pertains. In case of conflict, the present specification, including definitions, shall prevail.
[0074] While the present invention has been described in detail for purposes of illustration based on what are presently considered to be the most practical and preferred embodiments, it should be understood that such detail is for that purpose only and that the present invention is not limited to the disclosed embodiments, but on the contrary, is intended to cover modifications and equivalent arrangements within the scope of the appended claims. For example, it should be understood that the present invention contemplates that, to the extent possible, one or more features of any embodiment can be combined with one or more features of any other embodiment.
Claims
1. obtaining a video stream from a client device associated with a user, the video stream including video of the user's face; providing the video stream as an input to a first predictive model to obtain a first set of eye measurement parameters for the user's eye including a first pupil center, first iris data, and first eyelid data, the first set of eye measurement parameters being obtained at a first resolution; deconvolving the video stream based on the first eyelid data, the first pupil center, stimulus data and environmental data to obtain adjusted eyelid data, an adjusted pupil center and an adjusted pupil radius as part of the first eye measurement parameter set, the stimulus data including spatiotemporal and optical information of a video stimulus presented on a display associated with the client device and the environmental data including lighting conditions within an environment in which the client device is located; and providing the first ocular measurement parameter set as an input to a second predictive model to obtain a second ocular measurement parameter set at a second resolution, the second ocular measurement parameter set including (a) a second pupil center, (b) a second pupil radius, and (c) second iris data including an iris radius, an iris translation, and an iris rotation, the second resolution being greater than the first resolution; A method comprising:
2. 2. The method of claim 1, wherein the first resolution is measured in pixel values or other lower resolution values and the second resolution is measured in sub-pixel values.
3. Providing the input to the second predictive model comprises: training the second prediction model using a plurality of datasets to predict the second eye measurement parameter set at the second resolution, each dataset including a video of a specified user's face, a set of labels having a first set of values corresponding to the first eye measurement parameter set acquired at the first resolution, and a second set of values corresponding to the second eye measurement parameter set acquired at the second resolution; Including, 2. The method of claim 1.
4. 4. The method of claim 3, wherein the second set of values corresponding to the second set of eye measurement parameters is obtained using an eye tracking device configured to obtain the second set of eye measurement parameters at the second resolution for a video of the face of the specified user.
5. Deconvolving the video stream comprises: determining a point spread function of a deconvolution process based on the first eyelid data, the first pupil center, the stimulus data, and the environmental data; and deconvolving the video stream with the point spread function to obtain the adjusted eyelid data, the adjusted pupil center, and the adjusted pupil radius. Including, 2. The method of claim 1.
6. 6. The method of claim 5, wherein the step of deconvolving the video stream includes determining the adjusted eyelid data, the adjusted pupil center, and the adjusted pupil radius based on a maximum likelihood of prediction of the location of the first pupil center and eyelid boundary.
7. Determining the adjusted eyelid data includes: processing the first eyelid data using a maximum likelihood curve fitting operation to obtain a parabolic fit of an eyelid of the eye; and obtaining a plurality of coordinates along the parabolic fit as the adjusted eyelid data; Including, 6. The method of claim 5.
8. determining the adjusted pupil center, determining a plurality of locations for the first pupil center based on the maximum likelihood; generating a pupil shape for each of the plurality of positions based on a set of constraints and a set of assumptions; assigning a score to each of the plurality of shapes of the pupil; selecting a designated shape from the plurality of shapes based on a designated score of the designated shape that satisfies a criterion; and determining a center and a radius of the specified shape as the adjusted pupil center and the adjusted pupil radius, respectively; Including, 7. The method of claim 6.
9. providing the video stream to the first predictive model to obtain the first set of ocular measurement parameters, processing the video stream into a suitable format for the first predictive model to extract the first set of ocular measurement parameters prior to inputting the video stream into the first predictive model; Including, 2. The method of claim 1.
10. processing the video stream, Reducing noise, adjusting color and brightness of the video, and improving the resolution of the video Including, 10. The method of claim 9.
11. providing the video stream to the first predictive model to obtain the first set of ocular measurement parameters, obtaining user data associated with the user, the user data including optometric data of the eye of the user; and applying a correction to the first set of eye measurement parameters based on the user data. Further comprising:
2. The method of claim 1.
12. The method of claim 1 , wherein the first iris data includes: (a) a plurality of coordinates representing a shape of the iris; and (b) a first coordinate of an iris translation.
13. providing the first set of ocular measurement parameters to the second predictive model, processing the first iris data using a maximum likelihood based curve fitting operation to obtain an elliptical fit of the shape of the iris; and determining a first iris rotation and a first iris radius based on the ellipse fit; Including, 13. The method of claim 12.
14. The method of claim 1 , wherein the first eyelid data comprises a plurality of coordinates of a shape representing an eyelid of the eye.
15. The method of claim 1, wherein the first set of eye measurement parameters is acquired as time series data, a first set of values of the first set of eye measurement parameters is acquired for a first time point of the video, and a second set of values of the first set of eye measurement parameters is acquired for a second time point of the video.
16. The method of claim 1 , wherein the first set of eye measurement parameters is acquired in multiple coordinate systems.
17. 20. The method of claim 16, wherein the multiple coordinate systems include a device coordinate system having an origin located at a center of a display associated with the client device.
18. 17. The method of claim 16, wherein the plurality of coordinate systems includes a head coordinate system whose origin is centered between the eyes projected by a displacement in the z direction onto a plane defining the face of the user.
19. determining an additional ocular measurement parameter set based on the second ocular measurement parameter set, the additional ocular measurement parameter set including at least one of a pupil visible ratio, an iris visible ratio, a pupil coverage asymmetry, or an iris coverage asymmetry; 2. The method of claim 1, further comprising:
20. A non-transitory computer-readable medium, comprising: When executed by a computer, obtaining a video stream from a client device associated with a user, the video stream including video of the user's face; providing the video stream as an input to a first predictive model to obtain a first set of eye measurement parameters for the user's eye including a first pupil center, first iris data, and first eyelid data; deconvolving the video stream based on the first eyelid data, the first pupil center, stimulus data and environmental data to obtain adjusted eyelid data, an adjusted pupil center and an adjusted pupil radius as part of the first eye measurement parameter set, the stimulus data including spatiotemporal and optical information of a video stimulus presented on a display associated with the client device and the environmental data including lighting conditions within an environment in which the client device is located; and providing the first ocular measurement parameter set as an input to a second predictive model to obtain a second ocular measurement parameter set with a higher resolution than the first ocular measurement parameter set, the second ocular measurement parameter set including a second pupil center, a second pupil radius, an iris radius, an iris translation, and an iris rotation; instructions for causing the computer to carry out a method comprising: The non-transitory computer readable medium having stored thereon:
21. Providing the input to the second predictive model comprises: training the second prediction model using a plurality of datasets to predict the second eye measurement parameter set, each dataset including a video of a specified user's face, a set of labels having a first set of values corresponding to the first eye measurement parameter set acquired at a specified resolution, and a second set of values corresponding to the second eye measurement parameter set acquired at a resolution higher than the specified resolution; Including, 21. The computer-readable medium of claim 20.
22. 22. The computer-readable medium of claim 21, wherein the second set of values corresponding to the second set of eye measurement parameters is acquired using an eye tracking device configured to acquire the second set of eye measurement parameters at a resolution higher than the specified resolution.
23. Deconvolving the video stream comprises: determining a point spread function of a deconvolution process based on the first eyelid data, the first pupil center, the stimulus data, and the environmental data; and deconvolving the video stream with the point spread function to obtain the adjusted eyelid data, the adjusted pupil center, and the adjusted pupil radius. Including, 21. The computer-readable medium of claim 20.
24. 24. The computer-readable medium of claim 23, wherein deconvolving the video stream includes determining the adjusted eyelid data, the adjusted pupil center, and the adjusted pupil radius based on a maximum likelihood of prediction of the location of the first pupil center and eyelid boundary.
25. Determining the adjusted eyelid data includes: processing the first eyelid data using a maximum likelihood curve fitting operation to obtain a parabolic fit of an eyelid of the eye; and obtaining a plurality of coordinates along the parabolic fit as the adjusted eyelid data; Including, 25. The computer-readable medium of claim 24.
26. determining the adjusted pupil center, determining a plurality of locations for the first pupil center based on the maximum likelihood; generating a pupil shape for each of the plurality of positions based on a set of constraints and a set of assumptions; assigning a score to each of the plurality of shapes of the pupil; selecting a designated shape from the plurality of shapes based on a designated score of the designated shape that satisfies a criterion; and determining a center and a radius of the specified shape as the adjusted pupil center and the pupil radius, respectively; Including, 25. The computer-readable medium of claim 24.
27. 21. The computer-readable medium of claim 20, wherein the first set of eye measurement parameters is acquired as time series data, a first set of values of the first set of eye measurement parameters is acquired for a first time point of the video, and a second set of values of the first set of eye measurement parameters is acquired for a second time point of the video.
28. determining an additional ocular measurement parameter set based on the second ocular measurement parameter set, the additional ocular measurement parameter set including at least one of a pupil visible ratio, an iris visible ratio, a pupil coverage asymmetry, or an iris coverage asymmetry; 21. The computer readable medium of claim 20, further comprising:
29. A memory storing an instruction set; Executing the instruction set, obtaining a video stream from a client device associated with a user, the video stream including video of the user's face; training a first predictive model using a first plurality of datasets to output a first set of eye measurement parameters including first pupil center, first iris data, and first eyelid data of the user's eye at a first resolution, each dataset of the first plurality of datasets including a video of a specified user's face and the set of eye measurement parameters acquired at the first resolution; deconvolving the video stream based on stimulus data and environmental data to obtain adjusted eyelid data, an adjusted pupil center, and an adjusted pupil radius as part of the first eye measurement parameter set; and training a second predictive model using a second plurality of data sets to output a second set of eye measurement parameters at a second resolution higher than the first resolution, the second set of eye measurement parameters including a second pupil center, a second pupil radius, an iris radius, an iris translation, and an iris rotation, each data set including a video of a specified user's face, a set of labels having a first set of values corresponding to the first set of eye measurement parameters acquired at the first resolution, and a second set of values corresponding to the second set of eye measurement parameters acquired at the second resolution; a processor configured to cause the system to perform the method of The system comprising: