System and method for training eye state predictor
By combining feature, model and appearance-based learning methods, using eye state predictors and head-mounted device sensor data to train neural networks, the robustness and accuracy of head-mounted eye trackers when the environment and user appearance changes are solved, and efficient 3D eye state prediction is achieved.
Patent Information
- Application Number
- CN202280102687.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-21
- Publication Date
- 2025-07-25
AI Technical Summary
Existing head-mounted eye trackers are limited in robustness when dealing with ambient light and user appearance changes, and are difficult to accurately predict 3D eye status parameters.
A method is adopted to feed the first and second eye-related observations as input to the eye state predictor, and optimize the neural network model by training the loss function, output the 3D eye state, and train it in combination with the sensor data of the head-mounted device to achieve a combination of learning methods based on features, models and appearance.
Accurate and robust 3D eye state predictions during environmental changes and user appearance changes are achieved, improving the accuracy and robustness of predictions.
Smart Images

Figure CN120380439A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present invention relate to systems and methods for training an eye state predictor implemented as a neural network, particularly a convolutional neural network, and to methods for real-time prediction of 3D eye states of an object, particularly an object wearing a head-mounted device. Background Art
[0002] Over the past few decades, a wide variety of camera-based eye trackers have been proposed. Generally, methods for determining the gaze direction can be classified into methods that rely on the explicit extraction of features from eye images, and methods that do not rely on such explicit feature extraction, but instead accept the entire image of one eye (or both eyes) as input to some pre-trained algorithm, such as a machine learning algorithm, e.g., a trained neural network. The latter method is commonly referred to as an appearance-based method, in contrast to the former explicit feature-based method. Feature-based methods can be further divided into regression-based and model-based methods. Regression-based methods typically employ a polynomial mapping function that, after a person-specific calibration, is used to predict the gaze direction (typically 2D gaze coordinates on a screen) based on suitable eye image features, typically the 2D pupil center. Model-based methods fit a mathematical 3D eye model to the extracted eye image features (typically a series of pupil contours) and, following the 3D gaze direction, allow the estimation of other relevant parameters characterizing the three-dimensional (3D) eye state, such as, for example, the position of the eye center and the pupil radius. In all these methods, the explicit extraction of eye features, such as the pupil center, IR blink position (reflections actively generated by IR LEDs), or pupil contours, is typically performed by classical computer vision and image processing algorithms or by machine learning-based methods.
[0003] Many known head-mounted eye trackers suffer from the following drawbacks: Ambient stray light reflected on the eyes of the test user can negatively affect eye tracking functionality. In feature-based eye tracking methods, the camera monitoring the eyes of the test user may not be able to distinguish between the eye features explicitly used for tracking eye movements and features such as false reflections caused by environmental lighting conditions. Therefore, reliable eye tracking is often impaired by environmental conditions and unwanted stray light that disturbs the tracking mechanism. Thus, known head-mounted eye tracker devices often suffer from limited robustness when dealing with large variations in the environment and user appearance.
[0004] Head-mounted eye trackers that employ appearance-based (learning-based) gaze estimation methods have been shown to better handle such large variations in ambient light conditions and user appearance, giving very good accuracy even in uncalibrated setups. However, such methods require large amounts of labeled training data, including images of the eyes and corresponding gaze coordinates (e.g., within the image of a front-facing scene camera), and are also challenging to calibrate. Due to the lack of corresponding ground truth data, they typically also cannot (reliably) produce interesting 3D eye state parameters. This is because, in the context of eye tracking, neural networks are typically directly supervised for training only by using ground truth data of the gaze point or gaze direction (which can be obtained from dedicated data collection sessions) to perform end-to-end gaze prediction on the input 2D eye images, or by using 2D image features that are directly visible or annotatable on the 2D eye images to directly supervise training for performing eye feature extraction.
[0005] Although well-calibrated feature-based methods can still produce higher accuracy than appearance-based learning methods, in many use cases, such as head-mounted eye trackers, highly controlled and calibrated setups are difficult or even impossible to achieve. Thus, both gaze estimation paradigms have advantages and disadvantages.
[0006] Furthermore, model-based gaze estimation methods using head-mounted eye trackers based on the fitting of a mathematical 3D eye model to a series of pupil contours extracted from 2D eye images typically assume that the center of rotation of the eyeball is fixed in the eye camera coordinate system, which may only be approximately true due to the so-called head-mounted headset slippage (i.e., the inevitable movement of the head-mounted headset relative to the user's head during use), and may therefore require real-time eye model updates, as explained in WO 2020 / 244752 A1.
[0007] Accordingly, there is a need for further improvement in the detection of gaze direction and other eye state parameters. SUMMARY OF THE INVENTION
[0008] According to an embodiment of a method for training an eye state predictor implemented as a neural network, particularly as a convolutional neural network, the method includes feeding a first eye-related observation as input to the eye state predictor to determine a predicted 3D eye state of at least one eye of an object for an observation scenario as an output of the eye state predictor, the first eye-related observation relating to at least one eye of the object in and / or during the observation scenario. Feeding the predicted 3D eye state as input to a differentiable predictor to determine a prediction of at least one eye of the object for the observation scenario as an output of the differentiable predictor. Determining a training loss based on the prediction and at least one of the first eye-related observation and a second eye-related observation. The second eye-related observation also relates to at least one eye of the object in and / or during the observation scenario. The training loss is used to train the eye state predictor.
[0009] Thus, the eye state predictor (which is also referred to hereinafter as the eye state predictor model) is trained (based on the (one or more) training losses which may also be referred to as the (one or more) prediction losses) to output a corresponding 3D eye state when receiving the (one or more) eye-related observations.
[0010] At least the first eye-related observation typically includes an image relating to at least one eye of the object, particularly a corresponding eye image recorded by an eye camera, and shows at least a part of the left eye of the object (left eye image) or the right eye of the object (right eye image). More generally, the first eye-related observation typically includes a pair of corresponding left eye images and right eye images or a concatenated eye image, i.e., an eye image formed by concatenating the left eye image and the right eye image. The left eye image and the right eye image may particularly be provided by corresponding eye cameras of a head-mounted device worn by the object.
[0011] The observation period of the observation scenario may correspond to a normal video frame rate and / or be relatively short, e.g., less than 0.2 s, more generally less than 0.02 s. In these embodiments, the left eye image and the right eye image may correspond to images of corresponding video streams provided by the eye cameras.
[0012] While the eye-related observation typically includes at most one left eye image and at most one corresponding right eye image, and / or is determined based on at most one left eye image and at most one corresponding right eye image, the eye-related observation may also include a generally short sequence of left eye images and / or right eye images, and / or be determined based on a generally short sequence of left eye images and / or right eye images, e.g., a corresponding sequence having a length of at most 5 or 10.
[0013] Furthermore, the observation scenario may refer to and / or be represented by at least one of a (corresponding) observation time (observation date) and an observation identifier (ID) which typically depends on the observation time.
[0014] According to an embodiment, a method for training an eye state predictor includes feeding a first eye-related observation as an input to the eye state predictor to determine a predicted 3D eye state of at least one eye of an object during an observation time as an output of the eye state predictor, the first eye-related observation relating to at least one eye of the object during the observation time. Feeding the predicted 3D eye state as an input to a differentiable predictor to determine a prediction of at least one eye of the object during the observation time as an output of the differentiable predictor. Determining a training loss based on the prediction and at least one of a first eye-related observation and a second eye-related observation relating to at least one eye of the object during the observation time, in particular based on the prediction and one of the first eye-related observation and the second eye-related observation. The training loss is used to train the eye state predictor.
[0015] Hereinafter, the method for training an eye state predictor is also referred to as a training method.
[0016] The method for training an eye state predictor as explained herein allows combining the advantages of feature-based, model-based, and appearance-based learning methods to achieve accurate and robust 3D eye state prediction.
[0017] In particular, an eye state prediction model can be trained to receive a 2D (camera) image of an eye of an object (or a pair of corresponding images of the left and right eyes of the object), as in the appearance-based method mentioned above, but for a given use case, instead of merely producing a fixation point or direction, it outputs a generally complete 3D eye state, including a set or vector of eye state values (data) or consisting of a set or vector of eye state values (data), such as 3D eye ball center coordinates, 3D gaze direction, and 3D pupil size.
[0018] The term "3D eye state" as used herein is intended to describe a set of quantities or values, in particular a corresponding vector that describes and generally characterizes the (actual) three-dimensional state of at least one eye of an object (i.e., the left eye of the object (hereinafter also referred to as the first eye), the right eye of the object (hereinafter also referred to as the second eye), both eyes, and / or the fused (cyclopean) eye) at a given time.
[0019] A set consisting (only) of two-dimensional (2D) values directly visible in an eye image (such as, for example, the 2D pixel positions of an eye bounding box or the 2D pupil image (ellipse) attributes within a remotely captured image of an object's face) should not be understood as a "3D eye state".
[0020] The 3D eye state generally involves and / or includes one or more (usually two or three) 3D observables (physical quantities that can be measured) of at least one eye of the object.
[0021] (Predicted) 3D eye states can include any (measured / measurement-based) values that characterize physiological constants or transient parameters of the corresponding eye in 3D, in particular the 3D center of rotation of the eyeball, the 3D gaze direction, e.g., vectors characterizing the optical axis or the visual axis / line of sight and / or the pupil aperture size ("pupil size") in 3D.
[0022] (Predicted) 3D eye states typically include at least one of the following, typically include at least two of the following (e.g., three or even all): the predicted 3D (rotation) center of the eyeball of at least one eye, the 3D gaze direction of at least one eye, the 3D state of the eyelids of at least one eye, and the 3D state of the pupils of at least one eye.
[0023] Typically, the (predicted) 3D state of the pupil includes at least one of the predicted 3D pupil size of at least one eye, the predicted 3D pupil aperture of the iris of at least one eye, the predicted 3D pupil radius of the iris of at least one eye, and the predicted 3D pupil diameter of the iris of at least one eye.
[0024] The (predicted) 3D state of the eyelids of at least one eye can include at least one of eyelid shape, eyelid position, and eyelid closure percentage.
[0025] In addition, the predicted 3D eye state can be a monocular state, a pair of corresponding left and right monocular states, or a binocular state.
[0026] In embodiments involving 3D eye states consisting of the predicted 3D center of the eyeball of at least one eye, the 3D gaze direction of at least one eye, and the 3D state of the pupils of at least one eye, the predicted 3D eye state is typically one of the following: a 6-dimensional vector (characterizing a monocular 3D eye state), a 10-dimensional vector (characterizing a binocular 3D eye state), and a 12-dimensional vector (characterizing a binocular 3D eye state). Note that due to physiological reasons, the two monocular 3D eye states of an object are related to each other. Therefore, a 10-dimensional or 11-dimensional data set (vector) is typically sufficient to characterize the 3D eye state. In other words, compared to a combined data set (e.g., a vector) representing two corresponding left and right (monocular) 3D eye states, the binocular 3D eye state can be represented by a lower-dimensional data set (e.g., a vector) depending on the applicable (one or more) constraints.
[0027] During training, the training parameters of the eye state predictor are typically changed iteratively, i.e., the model parameters to be trained / parameters that the eye state predictor model learns by optimizing a loss function, in particular the weights of the connections between the artificial neurons of a neural network (NN).
[0028] The eye state predictor can be trained separately by machine learning (ML) and by using machine learning algorithms. In particular, corresponding ML optimization algorithms such as backpropagation can be used to train the eye state predictor. Thus, (at least one) loss function is used to determine the training loss and train the eye state predictor respectively, and the loss function involves, typically mathematically characterizes and / or provides a measure of the difference or distinction between (one or more) predictions of the eye state predictor (model) and (one or more) eye-related observations. Note that for the backpropagation that can be adopted, all steps in the computational graph (i.e., all computations starting from the algorithm input to the final loss function) must be differentiable.
[0029] The eye state predictor can be trained using a direct supervision training scheme or an indirect supervision training scheme.
[0030] Eye-related observations that can be used to train the eye state predictor can include further data related to at least one eye of the object.
[0031] In particular, eye-related observations for training can be determined using a head-mounted device worn by the object and providing eye cameras (typically a left eye camera, a right eye camera, and an optional scene camera) and / or additional sensors such as an inertial measurement unit as components. The data provided by these components during the observation scenario and / or at the observation time (also as post-processed data / derived data) can also be part of the corresponding eye-related observations.
[0032] Thus, further data (e.g., sensor readings of the head-mounted device) can be used to train the eye state predictor. Therefore, the accuracy and / or robustness of the 3D eye state prediction of the trained eye state predictor can be further improved.
[0033] The first eye-related observation and the second eye-related observation respectively used to train the eye state predictor can include the corresponding eye images of the (one or more) eyes of the object recorded by the corresponding eye cameras (of the head-mounted device) during the observation scenario and / or at the observation time as the main eye-related observations, and / or the corresponding secondary eye-related observations derived from the (one or more) eye images.
[0034] Secondary eye-related observations generally include visual features determined for and / or extracted from the corresponding eye images (more generally including several corresponding visual features such as edges, intermediate boundaries, and / or positions of anatomical features identified in the eye images), and / or visual feature representations of the eye images implemented to include the (one or more) visual features. Feature detection algorithms such as feature point detection, edge detection, and / or semantic segmentation can be used to determine the (one or more) visual features of the corresponding eye images.
[0035] In particular, the secondary eye-related observation can be a semantic segmentation of the corresponding eye image, which is also referred to hereinafter as a semantically segmented eye image.
[0036] Semantic segmentation (also known as image segmentation) can be described as clustering and / or partitioning a digital image into image segments and providing a semantically segmented image for the digital image consisting of segments that can completely cover the digital image. During semantic segmentation, objects and / or (segmentation) boundaries can be detected in the digital image. Additionally, image segmentation generally includes assigning labels to pixels in the digital image such that pixels with the same label share one or more characteristics. For example, pixels identified as belonging to the pupil in a digital eye image can be assigned a pupil label.
[0037] In particular, the first eye-related observation and / or the second eye-related observation can include, as the corresponding primary eye-related observation, the left-eye image of the left eye of the object and the right-eye image of the right eye recorded by the corresponding eye camera during the observation scenario and / or at the observation time.
[0038] Furthermore, the first eye-related observation and / or the second eye-related observation can include a secondary eye-related observation derived from the left-eye image and a secondary eye-related observation derived from the right-eye image, in particular, the semantic segmentation of the left-eye image (semantically segmented left-eye image) and the semantic segmentation of the right-eye image (semantically segmented right-eye image).
[0039] The semantic segmentation of the corresponding eye image can include at least one marker, which is typically selected from the list consisting of pupil markers, iris markers, sclera markers, eyelid markers, skin markers, and eyelash markers. Using such labels to train an eye state predictor can also improve the training results.
[0040] The training method can include determining, as an eye-related observation, a scene observation related to the field of view of an object's (at least one eye) during the observation scenario and / or at the observation time, typically using a scene camera of a head-mounted device, for example, as the second eye-related observation during the observation scenario and / or at the observation time.
[0041] The scene observation can include a scene image (recorded by the scene camera) as the primary eye-related observation, and / or (at least one) secondary eye-related observation derived from the scene image, in particular, the semantic segmentation of the scene image, a 3D fixation point typically measured in the scene camera coordinates, and a 2D fixation point typically measured in the scene camera image coordinates.
[0042] In addition, the first eye-related observation and / or the second eye-related observation may include at least one of the following as the respective main observation: the corneo-retinal resting potential of at least one eye of the object during the observation scenario and / or the observation time, the speed of the object's head during the observation scenario and / or the observation time, the acceleration of the object's head during the observation scenario and / or the observation time, the orientation of the object's head during the observation scenario and / or the observation time, the respective 2D fixation point during the observation scenario and / or the observation time, and the 3D fixation point during the observation scenario and / or the observation time.
[0043] The method for training an eye state predictor is typically performed iteratively.
[0044] Thus, multiple respective eye-related observations can (be determined using an eye camera and) be used for training (one or several times).
[0045] In addition, the respective eye-related observations of different objects can (be determined using an eye camera and) be used for training.
[0046] The observation scenario can be represented by an object ID and an observation time, or by an observation (identification) ID depending on the object ID and the observation time. In addition, the observation ID can depend on the ID of the hardware used to capture the image (hardware ID), in particular the device ID of a head-mounted device.
[0047] The determined eye-related observations can be buffered and / or stored in a (training) database (for eye-related observations), in particular before the actual training of the eye state predictor.
[0048] During the actual training (loop) of the eye state predictor, the database can be used to determine the first eye-related observation and the second eye-related observation respectively, typically a number of (first) eye-related observations to be fed as input to the eye state predictor.
[0049] For example, eye-related observations can typically be retrieved randomly from the database and fed to the eye state predictor, thereby outputting a corresponding prediction. The training loss can be determined based on the corresponding prediction and one or even both of the following two items: the retrieved eye-related observation and another eye-related observation retrieved from the database and corresponding to the same observation scenario (including the observation time).
[0050] In an embodiment where the differentiable predictor is the identity operator (identity function), predicting the 3D eye state and the corresponding 3D eye state (for the observation scenario and / or the observation time) can be used to determine the training loss. Alternatively, an operator that maps the input to a typically linearly scaled version of the input as output (which is also referred to hereinafter as a scaling operator) can be used to determine the training loss.
[0051] The first eye - related observation can be used to determine the corresponding 3D eye state, but it is different from the 3D eye state. Alternatively, the second eye - related observation can be used, or both the first eye - related observation and the second eye - related observation (or their corresponding parts) can be used to determine the corresponding 3D eye state. Further, an eye - related observation sequence can be used to determine the corresponding 3D eye state. In particular, a 3D geometric eye model can be used to determine the corresponding 3D eye state such that the corresponding 3D eye state conforms to the first eye - related observation (and / or the second eye - related observation or the eye - related observation sequence in the above alternatives).
[0052] Typically, a 3D geometric eye model that takes into account corneal refraction is used to determine the corresponding 3D eye state.
[0053] In other embodiments, the differentiable predictor is different from the identity operator and has a non - zero (non - constant) derivative with respect to the parameters of the eye state predictor.
[0054] In particular, the differentiable predictor can be configured to, in response to receiving a predicted 3D eye state as input from the eye state predictor, output a synthetic eye image or a pair of left and right synthetic eye images.
[0055] In this embodiment, the training loss is typically determined based on the corresponding semantic segmentation of the (one or more) synthetic eye images and the (one or more) eye images. Note that the (one or more) synthetic eye images can be considered as secondary eye - related observations and / or can be stored in a database.
[0056] Similarly, the left - eye image of the left eye of the object and the right - eye image of the right eye fed as input to the eye state predictor can be considered as the first eye - related observations, which can also be stored in a database.
[0057] (Non - identity) The differentiable predictor can in particular be based on or even implemented as a trained neural network (especially a generative neural network) and / or be implemented as a differentiable renderer, such as an approximate ray tracer, i.e., a differentiable ray tracer.
[0058] In one embodiment, the differentiable predictor is implemented as a neural network that is trained using a 3D eye model such as the LeGrand 3D eye model or the Navarro 3D eye model (see, for example, WO2020 / 244752A1 and the references [1], [2] cited therein) and any non - differentiable ray tracer configured to generate artificial images given a 3D eye model, a 3D eye state, and camera properties (i.e., the image resolution and camera intrinsic characteristics of the (one or more) eye cameras used).
[0059] In another embodiment, the differentiable predictor is implemented in a suitable manner as a ray-tracing algorithm designed to generate an eye image consistent with a 3D eye model (such as the LeGrand 3D eye model or the Navarro 3D eye model (see, for example, WO 2020 / 244752 A1 and the references [1], [2] cited therein)), or any non-differentiable ray-tracker configured to generate an artificial image given a 3D eye model, in order to achieve the differentiability of the generated image with respect to the internal parameters of the model.
[0060] In these embodiments, the training loss can be a region-based loss, in particular the Jaccard loss.
[0061] As used herein, the term "camera intrinsics" will describe the optical properties of a camera, in particular the imaging properties (imaging characteristics) of the camera are known and / or can be modeled using a corresponding camera model, which includes known intrinsic parameters (known intrinsic characteristics) of an eye camera that approximately produces an eye image. Typically, a pinhole camera model is used to model the eye camera. The known intrinsic parameters can include the focal length of the camera, the image sensor format of the camera, the principal point of the camera, the offset of the central image pixel of the camera, the shear parameters of the camera, and / or one or more distortion parameters of the camera.
[0062] In embodiments where only one differentiable predictor (e.g., a first differentiable predictor) is used, only one loss function (e.g., a first loss function) can be used to determine the (total) training loss.
[0063] In particular, the prediction output by the (first) differentiable predictor and the first eye-related observation or the second eye-related observation can be fed into the (first) loss function that outputs the (first) training loss.
[0064] In other embodiments, at least two differentiable predictors (e.g., two, three, or even more) can be used to determine the corresponding predictions of at least one eye of an object during an observation scenario and / or at an observation time, where the (total) training loss is determined based on each of the corresponding predictions. Thus, the training efficiency, prediction accuracy, and / or prediction can be further improved.
[0065] In particular, predicting the 3D eye state can be fed as input into a first differentiable predictor to determine a first prediction of at least one eye of an object during an observation scenario and / or at an observation time as the output of the first differentiable predictor (DP).
[0066] Similarly, predicting the 3D eye state can be fed as input to a second differentiable predictor to determine a second prediction of the respective eye of the object during and / or at the observation time as the output of the second differentiable predictor.
[0067] Furthermore, the first prediction and the first eye-related observation or the second eye-related observation can be fed as input to a first loss function to determine a first training loss as the output of the first loss function.
[0068] Similarly, the second prediction and one of the first eye-related observation, the second eye-related observation, and the third eye-related observation relating to the observation scenario can be fed as input to a second loss function to determine a second training loss as the output of the second loss function.
[0069] The first and second training losses can be used to train the eye state predictor and to change the parameters of the eye state predictor, respectively.
[0070] In particular, the (total) training loss determined as a function of the first training loss and the second training loss can be used to train the eye state predictor, for example as a usual weighted sum.
[0071] The (total) training loss may depend on at least one scenario-specific parameter.
[0072] In particular, both the eye state predictor and the differentiable predictor (or even two or more differentiable predictors) can receive at least one scenario-specific parameter as part of their respective inputs.
[0073] The eye state predictor and the differentiable predictor can receive a common scenario-specific parameter, a set of common scenario-specific parameters (relating to and / or characterizing the observation scenario or sequence of observation scenarios at different observation times under otherwise invariant conditions, in particular for the same object and the same hardware), or respective (one or more) scenario-specific parameters from the set of common scenario-specific parameters.
[0074] At least one scenario-specific parameter typically includes (at least one) object-specific parameter.
[0075] The (at least one) object-specific parameter can in particular be selected from a list consisting of: interpupillary distance (IPD), the angle or angles between the optical axis and the visual axis of the respective eye, a rotation operator (such as a rotation matrix) capturing the transformation between the capture optical axis and the visual axis, geometric parameters (such as spherical, aspherical, thickness, astigmatism, 3D topography, and radius, e.g., the spherical radius or angle-dependent radius of the cornea) relating to the shape and / or size of the cornea of the respective eye, the refractive index of at least one part of the respective eye, the iris radius of the respective eye, the pupil shape, and geometric parameters relating to the shape and / or size of the eyeball of the respective eye.
[0076] During the training of the eye state predictor, one or more of the situation-specific parameters can be determined and / or adjusted, in particular the respective object-specific parameters.
[0077] For example, the eye state predictor can receive the respective physiological values of the (one or more) respective object-specific parameters as initial inputs, such as the corneal radius of the respective eye, the iris radius of the respective eye, the IPD of the respective eye, or the angle or angles between the optical axis and the visual axis. During the process of training using eye-related observations of a particular object, the respective (one or more) object-specific parameters can be changed (as part of the optimization).
[0078] In particular, the training loss can be used to correct (or update) the situation-specific parameters.
[0079] After training, the (one or more) situation-specific parameters adapted (learned) can be output, stored (e.g., for later use as an input for predicting the 3D eye state of an object), further processed, and / or used as (one or more) measured values of the object (which would otherwise be laborious to measure).
[0080] In an embodiment in which the eye state predictor is trained with eye-related observations of different objects, the respective object IDs of the eye-related observations contribute to learning the (one or more) object-specific parameters.
[0081] In some embodiments, during training, only a subset of the situation-specific parameters is learned and / or adjusted.
[0082] The (one or more) object-specific parameters learned can later be used as inputs for real-time prediction of the 3D eye state of an object.
[0083] Furthermore, the trained eye state predictor can even be used to determine at least one object-specific parameter of a new object.
[0084] In addition, at least one situation-specific parameter can include (at least one) hardware-specific parameter (which can also be learned and / or adjusted during training).
[0085] The (at least one) hardware-specific parameter can be specifically selected from a list consisting of: the respective camera intrinsic characteristics, the relative camera extrinsic characteristics, and the attitude of the inertial measurement unit relative to at least one of the (eye and optional scene) cameras.
[0086] Typically, when worn by an object, the attitudes of the (one or more) cameras are fixed relative to each other and / or to the coordinate system of the head-mounted device.
[0087] However, cameras with limited mobility can also be used (e.g., only along a given axis, as in a VR headset, to adjust the individual pupil distance). In a device with some camera position adjustment capabilities (such as a VR headset), it may even be possible to implement alternative position "sensors" that determine, for example, the mutual distance / attitude of the measurement cameras.
[0088] Camera external characteristics such as known attitude information and internal characteristics such as focal length, center pixel offset, shear, and distortion parameters can be measured during the production of the head-mounted device and then linked to the (hardware) ID of the head-mounted device and thus used as training input (additional "learning aids"). As a result, both the training of the eye state predictor and the accuracy of the predicted 3D eye state output by the trained eye state predictor can be further improved.
[0089] According to an embodiment, the predicted 3D eye state is used to determine a third training loss that can be used to train the eye state predictor.
[0090] Specifically, the predicted 3D eye state and at least one case-specific parameter can be fed into a third loss function to determine the third training loss.
[0091] The third training loss can be determined based on the predicted 3D eye state and the (one or more) object-specific parameters.
[0092] Furthermore, the (total) training loss (the training loss to be used for training the predicted 3D eye state) can be determined as a function of the first training loss and the third training loss, or as a function of the first training loss, the second training loss, and the third training loss.
[0093] According to an embodiment of a method for real-time prediction of an object's 3D eye state, i.e., with a maximum delay of at most approximately 0.25 seconds or 0.1 seconds, the method (which is also referred to hereinafter as the prediction method) includes determining eye-related observations involving at least one eye of the object and feeding the eye-related observations as input into a trained eye state predictor (model) implemented as a neural network to generally determine the predicted 3D eye state in real time and / or instantaneously as the output of the trained eye state predictor.
[0094] Predicting the 3D eye state can particularly include a predicted 3D rotation center of the eyeball of at least one eye and a 3D gaze direction of the eyeball of at least one eye.
[0095] Typically, predicting the 3D eye state further includes a predicted 3D state of the pupil of at least one eye.
[0096] The predicted 3D state of the pupil can include at least one of the following: a predicted 3D pupil size of at least one eye, a predicted 3D pupil aperture of the iris of at least one eye, a predicted 3D pupil radius of the iris of at least one eye, and a predicted 3D pupil diameter of the iris of at least one eye.
[0097] Typically, the trained eye state predictor is trained according to the training method as explained herein.
[0098] In embodiments where the training loss depends on one or more object-specific parameters, the (one or more) object-specific parameters can be pre-determined for a new object, particularly as corresponding physiological values, as corresponding measured values, using a method for object-specific calibration (as explained herein) to determine a corresponding predicted value or any function of one or more of these values.
[0099] According to an embodiment of the method for object-specific parameter calibration, the method includes providing a trained eye state predictor as explained herein, feeding a (corresponding new) first eye-related observation (the current value of at least one object-specific parameter of the new object) as an input to the eye state predictor to determine a predicted 3D eye state of at least one eye of the new object for the (new) observation scenario as an output of the trained eye state predictor, the first eye-related observation relating to at least one eye of the new object in and / or during the observation scenario; feeding the predicted 3D eye state and the current value of at least one object-specific parameter of the new object as an input to a differentiable predictor to determine a prediction for at least one eye of the new object in and / or during the observation scenario as an output of the differentiable predictor; determining a training loss based on the prediction and at least one of the first eye-related observation and a second eye-related observation, the second eye-related observation relating to at least one eye of the new object in and / or during the observation scenario; and using the training loss to update the current value of the corresponding at least one object-specific parameter.
[0100] Typically, after a plurality of update cycles, the current value of the corresponding at least one object-specific parameter is output as the corresponding predicted value of the at least one object-specific parameter of the (corresponding object).
[0101] In addition, for the methods explained herein, an eye camera of a head-mounted device worn by an object is typically used to capture at least one eye image (at a given time).
[0102] More typically, two corresponding eye cameras of the head-mounted device are used to capture corresponding (left and right) eye images of the object (at one or more given times).
[0103] The head-mounted device is typically implemented as a glasses device.
[0104] However, the head-mounted device can also be implemented as an augmented reality (AR-) and / or virtual reality (VR-) device (AR / VR head-mounted device), in particular goggles, an AR head-mounted display, and a VR head-mounted display. For the sake of clarity, the head-mounted device is mainly described below with respect to a head-mounted glasses device.
[0105] (One or more) 3D eye states are typically determined relative to a coordinate system fixed to the (one or more) eye cameras and / or the head-mounted device.
[0106] For example, a Cartesian coordinate system defined by the (one or more) image planes of the (one or more) eye cameras can be used.
[0107] Points and directions can also be specified and / or converted in a device coordinate system, a head coordinate system, a world coordinate system, or any other suitable 3D coordinate system.
[0108] (The head-mounted) glasses device typically includes a glasses body configured such that it can be worn on the head of an object, for example, in the manner of wearing ordinary glasses. Thus, when worn by an object, the glasses device can be particularly at least partially supported by the nose region of the object's face. This usage state of the head-mounted (glasses) device on the object's face will be further defined as the "intended use" of the glasses device, where direction and position references (such as horizontal and vertical, parallel and perpendicular, left and right, front and back, up and down, etc.) relate to this intended use. Thus, left and right, up and down positions, and lateral positions such as front / forward and back / backward are to be understood from the normal perspective of the object. Similarly, this applies to horizontal and vertical orientations, where during the intended use, the object's head is in a normal, thus upright, non-inclined, non-drooping, and non-nodding position.
[0109] The glasses body (main body) typically includes a left eye opening and a right eye opening, which mainly have the function of allowing the object to view through these eye openings. The eye openings can be implemented as but are not limited to light-shielding objects, optical lenses, or non-optical transparent glasses, or implemented as a non-material optical path that allows light to pass through.
[0110] The spectacle body can form the eye openings by at least partially or completely separating these eye openings from the surrounding environment. In this case, the spectacle body serves as a frame for the optical openings. The frame does not necessarily need to form a complete and closed surrounding for the eye openings. Additionally, it is possible for the optical openings themselves to have a frame-like configuration, for example by providing a support structure with the help of a transparent glass. In the latter case, the spectacle device has a form similar to that of rimless glasses, where only the nose bridge / nasal part and the temple pieces are attached to the glass screen, so that the glass screen serves as both the overall frame and the optical opening at the same time.
[0111] Furthermore, the median plane of the spectacle body can be identified. In particular, the median plane describes the structural central plane of the spectacle body, where corresponding structural components or parts that are comparable or similar to each other are placed in a similar manner on each side of the median plane. When the spectacle device is in its intended use and correctly worn, the median plane coincides with the median plane of the object.
[0112] In addition, the spectacle body typically includes a nasal part, a left part, and a right part, where the median plane intersects the nasal part, and the corresponding eye openings are located between the nasal part and the corresponding side parts.
[0113] For orientation purposes, a plane perpendicular to the median plane should be defined, and the median plane is particularly vertically oriented, where the perpendicular plane does not necessarily be firmly located at a defined forward or backward position of the spectacle device.
[0114] The head-mounted device has an eye camera, which has a sensor arranged in or defining an image plane for capturing an image of the first eye of the object, that is, an image of the left or right eye of the object. In other words, the eye camera, which is also referred to as the camera and the first eye camera hereinafter, can be a left-eye camera or a right (near) eye camera. The eye camera typically has known camera intrinsic characteristics.
[0115] In addition, the head-mounted device can have another eye camera with known camera intrinsic characteristics for capturing an image of the second eye of the object, that is, an image of the right or left eye of the object. Hereinafter, the other eye camera is also referred to as the other camera and the second eye camera.
[0116] In other words, in a binocular setup, the head-mounted device can have left and right (eye) cameras, where the left camera is used to capture a (left) image or image stream of at least a part of the left eye of the object, and where the right camera captures a (right) image or image stream of at least a part of the right eye of the object.
[0117] Typically, the first and second eye cameras have the same or similar camera intrinsic characteristics (are of the same type, but can be calibrated individually).
[0118] However, the method explained in this document is also applicable to a monocular setup with only one (near) eye camera.
[0119] (One or more) eye cameras may be arranged at the spectacle body in the inner eye camera placement area and / or the outer eye camera placement area, in particular, where the areas are determined such that a proper picture of at least a part of the respective eye can be taken for the purpose of determining one or more eye state-related parameters; in particular, the camera is arranged in the bridge part and / or the side edge part of the spectacle frame such that the light field of the respective eye is not occluded by the respective camera. The light field is defined as being occluded if the camera forms a visibly distinct area / part in the light field, e.g., if the camera points from the boundary of the visible field to the field or projects into the field from the boundary. For example, the camera may be integrated into the frame of the spectacle body and thus be non-occluding. In the context of the present invention, a limitation of the visual field caused by the spectacle device itself, in particular by the spectacle body or the frame, is not considered an obstacle to the light field.
[0120] Furthermore, the head-mounted device may have illumination components for illuminating the left eye and / or the right eye of the object, especially if the light conditions in the environment of the spectacle device are not optimal.
[0121] Furthermore, the head-mounted device may be provided with a scene camera for taking an image of the field of view (FOV) of the object wearing the head-mounted device, e.g., an integrated scene camera typically arranged in the middle plane.
[0122] The prediction method may be at least partially controlled and / or executed by the computing and control unit of the head-mounted device.
[0123] Alternatively or additionally, an accompanying device (in particular a mobile accompanying device connected to the head-mounted device, such as a smartphone, a tablet computer, or a laptop computer, or a desktop computer) functionally connected to the computing and control unit (or only the control unit) of the head-mounted device of the system via, for example, a wired or wireless network (TCP / IP) connection and / or a USB connection may supervise, control, and / or execute the prediction method as explained in this document.
[0124] Generally, the computing system (hardware) that executes the training method and / or the prediction method as explained in this document includes one or more processors (in particular one or more CPUs, GPUs, and / or DSPs) and a neural network software module including instructions that, when executed by at least one of the one or more processors, implement an instance of a neural network.
[0125] According to an embodiment of the system, in particular a training system for an eye state predictor, the system comprises: a head-mounted device comprising at least one eye camera configured to generate an eye image of at least a part of the (at least one) eye of an object wearing the head-mounted device; and a computing system connectable to the at least one eye camera for receiving the eye image and configured to: generate a first eye-related observation of the (at least one) eye of the object in and / or during the observation scenario based on the eye image received from the at least one eye camera and relating to the observation scenario; run an eye state predictor implemented as a neural network; feed the first eye-related observation as an input to the eye state predictor to determine a predicted 3D eye state of the (at least one) eye of the object in and / or during the observation scenario; feed the predicted 3D eye state as an input to a differentiable predictor to determine a prediction of the (at least one) eye of the object in and / or during the observation scenario as an output of the differentiable predictor; determine a training loss based on the prediction and at least one of the first eye-related observation and a second eye-related observation, the second eye-related observation also relating to the eye of the object in and / or during the observation scenario; and use the training loss to train the eye state predictor.
[0126] Typically, the computing system is configured to determine the second eye-related observation differently than the first eye-related observation.
[0127] The head-mounted device typically comprises a respective eye camera for each eye of the object.
[0128] Furthermore, the first eye-related observation can be generated based on the respective (left and right) eye images of the left and right eyes of the object during the observation scenario and / or at the observation time at which the eye image is generated.
[0129] Similarly, the second eye-related observation, even if determined differently, can be generated based on the respective (left and right) eye images.
[0130] Furthermore, the head-mounted device can comprise a scene camera configured to generate a scene image relating to the field of view of the object wearing the head-mounted device.
[0131] Furthermore, the computing system can be configured to host or access a database for eye-related observations.
[0132] Thus, the training of the eye state predictor (model), which is typically performed iteratively, can be facilitated.
[0133] In particular, the computing system can be configured to perform the training method as explained herein.
[0134] The (actual) training period of the eye state predictor is typically performed after generating multiple eye-related observations and / or by sufficiently powerful computing hardware, such as a corresponding (remote) server and / or cloud-based architecture.
[0135] The computing system can also be configured to perform the prediction method as explained herein.
[0136] Compared to the training method, the prediction method has a lower computational intensity. Thus, the prediction method can even be controlled and / or executed by the controller of the head-mounted device (which typically provides the computing and control unit) (the controller is typically configured to run the trained eye state predictor (instance) implemented as a neural network and / or can even be integrated into the head-mounted device), by a connected companion device such as a smartphone, or by the (computing and) control unit of the head-mounted device and the companion device.
[0137] As used in this specification, the term "neural network" (NN) is intended to describe an artificial neural network (ANN) or connectionist system that includes multiple connected units or nodes called artificial neurons. The output signal of an artificial neuron is computed by a (non-linear) activation function of the sum of its (one or more) input signals. The connections between artificial neurons typically have respective weights (gain factors for the (one or more) output signals being passed) that are adjusted during one or more learning phases. Other parameters of the NN that may or may not be modified during learning can include parameters of the activation function of the artificial neurons, such as thresholds. Generally, artificial neurons are organized into layers that are also called modules. The most basic NN architecture, called a "multi-layer perceptron", is a sequence of so-called fully connected layers. A layer consists of multiple different units (neurons), each computing a linear combination of the inputs, followed by a non-linear activation function. Different layers (of neurons) can perform different kinds of transformations on their respective inputs. A neural network can be implemented in software, firmware, hardware, or any combination thereof. During one or more learning phases, machine learning methods, particularly supervised, unsupervised, or semi-supervised (deep) learning methods, can be used. For example, deep learning techniques, particularly gradient descent techniques such as backpropagation, can be used to train a (feedforward) NN with a hierarchical architecture. Modern computer hardware such as GPUs makes backpropagation efficient for multi-layer neural networks. A convolutional neural network (CNN) is a feedforward artificial neural network that includes an input (neural network) layer, an output (neural network) layer, and one or more hidden (neural network) layers arranged between the input layer and the output layer. The CNN is characterized by using convolutional layers to perform the mathematical operation of convolving the input with a kernel. The hidden layers of a CNN can include convolutional layers and optional pooling layers (for downsampling the output of the previous layer before feeding it into the next layer), fully connected layers, and normalization layers. At least one of the hidden layers in the CNN is a convolutional neural network layer, also called a convolutional layer hereinafter. A typical convolutional kernel size is, for example, 3×3, 5×5, or 7×7. Compared with fully connected layers, using one or more convolutional layers can help compute recurrent features in the input more efficiently. Therefore, the memory footprint can be reduced and the performance can be improved. Due to the shared weight architecture and the translational invariance property, the CNN is also called a shift-invariant or space-invariant artificial neural network (SIANN). Hereinafter, the term "neural network model" is intended to describe the set of data that defines a neural network operable in software and / or hardware. The model typically includes data related to the NN architecture (particularly the network structure, including the arrangement of neural network layers, the sequence of information processing in the NN), and data representing or consisting of NN parameters (particularly the connection weights within fully connected layers and the kernel weights within convolutional layers).During the training phase, the network learns to map (one or more) inputs ((one or more) eye-related observations) to (one or more) corresponding output results ((one or more) 3D eye states) with sufficient accuracy.
[0138] Other embodiments include corresponding computer systems, (non-transitory) computer-readable storage media or devices, and / or computer programs recorded on one or more computer-readable storage media or computer storage devices, each configured to perform the processes of the methods described herein.
[0139] A system of one or more computers and / or a system including one or more computers can be configured to perform particular operations or processes by virtue of software, firmware, hardware, or any combination thereof installed on the one or more computers, which in operation can cause the system to perform the processes. One or more computer programs can be configured to perform particular operations or processes by virtue of including instructions that, when executed by one or more processors of the system, cause the system to perform the processes.
[0140] Those skilled in the art will recognize additional features and advantages upon reading the following detailed description and upon viewing the drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0141] The components in the drawings are not necessarily to scale, emphasis instead being placed upon illustrating the principles of the invention. Further, in the drawings, like reference numerals denote corresponding parts. In the drawings:
[0142] Figure 1A and Figure 1B corresponding flowcharts illustrate methods for training an eye state predictor according to an embodiment;
[0143] Figure 1C A flowchart illustrates a method for real-time predicting the 3D eye state of an object according to an embodiment;
[0144] Figure 1D illustrates the 3D eye state of an object according to an embodiment;
[0145] Figure 2A 、 2B corresponding flowcharts illustrate methods for training an eye state predictor according to an embodiment;
[0146] Figure 2C A flowchart illustrates a method for training an eye state predictor according to an embodiment;
[0147] Figure 3A 、 3B corresponding flowcharts illustrate methods for training an eye state predictor according to an embodiment;
[0148] Figure 4A 、 4B Figures 4C illustrate corresponding flowcharts of a method for training an eye state predictor according to an embodiment;
[0149] Figure 5A illustrates a flowchart of an object - specific parameter calibration method according to an embodiment;
[0150] Figure 5B illustrates a flowchart of a method for training an eye state predictor according to an embodiment; and
[0151] Figure 5C illustrates a perspective view of a system including a head - mounted device according to an embodiment. DETAILED DESCRIPTION
[0152] In the following detailed description, reference is made to the accompanying drawings, which form a part hereof, and in which are shown by way of illustration specific embodiments in which the invention may be practiced. In this regard, directional terms such as "top", "bottom", "front", "rear", "frontward", "rearward", etc. are used with reference to the orientation of the described (one or more) figures. Since the components of the embodiments can be positioned in many different orientations, the directional terms are for illustrative purposes and are in no way limiting. It is to be understood that other embodiments may be utilized and structural or logical changes may be made without departing from the scope of the invention. Accordingly, the following detailed description is not to be taken in a limiting sense, and the scope of the invention is defined by the appended claims.
[0153] Reference is now made in detail to various embodiments, one or more examples of which are illustrated in the accompanying drawings. Each example is provided by way of explanation and is not meant as a limitation of the invention. For example, features illustrated or described as part of one embodiment can be used on or in combination with other embodiments to yield yet a further embodiment. It is intended that the invention include such modifications and variations. The examples are described in specific language which should not be construed as limiting the scope of the appended claims. The drawings are not to scale and are for illustrative purposes only. For clarity, the same elements or manufacturing steps have been designated by the same reference numerals in different drawings if not otherwise specified.
[0154] Reference Figure 1A and 1B , a method 1000 for training an eye state predictor ESP implemented as a neural network is explained.
[0155] In block 1200 of method 1000 , a first eye-related observation ERO1 is fed as input to an eye state predictor ESP to be trained, which outputs an eye state predictor corresponding to the eye state observed at observation time t k , where k is an integer index. In block 1100, if the training method 1000 is repeatedly and / or iteratively performed (e.g., by Figure 1B ), then the time t can be randomly selected, for example k and index k, and retrieve eye-related observations EROs from a database pre-populated with eye-related observations EROs.
[0156] Alternatively or additionally, the first eye-related observation ERO1(t k )(and optionally a second eye-related observation ERO2(t k )) can be determined using at least one eye camera or any other eye-related data providing component of a head-worn device worn by the subject during execution of the training method 1000, such as a scene camera of the head-worn device.
[0157] At least one of the eye-related observations EROs for a respective observation situation typically comprises as a respective primary eye-related observation pERO a plurality of eye-related observations, typically taken by the respective eye camera during the observation situation and / or at the observation time t k The left eye image P of the left eye of the subject recorded in l and the right eye image P of the right eye r , and / or from the left eye image P l and the right eye image P r The corresponding secondary eye-related observation sERO derived, such as the left eye image P l Semantic segmentation S l and the right eye image P r Semantic segmentation S r .
[0158] Furthermore, the eye-related observations EROs for the observation situation may include as a primary eye-related observation pERO a scene image Ps, typically recorded by a scene camera, and / or at least one corresponding secondary eye-related observation derived from the scene image Ps, in particular a derived 2D gaze point or even a derived 2D gaze point. S Semantic segmentation S S As a secondary eye-related observation.
[0159] In an exemplary embodiment, as indicated by brackets "{}", the predicted 3D eye state 3DES is a predicted 3D center E of an eyeball of at least one eye of a human subject.C and the predicted 3D gaze direction E of at least one eye G and the predicted 3D state E of the pupil of at least one eye P constitute a set (e.g., a vector).
[0160] In other embodiments, the predicted 3D eye state 3DES may additionally (or instead of, for example, the predicted 3D state E of the pupil P ) include the predicted 3D state of the eyelid of at least one eye.
[0161] As Figure 1D illustrated, the predicted 3D eye state 3DES may be one of the following: the monocular state 3DES of the left eye of the object ml (having the predicted 3D center E of the eyeball of the left eye of the object Cl , the predicted 3D gaze direction E of the left eye of the object Gl and the predicted 3D pupil state E of the left eye of the object Pl ), the monocular state 3DES of the right eye of the object mr (having the predicted 3D center E of the eyeball of the right eye of the object Cr , the predicted 3D gaze direction E of the right eye of the object Gr and the predicted 3D pupil state E of the right eye of the object Pr ), the monocular state 3DES of the (virtual) fused eye of the object mc (having the predicted 3D center E of the eyeball of the fused eye of the object Cc , the predicted 3D gaze direction E of the fused eye of the object Gc and the predicted 3D pupil state E of the fused eye of the object Pc ), and the binocular state 3DES b (having the corresponding monocular 3D states of the left and right eyes of the object).
[0162] Although the (right, left, and fused eye) monocular states 3DES ml , 3DES mr , 3DES mc can be represented by a six-dimensional vector, the binocular state 3DES b can be represented by a ten-dimensional vector or a twelve-dimensional vector.
[0163] As Figure 1A and Figure 1B further illustrated, in the subsequent block 1300 of Figure 1B , the predicted 3D eye state 3DES ({E C , E G , E P}) can be fed as input into a differentiable predictor DP that outputs a prediction П related to at least one eye of an object and corresponding to an observation situation. The differentiable predictor DP can have one or more non-zero derivatives with respect to the input observables that are its variables.
[0164] In a subsequent block 1400, the prediction П of at least one eye of the object in and / or during the same observation situation t k and the first eye-related observation ERO1(t k )( Figure 1A the dashed arrow in) and / or the second eye-related observation ERO2(t k )(usually one of ERO1(t k ) and ERO2(t k ) are fed into a loss function LF that outputs a corresponding training loss Δ for training an eye state predictor ESP in block 1500.
[0165] For example, the first eye-related observation ERO1 can be the primary eye-related observation, and the loss function LF can receive the prediction П and the secondary eye-related observation derived from the primary eye-related observation or a different primary eye-related observation as the second eye-related observation ERO2(t k ).
[0166] Alternatively, N eye-related observations can be fed into an eye state predictor ESP that outputs M predictions that are compared with L other eye-related observations to calculate a loss Δ (where N, M, and L are positive integers, each of which can be greater than 1). Generally, 1 < N < 100, M <= N, and / or L < N.
[0167] In one example, 20 > N > 1, M = 1, and L = M (or for example L <= N). Thus, a prediction П can be determined for a short sequence of N eye-related observations, for example as an average prediction П, and compared with a representative (other) eye-related observation for the same time interval (L = M = 1), or with an average eye-related observation determined using L > 1, for example L = N (other) eye-related observations.
[0168] After multiple training cycles, the resulting trained eye state predictor (model) tESP can be used to predict the 3D eye state of the same or different objects in real time.
[0169] As Figure 1CAs illustrated, for the exemplary prediction method 9000, the (new) newly determined eye-related observation ERO* can be input into the trained eye state predictor tESP, thereby outputting the corresponding predicted 3D eye state 3DES*.
[0170] In particular, method 9000 can be performed for an object wearing a head-mounted device of the same type as the head-mounted device used for training method 1000.
[0171] Furthermore, method 9000 can be performed several times and / or substantially continuously when the object is wearing the head-mounted device and, for example, viewing a real 3D scene, a 3D representation of a real scene, or a virtual 3D scene (presented on a screen).
[0172] Furthermore, the predicted 3D eye state(s) 3DES* can be further processed, particularly for determining the intent of the object and / or interacting with the object, such as presenting additional information (audibly and / or visually) and altering the 3D representation of the real scene or the virtual 3D scene.
[0173] Figure 2A and Figure 2B illustrates the training method 2000. The training method 2000 is similar to the training method 1000 explained above with reference to Figure 1A 、 1B and includes blocks 2100 to 2500, each of which is generally similar to the corresponding block of blocks 1100 to 1500 of method 1000. However, the training method 2000 is more specific.
[0174] In an exemplary embodiment, the differentiable predictor of block 2300 is the identity operator I.
[0175] Thus, the predicted 3D eye state 3DES({E C ,E G ,E P}) determined in block 2200 is forwarded as the prediction П in block 2300 and used as an input to the loss function LF in block 2400.
[0176] As illustrated by the dashed arrows in Figure 2A 、 2B , using the identity operator in block 2300 can also be considered as bypassing block 2300 and using the predicted 3D eye state 3DES({E C ,E G ,E P}) which was determined as the output of the eye state predictor ESP in block 2200 upon receiving the first eye-related observation ERO1(t k ) as the prediction П and the input to the loss function LF in block 2400, respectively.
[0177] According to Figure 2C the embodiment illustrated in, a method 2000' for training an eye state predictor ESP includes: in block 2100, feeding a first eye-related observation ERO1 as an input to the eye state predictor ESP to determine a predicted 3D eye state 3DES{E C ,E G ,E P} as an output of the eye state predictor ESP, which may relate to an observation time t k and / or represented by the observation time t k , where the first eye-related observation relates to at least one eye of the object in and / or during the observation scenario; in block 2400', determining a training loss Δ based on the first eye-related observation ERO1 and a second eye-related observation, the second eye-related observation also relating to at least one eye of the object for the observation scenario (in and / or during); and in block 2500, using the training loss (Δ) to train the eye state predictor ESP.
[0178] In Figures 2A to 2C the embodiment, a corresponding loss function LF generally receives a corresponding 3D eye state {E C ,E G ,E P} determined differently compared to the 3D eye state ({E C ,E G ,E P}) as a second input (second eye-related observation), but also relating to at least one eye of the object for the observation scenario (in and / or during).
[0179] As Figure 2B further shown in, the method 2000 (direct supervision training) may be performed until the training loss Δ (or a running average of the training loss Δ) is below a predetermined threshold Δ th .
[0180] Referring to Figure 3A of FIG. 3B, a training method 3000 is explained. The method 3000 is also similar to the training method 1000 explained above with reference to Figure 1A , 1B and includes blocks 3100 to 3500, each of which is generally similar to the corresponding block 1100 to 1500 of the method 1000. However, the training method 3000 is more specific.
[0181] In Figure 3A , 3BIn an exemplary embodiment, two different differentiable predictors DP, DP1 are used in respective boxes 3300, 3301 to determine respective predictions P, P1 regarding at least one eye of an object during an observation scenario t k during the observation scenario t l .
[0182] Furthermore, in box 3400, a first prediction П and a second eye-related observation ERO2 regarding the same observation scenario t k can be fed as inputs to a first loss function LF to determine a first training loss Δ as the output of the first loss function LF
[0183] In an alternative, a first eye-related observation and a first prediction П are fed as inputs to a first loss function LF
[0184] Similarly, in box 3401, a second prediction П1 and a third eye-related observation regarding the same observation scenario t k can be fed as inputs to a second loss function LF1 to determine a second training loss Δ1 as the output of the second loss function LF1
[0185] For example, a differentiable renderer can be used in box 3300 to determine a synthetic left-eye image SI l and a synthetic right-eye image SI r (as a first prediction P), and in box 3400, they are respectively compared with the semantic segmentations S l and S r of the left-eye image P l and the right-eye image P r for determining a first loss Δ, where the left-eye image P l and the right-eye image P r are part of or even form the first eye-related observation
[0186] Furthermore, in box 3401, 2D gaze predictions G l and G r for the left and right eyes respectively can be determined as the second prediction П1 and compared with the corresponding 2D gaze values or labels L l and L r that have been independently determined
[0187] The first training loss Δ and the second training loss Δ1 can be used to train an eye state predictor ESP
[0188] Generally, in box 3450, a total training loss Δ2 is determined as a function f(Δ, Δ1) of the first training loss Δ and the second training loss Δ1, for example as a weighted average, and in box 3500, training can be performed based on the total training loss Δ2
[0189] Reference Figure 4A 、 4B 、4C explain the training method 4000. The method 4000 is also similar to the training method 1000 explained above with reference Figure 1A 、 1B and includes blocks 4100, 4200, 4300, 4400, 4500, each of which is generally similar to the corresponding blocks 1100 to 1500 of method 1000. However, the eye state predictor 3DES is trained separately according to the training method 4000 and determining the training losses Δ, Δ' depends on at least one case-specific parameter PAR.
[0190] In particular, determining the predicted 3D eye state 3DES in block 4200 and / or determining the prediction П in block 4300 as an output of the differentiable predictor DP can depend on one or more case-specific parameters PAR. The (one or more) case-specific parameters PAR used in blocks 4200, 4300 can be (at least partially) the same, but can also be completely different depending on the training settings.
[0191] For example, the differentiable predictor DP can receive the predicted 3D eye state 3DES (e.g., having the predicted 3D center E Figure 4C of the left eye eyeball of the object as shown in Cl , the predicted 3D gaze direction E Gl of the left eye, the predicted 3D pupil state E Pl of the left eye, the predicted 3D center E Cr of the right eye eyeball, the predicted 3D gaze direction E Gr of the right eye, and the predicted 3D pupil state EP r of the right eye for the binocular state 3DES b ) and hardware-specific parameters (such as the corresponding camera intrinsic characteristics, relative camera extrinsic characteristics, and the attitude of the inertial measurement unit of the object's head relative to at least one of the cameras) as inputs. Thus, the determination of the prediction P that can be determined as the synthetic eye image in block 4300 can be facilitated.
[0192] In addition, an additional loss function LF' (also referred to as the third loss function) can be used to determine an additional (or third) loss Δ' depending on the predicted binocular state 3DES b and (one or more) object-specific parameters (such as the interpupillary distance IPD and the angle between the optical axis and the visual axis).
[0193] For example, optionally depending on (one or more) object-specific parameters, the loss Δ' can be a symmetry loss, the predicted binocular state 3DES bA measure of the relationship between the predicted left-eye 3D state and the predicted right-eye 3D state or the deviation of the expected distance.
[0194] For example, the loss Δ' can depend on or even correspond to the IPD loss, i.e., the deviation of the IPD according to the predicted binocular state 3DES b from the physiological (average) IPD of the object or the measured IPD.
[0195] Similarly, the loss Δ' can depend on or even correspond to a loss (gaze angle symmetry loss) that is a deviation of the (expected) relationship between the 2D or 3D gaze angles of the left and right eyes, a loss that involves the difference between the pupils of the left and right eyes (such as pupil size difference, deviation of the (expected) 2D or 3D pupil orientation relationship with the pupils of the left and right eyes (pupil symmetry loss)), or any (other) symmetry loss of the predicted binocular state 3DES b .
[0196] Furthermore, the training of the eye state predictor ESP can be based on a total training loss Δ2, which is determined as a function g of two partial losses Δ, Δ', as also Figure 4C shown.
[0197] Furthermore, in particular, one or more object-specific parameters PAR (such as corneal radius, iris radius, etc.) may change during training.
[0198] For example and as indicated by the dashed arrow in Figure 4A , even the (one or more) object-specific parameters PAR can be learned by optimizing the loss function LF.
[0199] Regarding Figure 5A , a flowchart of a method 8000 for object-specific parameter calibration is explained.
[0200] In the first block 8100, a trained eye state predictor as explained herein is provided. For example, the trained eye state predictor can be obtained as explained above regarding Figures 4A - 4C .
[0201] In the subsequent block 8200, the corresponding first eye-related observation is fed as input to the trained eye state predictor to determine the predicted 3D eye state of at least one eye of a new object for a new observation scenario as the output of the eye state predictor. The first eye-related observation relates to at least one eye of the new object in and / or during the new observation scenario.
[0202] In the subsequent block 8300, the predicted 3D eye state and the current value of at least one object-specific parameter of the new object are fed as input to a differentiable predictor, for example Figure 4AA differentiable predictor LF1 to determine a prediction of at least one eye of a new object in and / or during a new observation scenario as an output of the differentiable predictor.
[0203] In a subsequent block 8400, a training loss is determined based on the prediction and at least one of a corresponding first eye-related observation and a corresponding second eye-related observation of at least one eye of the new object in and / or during the new observation scenario.
[0204] In a subsequent block 8600, the training loss is used to update a current value of at least one object-specific parameter.
[0205] The current value of at least one object-specific parameter can be updated according to an optimization technique.
[0206] Thereafter, method 8000 can return to block 8200.
[0207] After multiple cycles, the current updated value of at least one object-specific parameter can be output as a corresponding predicted value of at least one object-specific parameter and / or stored in a database.
[0208] Reference Figure 5B , explains the training method 5000. Method 5000 is also similar to the training method 1000 referenced above Figure 1A , 1B , and includes blocks 5100, 5200, 5300, 5400, 5500, each of which is generally similar to the corresponding blocks 1100 to 1500 of method 1000. However, the training method 5000 is more specific.
[0209] In a first block 5100, a left-eye image P of an object for an observation scenario is determined l and a right-eye image P r , for example, captured using a corresponding eye camera or retrieved from a database.
[0210] In a subsequent block 5200, an eye state predictor is used to determine a predicted 3D eye state 3DES{E l , E r}, E C , E G , E P} of the left-eye image P
[0211] and the right-eye image P l and a synthetic left-eye image SI r is determined for the predicted 3D eye state 3DES (as a prediction Π obtained from the predicted 3D eye state 3DES).
[0212] In addition, the semantic segmentation S of the left-eye image P is determined in another block 5350 after the block 5100 l and the semantic segmentation S of the right-eye image P l . r . r
[0213] Based on the comparison between the synthetic eye image SI l , SI r and the semantic segmentation S l , S r , the training loss Δ is determined in the block 5400.
[0214] In a subsequent block 5500, the training loss Δ can be used to train an eye state predictor (model), such as an NN, in particular a CNN using a machine learning algorithm (in particular a corresponding optimization algorithm).
[0215] Thereafter, the method 5000 can return to the block 5100, as indicated by the dashed arrow.
[0216] For example, determining the training loss Δ can be based on a first comparison between the synthetic left-eye image SI l and the semantic segmentation S of the left-eye image P l , and / or based on a second comparison between the synthetic right-eye image SI l and the semantic segmentation S of the right-eye image P r , and typically based on the first comparison and the second comparison. r . r
[0217] Alternatively or more typically, in addition to determining the training loss Δ in the block 5400, an additional training loss Δ' can be determined in the block 5450 based on the comparison between the synthetic left-eye image SI l and the synthetic right-eye image SI r , and used to train the eye state predictor in the block 5500.
[0218] Figure 5C The system 500 for performing the method explained herein is illustrated, in particular the training method 1000 - 5000.
[0219] In an exemplary embodiment, the system 500 includes a head-mounted device 100 implemented as a glasses device. Thus, the frame of the glasses device 100 has a front part 114 surrounding the left-eye opening and the right-eye opening. The bridging part of the front part 114 is arranged between the eye openings. In addition, the left temple 113 and the right temple 123 are attached to the front part 114.
[0220] The exemplary camera module 140 is received in the bridge portion and is arranged on the wearer side of the bridge portion. Channel openings for the scene camera 160 of the module 140 and the field of view (FOV) of the scene camera 160 are formed in the bridge portion.
[0221] The scene camera 160 is typically arranged centrally, i.e., at least near the central vertical plane between the left eye opening and the right eye opening and / or near the (one or more) (expected) eye midpoints of the human subject (user) wearing the head-mounted device 100. The latter also promotes a compact design. In addition, the influence of parallax errors on gaze prediction can be significantly reduced in this way.
[0222] In addition, in an exemplary embodiment, the scene camera 160 can define a Cartesian coordinate system x, y, z and have an optical axis that is at least substantially arranged in the central vertical plane (to reduce parallax errors), arranged in the central x, z plane, and / or pointing in the x direction.
[0223] The legs 134, 135 of the module 140 can be at least substantially complementary to the frame below the bridge portion such that the eye openings are at least substantially surrounded by the frame and the material of the module 100.
[0224] The right-eye camera 150 for taking an image of the user's right eye is arranged in the right leg portion 135. The right-eye image can be considered eye-related observation data and can form at least a part of the corresponding (primary) eye-related observation.
[0225] Similarly, the left-eye camera 150 for taking an image of the user's left eye can be arranged in the left leg portion 134 (for providing eye-related observation data).
[0226] In an exemplary embodiment, the head-mounted device 500 is additionally provided with an inertial measurement unit 170 for measuring the movement and / or orientation of the head-mounted device 500 and the human subject wearing the head-mounted device 500, respectively.
[0227] Both the scene camera 160 and the inertial measurement unit 170 can provide corresponding (primary) eye-related observation data for the user.
[0228] As indicated by the dotted arrows in Figure 5C the computing systems 200, 300 of the system 500 can be connected to the scene camera 160 for receiving scene images, can be connected to the (one or more) eye cameras 150 for receiving eye images, and can be connected to the inertial measurement unit 170 for receiving data related to the movement and / or orientation of the head-mounted device 500.
[0229] The computing systems 200, 300 are configured to perform the methods explained herein.
[0230] For this purpose, computing systems 200, 300 generally include one or more processors and a (corresponding) non-transitory computer-readable storage medium including instructions that, when executed by the one or more processors, cause system 500 to perform the methods as explained herein.
[0231] In an exemplary embodiment, computing systems 200, 300 include several interconnected parts, namely a first computing unit 200 and a second computing unit 300.
[0232] While the first computing unit 200 is generally a local unit or system and / or is configured to perform the prediction method as explained herein, the second computing unit 300 can be a remote unit or system (e.g., even cloud-based) and / or is generally configured to perform the actual training steps (cycles) of the training method as explained herein (after receiving eye-related observation data via the first computing unit 200).
[0233] The second computing unit 300 can be configured to determine eye-related observations from the received eye-related observation data and even host and / or manage a database of eye-related observations.
[0234] The first computing unit 200 can be implemented as a controller and can even be at least partially arranged within the housing of module 140.
[0235] However, the first computing system 200 can also be at least partially provided by an accompanying device of the controller that can be (and is) connected to the head-mounted device 100 (e.g., via a USB connection (e.g., one of the temple pieces can provide a corresponding plug or socket), in particular a mobile accompanying device such as a smartphone, a tablet computer, or a laptop computer).
[0236] In addition, the second computing unit 300 can be configured to upload the trained eye state predictor tESP to the second computing unit 200.
[0237] According to an embodiment of a method for training an eye state predictor model, which can be implemented as a neural network, particularly as a convolutional neural network, the method includes feeding a first eye-related observation as an input to the eye state predictor to determine a predicted 3D eye state of at least one eye of an object for an observation scenario (which can be represented by an observation ID or an observation time) as an output of the eye state predictor model, where the first eye-related observation relates to at least one eye of the object during the observation scenario and / or includes first eye-related observation data relating to at least one eye of the object during the observation scenario. The predicted 3D eye state is fed as an input to a differentiable predictor to determine a prediction for at least one eye of the object during the observation scenario as an output of the differentiable predictor. Based on the prediction and at least one of the first eye-related observation and a second eye-related observation (relating to the observation scenario and / or including second eye-related observation data relating to at least one eye of the object during the observation scenario and different from the first eye-related observation data), particularly based on the prediction and one of the first eye-related observation and the second eye-related observation, a (current) training loss is determined. The training loss is used to train the eye state predictor model and / or change the training parameters of the eye state predictor. In particular, the eye state predictor model can be trained respectively by machine learning and using machine learning algorithms (particularly corresponding optimization algorithms).
[0238] Although various exemplary embodiments of the present invention have been disclosed, it will be apparent to those skilled in the art that various changes and modifications can be made, which will achieve some of the advantages of the present invention without departing from the spirit and scope of the present invention. It will be obvious to those of ordinary skill in the art that other components performing the same functions can be appropriately substituted. It should be mentioned that features explained with reference to a particular drawing can be combined with features of other drawings, even in those cases where this is not explicitly mentioned therein. Such modifications to the inventive concept are intended to be covered by the appended claims.
[0239] Although processes may be depicted in the drawings in a particular order, if not otherwise stated, this should not be construed as requiring such operations to be performed in the particular order shown or in a sequential order to obtain the desired result. In some cases, multitasking and / or parallel processing may be advantageous.
[0240] For ease of description, spatially relative terms such as "below", "beneath", "lower", "above", "upper", etc. are used to explain the positioning of one element relative to a second element. These terms are intended to encompass different orientations of the device in addition to the different orientations depicted in the figures. Additionally, terms such as "first", "second", etc. are also used to describe various elements, regions, parts, etc., and are not intended to be limiting. Throughout the description, like terms refer to like elements.
[0241] As used herein, the terms "having", "containing", "including", "comprising", etc. are open terms indicating the presence of the stated element or feature but not excluding additional elements or features. The articles "a", "an", and "the" are intended to include the plural and the singular unless the context clearly indicates otherwise.
[0242] Considering the variations and applications of the above scope, it should be understood that the present invention is not limited by the foregoing description nor by the accompanying drawings. Instead, the present invention is limited only by the appended claims and their legal equivalents. Reference numerals 100 Head-mounted device 114 Front part of the frame 113, 123 Temple arms 140 Camera module 150 (Right) eye camera 160 Scene camera 170 Inertial measurement unit 200, 300 Computing system / controller / computing unit / accompanying device 500 System 1000 - 1430 Method / method steps 3DES{E C ,E G ,E P} 3D eye state DP, DP1 Differentiable predictor ERO, ERO1, P l ,P r Eye-related observation ESP Eye state predictor (model) pERO Primary eye-related observation sERO Secondary eye-related observation P l ,P r Eye image S l ,S r P l ,P r Semantic segmentation of Δ, Δ1, Δ2, Δ’ training losses Ps scene image Ss P S semantic segmentation of Π, Π1 predictions 10005500 method, method steps
Claims
1. A method (1000, 2000, 3000, 4000, 5000) for training an eye state predictor (ESP), the eye state predictor being implemented as a neural network (NN), particularly as a convolutional neural network (CNN), the method comprising: - Feed a first eye-related observation (ERO, ERO1, P l , P r ) as an input feed (1200, 2200, 3200, 4200, 5200) to the eye state predictor (ESP) to determine a predicted 3D eye state (3DES, {E k ) of at least one eye of the object for an observation scenario (t C , E G , E P ) as an output of the eye state predictor (ESP), the first eye-related observation (ERO, ERO1, P l , P r ) relating to the at least one eye of the object in and / or during the observation scenario (t k ); - Feed the predicted 3D eye state ({E C , E G , E P}) as input (1300, 2300, 3300, 4300, 5300) to a differentiable predictor (DP, DP1) to determine the prediction (П, П1) of the at least one eye of the object in and / or during the observation scenario (t k ) as the output of the differentiable predictor (DP); - Based on the prediction (П, П1) and at least one of the first eye-related observations (ERO, ERO1, P l , P r ), and the second eye-related observation, determine (1400, 2400, 3400, 3401, 4400, 4500) a training loss (Δ, Δ1, Δ2, Δ’), where the second eye-related observation (ERO2) relates to the at least one eye of the subject in and / or during the observation scenario (t k ); and - Using (1500, 2500, 3500, 4500, 5500) the training loss (Δ, Δ1, Δ2, Δ’) to train the eye state predictor (ESP).
2. The method (1000, 2000, 3000, 4000, 5000) according to claim 1, wherein, The observed situation (t k ) relates to a corresponding observation time (t k ) and / or is represented by at least one of the observation time (t k ) and an observation ID that typically depends on the observation time (t k ), wherein at least one of the first eye-related observation (ERO, P l , P r ) and the second eye-related observation (ERO2) includes, as a primary eye-related observation (pERO), an eye image (P k ) of the eyes of the object recorded by an eye camera during the observed situation and / or within the time (t l , P r ), and at least one of the corresponding secondary eye-related observations (sERO) derived from the eye image (P l , P r ), the corresponding secondary eye-related observation (sERO) typically including at least one visual feature extracted from the eye image (P l , P r ) and / or being implemented as a visual feature representation of the eye image (P l , P r ) including the at least one visual feature, in particular being implemented as a semantic segmentation (S l , S r ) of the eye image (P l , S r ).
3. The method (1000, 2000, 3000, 4000, 5000) according to claim 1 or 2, wherein, The first eye-related observation (ERO, P l , P r ) and at least one of the second eye-related observations (ERO2) includes, as the respective primary eye-related observation (pERO), the left-eye image (P k ) of the left eye of the object and the right-eye image (P l ) of the right eye of the object, recorded by the respective eye camera during the observation scenario and / or within the observation time (t r ), and / or wherein at least one of the first eye-related observations (ERO, P l , P r ) and the second eye-related observations (ERO2) includes the corresponding secondary eye-related observations (sERO) derived from the left-eye image and the right-eye image (P l , P r ), in particular the semantic segmentation (S l ) of the left-eye image (P l ) and the semantic segmentation (S r ) of the right-eye image (P r ).
4. The method (1000, 2000, 3000, 4000, 5000) according to claim 2 or 3, wherein, The related eye image (P l , P r ) and the semantic segmentation (S l , S r ) include at least one label (L l , L r ) selected from the list consisting of pupil label, iris label, sclera label, eyelid label, skin label, and lash label.
5. The method (1000, 2000, 3000, 4000, 5000) according to any one of the preceding claims, further comprising: - Usually, a scene camera (160) determines (1100, 2100, 3100) a scene observation (P s ), and the scene observation (P s ) involves the field of view of at least one eye of the object as an eye-related observation (ERO) during the observation scenario and / or within the observation time (t k ).
6. The method (1000, 2000, 3000, 4000, 5000) according to claim 5, wherein, The scene observation (P s ) includes, as a primary eye-related observation (pERO), a scene image (P s ) typically recorded by the scene camera and at least one of at least one corresponding secondary eye-related observation derived from the scene image (P s ).
7. The method (1000, 2000, 3000, 4000, 5000) according to any one of claims 2 to 6, wherein, The at least one corresponding secondary eye-related observation (sERO) includes the semantic segmentation (S s ) of the scene image (P l , S r ), at least one of a 3D fixation point typically measured in scene camera coordinates and a 2D fixation point typically measured in scene camera image coordinates.
8. The method (1000, 2000, 3000, 4000, 5000) according to any one of the preceding claims, wherein, The first eye-related observation (ERO, P l , P r ) and at least one of the second eye-related observations include at least one of the following as the respective main observation: the corneo-retinal resting potential of at least one eye of the object during the observation scenario and / or during the observation time (t k ), the velocity of the head of the object during the observation scenario and / or during the observation time (t k ), the acceleration of the head of the object during the observation scenario and / or during the observation time (t k ), the orientation of the head of the object during the observation scenario and / or during the observation time (t k ), the respective 2D fixation point during the observation scenario and / or during the observation time (t k ) and the 3D fixation point during the observation scenario and / or during the observation time (t k ).
9. The method (1000, 2000, 3000, 4000, 5000) according to any one of the preceding claims, wherein, The second eye-related observation (ERO2) is different from the first eye-related observation (ERO1), wherein the first eye-related observation (ERO1) and the second eye-related observation (ERO2) are used to determine the training loss (Δ, Δ1, Δ2, Δ’), wherein the first eye-related observation (ERO1) includes or even is the primary eye-related observation (pERO), and / or wherein the second eye-related observation includes or even is the secondary eye-related observation (sERO).
10. The method (1000, 2000, 3000, 4000, 5000) according to any one of the preceding claims, wherein, Using a head-mounted device including the eye camera to determine the corresponding eye-related observation, the head-mounted device generally including at least one of a left eye camera, a right eye camera (150), a scene camera (160), and an inertial measurement unit (170).
11. The method (1000, 2000, 3000, 4000, 5000) according to any one of the preceding claims, further comprising at least one of the following: - Using an eye camera (150) to determine (1100, 2100, 3100) at least the first eye-related observation (ERO, P k ) during the observation scenario and / or within the observation time (t l , P r ), typically multiple eye-related observations (ERO, P k ) at respective observation times (t l , P r ); - Store the first eye-related observation (ERO, P l , P r ), typically the multiple eye-related observations (ERO, P l , P r ) in a database; - Using the database to determine the first eye-related observation; - Using the database to determine the second eye-related observation; - Feed a number of eye-related observations (P l , P r , S s , S l , S r ) as input feeds (1200, 2200, 3200) to the eye state predictor (ESP) to determine the corresponding predicted 3D eye states ({E C , E G , E P ) of the at least one eye.
12. The method (1000, 2000, 3000, 4000, 5000) according to any one of the preceding claims, wherein, Determine a plurality of corresponding eye-related observations (ERO, P l , P r ), wherein said corresponding eye-related observations (ERO, P k ) are determined for different objects and / or for a number of corresponding observation times (t l ), P r ), and / or wherein the method (1000, 2000, 3000, 4000, 5000) is performed iteratively.
13. The method (1000, 2000, 3000, 4000, 5000) according to any one of the preceding claims, wherein, The predicted 3D eye state ({E C 、E G 、E P}) includes at least one of the following, typically three or even all: the predicted 3D center (E C ) of the eyeball of the at least one eye, the predicted 3D gaze direction (E G ) of the at least one eye, the predicted 3D state of the eyelid of the at least one eye and the predicted 3D state of the pupil of the at least one eye (E P ) 14. The method (1000, 2000, 3000, 4000, 5000) according to claim 13, wherein, The 3D state of the pupil includes at least one of the following: the predicted 3D pupil size of the at least one eye, the predicted 3D pupil aperture of the iris of the at least one eye, the predicted 3D pupil radius of the iris of the at least one eye, and the predicted 3D pupil diameter of the iris of the at least one eye.
15. The method (1000, 2000, 3000, 4000, 5000) according to claim 13 or 14, wherein, The 3D state of the eyelid of the at least one eye includes at least one of eyelid shape, eyelid position, and eyelid closure percentage.
16. The method (1000, 2000, 3000, 4000, 5000) according to any one of the preceding claims, wherein, The predicted 3D eye state ({E C 、E G 、E P}) is a monocular state (3DES m ), a pair of corresponding left and right monocular 3D eye states, or a binocular state (3DES b ), where the predicted 3D eye state ({E C 、E G 、E P}) includes and / or is one of a 6-dimensional vector, a 10-dimensional vector, an 11-dimensional vector, and a 12-dimensional vector, and / or where the binocular state (3DES b ) is represented by a data set having a lower dimension than the sum of the dimensions of the corresponding left and right monocular 3D eye states.
17. The method (2000) according to any one of the preceding claims, wherein The differentiable predictor (DP) is the identity operator (I), and wherein the predicted 3D eye states ({E C , E G , E P}) and the corresponding 3D eye states ({E k ) within the observed scenario and / or the observation time (t C , E G , E P}’) are used to determine the training losses (Δ, Δ1, Δ2, Δ’).
18. The method (1000, 2000, 3000, 4000, 5000) according to claim 17, wherein, Using the first eye-related observation (ERO1, P C , E G , E P ) different from the 3D eye state ({E l , P r ) or using the second eye-related observation (ERO2) to determine the corresponding 3D eye state ({E C , E G , E P}’).
19. The method (1000, 2000, 3000, 4000, 5000) according to claim 18, wherein, Using a 3D geometric eye model, particularly a 3D geometric eye model that takes into account corneal refraction, to determine the corresponding 3D eye state ({E C , E G , E P}’).
20. The method (1000, 3000, 4000, 5000) according to any one of claims 1 to 16, wherein, The differentiable predictor (DP) is different from the identity operator (I).
21. The method (1000, 2000, 3000, 4000, 5000) according to any one of claims 1 to 16 and 20, wherein, The differentiable predictor (DP) is configured to output a synthetic eye image (SI l , SI r ) in response to receiving the input, and the training loss (Δ) is determined based on the synthetic eye image (SI l , SI r ) and the corresponding visual feature representations extracted from the eye image (P l , P r ), in particular the corresponding semantic segmentation (S l , S r ) of the eye image (P l , S r ).
22. The method (1000, 2000, 3000, 4000, 5000) according to any one of claims 1 to 16, 20 and 21, wherein, The differentiable predictor (DP) is implemented as a generative neural network or a differentiable renderer, such as a differentiable approximate ray tracing algorithm.
23. The method (1000, 2000, 3000, 4000, 5000) according to any one of the preceding claims, wherein, At least one loss function (LF, LF1, LF2, LF') is used to determine the training loss (Δ, Δ1, Δ2, Δ'), where at least two differentiable predictors (DP, DP1) are used to determine the at least one eye of the object during the observation scenario and / or the corresponding prediction (П, П1) within the observation time (t k )), where the training loss (Δ, Δ1, Δ2, Δ') is determined based on each of the corresponding predictions (П, П1), where the predicted 3D eye state ({E C , E G , E P}) is fed as input to a first differentiable predictor (DP) to determine the first prediction (П) of the at least one eye of the object during the observation scenario and / or within the observation time (t k ) as the output of the first differentiable predictor (DP), where the predicted 3D eye state ({E C , E G , E P}) is fed as input to a second differentiable predictor (DP1) to determine the second prediction (П1) of the at least one eye of the object during the observation scenario and / or within the observation time (t k ) as the output of the second differentiable predictor (DP1), where the first prediction (П) and the first eye-related observation (ERO1) or the second eye-related observation (ERO2) are fed as input to a first loss function (LF) to determine a first training loss (δ) as the output of the first loss function (LF); where the second prediction (П1) and the first eye-related observation (ERO1), the second eye-related observation (ERO2) or the third eye-related observation (ERO3) are fed as input to a second loss function (LF1) to determine a second training loss (Δ1) as the output of the second loss function (LF1), where the second training loss (Δ1) is used to train the eye state predictor (ESP), where the first training loss (Δ) and the second training loss (Δ1) are used to train the eye state predictor (ESP), and / or where the training loss is determined as a function of the first training loss (Δ) and the second training loss (Δ1).
24. The method (1000, 2000, 3000, 4000, 5000) according to any one of the preceding claims, wherein, The training loss (Δ, Δ1, Δ2, Δ’) depends on at least one case-specific parameter (PAR), and / or wherein the training loss (Δ, Δ1, Δ2, Δ’) is used to correct and / or learn at least one case-specific parameter (PAR).
25. The method (1000, 2000, 3000, 4000, 5000) according to claim 24, wherein, Both the eye state predictor (ESP) and the differentiable predictor (DP) receive the at least one case-specific parameter (PAR) as part of the input.
26. The method (1000, 2000, 3000, 4000, 5000) according to claim 24 or 25, wherein, The at least one case-specific parameter (PAR) includes at least one object-specific parameter.
27. The method (1000, 2000, 3000, 4000, 5000) according to claim 26, wherein, The at least one object-specific parameter is selected from the list consisting of: interpupillary distance (IPD), the angle between the optical axis and the visual axis of the corresponding eye, a rotation operator capturing the transformation between the optical axis and the visual axis, geometric parameters related to the shape and / or size of the cornea of the corresponding eye such as spherical, aspherical, thickness, astigmatism, 3D topographies, and radius, the refractive index of at least one part of the corresponding eye, the iris radius of the corresponding eye, the pupil shape, and geometric parameters related to the shape and / or size of the eyeball of the corresponding eye.
28. The method (1000, 2000, 3000, 4000, 5000) according to any one of claims 24 to 27, wherein, The at least one situation-specific parameter (PAR) includes hardware-specific parameters.
29. The method (1000, 2000, 3000, 4000, 5000) according to claim 28, wherein, The at least one hardware-specific parameter is selected from the list consisting of: the corresponding camera intrinsic characteristics, the relative camera extrinsic characteristics, and the attitude of the inertial measurement unit relative to at least one of the cameras.
30. The method (1000, 2000, 3000, 4000, 5000) according to any one of the preceding claims, further comprising at least one of the following: - Use the predicted 3D eye state ({E C , E G , E P}) to determine a third training loss (Δ'); - Based on the predicted 3D eye state ({E C , E G , E P}) and the at least one situation - specific parameter (PAR), use a third loss function (LF’) to determine a third training loss (Δ’); and - using (1500, 2500, 3500) the third training loss (Δ’) to train the eye state predictor (ESP), and / or wherein the training loss (Δ, Δ1, Δ2, Δ’) is determined as a function (g) of the first training loss (Δ) and the third training loss (Δ’), for example as a function of the first training loss (Δ), the second training loss (Δ1), and the third training loss (Δ’).
31. A method (8000) for calibrating object-specific parameters, the method comprising: - providing (8100) a trained eye state predictor (tESP), the trained eye state predictor (tESP) being trained according to any one of the preceding claims; and performing the following steps, typically several times: o Feed the corresponding first eye-related observation (ERO, ERO1, P l , P r ) as an input feed (8200) to the trained eye state predictor (tESP) to determine the predicted 3D eye state (3DES, {E k ) of at least one eye of the new object for the observation scenario (t C , E G , E P}) as the output of the trained eye state predictor (tESP), the first eye-related observation (ERO, ERO1, P l , P r ) relating to said at least one eye of said new object in and / or during said observation situation (t k ) o Feed the predicted 3D eye state ({E C , E G , E P}) and the current value of at least one object-specific parameter of the new object as inputs (8300) to a differentiable predictor (DP, DP1) to determine the prediction (П, П1) of the at least one eye of the new object in and / or during the observation scenario (t k ) as the output of the differentiable predictor (DP); o based on the prediction (П, П1) and the first eye-related observation (ERO, ERO1, P l , P r ), and at least one of the corresponding second eye-related observations, determine (8400) a training loss (Δ, Δ1, Δ2, Δ'), wherein the corresponding second eye-related observation (ERO2) relates to said at least one eye of said new object in and / or during said observation scenario (t k ); and / or said at least one eye of said new object during and o using (8600) the training loss (Δ, Δ1, Δ2, Δ’) to update the current value of the at least one object-specific parameter.
32. A method (9000) for real-time predicting the 3D eye state (3DES) of an object, the method comprising: - determining an eye-related observation (ERO*) related to at least one eye of the object; and - Feed the eye-related observation (ERO*) as an input to the trained eye state predictor (tESP) implemented as a neural network (NN) to determine (5300) a predicted 3D eye state (3DES*, {E C , E G , E P}) as an output of the trained eye state predictor (tESP).
33. The method (9000) according to claim 32, wherein, The predicted 3D eye state (3DES*, {E C , E G , E P ) includes a predicted 3D rotation center of the eyeball of the at least one eye and a 3D gaze direction of the eyeball of the at least one eye.
34. The method (9000) according to claim 33, wherein, The predicted 3D eye state (3DES*, {E C , E G , E P ) includes the predicted 3D state of the pupil of the at least one eye.
35. The method (9000) according to claim 34, wherein, The predicted 3D state of the pupil includes at least one of the predicted 3D pupil size of the at least one eye, the predicted 3D pupil aperture of the iris of the at least one eye, the predicted 3D pupil radius of the iris of the at least one eye, and the predicted 3D pupil diameter of the iris of the at least one eye.
36. The method (9000) according to any one of claims 32 to 35, wherein, The trained eye state predictor (tESP) is trained according to the method according to any one of claims 1 to 30, and / or the method includes training an eye state predictor (ESP) according to the method according to any one of claims 1 to 30 to obtain the trained eye state predictor (tESP).
37. The method (9000) according to any one of claims 32 to 36, wherein, The eye camera (150) of the head-mounted device (100) worn by the object is used to capture the at least one eye image (P l , P r ), the head-mounted device (100) is generally implemented as a glasses device, and / or wherein the method is at least partially controlled and / or executed by the computing and control unit of the head-mounted device.
38. The method (9000) according to any one of claims 32 to 37, wherein, The training losses (Δ, Δ1, Δ2, Δ') depend on at least one object-specific parameter (PAR), and wherein the object-specific parameter (PAR) is predetermined for the new object, in particular as a corresponding physiological value, as a corresponding measured value, using any function for determining a corresponding predicted value or one or more of the values according to the method of claim 31.
39. A system (500) comprises: - Head-mounted device (100), the head-mounted device (100) includes at least one eye camera (150), the eye camera (150) is configured to generate an eye image (P l , P r ) of at least a part of the eyes of the object wearing the head-mounted device and - a computing system (200, 300) connectable to the at least one eye camera (150) for receiving the eye images and configured to: o Based on the observed situation (t k ) and from the at least one eye camera (150) Received eye image (P l , P r ), generating a first eye-related observation (ERO1, P k ) of the eye of the object in and / or during the observed situation (t l , P r ); o run an eye state predictor (ESP) implemented as a neural network (NN); o Feed the first eye-related observation (ERO1, P l , P r ) as an input feed (1200, 2200, 3200) to the eye state predictor (ESP) to determine the predicted 3D eye state (3DES, {E k ) of the eyes of the subject in and / or during the observation scenario (t C , E G , E P}); o Feed the predicted 3D eye state ({E C , E G , E P}) as input (1300, 2300, 3300) to a differentiable predictor (DP, DP1) to determine the prediction (P, k ) of the eyes of the object in and / or during the observation scenario (t P1) as an output of the differentiable predictor (DP); o based on the prediction (P, P1) and the first eye-related observation (ERO1, P l , P r ), and at least one of the second-eye related observations, determine (1400, 2400, 3400, 3401) Training losses (Δ, Δ1, Δ2, Δ'), the second eye-related observation (ERO2) relates to the eyes of the object in and / or during the observation scenario (t k ) and the computing system (200, 300) is generally configured to determine the second eye-related observation (ERO2) different from the first eye-related observation (ERO1); and o use (1500, 2500, 3500) the training losses (Δ, Δ1, Δ2, Δ') to train the eye state predictor (ESP).
40. The system (500) according to claim 39, wherein, The head-mounted device (100) includes respective eye cameras (150) for each eye of the subject, wherein the first eye-related observation (ERO1, P l , P r ) is based on the eye images (P l , P r ) generated during the observation scenario and / or at the observation time (t k ) at which the eye images (P l , P r ) are generated for each eye, wherein the head-mounted device (100) includes a scene camera (160) configured to generate a scene image related to the field of view of the subject wearing the head-mounted device, and / or wherein the computing system is configured to o Hosting or accessing a database of the eye-related observations (ERO, P l , P r ); and / or o perform the method (1000-5000, 8000) according to any one of claims 1 to 31.
41. A computer program product or computer-readable storage medium comprising instructions which, when executed by one or more processors of a computing system, cause the computing system to perform any of the steps of the method (1000-5000, 8000) according to any one of claims 1 to 31.
Citation Information
Patent Citations
Devices, systems and methods for predicting gaze-related parameters
WO2020244752A1