Object Observation and Tracking in Images Using an Encoder-Decoder Model
A composite model using convolutional and linear networks trained with real and synthetic data addresses the challenge of limited training data in augmented reality gaze tracking, ensuring accurate and efficient gaze prediction.
Patent Information
- Application Number
- JP2023579020
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2021-06-22
- Filing Date
- 2022-06-22
- Publication Date
- 2025-07-02
- Estimated Expiration
- 2042-06-22
AI Technical Summary
Existing machine learning models for line of sight tracking in augmented reality applications face challenges due to limited actual training data and variations in human eye features, leading to inaccurate predictions and impaired robustness.
A composite model comprising a first convolutional neural network for feature extraction, a second neural network for segmentation, and a third linear network for gaze prediction is trained using both real and synthetic data, then optimized to reduce complexity and resource usage, allowing for accurate gaze prediction in augmented reality systems.
The model achieves robust and accurate real-time gaze tracking with reduced computational requirements, effectively handling variations in human eye features and improving prediction accuracy.
Smart Images

Figure 0007701995000001 
Figure 0007701995000002 
Figure 0007701995000003
Abstract
Description
Technical Field
[0001] Cross - Reference to Related Applications This application is a continuation of U.S. Non - Provisional Patent Application No. 17 / 304,505, entitled "Object Observation Tracking in Images Using an Encoder - Decoder Model", filed on Jun. 22, 2021, the entire disclosure of which is incorporated herein by reference and for which priority is claimed.
[0002] Field Embodiments relate to determining where a user is looking to track the user's line of sight or eyes for an augmented reality device.
Background Art
[0003] Background Machine learning techniques using convolutional networks can be used for predicting and tracking the line of sight (e.g., eye direction) in augmented reality (AR) applications. A convolutional network can successively apply a plurality of convolutional layers and pooling layers. A convolutional network can start from an (N×N) high - resolution image and generate a spatially pooled feature map of dimension N / m × N / m × F, where F is the number of feature channels.
Summary of the Invention
[0004] Summary Embodiments relate to using a machine learning model (e.g., an encoder - decoder model, a convolutional neural network (CNN), a linear network, etc.) to track the observation of an object within an image (e.g., eye - gaze tracking). These machine - learned models can also be used to predict the characteristics of data for which the model is trained in a first configuration and used in a second configuration.
[0005] In a general aspect, a device, a system, a non-transitory computer-readable medium (storing computer-executable program code executable on a computer system), and / or a method, in a training phase, includes training a gaze prediction model including a first model and a second model, the first model and the second model being configured to cooperate to predict segmentation data based on training data, the method further includes training a third model together with the first model and the second model, the third model being configured to predict training characteristics using an output of the first model based on training data, and the method, in an operation phase, further includes receiving operation data and predicting operation characteristics using the trained first model and the trained third model, and can execute a process involving a method.
[0006] An embodiment can include one or more of the following features. For example, the training data can include an image of an eye, the predicted segmentation data can include an eye region, and the training characteristic can be an eye gaze. The motion data can include an image of an eye captured using an augmented reality (AR) user device, and the motion characteristic can be an eye gaze. Training the gaze prediction model can include generating a first feature map based on training data using a first model, generating a second feature map based on the first feature map using a second model, predicting segmentation data based on the second feature map, generating a loss associated with the predicted segmentation data, and training at least one of the first model and the second model based on the loss and a loss associated with the training characteristic. Training the gaze prediction model can include generating a first feature map based on training data using a first model, predicting a training characteristic based on the first feature map using a third model, generating a loss associated with the predicted training characteristic, and training the third model based on the loss and a loss associated with the segmentation data. Predicting the motion characteristic can include generating a feature map based on the motion data using the first model and predicting the motion characteristic based on the feature map using the third model. The first model can be a first convolutional neural network (CNN), the second model can be a second CNN including at least one skip connection from the first CNN, and the third model can be a linear neural network. The second model can be removed from the gaze prediction model used in the operation phase. Training the gaze prediction model can include changing at least one of the parameters, features, and feature characteristics associated with at least one of the first model, the second model, and the third model.The method can further include, in a calibration phase before an operation phase, training at least one of a first model, a second model, and a third model based on user data captured using an AR user device.
[0007] In other general aspects, a device, a system, a non-transitory computer-readable medium (storing computer-executable program code executable on a computer system), and / or a method can include generating a first feature map based on training data using a first model; generating a second feature map based on the first feature map using a second model; predicting segmentation data based on the second feature map; generating a first loss associated with the predicted segmentation data; predicting training characteristics based on the first feature map using a third model; generating a second loss associated with the predicted training characteristics; and training at least one of the first model, the second model, and the third model based on the first loss and the second loss. A process involving the method can be executed.
[0008] An embodiment can include one or more of the following features. For example, the training data can include an image of an eye, the predicted segmentation data can include an eye region, and the training characteristic can be an eye gaze. The first model can be a first convolutional neural network (CNN), the second model can be a second CNN including at least one skip connection from the first CNN, and the third model can be a linear neural network. In the method according to claim 11, the second model can be removed for use in an operating phase. Training at least one of the first model, the second model, and the third model can include changing at least one of the parameters, features, and feature characteristics associated with at least one of the first model, the second model, and the third model. The segmentation data can be a process including separating data associated with a second feature map into separate groups. Predicting the segmentation data can include a detection process and a suppression process.
[0009] In yet another general aspect, a device, a system, a non-transitory computer-readable medium (storing computer-executable program code executable on a computer system), and / or a method includes generating a feature map based on operation data using a trained first model and predicting a characteristic based on the feature map using a trained second model, wherein training the first model and the second model includes using a third model to predict segmentation data during training and further includes removing the third model for predicting the characteristic, and can execute a process involving a method.
[0010] Embodiments can include one or more of the following features. For example, the motion data can include an image of the eye captured using an augmented reality (AR) user device, and the motion characteristic can be the line of sight of the eye. The method can further include further training at least one of a trained first model, a trained second model, and a third model based on user data captured using an AR user device.
[0011] Exemplary embodiments will become more fully understood from the following detailed description given herein and the accompanying drawings, wherein like elements are represented by like reference numerals throughout the detailed description and the accompanying drawings, which are given by way of illustration only and thus are not limiting of the exemplary embodiments.
Brief Description of the Drawings
[0012]
Figure 1
Figure 2A
Figure 2B
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Modes for Carrying Out the Invention
[0013] It should be noted that these figures are intended to illustrate the general characteristics of the methods, structures, and / or materials utilized in certain exemplary embodiments and to supplement the written description provided below. However, these drawings are not to scale and may not accurately reflect the exact structural or performance characteristics of a given embodiment, and should not be construed as defining or limiting the range of values or characteristics encompassed by the exemplary embodiments. For example, the relative thicknesses and positions of regions and / or structural elements may be reduced or exaggerated for clarity. The use of like or the same reference numerals in the various drawings is intended to indicate the presence of like or the same elements or features.
[0014] Detailed Description The disclosed technology provides object observation (e.g., line of sight) prediction and tracking for objects in recently acquired (e.g., captured) images with higher accuracy and easier customization than conventional systems. Machine learning techniques using convolutional networks can be used for data characteristic prediction, line of sight (e.g., eye direction) tracking, and / or the like in augmented reality (AR) applications and other live data capture systems. In these convolutional networks, a large amount of training data may be required. In some scenarios, such as line of sight tracking applications, the supply of actual training data may be limited. Thus, in embodiments, actual training data may be augmented with synthetic training data, e.g., computer-generated images. However, some object observation tasks, including line of sight tracking, may include many variations and special cases due to differences in human eyes (such as iris texture, pupil shape / size, etc.), makeup, changes in skin color / texture, etc. These variations are difficult to capture with synthetic (e.g., computer-generated) data and are limited in real data. As a result, the robustness of the trained models, including AR systems, may be impaired by inappropriate (inaccurate) predictions. Eye line of sight or eye line of sight tracking can include determining or specifying a point in space. The point in space may represent the point that the user of the system is looking at. The point in space may represent the direction in which the user is looking (e.g., the direction relative to the surface of the user's eye). The point in space can have coordinates (e.g., Cartesian coordinates or x, y coordinates) and / or depth (such as z coordinate). The point in space can be a measurement variable that can be determined using elements (such as cameras) of the AR system.
[0015] The example embodiments can generate a robust yet less complex eye-tracking model that enables real-time operation on an AR system and can accurately handle the aforementioned changes. Machine learning techniques using models, algorithms, networks, neural networks, and / or convolutional neural networks (CNNs) can be used in the eye-tracking system. These models can be trained in a first configuration and used operationally in a second configuration. The first model configuration can be advantageous for highly robust training using limited actual training data, and the results of the training can be applied to the second configuration, which can be advantageous in terms of operation or use cases by reducing resource (e.g., processor, memory, and / or the like) utilization while still generating accurate predictions.
[0016] FIG. 1 shows a block diagram of an eye-tracking system according to an exemplary embodiment. As shown in FIG. 1, the eye-tracking system can include a first computing device 105 and a second computing device 120. The first computing device 105 can be an example of the computing device 700 or 750 shown in FIG. 7. The second computing device can be an example of the computing device 790 shown in FIG. 7. The first computing device 105 can include a model trainer 110 block and a model modifier 115 block, and the second computing device 120 can include a model implementer 125 block. The first computing device 105 can be associated with a product manufacturer and can be implemented, for example, on a server, a network-connected computer, a mainframe computer, a local computer, and the like. The second computing device 120 can be a user device and can be implemented, for example, on an AR headset (e.g., AR device 130), a mobile device, a laptop device, a mobile phone, a personal computer, and the like. The second computing device 120 can have limited computing resources compared to the first computing device 105.
[0017] The model trainer 110 can be configured to train a model (e.g., a statistical model, a neural network (e.g., a CNN), a linear network, an encoder-decoder network, etc.). During training, the model can be configured to predict object characteristics (such as a line of sight) and segments (or segmentations). In the prediction of object characteristics (such as a line of sight), the coordinates of the object characteristics (such as a line of sight) can be predicted. For example, the line-of-sight prediction can predict which coordinates of an image a human user is looking at. As another example, there can be an estimation of a body pose. For example, the coordinates of the joints of a human body can be estimated while at the same time segmenting an object of interest (such as a leg). At runtime, coordinates can be output without segmentation in order to save computation time. The model trainer 110 can generate a model trained using training data.
[0018] As shown in FIG. 1, the training data can be a plurality of training images 15. The training images 15 can be used to train a model for line-of-sight prediction and segmentation prediction for each training image. For example, the model can include a plurality of weights that are modified after each training iteration until the loss is minimized and / or until the change in one or more losses associated with the model is minimized.
[0019] The model modifier 115 can be configured to modify a model trained using the model trainer 110. For example, the trained model can include two or more models (referred to as a composite model), and the model modifier 115 can operably disconnect (or remove) at least one of the two or more models from the composite model. In an exemplary embodiment, the trained model is a composite model (e.g., two or more models trained together, co-trained, and / or optimized together) that can include a first model (sometimes referred to as an inference or encoder), a second model (sometimes referred to as a segmentation model or decoder), and a third model configured to predict object characteristics (such as a line of sight). The first model can be configured to generate (or infer) a feature map. The second model can be configured to predict a segmentation from the feature map. The third model can be configured to predict a characteristic (such as a line of sight) from the feature map. The segmentation can only be used when training a composite model (e.g., the model trainer 110). Thus, the model modifier 115 can be configured to modify the trained composite model by operably disconnecting (or removing) the second model from the composite model. Operably disconnecting (or removing) the second model can also include having a first trained composite model with the second model and a second trained composite model without the second model that holds the trained data (e.g., parameters, weights, etc.).
[0020] The model corrector 115 can also be configured to correct parameters related to a trained model (or composite model). The model corrector 115 can also use a plurality of training images 15 to adjust (e.g., further correct the parameters of) the trained model. The adjustment of the trained model can be performed after the trained model has been corrected. For example, adjusting the aforementioned trained composite model can include performing a training operation using the first model and the third model.
[0021] The model implementer 125 (e.g., associated with the second computing device 120) can include a trained (and optionally adjusted) model (e.g., a composite model or a first model and a third model). Thus, it can be said that the model implementer 125 uses the model in an operation or gaze prediction stage or mode. The model implementer 125 can be configured to determine (e.g., predict) the gaze 10 associated with the image 5. The image 5 can be captured by a computing device such as the AR device 130. Thus, the gaze 10 can be associated with the user of the AR device 130. In an exemplary embodiment, the output of the first network is input to the third network, and the third network can predict the gaze 10. The model implementation unit 125 can be configured to calibrate (e.g., further correct the weights) the model based on an image associated with the user of the AR device 130. In this embodiment, the model implementer 125 can include both the trained model and the corrected trained model. In this embodiment, the trained model can be further trained by the user (or an engineer cooperating with the user), and the parameters and / or weights associated with the further trained model can be used by the corrected trained model. In some embodiments, the model implementer 125 can access, for calibration, a model trained, for example, in the first computing device 105.
[0022] The training images can be real data (e.g., an image of an eye) and / or synthetic data (e.g., a computer-generated image of an eye). The real data and / or synthetic data can be in the form of (image[i], gaze_direction[i], i = 1, ... K), where the direction is represented by x, y coordinates. The real data and / or synthetic data can also be in the form of (image[i], gaze_direction[i], segmentation[i], i = 1, ... K), where the i-th segmentation data (associated with the i-th image) contains a plurality of data points each identifying a region of the eye (e.g., pupil, iris, sclera, etc.). The segmentation labels can be crowdsourced or generated through ML techniques with human assistance in cases where the calculation is complex. Since it takes time to acquire real data, it may be unrealistic to acquire sufficient real data to robustly train a model related to the gaze tracking system. To address this problem, an exemplary embodiment can train a first model (or a first composite model) that can be robustly trained using limited real data and synthetic data. This first model can then be modified to a second model (or a second composite model) for use in a user device including an operational gaze tracking system. The second model can use fewer processing resources than the first model, whereby the second model is typically suitable for the resources associated with the user device.
[0023] Accurate gaze prediction (e.g., x, y coordinates in an image) can depend on obtaining accurate information related to the pupil, the pupil center, and / or other eye regions. Segmentation data can be a factor in obtaining accurate gaze prediction. Thus, in an example embodiment, the gaze prediction model can be combined with a segmentation model as a composite model (e.g., a gaze and segmentation prediction model). The segmentation portion of the composite model can be used during training (e.g., by model trainer 110) to normalize and improve the accuracy of the gaze prediction portion of the composite model. An example embodiment can include the composite model as an inference segmentation model that can be split or decomposed (e.g., by model modifier 115) into a multi-resolution encoder network (e.g., for gaze tracking using model implementer 125), a multi-resolution encoder-decoder network (e.g., for training and / or calibration) that can be split or decomposed (e.g., by model modifier 115). In other words, the inference segmentation model can be split or decomposed (e.g., by model modifier 115) into an inference model (e.g., for gaze tracking using model implementer 125).
[0024] In the operation phase or the gaze tracking mode, the inference model (e.g., a multi-resolution encoder-decoder network with the decoder network removed) can operate on the feature maps used to generate gaze predictions or gaze outputs (e.g., gaze direction, gaze point, etc.). In this mode, the inference model may not include skip connections and other calculations necessary to realize the segmentation model output of the inference segmentation model. The complexity of the inference model is significantly lower than that of the inference segmentation model, but the inference model includes the same training parameters and / or weights. Therefore, the parameters and / or weights identified during the training phase of the inference segmentation model can be used in the inference model to obtain gaze predictions or gaze outputs. In other words, gaze predictions using the inference segmentation model (e.g., during the training phase) should be the same as gaze predictions using the inference model (e.g., during the implementation or operation phase).
[0025] In some embodiments, the inference segmentation model can be used in a calibration phase or a calibration mode. For example, the inference segmentation model can be fine-tuned for a specific user. The calibration mode can be implemented by the model corrector 115 and / or the model implementer 125. The operation of the first computing device 105 can be referred to as the training phase or the training mode, and the operation of the second computing device 120 can be referred to as the operation phase or the operation mode.
[0026] With reference to FIG. 2A, the training stage, for example, executed by the model trainer 110 of FIG. 1 can be further described, and with reference to FIG. 2B, the operation stage, for example, executed by the model implementer 125 of FIG. 1 can be further described. FIG. 2A shows a block diagram of the training of a gaze tracking system according to an exemplary embodiment. As shown in FIG. 2A, the gaze tracking system can include, in a training stage 205, an inference and segmentation model 210 block, a gaze model 215 block, a loss 220 block, and a trainer 225 block.
[0027] The inference and segmentation model 210 (shown in more detail in FIG. 3) can be configured to predict the segmentation of an image. The segmentation or segmentation data can be used to identify (or assist in identifying) objects or regions such as regions of the eye (e.g., the pupil, iris, sclera, etc.). The segmentation (such as the region of the eye) can be used to predict (or assist in predicting) the line of sight. The predicted segmentation can be used to train the inference and segmentation model 210 and / or the line of sight model 215. The inference and segmentation model 210 can include two parts, an inference part and a segmentation part. The inference part of the inference and segmentation model 210 can be a neural network (e.g., the first neural network in a line of sight tracking system). The inference part of the inference and segmentation model 210 can be referred to as an encoder. The inference part obtains an image and outputs a feature map corresponding to the image. Each layer of the network extracts features associated with the input image. Each layer uses one or more convolution functions, ReLu functions, and pooling functions to build higher-order features. The feature map (also called the activation map) can be the final output or the output of the last layer of the network. The segmentation part of the inference and segmentation model 210 can be a neural network (e.g., the second neural network in a line of sight tracking system). The segmentation part of the inference and segmentation model 210 can be referred to as a decoder.
[0028] The gaze model 215 (see FIGS. 3 and 4 for further details) can be configured to predict gaze, for example, using a layered convolutional neural network without sparse constraints. As described above, the inference and segmentation model 210 can include a first model (e.g., an inference part or an encoder) and a second model (e.g., a segmentation part or a decoder). The gaze model 215 may be a third model where the output of the first model (e.g., a feature map) can be an input to the third model. In other words, the data generated by the inference part or the encoder part of the encoder-decoder CNN can be used to predict gaze by the gaze model or a layered convolutional neural network without sparse constraints. As mentioned herein, the gaze tracking system can be trained using segmentation prediction that uses a segmentation model (e.g., a decoder) based on the feature map of the inference model (e.g., an encoder) within the inference segmentation model (e.g., an encoder-decoder CNN) and gaze prediction based on the feature map of the inference model. By including segmentation prediction during training, the training of the gaze model is improved and the gaze prediction during use becomes more accurate.
[0029] The loss 220 and the trainer 225 can be used to implement the training and optimization process. The training and optimization process can be configured to generate a loss based on a loss function (of the loss 220) and a comparison (by the trainer 225) of the predicted segmentation and predicted gaze with the ground truth data. For example, the loss 220 can be calculated as (Loss g +λLoss s ), where Loss g is the gaze loss, Loss s is the segmentation loss, and λ can be a Lagrange multiplier. Then, the trainer 225 can use the resulting loss (e.g., (Loss g +λLoss s)) can be configured to minimize. Generally, ground truth is data associated with training data (e.g., training image 15). In other words, each training image 15 can have associated segmentation prediction data and line-of-sight direction data (e.g., developed by a user, developed by a proven line-of-sight algorithm, etc.). In one embodiment, the line-of-sight loss is the difference between the prediction of the line-of-sight model 215 in the training image and the line-of-sight direction data in the training image, and the segmentation loss is the difference between the segmentation prediction of the inference and segmentation model 210 and the segmentation prediction data in the training image. In another example of an embodiment, the inference and segmentation model 210 can be configured to predict the line of sight based on the segmentation prediction. In such an embodiment, calculating the loss 220 can include comparing the line of sight predicted by a second neural network (or decoder) based on the segmentation with the line of sight predicted by a third neural network (or a convolutional neural network without sparse constraints).
[0030] The loss 220 can be a numerical value indicating how bad the model's prediction was for a single example (e.g., training image 15). If the model's prediction (such as segmentation prediction and / or line-of-sight prediction) is perfect, the loss is zero; otherwise, the loss increases. The goal of training (e.g., trainer 225) is to find (e.g., change, adjust, or correct) parameters (e.g., weights and biases) that result in a low loss on average across all training examples. The loss algorithm can be mean squared loss, mean squared error loss, etc. Training can include modifying parameters associated with at least one of the inference and segmentation model 210 (e.g., the first model and the second model), or the line-of-sight model 215 (e.g., the third model) used to predict the line of sight based on the result of the loss 220.
[0031] Modifying the first model, the second model, and / or the third model can include changing features, feature characteristics (e.g., important features or the importance of features), and / or parameters. Non-limiting examples of features, feature characteristics, and parameters include proposals of bounding boxes, aspect ratios, data augmentation options, loss functions, depth multipliers, number of layers, image input sizes (e.g., normalization), anchor boxes, positions of anchor boxes, number of boxes per cell, feature map sizes, convolutional parameters (e.g., weights), etc. The parameters can be coefficients of the model, and they are selected by the model itself. In other words, the algorithm optimizes these coefficients (according to a given optimization strategy) during learning and returns an array of parameters that minimizes the error. Some parameters (sometimes called hyperparameters) can be elements set manually. For example, the model may not update some parameters (such as during training) according to the optimization strategy.
[0032] The training and optimization process executed by the trainer 225 can be configured based on the desired trade-off between the computational time spent and the quality of the desired result. Generally, since the accuracy improves approximately logarithmically with the number of iterations used during the training process, it may be preferable to use an automatic threshold to stop further optimization. If the quality of the result is prioritized, for example, by calculating the mean squared error, the automatic threshold can be set to a predetermined value of the reconstruction error, but other methods can also be used. The automatic threshold can be set to limit the training and optimization process to a predetermined number of iterations. In some embodiments, these two factors can be used in combination.
[0033] FIG. 2B shows a block diagram of using a gaze tracking system according to an exemplary embodiment. As shown in FIG. 2B, the gaze tracking system can include an inference model 235 block, a gaze model 215, and a gaze 240 block. As described throughout this specification, a trained model (such as an encoder-decoder CNN) can be modified to operatively disconnect (or remove) a portion of the model used for segmentation, such as the decoder portion. This modified model is used in the operation stage 230 and can significantly reduce complexity and processing resource usage. The inference model 235 can be consistent with the inference model of the trained inference and segmentation model 210.
[0034] Therefore, the inference model 235 (see FIG. 4 for further details) can be configured to generate a feature map associated with an image (e.g., image 5). The gaze model 215 (see FIG. 4 for further details) can be configured to predict a gaze 240 based on the feature map. The gaze 240 can be used to determine the gaze of a user of an AR device (e.g., AR device 130).
[0035] As described above, generating accurate gaze predictions can include obtaining accurate information related to the pupil, pupil center, and / or other eye regions. Segmentation data can be a factor in obtaining accurate gaze. When a gaze prediction model can be combined with a segmentation model as a composite model (such as a gaze and segmentation prediction model), the segmentation part of the composite model can be used during training (e.g., by model trainer 110) to normalize and improve the accuracy of the gaze prediction part of the composite model. Example embodiments can include a composite model as an inference segmentation model that can be split or decomposed (e.g., for training and / or calibration) into a multi-resolution encoder-decoder network (e.g., for gaze tracking using model implementer 125) (e.g., by model modifier 115). In other words, the inference segmentation model can be split or decomposed into an inference model (e.g., for gaze tracking using model implementer 125) (e.g., by model modifier 115).
[0036] For example, at the coarsest resolution of the inference part of the inference segmentation model, a gaze model can be added to generate or predict a gaze output. During the training phase, the inference segmentation model can be utilized to generate an input to the gaze model and a segmentation prediction. The gaze model can be used to generate a gaze prediction. The gaze prediction and the segmentation prediction can be input into a loss function to utilize the advantages of segmentation regularization to train the inference segmentation model and / or the gaze model. Regularization is a type of regression that restricts / regularizes the estimated values of coefficients (such as weights and biases) between zero and an upper limit, e.g., 1. In other words, this technique prevents the learning of more complex or flexible models in order to avoid the risk of overfitting.
[0037] FIG. 3 shows a block diagram of an exemplary inference segmentation model (e.g., a multi-resolution encoder-decoder network). FIG. 3 is an example of the inference and segmentation model 210 and the gaze model 215 of FIG. 2A. The inference model and the segmentation model can each include at least one convolutional layer or convolution. For example, as shown in FIG. 3, the inference model 365 can include four convolutional layers each including three convolutions. The first convolutional layer 370-1 includes convolutions 305-1, 305-2, 305-3. The second convolutional layer 370-2 includes convolutions 310-1, 310-2, 310-3. The third convolutional layer 370-3 includes convolutions 315-1, 315-2, 315-3. The fourth convolutional layer 370-4 includes convolutions 320-1, 320-2, 320-3. Further, as shown in FIG. 3, the segmentation model 375 can include four convolutional layers each including three convolutions. The first convolutional layer 380-1 includes convolutions 325-1, 325-2, 325-3. The second convolutional layer 380-1 includes convolutions 330-1, 330-2, 330-3. The third convolutional layer 380-3 includes convolutions 335-1, 335-2, 335-3. The fourth convolutional layer 380-4 includes convolutions 340-1, 340-2, 340-3. The inference segmentation model can include skip connections (e.g., the output of convolution 305-3 can communicate as an input to convolution 340-1) and other calculations necessary to realize the segmentation model output. The skip connection can be configured to skip one or more layers within the neural network. Thus, the skip connection supplies the output of one layer as an input to the next layer after the skipped layer.
[0038] The convolutional layer (e.g., convolutional layers 370-1, 370-2, 370-3, 370-4, 380-1, 380-2, 380-3, and / or 380-4) or convolution can be configured to extract features from the image 5. The features can be based on, for example, color, frequency domain, edge detector, etc. The convolution can have a filter (sometimes called a kernel) and a stride. For example, the filter can be a 1×1 filter with a stride of 1 that results in an output of a cell based on a combination (e.g., addition, subtraction, multiplication, etc.) of the features of the cells at the positions of the M×M grid for each channel (or 1×1×n in the case of conversion to n output channels, and the 1×1 filter is sometimes called pointwise convolution). In other words, a feature map having a depth or number of channels greater than 1 is combined into a feature map having a single depth or number of channels. The filter can be a 3×3 filter with a stride of 1 that results in an output with fewer cells within / for each channel of the M×M grid or feature map. The output can have the same depth or number of channels (e.g., a 3×3×n filter, where n = the number of channels sometimes called depth or depth unit filters) or a reduced depth or number of channels (e.g., a 3×3×k filter, where k < the depth or number of channels). Each channel, depth, or feature map can have an associated filter. Each associated filter can be configured to emphasize different aspects of the channel. In other words, different features can be extracted from each channel based on the filter (this is sometimes called a depthwise separable filter). Other filters are also within the scope of this disclosure.
[0039] Another type of convolution can be a combination of two or more convolutions. For example, the convolution can be made separable in depth units and point units. This can include, for example, a two-step convolution. The first step can be a convolution in depth units (e.g., 3×3 convolution). The second step can be a convolution in point units (e.g., 1×1 convolution). The depth unit and point unit convolutions can be separable convolutions in that different filters (e.g., filters for extracting different features) can be used for each channel or each depth of the feature map. In an example embodiment, the point unit convolution can transform the feature map to include c channels based on the filter. For example, an 8×8×3 feature map (or image) can be transformed to an 8×8×256 feature map (or image) based on the filter. In some embodiments, two or more filters can be used to transform the feature map (or image) to an M×M×c feature map (or image).
[0040] The convolution can be linear. A linear convolution is described as being linear time-invariant (LTI) with respect to the input. The convolution can also include a rectified linear unit (ReLU). The ReLU is an activation function that normalizes the LTI output of the convolution and limits the normalized output to a maximum. The ReLU can be used to accelerate convergence (e.g., more efficient computation).
[0041] In an example embodiment, a combination of depth - unit convolutions and separable convolutions of depth and point units can be used. Each convolution can be configurable (e.g., configurable features, strides, and / or depth). For example, the inference model 365 can include convolutions 305 - 1, 305 - 2, 305 - 3, 310 - 1, 310 - 2, 310 - 3, 315 - 1, 315 - 2, 315 - 3, 320 - 1, 320 - 2, and 320 - 3 that can convert the image 5 into a first feature map. The segmentation model 375 can include convolutions 325 - 1, 325 - 2, 325 - 3, 330 - 1, 310 - 2, 330 - 3, 335 - 1, 335 - 2, 335 - 3, 340 - 1, 340 - 2, and 340 - 3 that can incrementally convert the first feature map into a second feature map. This incremental conversion can generate bounding boxes (regions of the feature map or grid) of various sizes, enabling the detection of objects of many sizes. Each cell can have at least one associated bounding box. In an example embodiment, as the grid (e.g., the number of cells) increases, the number of bounding boxes per cell decreases. For example, the largest grid can use three bounding boxes per cell, and the smallest grid can use six bounding boxes per cell.
[0042] The second feature map can be used to predict segmentation 20. The prediction of the segmentation can include detection (e.g., using a detection layer not shown) and suppression (e.g., a suppression layer not shown). Detection can include using data associated with each bounding box associated with the second feature map (e.g., at least the output of convolution 340-3). The data can be associated with features within the bounding box. The data can indicate an object within the bounding box (the object can be no object or a part of an object). The object can be identified by its features. The data can, cumulatively, be called a class or classifier. The class or classifier can be associated with the object. The data (such as a bounding box) can also include a confidence score (such as a numerical value from 0 to 1).
[0043] After detection, the results can include multiple classifiers that indicate the same object. In other words, an object (or part of an object) can be present within multiple overlapping bounding boxes. However, the confidence scores in each classifier can be different. For example, a classifier that identifies a part of an object may have a lower confidence score than a classifier that identifies the complete (or substantially complete) object. Detection can further include discarding bounding boxes without an associated classifier. In other words, in detection, bounding boxes that do not contain an object can be discarded.
[0044] Suppression can include sorting the bounding boxes based on the confidence score and selecting the bounding box with the highest score as the classifier that identifies the object. The suppression layer can repeat the sorting and selection process for each bounding box having the same or substantially similar classifier. As a result, the suppression layer can include data (e.g., a classifier) that identifies each object within the input image.
[0045] In an augmented reality (AR) eye tracking application, the eye region to be specified can be limited to the eye region captured by the AR application (e.g., within an image). For example, in an exemplary embodiment, a trained ML model can be used to identify any possible eye region (such as the texture of the iris, the shape / size of the pupil, etc.) to assist in determining or verifying the user's line of sight.
[0046] The exemplary embodiment can include a third model 385. The third model 385 can use the first feature map (e.g., the output of convolution 320-3) to predict the line of sight 10. As shown in FIG. 3, the third model 385 can include multiple layers within a convolutional neural network without sparsity constraints. The layered neural network can include three layers 350, 355, 360. Each layer 350, 355, 360 can be formed from a plurality of neurons 345. In this embodiment, no sparsity constraint is applied. Thus, all neurons 345 in each layer 350, 355, 360 are network-connected to all neurons 345 in any adjacent layer 350, 355, 360.
[0047] The third model 385 shown in FIG. 3 can be computationally less complex due to the small number of neurons 345 and layers. In other words, the computational complexity can be associated with the number of neurons 345. Using an initial sparsity condition, the computational complexity of the neural network can be reduced. For example, when the neural network is functioning as an optimization process, the neural network approach can process high-dimensional data by restricting the number of connections between neurons and / or between layers. Additionally, the convolutional neural network can also utilize pooling or max pooling to reduce the dimension (and thus the complexity) of the data flowing through the neural network. Other approaches for reducing the computational complexity of the convolutional neural network can also be used.
[0048] FIG. 4 shows a block diagram of an inference model (e.g., a multi-resolution encoder network) according to an exemplary embodiment. This inference model is an example of the inference model 235 of FIG. 2B. In this embodiment, a segmentation model (such as a decoder) is not included. Therefore, this embodiment is used to predict the line of sight without predicting the segmentation. By not predicting the segmentation, the processing of the image for predicting the line of sight is simplified and significantly speeded up (e.g., fewer processor cycles are used).
[0049] The inference model (e.g., an encoder) can include at least one convolutional layer or convolution. For example, as shown in FIG. 4, the inference model 365 can include four convolutional layers, each including three convolutions. The first convolutional layer 370-1 includes convolutions 305-1, 305-2, 305-3, the second convolutional layer 370-2 includes convolutions 310-1, 310-2, 310-3, the third convolutional layer 370-3 includes convolutions 315-1, 315-2, 315-3, and the fourth convolutional layer 370-4 includes convolutions 320-1, 320-2, 320-3. The inference model 365 of FIG. 4 is the same as the inference model 365 described in FIG. 3. In some embodiments, the inference model 365 of FIG. 4 is different from the inference model 365 described in FIG. 3 in that the inference model 365 of FIG. 4 does not include skip connections.
[0050] The second model 385 (the third model 385 in FIG. 3) can predict the line of sight 10 using the first feature map (e.g., the output of convolution 320-3). As shown in FIG. 4, the second model 385 can include multiple layers within a convolutional neural network without sparsity constraints. The layered neural network can include three layers 350, 355, and 360. Each layer 350, 355, 360 can be formed from a plurality of neurons 345. In this embodiment, no sparsity constraints are applied. Thus, all neurons 345 in each layer 350, 355, 360 are network-connected to all neurons 345 in any adjacent layer 350, 355, 360. The example of the neuron network shown in FIG. 4 is not computationally complex because the number of neurons 345 and layers is small. In other words, the computational complexity can be associated with the number of neurons 345. Additionally, the convolutional neural network can also use pooling or max pooling to reduce the dimensionality (and thus the complexity) of the data flowing through the neural network. Other approaches can also be used to reduce the computational complexity of the convolutional neural network.
[0051] The parameters of the neural networks shown in FIGS. 3 and 4 can be selected to minimize the complexity of the models (e.g., encoders and decoders) while maintaining the overall accuracy. For example, unlike a typical multi-resolution encoder-decoder network model, an example of an encoder can have different asymmetric parameters when compared to the decoder. For example, if each level of the decoder can use two convolutional layers, the corresponding layer of the encoder can use a single convolutional layer, a smaller convolutional size, etc.
[0052] The above embodiments have been described using a gaze tracking system for an AR user device. However, in other embodiments used for object identification, benefits can be obtained from the training and use of the neural network described above. These embodiments can include a system for identifying an object (e.g., eye features in a gaze tracking system) and an attribute associated with the object (e.g., pupil position / direction in a gaze tracking system). For example, a system for identifying variations in the direction of movement (e.g., the direction of movement of a vehicle, a human (or part of a human)), weather (e.g., wind, snow, rain, clouds, etc.), product features in manufacturing (e.g., color, size, etc.) can utilize the techniques described herein.
[0053] In a training stage or training mode, an example embodiment can include training a multi-resolution encoder-decoder network including a first neural network and a second neural network, where the first neural network and the second neural network are configured to predict segmentation data based on training data, and the example embodiment can also include training a third neural network (e.g., a layered neural network) together with the first neural network and the second neural network, where the third neural network is configured to predict training characteristics (e.g., gaze) based on training data.
[0054] FIG. 5 shows a method of training a gaze prediction model according to an exemplary embodiment. As shown in FIG. 5, in step S505, an inference segmentation model (e.g., a multi-resolution encoder-decoder network) including a first model (e.g., a first neural network) and a second model (e.g., a second neural network) is configured. For example, the multi-resolution encoder-decoder network can be configured using a CNN as an encoder and a CNN as a decoder. Alternatively, the multi-resolution encoder-decoder network can be selected from the data stores of encoder-decoder networks.
[0055] In step S510, a third model (e.g., a third neural network) is configured. For example, the third model can be a network including multiple layers of a layered neural network, a deep neural network, a dense network, a convolutional neural network without sparse constraints, etc. The third model can also be selected from the data stores of neural networks. In step S515, the output of the first model is used as the input to the third model. For example, the output of the first model can be communicatively coupled directly or indirectly to the input to the third model. In step S520, training data is received. For example, a plurality of actual (captured) images or synthetic (e.g., computer-generated) images can be received. The plurality of real images and synthetic images include labels for use as ground truth during the training process.
[0056] In step S525, a first feature map is generated based on the training data using the first model. In step S530, a second feature map is generated based on the training data using the second model. For example, the feature map can be generated by applying a filter or a feature detector to the input image or the feature map output of a previous layer of the neural network.
[0057] In step S535, segmentation data is predicted based on the second feature map. For example, segmentation can include dividing data into individual groups. Segmentation can be a progression from a rough inference to a detailed inference. Segmentation can be based on classification that includes predicting an input image or a part of the input image. Segmentation can include location identification / detection that can identify classes (such as eye features) and information regarding the spatial positions of those classes. Further, the detailed inference can include performing a dense prediction that infers the labels of all pixels such that each pixel is labeled with the class of the surrounding object or region.
[0058] In step S540, characteristics are predicted based on the first feature map. For example, the third model can process the first feature map to generate a predicted output (such as characteristics). The characteristics can be gaze prediction. The gaze prediction can be the x, y position on a device (such as the display screen of an AR user device) or the direction in the real-world environment. The first feature map can include tokens representing the gaze and other features characterizing the gaze. In this example, the predicted output can be the likelihood that the eyes have a specific gaze (such as a specific x, y position).
[0059] In step S545, a training loss is generated based on the loss associated with the predicted characteristics and the loss associated with the predicted segmentation. For example, the training loss can be calculated as (Loss g + λLoss s ), where Loss g is the gaze loss and Loss sis the segmentation loss, and λ can be the Lagrange multiplier. Generally, the ground truth is the data associated with the training data. In other words, each training image can have associated segmentation prediction data and gaze prediction data (e.g., labels developed by the user, labels developed by a proven gaze algorithm, etc.). The loss can be calculated as squared loss, mean squared error loss, etc. In another example of an embodiment, the gaze predicted by a second neural network (or decoder) based on the segmentation can be compared with the gaze predicted by a third neural network (or a convolutional neural network without sparse constraints).
[0060] In step S550, the first model, the second model, and / or the third model are modified based on the training loss (e.g., to minimize (Loss g + λLoss s ). For example, the goal of training is to find a set of parameters (e.g., parameters, weights, and / or biases) that result in less loss on average across all training examples. Training can include modifying the parameters associated with at least one of the first model, the second model, and the third model used to generate gaze predictions based on the calculated training loss. Finding a set of parameters can include changing features and / or characteristics of the features (such as important features or the importance of features), and the parameters can include box proposals, aspect ratios, data augmentation options, loss functions, depth multipliers, number of layers, image input size (e.g., normalization), anchor boxes, positions of anchor boxes, number of boxes per cell, feature map size, convolutional parameters (e.g., weights), etc.
[0061] In an operation stage or operation mode, an exemplary embodiment can include receiving operation data and predicting an operation characteristic (e.g., line of sight) using a trained first model (e.g., an inference model, an encoder, etc.) and a trained second model (e.g., a line of sight prediction model, a layered neural network, etc.). In an exemplary embodiment, user data captured using an AR user device can be used in a pre-operation training operation. The user data can be data associated with a user who uses the AR user device. This user data can be used to fine-tune a model trained for a specific user. By this pre-operation training operation, the model trained for a specific user can be improved.
[0062] FIG. 6 shows a method of using a trained line of sight prediction model according to an exemplary embodiment. As shown in FIG. 6, in step S605 (the dashed line indicates that step S605 is optional), the trained line of sight prediction mode is calibrated for the user of the AR device. For example, while the user is using the AR device, a training process (e.g., similar to the method described with respect to FIG. 5) can be executed. In this embodiment, the training data can be an image generated by the AR device while the user is using the AR device. In step S610, operation data is received. For example, the AR device can use a camera to capture an image of the user's eye. The captured image is the operation data.
[0063] In step S615, using a trained multi-resolution encoder-decoder network trained using a first neural network and a second neural network whose output of the first neural network is used as an input to a third neural network, characteristics related to the operation data are predicted. The prediction is made by removing the second neural network. For example, the neural network in FIG. 3 can be used when training the encoder-decoder network of the gaze tracking system, and the neural network in FIG. 4 can be used when operating the gaze tracking system.
[0064] Figure 7 shows examples of a computer device 700 and a mobile computer device 750, which can be used with the techniques described herein (e.g., as a first computing device 105 for implementing the techniques described herein, a second computing device 120, and for implementing other resources). The computing device 700 includes a processor 702, a memory 704, a storage device 706, a high-speed interface 708 connected to the memory 704 and a high-speed expansion port 710, and a low-speed interface 712 connected to a low-speed bus 714 and the storage device 706. Each component 702, 704, 706, 708, 710, and 712 is interconnected using various buses and can be mounted on a common motherboard or in other manners as required. The processor 702 can process instructions for execution within the computer device 700, and the instructions include instructions stored in the memory 704 or the storage device 706 for displaying graphic information for a GUI on an external input / output device such as a display 716 coupled to the high-speed interface 708. In other embodiments, multiple processors and / or multiple buses can be used, along with multiple memories and memory types, as required. Also, multiple computing devices 700 can be connected, in which case each device provides a portion of the required operations (e.g., as a server bank, a group of blade servers, or a multiprocessor system).
[0065] The memory 704 stores information within the computing device 700. In one embodiment, the memory 704 is one or more volatile memory units. In other embodiments, the memory 704 is one or more non-volatile memory units. The memory 704 can also be another form of computer-readable medium, such as a magnetic disk or an optical disk.
[0066] The memory device 706 can provide large-capacity storage for the computing device 700. In one embodiment, the memory device 706 may be a computer-readable medium, such as, for example, a floppy disk device, a hard disk device, an optical disk device, or a tape device, a flash memory or other similar solid-state memory device, or an array of devices including a device within a storage area network or other configuration, or may include these. The computer program product can be tangibly incorporated into an information carrier. The computer program product can also include instructions that, when executed, perform one or more of the methods as described above. The information carrier is a computer or machine-readable medium such as the memory 704, the memory device 706, or the memory on the processor 702.
[0067] The high-speed controller 708 manages operations that consume a large amount of bandwidth in the computing device 700, while the low-speed controller 712 manages operations that consume less bandwidth. Such an assignment of functions is merely an example. In one embodiment, the high-speed controller 708 is coupled to the memory 704, the display 716 (e.g., via a graphics processor or accelerator), and to a high-speed expansion port 710 that can receive various expansion cards (not shown). In an embodiment, the low-speed controller 712 is coupled to the memory device 706 and a low-speed expansion port 714. The low-speed expansion port, which may include various communication ports (e.g., USB, Bluetooth, Ethernet, wireless Ethernet), may be coupled to one or more input / output devices such as a keyboard, a pointing device, a scanner, or to a networking device such as a switch or router via, for example, a network adapter.
[0068] As shown in this figure, computing device 700 can be implemented in many different forms. For example, computing device 700 may be implemented as a standard server 720, or may be implemented multiple times within a group of such servers. Computing device 700 can also be implemented as part of a rack server system 724. Further, computing device 700 may be implemented in a personal computer such as a laptop computer 722. Alternatively, the components of computing device 700 can be combined with other components within a mobile device (not shown) such as device 750. Each of such devices can include one or more computing devices 700, 750, and the entire system may be composed of a plurality of computing devices 700, 750 that communicate with each other.
[0069] Computing device 750 includes, among other components, a processor 752, a memory 764, input / output devices such as a display 754, a communication interface 766, and a transceiver 768. Device 750 may also be provided with a storage device such as a microdrive or other device to provide additional storage. Each of the components 750, 752, 764, 754, 766, and 768 are interconnected using various buses, and some of the components may be implemented on a common motherboard or in other ways as needed.
[0070] Processor 752 can execute instructions within computing device 750, including instructions stored in memory 764. The processor can be implemented as a chipset of chips including separate multiple analog and digital processors. The processor can adjust other components of device 750, such as, for example, the user interface, applications executed by device 750, and control of wireless communication by device 750.
[0071] The processor 752 can communicate with the user via a control interface 758 and a display interface 756 coupled to a display 754. The display 754 may be, for example, a TFT LCD (Thin Film Transistor Liquid Crystal Display), and an LED (Light Emitting Diode) or OLED (Organic Light Emitting Diode) display, or other suitable display technology. The display interface 756 can include appropriate circuitry for driving the display 754 to present graphic information and other information to the user. The control interface 758 can receive commands from the user, convert them, and send them to the processor 752. Further, an external interface 762 that communicates with the processor 752 may be provided to enable short-range communication between the device 750 and other devices. The external interface 762 can provide, for example, wired communication in some embodiments, wireless communication in other embodiments, or use multiple interfaces.
[0072] Memory 764 stores information within computing device 750. Memory 764 can be implemented as one or more of one or more computer-readable media, one or more volatile memory units, or one or more non-volatile memory units. An extended memory 774 may be further provided. The extended memory 774 may be connected to device 750 via an expansion interface 772, and the expansion interface 772 may include, for example, a SIMM (Single In-line Memory Module) card interface. Such extended memory 774 can provide additional storage space for device 750, or can store applications or other information for device 750. Specifically, the extended memory 774 can include instructions for executing or supplementing the aforementioned processes, and can also include secure information. Thus, for example, the extended memory 774 may be provided as a security module for device 750 and may be programmed with instructions that enable secure use of device 750. Further, a secure application may be provided via the SIMM card, along with additional information, such as placing specific information on the SIMM card in a non-hackable manner.
[0073] The memory may include, for example, flash memory and / or NVRAM memory, as described below. In one embodiment, a computer program product is tangibly incorporated in an information carrier. The computer program product includes instructions that, when executed, perform one or more of the methods as described above. The information carrier is a computer or machine-readable medium such as memory 764, extended memory 774, or memory on processor 752, which may be received via, for example, transceiver 768 or external interface 762.
[0074] Device 750 can communicate wirelessly via a communication interface 766 that can include a digital signal processing circuit as needed. The communication interface 766 can provide communication under various modes or protocols, such as GSM voice calls, SMS, EMS, or MMS messaging, CDMA, TDMA, PDC, WCDMA (registered trademark), CDMA2000, or GPRS. Such communication can be performed, for example, via a radio frequency transceiver 768. Additionally, short-range communication can be performed, such as using Bluetooth, Wi-Fi, or other such transceivers (not shown). Further, a GPS (Global Positioning System) receiver module 770 can provide additional navigation and location-related wireless data to device 750, and this data can be appropriately used by applications running on device 750.
[0075] Device 750 can also communicate using an audio codec 760, which can receive information spoken by the user and convert it into usable digital information. The audio codec 760 can similarly generate audible sound for the user, for example, via a speaker within the handset of device 750. Such sound can include sound from a voice call, recorded sound (such as a voice message, music file, etc.), and even sound generated by an application operating on device 750.
[0076] As shown in this figure, computing device 750 can be implemented in many different forms. For example, the computing device may be implemented as a mobile phone 780. Also, the computing device may be implemented as part of a smartphone 782, a personal digital assistant, or other similar mobile devices.
[0077] The various implementations of the systems and techniques described in this specification can be realized in digital electronic circuitry, integrated circuitry, specially designed ASICs (application specific integrated circuits), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include an implementation in one or more computer programs executable and / or interpretable on a programmable system that couples data and instructions to and from a memory system, at least one input device, and at least one output device and includes at least one programmable processor which can be either special purpose or general purpose.
[0078] These computer programs (also known as programs, software, software applications, or code) include machine instructions for a programmable processor and can be implemented in high-level procedural and / or object-oriented programming languages, and / or in assembly / machine language. As used herein, the terms "machine-readable medium" and "computer-readable medium" refer to any computer program product, apparatus, and / or device (e.g., magnetic disks, optical disks, memory, programmable logic device (PLD)) used to provide machine instructions and / or data to a programmable processor, including a machine-readable medium that receives machine instructions as a machine-readable signal. The term "machine-readable signal" refers to any signal used to provide machine instructions and / or data to a programmable processor.
[0079] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having a display device (LED (light-emitting diode), OLED (organic LED), or LCD (liquid crystal display) monitor / screen) for displaying information to the user, as well as a keyboard and a pointing device (such as a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide similar interaction with the user. For example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback), and the input from the user can be received in any form, such as acoustic, voice, or tactile input.
[0080] The systems and techniques described herein can be implemented in a computing system that includes back-end components (such as a data server), or middleware components (such as an application server), or front-end components (such as a client computer having a graphical user interface or a web browser through which a user can interact with an implementation of the systems and techniques described herein), or any combination of such back-end, middleware, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (such as a communication network). Examples of communication networks include local area networks ("LANs"), wide area networks ("WANs"), and the Internet.
[0081] A computing system can include clients and servers. Clients and servers are generally remote from each other and typically interact through a communication network. The relationship between a client and a server arises from computer programs that are executed on respective computers and have a client-server relationship with each other.
[0082] In some embodiments, the computing device shown in this figure can include sensors that interface with an AR headset / HMD device 790 to generate an extended environment for viewing content inserted into the physical space. For example, one or more sensors included in computing device 750 or other computing devices shown in this figure can provide input to AR headset 790 or, generally, to the AR space. Sensors include, but are not limited to, touchscreens, accelerometers, gyroscopes, pressure sensors, biometric sensors, temperature sensors, humidity sensors, ambient light sensors, etc. Computing device 750 can use the sensors to determine the absolute position and / or detected rotation of the computing device within the AR space, which can then be used as input to the AR space. For example, computing device 750 can be incorporated into the AR space as a virtual object such as a controller, laser pointer, keyboard, weapon, etc. By positioning the computing device / virtual object by the user when incorporated into the AR space, the user may be able to position the computing device to view the virtual object in a specific manner within the AR space. For example, if the virtual object represents a laser pointer, the user can operate the computing device as if it were an actual laser pointer. The user can use the device in a manner similar to using a laser pointer, such as moving the computing device left and right, up and down, or in a circular motion. In some embodiments, the user can use the virtual laser pointer to target a target position.
[0083] In some embodiments, one or more input devices included on or connected to computing device 750 can be used as input to the AR space. Input devices can include, but are not limited to, touchscreens, keyboards, one or more buttons, trackpads, touch pads, pointing devices, mice, trackballs, joysticks, cameras, microphones, headsets or buds with input capabilities, game controllers, or other connectable input devices. Interacting the user with the input devices included in computing device 750 when the computing device is incorporated into the AR space can cause specific actions to occur within the AR space.
[0084] In some embodiments, the touchscreen of computing device 750 can be rendered as a touchpad in the AR space. The user can interact with the touchscreen of computing device 750. The interaction can be rendered, for example, in the AR headset 790 as movement on a touchpad rendered within the AR space. The rendered movement can control virtual objects within the AR space.
[0085] In some embodiments, one or more output devices included in computing device 750 can provide output and / or feedback to the user of the AR headset 790 within the AR space. The output and feedback can be visual, tactile, or auditory. Output and / or feedback can include, but are not limited to, vibration, on / off or flashing and / or strobe of one or more lights or strobes, alarm sounding, chime playing, song playing, and audio file playing. Output devices can include, but are not limited to, vibration motors, vibration coils, piezoelectric devices, electrostatic devices, light emitting diodes (LEDs), strobes, speakers, etc.
[0086] In some embodiments, computing device 750 can appear as another object in a computer-generated 3D environment. Interaction by a user with computing device 750 (e.g., rotating, shaking, touching the touch screen, swiping a finger on the touch screen) can be interpreted as interaction with an object in the AR space. In an example of a laser pointer in the AR space, computing device 750 appears as a virtual laser pointer in a computer-generated 3D environment. When the user operates computing device 750, the user in the AR space sees the movement of the laser pointer. The user receives feedback on computing device 750 or on AR headset 790 from the interaction with computing device 750 in the AR environment. The user's interaction with the computing device may be converted into interaction with a user interface generated in the AR environment for a controllable device.
[0087] In some embodiments, computing device 750 may include a touch screen. For example, the user can interact with the touch screen to interact with a user interface for a controllable device. For example, the touch screen can include user interface elements such as sliders that can control the properties of a controllable device.
[0088] Computing device 700 is intended to represent various forms of digital computers and devices including, but not limited to, laptops, desktops, workstations, personal digital assistants, servers, blade servers, mainframes, and other suitable computers. Computing device 750 is intended to represent various forms of mobile devices such as personal digital assistants, cellular phones, smartphones, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are intended to be illustrative only and are not intended to limit the implementation of the invention described and / or claimed in this document.
[0089] Numerous embodiments have been described. Nevertheless, it can be understood that various modifications can be made without departing from the spirit and scope of this specification.
[0090] Furthermore, the logical flows shown in the figures do not require the specific order, or series of orders, shown to achieve the desired result. In addition, steps can be provided to the described flows or removed from the described flows, and other components can be added to or removed from the described system. Accordingly, other embodiments are also included within the scope of the claims.
[0091] In addition to the above description, it can be provided to the user to control whether the systems, programs, or functions described herein enable the collection of user information (e.g., information regarding the user's social network, social actions or activities, occupation, user preferences, or the user's current location), and when the collection is performed, and whether the user can select whether the user is sent content or communication from the server. Furthermore, certain data may be processed in one or more ways before being stored or used so that information that can identify an individual is removed. For example, the user's identity information may be processed so that information that can identify the user's individual cannot be identified, or the user's geographical location may be generalized at the location where the location information is obtained so that the user's specific location cannot be determined (to the city, postal code, state level, etc.). Accordingly, the user can control what information is collected regarding the user, how that information is used, and what information is provided to the user.
[0092] While specific features of the described embodiments are shown as described herein, many modifications, substitutions, changes, and equivalents will readily occur to those skilled in the art. Accordingly, it should be understood that the appended claims are intended to cover all such modifications and changes that fall within the scope of the implementation. These are presented by way of example only and not by way of limitation, and it should be understood that various changes are possible in form and detail. Any part of the apparatus and / or method described herein can be combined in any combination, except mutually exclusive combinations. The implementations described herein can include various combinations and / or sub-combinations of the functions, components, and / or features of the different implementations described.
[0093] Exemplary embodiments can include various modifications and alternative forms, but those embodiments are shown in the drawings by way of example and described in detail herein. However, it should be understood that there is no intention to limit the exemplary embodiments to the specific forms disclosed, and conversely, the exemplary embodiments are intended to cover all modifications, equivalents, and alternatives that fall within the scope of the claims. Throughout the description of the drawings, like numbers refer to like elements.
[0094] Some of the above exemplary embodiments are described as a process or method shown as a flowchart. In the flowchart, the operations are described as a sequential process, but many of the operations can be executed in parallel, concurrently, or simultaneously. Further, the order of the operations may be rearranged. The process may end when the operations are completed, but there may be additional steps not included in the figure. The process may correspond to a method, function, procedure, subroutine, subprogram, etc.
[0095] The methods discussed above are, in part, illustrated by flowcharts and can be implemented by hardware, software, firmware, middleware, microcode, hardware description language, or any combination thereof. When implemented in software, firmware, middleware, or microcode, the program code or code segments for performing the necessary tasks can be stored in a machine or computer-readable medium such as a storage medium. The processor can execute the necessary tasks.
[0096] The specific details of the structures and functions disclosed herein are merely representative for the purpose of describing exemplary embodiments. However, the exemplary embodiments may be embodied in many alternative forms and should not be construed as limited to only the embodiments described herein.
[0097] In this specification, terms such as first, second, etc. may be used to describe various elements, but it should be understood that these elements should not be limited by these terms. These terms are only used to distinguish one element from another. For example, without departing from the scope of the exemplary embodiments, the first element can be called the second element, and similarly, the second element can be called the first element. As used herein, the term "and / or" includes any and all combinations of one or more of the associated listed items.
[0098] When an element is referred to as being connected or coupled to another element, it should be understood that it can be directly connected or coupled to the other element, or intervening elements may be present. In contrast, when an element is called directly connected or directly coupled to another element, there are no intervening elements. Other words used to describe the relationship between elements should be interpreted in the same way (e.g., between and directly between, adjacent and directly adjacent, etc.).
[0099] The terms used in this specification are for the sole purpose of describing particular embodiments and are not intended to limit the exemplary embodiments. As used herein, the singular forms "a," "an," and "the" are to be construed to include the plural forms as well, unless the context clearly dictates otherwise. Further, as used herein, the terms "comprises," "comprising," "includes," and / or "including" specify the presence of the stated features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.
[0100] Also, in some alternative implementations, it should be noted that the described functions / operations may occur out of the order described in the figures. For example, two figures shown consecutively may actually be executed simultaneously, or sometimes in the reverse order depending on the relevant functions / operations.
[0101] Unless otherwise defined, all terms (including technical and scientific terms) used herein have the same meaning as commonly understood by one of ordinary skill in the art to which the exemplary embodiments belong. Further, terms such as those defined in commonly used dictionaries are to be construed to have a meaning that coincides with their meaning in the context of the relevant art and are not to be construed in an idealized or overly formal sense unless expressly so defined herein.
[0102] Some of the above exemplary embodiments and corresponding detailed descriptions are presented from the perspective of algorithms and symbolic representations of operations on software or data bits in a computer memory. These descriptions and representations are for those skilled in the art to effectively communicate the content of their research to other skilled persons. The terms used herein and algorithms as commonly used terms are considered to be a consistent series of steps leading to a desired result. These steps are steps that require physical operations on physical quantities. Although not necessarily so, usually these quantities take the form of optical, electrical, or magnetic signals capable of being stored, transferred, combined, compared, and other operations. Mainly for reasons of general use, it may be convenient to refer to these signals as bits, values, elements, symbols, characters, terms, numbers, and the like.
[0103] In the above exemplary embodiments, references to operations and symbolic representations of operations (e.g., in the form of a flowchart) that can be implemented as program modules or functional processes include routines, programs, objects, components, data structures, etc. that can perform specific tasks or implement specific abstract data types and can be described and / or implemented using existing hardware in existing structural elements. Such existing hardware can include one or more central processing units (CPUs), digital signal processors (DSPs), application specific integrated circuits, field programmable gate array (FPGA) computers, and the like.
[0104] However, it should be noted that all of these terms and similar terms are merely associated with appropriate physical quantities and are nothing more than convenient labels applied to these quantities. Unless otherwise specified or apparent from the discussion, terms such as processing or calculation, or terms such as calculation or determination in display, etc., operate on data represented as physical quantities, electronic quantities in the registers and memories of a computer system, and convert them into other data similarly represented as physical quantities in the memories and registers of the computer system, or other such information storage devices, transmission devices, or display devices. This refers to the operations and processes of a computer system or a similar electronic computing device.
[0105] Also, it should be noted that the software-implemented aspects of the exemplary embodiments are usually encoded on some form of non-transitory program storage medium or implemented on some kind of transmission medium. The program storage medium may be magnetic (e.g., a floppy disk or a hard drive) or optical (e.g., a compact disc read-only memory i.e. CDROM), and may be read-only or random access. Similarly, the transmission medium may be a twisted pair wire, a coaxial cable, an optical fiber, or any other suitable transmission medium known in the art. The exemplary embodiments are not limited by these aspects of any given implementation form.
[0106] Finally, the appended claims describe specific combinations of features described herein, but the scope of the present disclosure is not limited to the specific combinations claimed below. Instead, it should be noted that the scope is extended to include any combination of features or embodiments disclosed herein, whether or not that specific combination is specifically recited in the appended claims at the present time.
Claims
1. A method, comprising: In a training stage, training a gaze prediction model including a first model and a second model, wherein the first model and the second model are configured to cooperate to predict segmentation data based on training data; The method further includes training a third model together with the first model and the second model, wherein the third model is configured to predict training characteristics using the output of the first model based on the training data; In an operation stage, the method includes receiving operation data and predicting operation characteristics using the trained first model and the trained third model; Training the gaze prediction model includes predicting the segmentation data based on a feature map and training at least one of the first model and the second model based on a loss related to the predicted segmentation data and a loss related to the training characteristics. A method according to claim 1, wherein the training data includes an eye image.
2. A method, comprising: In a training stage, training a gaze prediction model including a first model and a second model, wherein the first model and the second model are configured to cooperate to predict segmentation data based on training data; The method further includes training a third model together with the first model and the second model, wherein the third model is configured to predict training characteristics using the output of the first model based on the training data; In an operation stage, the method includes receiving operation data and predicting operation characteristics using the trained first model and the trained third model; Training the gaze prediction model includes predicting the training characteristics based on a feature map using the third model and training the third model based on a loss related to the predicted training characteristics and a loss related to the segmentation data.
3. The training data includes an eye image. The predicted segmentation data includes the region of the eye, The training characteristic is the line of sight of the eye, The method according to claim 1 or claim 2.
4. The motion data includes an image of the eye captured using an augmented reality (AR) user device, The motion characteristic is the line of sight of the eye, The method according to claim 1 or claim 2.
5. Predicting the motion characteristic includes: Generating a feature map based on the motion data using the first model, Predicting the motion characteristic based on the feature map using the third model The method according to claim 1 or claim 2.
6. The first model is a first convolutional neural network (CNN), The second model is a second CNN including at least one skip connection from the first CNN, The third model is a linear neural network, The method according to claim 1 or claim 2.
7. The second model is configured to predict a segmentation from the feature map generated by the first model, In the motion stage, the second model is removed from the line of sight prediction model, The method according to claim 1 or claim 2.
8. Training the line of sight prediction model includes changing at least one of the parameters, features, and feature characteristics related to at least one of the first model, the second model, and the third model. The method according to claim 1 or claim 2.
9. Before the motion stage, in a calibration stage, training at least one of the first model, the second model, and the third model based on user data captured using an AR user device The method according to claim 1 or claim 2 further includes.
10. Generating a first feature map based on training data using a first model, Generating a second feature map based on the first feature map using a second model, Predicting segmentation data based on the second feature map, Generating a first loss related to the predicted segmentation data, Predicting training characteristics based on the first feature map using a third model, generating a second loss related to the predicted training characteristics; generating a training loss based on the first loss and the second loss; training at least one of the first model, the second model, and the third model based on the training loss A method comprising. **Claim 11** The training data includes an image of an eye, The predicted segmentation data includes the region of the eye, The training characteristic is the line of sight of the eye, The method according to claim 10. **Claim 12** The first model is a first convolutional neural network (CNN), The second model is a second CNN including at least one skip connection from the first CNN, The third model is a linear neural network, The method according to claim 10 or claim 11. **Claim 13** The second model is configured to predict a segmentation from a feature map generated by the first model, In the operating phase, the second model is removed, The method according to claim 10 or claim 11. **Claim 14** Training at least one of the first model, the second model, and the third model includes changing at least one of the parameters, features, and feature characteristics associated with at least one of the first model, the second model, and the third model. The method according to claim 10 or claim 11. **Claim 15** Predicting the segmentation data is a process including separating data related to the second feature map into separate groups. The method according to claim 10 or claim 11. **Claim 16** Predicting the segmentation data includes a detection process and a suppression process. The method according to claim 10 or claim 11. **Claim 17** A method implemented in a system using a first model, a second model, and a third model, comprising: generating a feature map based on operation data using the trained first model; predicting a characteristic based on the feature map using the trained third model; The training of the first model and the third model is using the second model to predict segmentation data during the training; generating a training loss using a loss associated with the segmentation data and a loss associated with the characteristic; modifying the first model and the third model based on the training loss, including; the method includes; further comprising removing the second model for predicting the characteristic; method.
18. The motion data includes an image of an eye captured using an augmented reality (AR) user device, The method includes predicting motion characteristics using the trained first model and the trained third model, The motion characteristic is the line of sight of the eye, The method according to claim 17.
19. The method according to claim 17 or claim 18, further comprising further training at least one of the trained first model, the trained third model, and the second model based on user data captured using an AR user device.
Citation Information
Patent Citations
Classification of source data by neural network processing
US20190273510A1
Geometrically constrained, unsupervised training of convolutional autoencoders for extraction of eye landmarks
US20210034836A1
Eye tracking and gaze estimation using off-axis camera
WO2021034961A1