Methods for eye tracking

The eye tracking method implemented by computers uses image processing and artificial neural network technology to solve the problem of low accuracy in the absence of dedicated hardware in the existing technology, realize high-accuracy viewpoint positioning, and simplify user operations.

JP7673091B2Active Publication Date: 2025-05-08イリスボンド クラウドボンディング エスエレ
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2022559517
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2020-04-09
Filing Date
2021-02-17
Publication Date
2025-05-08
Estimated Expiration
2041-02-17

AI Technical Summary

Technical Problem

In the absence of dedicated hardware components, existing eye tracking technology has limited accuracy and depends on user calibration, which is cumbersome and impractical.

Method used

Using computer-implemented methods, by acquiring images and positioning facial landmarks, building line-of-sight vectors, and using artificial neural network (ANN) and support vector regression (SVR) algorithms, the viewpoint position on the screen is accurately determined.

Benefits of technology

Without relying on dedicated hardware, high accuracy of viewpoint positioning is achieved, with the accuracy of less than 1 degree, reducing the steps of user calibration and improving the practicality and popularity of the method.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007673091000069
    Figure 0007673091000069
  • Figure 0007673091000070
    Figure 0007673091000070
  • Figure 0007673091000071
    Figure 0007673091000071
Patent Text Reader

Abstract

The present invention relates to a computer-implemented method for locating a viewpoint on a screen (151). The method comprises the steps of: initiating (210) acquisition of an image (300); and initiating (220, 230) locating a first facial landmark location (301) and a second facial landmark location (302) in the image (300). The method further comprises the step of initiating (240) selection of a region of interest (310) in the image (300), wherein the selection is performed by using the landmark locations (301, 302). The method also comprises the step of initiating (250) construction of a gaze vector, wherein the construction of the gaze vector is performed using an artificial neural network that uses the first region of interest (310) as input. Furthermore, the method comprises a step (260) of initiating a positioning of a viewpoint on the screen (151), wherein the positioning of the viewpoint is performed using a line-of-sight vector.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present invention refers to the field of eye-tracking methods and devices for locating a position of a gaze point on a screen, e.g., on a display unit of a computing device. According to the present invention, the computing device may be, for example, a personal computer, a tablet, a laptop, a smartphone, a video game console, a camera, a head-mounted display (HMD), a smart TV, etc. [Background technology]

[0002] Eye-tracking methods are widely used in human-computer interaction. A computer program implementing the method is capable of tracking a user's gaze, and thus enabling input to be provided to a computing device without using traditional input devices (such as a keyboard, a mouse device, a touchpad, etc.) by simply looking at a specific location on a display unit of the computing device. For example, a user may provide input to a computer's graphical user interface (GUI) without having to use his or her hands, thereby enabling, for example, a user with impaired mobility to successfully interact with a computer.

[0003] Methods for eye-tracking are also used in the assembly or maintenance of complex systems. Operators performing such assembly or maintenance tasks often use HMDs, especially augmented reality HMDs that use computer graphics techniques to augment the operator's natural vision. HMDs implementing eye-tracking methods can be controlled hands-free by the operator, thus leaving both hands available for the task.

[0004] Eye-tracking methods are also important in the automotive industry. For example, these methods can be implemented in driver assistance systems, which make it possible to track the driver's gaze to see if he / she is paying attention to the road, for example, if he / she is looking through the windshield of the car. For example, a driver assistance system implementing an eye-tracking method can detect if the driver is looking at the screen of the vehicle rear-view camera and activate said camera only when needed, in other words only when the driver is looking at the screen.

[0005] Methods for eye-tracking may also enable hands-free interaction between the driver and the vehicle's software, so that the driver may give commands to the vehicle's software without taking her / his hands off the steering wheel. The driver may thereby command the software to perform specific activities, such as adjusting the intensity of the vehicle lights, locking / unlocking the doors, controlling the vehicle speed, etc., simply by looking in a particular direction.

[0006] Gaze-tracking methods for locating the position of the gaze point on a screen are known in the art. Known gaze-tracking methods can achieve relatively high accuracy only if they rely on dedicated hardware components, such as infrared (IR) cameras, wearable hardware components, and eye-tracking glasses, or on a calibration procedure that is user-dependent. For example, methods using IR cameras can reach an accuracy of about 0.5°. However, such dedicated hardware components are relatively expensive and do not exist in mainstream computing devices, such as laptops and smartphones. Moreover, wearable hardware components can be uncomfortable to use and hinder the user's mobility. Alternatively, the calibration procedure is time-consuming, limiting the usefulness of devices implementing known methods for gaze tracking, and thus the practicality of the devices.

[0007] In general, known eye-tracking methods suffer from limited accuracy under real-world operating conditions, e.g., in the absence of dedicated hardware components and / or under circumstances characterized by relatively wide variability in eye appearance, lighting, head pose, camera technical specifications, image quality, etc. Summary of the Invention

[0008] These problems are at least partly solved by the invention of the present application, which relates to a computer-implemented method according to claim 1, to a device according to claim 14, to a computer program product according to claim 15 and to a computer-readable storage medium according to claim 16. Embodiments of the invention are the subject of the dependent claims.

[0009] The present invention provides a computer-implemented method for locating a first viewpoint on a screen, the method comprising: Initiating acquisition of at least a first image; Initiating locating a first facial landmark location of a first facial landmark in a first image; ● commencing locating a second facial landmark location of a second facial landmark in the first image; initiating a selection of a first region of interest in a first image, wherein the selection of the first region of interest is performed by using at least a first facial landmark location and a second facial landmark location; - initiating construction of a first gaze vector, the construction of the first gaze vector being performed using at least an artificial neural network, the artificial neural network using at least the first region of interest as an input; - commencing locating a first viewpoint on a screen, the locating of the first viewpoint being performed using at least a first line of sight vector; The present invention relates to a computer-implemented method comprising at least the steps of:

[0010] The screen may be a concave or convex surface. In particular, the screen may be a substantially flat surface, e.g. a panel, such as a canvas, a glass panel and / or a windshield of a vehicle. The screen may be a display unit of a computing device. For example, the screen may be a monitor or a screen of the computing device, e.g. a substantially flat area of ​​the computing device on which a GUI and / or data, in particular data in the form of an image, is displayed. A point on the screen may be represented by two-dimensional screen coordinates in a two-dimensional reference frame defined on the screen. The screen coordinates may in particular be Cartesian or polar coordinates. For example, the screen location of a point on the screen is represented by two-dimensional screen coordinates (a,b) with respect to a screen reference frame centered on the top left corner of the screen.

[0011] According to the invention, the image may be a vector image or a two-dimensional grid of pixels, for example a rectangular grid of pixels. In particular, the location of a pixel in an image may be uniquely determined by the two-dimensional image coordinates of the location in the image, said coordinates representing the location of said pixel in the two-dimensional grid of pixels. The two-dimensional image coordinates may be Cartesian or polar coordinates with respect to a two-dimensional reference frame in the plane of the image, for example in a plane comprising the grid of pixels. For example, the two-dimensional image coordinates of a pixel are the coordinates of the pixel in the image plane reference frame of the first image.

[0012] In particular, the first image is a first two-dimensional grid of pixels, e.g., a first rectangular grid of pixels. The entries of the first two-dimensional grid may be arranged in columns and rows, and may be enumerated in ascending order, with each column and each row associated with a column number and a row number, respectively. In particular, the location of each pixel in the first image is represented by the row number of the row to which the pixel belongs. TIFF0007673091000001.tif11170 and the column number of the column to which the pixel belongs The two-dimensional image coordinates of the pixel in the first image can be uniquely determined by the two-dimensional vector TIFF0007673091000003.tif12170. For example, the two-dimensional image coordinates of a pixel in the first image are the coordinates of the pixel in the image plane reference frame of the first image.

[0013] An image, for example a first image, may be encoded by at least a bitmap. A bitmap encoding an image or a portion of an image may comprise, for example consist of, an array of bits specifying the color of each pixel of said image or portion of said image. The bitmap may be palette indexed such that the entries of the array are indexed onto a color table. The entries of the array may store bits that encode the color of a pixel. In particular, the bitmap may comprise, for example consist of, a dot matrix data structure representing a two-dimensional grid of pixels. The bitmap may further comprise information relating to the number of bits per pixel, the number of pixels per row of the two-dimensional grid of pixels and / or the number of pixels per column of said rectangular grid. An image viewer may use the information encoded in the bitmap to render the image or portion of the image on a screen of a computing device, for example a computing device performing the method of the invention.

[0014] The image, for example a first image, may be stored, in particular temporarily, in a primary memory and / or in a secondary memory of a computing device, for example a computing device performing the method of the invention. According to the invention, the image may be acquired by accessing the memory in which said image is stored. Alternatively, or in conjunction with the above, the acquisition of the image may be performed by capturing said image with a recording device, for example a photo and / or video recording device, such as a photo or video camera. The photo and / or video recording device may be integrated into the computing device, in particular a computing device performing the method of the invention. The captured image may then be stored in the primary and / or secondary memory of the computing device and may be accessed to locate the facial landmarks and / or to select the region of interest.

[0015] A facial landmark is a point in a human face that specifically marks a characteristic anatomical region of a human face in general. For example, a facial landmark may be the tip of the nose, the right edge of the mouth, or the left edge of the mouth. Similarly, a facial landmark may be a point on the eyebrow or on the lip that marks the eyebrow and lip, respectively, along with other landmarks.

[0016] The facial landmarks may for example be eye landmarks. The eye landmarks are in particular points of the eye which together with other eye landmarks mark the shape of the eye. For example, the eye landmarks may be the left or right edge of the eye, a point of the eyelid or the centre of the eyeball. The eye landmarks may be iris landmarks. In particular, the iris landmarks are points of the iris which together with other iris landmarks mark the shape of the iris. For example, the iris landmark is the centre of the iris.

[0017] The location in the image of the facial landmark, for example the first facial landmark location and / or the second facial landmark location, is in particular the location in the image of a representation of said landmark in the image. For example, if in the image the facial landmark is represented by a set of pixels, the facial landmark location may be the location in the image of a reference pixel of said set of pixels. In this way, the facial landmark location is uniquely represented by the location of this reference pixel, for example the two-dimensional image coordinates of the reference pixel.

[0018] The localization of the first and / or the second facial landmark in the first image may be performed using the algorithm disclosed in the paper "One Millisecond Face Alignment with an Ensemble of Regression Trees" by V. Kazemi et al., DOI: 10.1109 / cvpr.2014.241, hereinafter referred to as "First Location Algorithm", which comprises an ensemble of regression trees.

[0019] In particular, the first location algorithm allows locating a set of n0 facial landmarks, said set comprising at least a first landmark and a second landmark, for example, n0 being between 2 and 194, in particular between 30 and 130. Moreover, n0 may be between 50 and 100, more particularly equal to either 6 or 68.

[0020] The three-dimensional location of a facial landmark, for example, an eyeball center or an iris center, is the position of said landmark in three-dimensional space. For example, the three-dimensional location of a facial landmark is represented by a three-dimensional vector representing the three-dimensional coordinate of this landmark in a camera reference frame. In particular, said three-dimensional coordinate can be obtained from the two-dimensional image coordinate of the landmark in the image plane reference frame of the first image via the inverse of a camera matrix related to the camera that acquires the first image.

[0021] The first region of interest (ROI) may comprise at least the first eye or a portion of the first eye. The first ROI is represented by a set of pixels of the first image. In particular, the first ROI is a two-dimensional grid of pixels of the first image, more particularly, a rectangular grid of pixels of the first image. The first ROI may be encoded by at least a first bitmap.

[0022] For example, if the first facial landmark is the tip of the nose and the second landmark is the left eyebrow point, the first ROI may consist of a row of pixels between the row of the second landmark and the row of the first landmark. Moreover, if the first and second facial landmarks are the left and right edges of the left eye, respectively, the first ROI may be an integer C1=C FL1 -E to integer C2=C FL2 +E, and C FL1 and C FL2 and E are the column numbers of the first and second facial landmarks, respectively. For example, E is a number in the range between 5 and 15, particularly between 8 and 12. More particularly, the integer E may be equal to 10.

[0023] The location of a pixel in the first ROI may be uniquely determined by two-dimensional coordinates representing the location of said pixel in a two-dimensional grid of pixels representing the first ROI. Moreover, the location of a pixel in the first two-dimensional grid may be expressed by Cartesian or polar coordinates with respect to a two-dimensional reference frame in the plane of the first image.

[0024] The selection of the first ROI may comprise storing information about the first ROI. The information about the first ROI may comprise information about the color of the pixels of the first ROI and the location of said pixels in the first ROI. The information about the first ROI is in particular in a first bitmap. The selection of the first ROI may include storing data comprising memory addresses of bits that store the information about the first ROI. For example, said bits may be arranged in a bitmap that encodes the first image. The data and / or the information about the first ROI may be stored, for example temporarily, in a primary memory and / or a secondary memory of a computing device, for example a computing device performing the method of the present invention.

[0025] Any structure format can be used to encode the information about the first ROI, as long as the information can be retrieved and correctly interpreted. For example, the information about the location of the pixels of the first ROI can specify the location of some of the pixels in the first ROI, as long as the information is sufficient to correctly obtain the location of each of the pixels of the first ROI. For example, if the first ROI is a rectangular grid of the first image, the information about the location of the vertices of the grid is sufficient to obtain the location of each of the pixels of the first ROI.

[0026] In particular, the first gaze vector estimates the three-dimensional direction in which the eye contained in the first ROI is looking. The first gaze vector can be a three-dimensional unit vector that can be expressed in Cartesian, spherical or cylindrical coordinates. For example, the first gaze vector can be expressed in spherical coordinates relative to the three-dimensional location of the eyeball center or the iris center of the eye contained in the first ROI. In this case, the first gaze vector can be expressed by polar angle and azimuth angle.

[0027] In particular, an artificial neural network (ANN) is a computational model comprising a number of interconnected nodes that map ANN inputs to ANN outputs, and each node maps inputs to outputs. In particular, the nodes of an ANN are interconnected with each other such that, for each node, the input of said each node comprises the output of another node and / or the output of said each node is part of the input of another node. For example, the output of a generic node of an ANN is part of the ANN output and / or of the input of another node. In particular, the node input of a generic node of an ANN comprises one or more data items, each data item being either an output of another node or a data item of an ANN input.

[0028] For example, each node of the ANN may map the node's inputs to the node's outputs using an activation function that may depend on the node. In general, the activation function of a node may depend on one or more weights that weight the data items at the node's input.

[0029] In particular, the output of a node of an ANN may depend on a threshold, for example whether the value of an activation function evaluated at the input of said node is greater than, equal to, or less than a threshold.

[0030] For example, the ANN may be a VGG-16 neural network or a MnistNet neural network. The ANN may be a convolutional neural network, such as AlexNet. For example, the ANN may be part of a generative adversarial network.

[0031] The values ​​of the weights of the ANN may be obtained by training the ANN with at least a training dataset. During training, the values ​​of the weights are typically iteratively adjusted to minimize the value of a cost function that depends on the weights of the ANN, the ANN inputs, the ANN outputs, and / or the biases. For example, the training dataset may be an MPIIGaze or gaze capture dataset or a synthetic dataset, such as the SynthesEyes dataset or the UnityEyes dataset.

[0032] The performance of the ANN may be improved by using data augmentation during training. For example, the training data set may be expanded by augmenting at least a portion of the data of the training data set. Data augmentation may be performed by translating and / or rotating at least some of the images in the training data set. At least some of the images of the training data set may be augmented by changing the intensity of the images and / or by adding lines or obstacles to the images.

[0033] The ANN input may include information about a number of pixels of the input image or a portion of the input image. For example, the ANN input may include information about the location and color of said pixels. In particular, the information about the location of a pixel may be encoded in the two-dimensional coordinates of the pixel with respect to a two-dimensional reference frame in the plane of the input image.

[0034] When constructing the first gaze vector, the ANN uses the first ROI as an input. In this case, in particular, the ANN input comprises information about the positions and colors of the pixels of the first ROI, and more particularly, the ANN input may comprise or consist of the first bitmap. The ANN output may comprise information characterizing the first gaze vector, for example, the spherical, cylindrical or Cartesian coordinates of the first gaze vector. In this case, the ANN constructs the first gaze vector.

[0035] The three-dimensional location of the first viewpoint is in particular the intersection between the screen and the first line of sight. The first line of sight is in particular a line that intersects with the three-dimensional location of the ocular center or iris center of the eye included in the first ROI and is parallel to the first line of sight vector. The localization of the first viewpoint on the screen can be obtained by modeling the screen in terms of a plane (hereinafter also referred to as the "screen plane") and constructing the first viewpoint as the intersection between said plane and the first line of sight. For example, the three-dimensional coordinates of the location of the first viewpoint in a given reference frame (e.g., the camera reference frame) TIFF0007673091000004.tif9170 is Given by TIFF0007673091000005.tif13170, TIFF0007673091000006.tif9170 is the first line of sight vector, TIFF0007673091000007.tif9170 is the 3D coordinates of the 3D location of the ocular center or iris center of the eye contained in the first ROI. TIFF0007673091000008.tif9170 is perpendicular to the screen plane, TIFF0007673091000009.tif12170 is the 3D coordinates of the 3D location of a reference point on the screen surface. For example, this reference point could be the top left corner of the screen.

[0036] 3D coordinates of the first viewpoint with respect to a 3D reference frame centered on a reference point on the screen surface TIFF0007673091000010.tif10170 is The screen coordinates relative to the reference point of the screen can be obtained by appropriately rotating the reference frame around the reference point to obtain a further three-dimensional reference frame. In the further three-dimensional reference frame, the three-dimensional coordinates of the first viewpoint are given by is given by TIFF0007673091000012.tif9170, where: TIFF0007673091000013.tif10170 are the 2D screen coordinates of the first eye point relative to the screen reference frame centered on the screen's reference point.

[0037] Screen coordinates are typically expressed in units of length, such as centimeters, and may be converted to units of pixels as follows: TIFF0007673091000014.tif18170

[0038] The selection of the first ROI improves the selection of inputs for the ANN, thereby resulting in more accurate gaze vector construction and reducing the processing load of the ANN. The selection of the first ROI and the construction of the first gaze vector using the ANN synergistically interact with each other to improve the accuracy of the method under real-world operating conditions, particularly in the absence of dedicated hardware components. The method of the present invention can achieve an accuracy of less than 1° under a wide range of operating conditions.

[0039] According to an embodiment of the method of the present invention, the artificial neural network detects in the first ROI at least a first eye landmark location of the first eye landmark and a second eye landmark location of the second eye landmark, in particular the eye landmarks detected by the ANN in the step of constructing the first gaze vector make it possible to reconstruct the eye depicted in the first ROI.

[0040] In particular, the first ocular landmark and the second ocular landmark are the ocular center and the iris center of the eye included in the first ROI, respectively. The three-dimensional coordinates of the eye center in the camera reference frame (hereinafter also referred to as "ocular coordinates") can be constructed from the two-dimensional image coordinates of the first ocular landmark location in the first ROI using the camera matrix of the camera that acquires the first image. Similarly, the three-dimensional coordinates of the iris center in the camera reference frame (hereinafter also referred to as "iris coordinates") can be constructed from the two-dimensional image coordinates of the second ocular landmark location in the first ROI using the above-mentioned camera matrix. In particular, the first gaze vector can be constructed as the difference between the iris coordinates and the ocular coordinates.

[0041] In this case, the first gaze vector can be constructed by using basic algebraic operations that manipulate the two landmark locations. In this way, the complexity of the ANN can be reduced and the computational load of the method is decreased.

[0042] For example, in this embodiment, the ANN output comprises information about the first ocular landmark and the second ocular landmark in the first image and / or in the first ROI. In particular, the ANN output comprises two-dimensional image coordinates of the first ocular landmark and the second ocular landmark in the first image and / or in the first ROI.

[0043] The ANN output may comprise at least a heatmap associated with the eye landmark. In particular, the heatmap associated with the eye landmark is an image that represents a pixel-by-pixel confidence of the location of said landmark by using the color of the pixels of the eye landmark. In particular, the pixels of the heatmap correspond to pixels of the first ROI and / or the first image. For example, this correspondence may be implemented by using a mapping function, e.g., an isomorphic mapping, that maps each pixel of the heatmap onto a pixel of the first ROI and / or the first image.

[0044] The color of a pixel of the heatmap encodes information about the probability that the eye landmark associated with the heatmap is located at the pixel of the first ROI that corresponds to said pixel of the heatmap.

[0045] In particular, a pixel of the heatmap associated with an eye landmark encodes a per-pixel confidence, e.g., likelihood or probability, that the landmark is located at the pixel of the first ROI associated with that pixel of the heatmap, e.g., the darker the pixel of the heatmap, the more likely it is that an eye landmark is located at the pixel of the first ROI that corresponds to that pixel of the heatmap.

[0046] For example, the location of the eye landmark in the first ROI may be the location of a pixel of the first ROI that corresponds to a pixel in the first region of the heatmap associated with the eye landmark. In particular, the first region is the region of the heatmap in which the eye landmark is most likely to be located, e.g., in which the per-pixel confidence is the largest. In this case, in particular, the location of the eye landmark in the first ROI is detected by using a pixel in the first region of the heatmap associated with the landmark and the mapping function described above.

[0047] In one embodiment of the present invention, the ANN further detects at least eight eye boundary landmark locations in the first ROI ranging from the first eye boundary landmark location to an eighth eye boundary landmark location, in particular, the eye boundary landmarks are points of the outer boundary of the eye, e.g., of the eyelid or of the edge of the eye.

[0048] In a further embodiment of the method of the present invention, the ANN detects at least eight iris boundary landmark locations in the first ROI ranging from a first iris boundary landmark location to an eighth iris boundary landmark location, for example, the iris boundary landmarks being points of the corneal limbus of the eye, in other words, the iris-scleral boundary.

[0049] For example, the ANN output comprises a first heatmap, a second heatmap, and 16 further heatmaps ranging from third to eighteen. In particular, each of the third to tenth heatmaps encodes a pixel-by-pixel confidence of the location of one of the first to eighth eye boundary landmarks in such a manner that a different eye boundary landmark is associated with a different heatmap. Moreover, each of the eleventh to eighteenth heatmaps may encode a pixel-by-pixel confidence of the location of one of the first to eighth iris boundary landmarks in such a manner that a different iris boundary landmark is associated with a different heatmap.

[0050] The locations in the first ROI of the iris center, the eyeball center, the eight eye boundary landmarks, and the eight iris boundary landmarks can be obtained by processing the 18 heatmaps mentioned above with a soft-argmax layer.

[0051] For example, the ANN may be implemented using the following cost function: TIFF0007673091000015.tif13170, M j (s) is the value at pixel s of the jth heatmap calculated by the ANN by using the training ANN input IN of the training dataset as input. TIFF0007673091000016.tif9170 is the value of the jth heatmap of the ground truth associated with the training ANN input IN. The scale factor λ lies between 0.1 and 10, in particular between 0.5 and 5. More precisely, the factor λ is equal to 1. For example, N h can be between 1 and 20, and more particularly can be equal to 2 or 18.

[0052] According to another embodiment of the invention, the construction of the gaze vector is performed using a support vector regression (SVR) algorithm, which uses at least the first and second ocular landmark locations as input, which makes it possible in particular to improve the accuracy of the gaze location without having to rely on an eyeball model of the eye.

[0053] The SVR algorithm may allow for constructing the first gaze vector by estimating the eye pitch and yaw and / or the polar and azimuth angles of the spherical coordinates of the first gaze vector. In particular, the SVR algorithm uses as input a feature vector comprising a first eye landmark location and a second eye landmark location. The feature vector may further comprise eight iris boundary landmark locations and eight eye boundary landmark locations.

[0054] The feature vector may be a vector connecting an eyeball center location in the first image or first ROI to an iris center location in the first image or first ROI, and may also comprise a two-dimensional forward gaze vector directed toward the latter location. The presence of the forward gaze vector significantly increases the accuracy of the SVR algorithm, especially when the algorithm is trained by using a relatively small number of training samples, for example, about 20 training samples.

[0055] To improve the accuracy of gaze estimation, at least one component of the feature vector, for example each of the components of the feature vector, may be normalized to the eye width, in other words the distance between the leftmost location and the rightmost location of the eye contained in the first image or the first ROI. Moreover, the facial landmark locations in the feature vector and, if present, the previous two-dimensional gaze vector may be expressed in Cartesian or polar coordinates with respect to a two-dimensional reference frame in the plane of the first image, the origin of said reference frame being the eyeball center location.

[0056] An SVR algorithm may be trained in a user-independent manner by using a relatively large set of training feature vectors obtained from images containing different people. An SVR algorithm may be trained in a user-dependent manner by using a relatively small set of training feature vectors obtained from training images containing a particular user. For example, training images and corresponding training feature vectors may be obtained by asking a user to look at predefined locations on a screen and by obtaining (e.g., capturing) images depicting the user's face while the user is looking at these locations.

[0057] In one embodiment of the method according to the invention, the artificial neural network is a Hourglass Neural Network (HANN).

[0058] HANNs allow for gathering information across different scales of an input image or of parts of an input image. In particular, HANNs comprise one or more hourglass modules that map input features to output features. In particular, the hourglass modules are stacked in series such that the output features of an hourglass module form the input features of another hourglass module. The nodes of the hourglass modules are arranged in layers and configured in such a way that convolutional and max-pooling layers modify the input features by reducing the resolution. At each max-pooling step, the module branches and applies further convolutions at the original pre-pooled resolution. After reaching a minimum resolution, the hourglass modules increase the resolution by upsampling and by combining features across scales.

[0059] The use of the Hourglass artificial neural network in constructing the gaze vector, typically reserved for human pose estimation, results in a surprising improvement in the accuracy of the location of the first gaze point on the screen.

[0060] According to a further embodiment of the present invention, the first facial landmark is a third eye landmark and / or the second facial landmark is a fourth eye landmark, in which case the first facial landmark location and the second facial landmark location are the third eye landmark location and the fourth eye landmark location, respectively.

[0061] The detection of the two eye landmarks makes it possible to improve the accuracy of the selection of the first ROI, thereby improving the quality of the inputs supplied to the input of the ANN and, ultimately, the accuracy of the construction of the first gaze vector.

[0062] A further embodiment of the method of the present invention comprises the step of initiating locating a fifth eye landmark location of the fifth eye landmark, a sixth eye landmark location of the sixth eye landmark, a seventh eye landmark location of the seventh eye landmark, and an eighth eye landmark location of the eighth eye landmark in the first image, in this embodiment, the selection of the first ROI is performed by using the third eye landmark location, the fourth eye landmark location, the fifth eye landmark location, the sixth eye landmark location, the seventh eye landmark location, and the eighth eye landmark location.

[0063] In this case, for example, a first ROI may consist of pixels having column numbers between integers C3 and C4 and row numbers between integers R3 and R4, where C3, C4, R3 and R4 have the following relationship: C3 = min(S e,x )-E, C4=max(S e,x )+E, R3=min(S e,y )-E, R4=max(S e,y )+E (4) Meet the S e,x ={x e,3 ,x e,4 ,x e,5 ,x e,6 ,x e,7 ,xe,8} and S e,y ={y e,3 ,y e,4 ,y e,5 ,y e,6 ,y e,7 ,y e,8}. The coordinates (x e,3 ,y e,3 ), (x e,4 ,y e,4 ), (x e,5 ,y e,5 ), (x e,6 ,y e,6 ), (x e,7 ,y e,7 ), and (x e,8 ,y e,8 ) denote the 2D image coordinates of the six ocular landmark locations used to select the first ROI. The expressions min(S) and max(S) denote the minimum and maximum values ​​of the general set S, respectively.

[0064] Surprisingly, the selection of the first ROI using the six ocular landmark locations results in an improvement in the accuracy of the construction of the first gaze vector using the ANN.

[0065] One embodiment of the method of the present invention comprises the steps of: Initiating construction of a head pose estimation vector, where construction of the head pose estimation vector is performed by using at least a first facial landmark location and a second facial landmark location. The method further includes the steps of: The localization of the first viewpoint on the screen is based on the head pose estimation vector.

[0066] The head pose estimation vector is in particular a vector that comprises at least the information needed to calculate the yaw, pitch and roll of the face contained in the first image with respect to a reference face position.

[0067] The three-dimensional location of the eyeball center or iris center of the eye contained in the first ROI. TIFF0007673091000017.tif9170 is the head pose estimation vector In particular, the head pose estimation vector can be expressed as: It is written as TIFF0007673091000019.tif11170, where: TIFF0007673091000020.tif9170 and TIFF0007673091000021.tif8170 is a 3x3 rotation matrix and translation vector that respectively transforms 3D coordinates in the 3D world reference frame to 3D coordinates in the 3D camera reference frame. For example, if the location of the eye center or iris center of the first eye in the first image is represented by the 2D image coordinates (x0,y0), then the 3D vector TIFF0007673091000022.tif8170 is the following, in terms of the inverse of the camera matrix C, i.e. It is written as TIFF0007673091000023.tif19170. x and f y are the focal lengths in the x and y directions, respectively, and (c x ,c y ) is the optical center. The constant s is the global scale. The rotation matrix TIFF0007673091000024.tif8170 Translation vector The TIFF0007673091000025.tif8170 and / or the overall scale s may be calculated, inter alia, by solving the Perspective n-Point (PnP) problem, which may be formulated as a problem of estimating the pose of a face contained in a first image, given the 3D coordinates in a world reference frame and the 2D image coordinates of each facial landmark of a set of n1 facial landmarks of said face.

[0068] The set of n1 facial landmarks comprises a first facial landmark and a second facial landmark. For example, n1 is between 2 and 80, particularly between 3 and 75. Moreover, n1 may be between 4 and 70, more particularly equal to either 6 or 68. For example, the set of n1 facial landmarks is a subset of the set of n0 facial landmarks, particularly, the set of n0 facial landmarks coincides with the set of n1 facial landmarks.

[0069] A solution to the PnP problem can be obtained by using the P3P method, the efficient PnP (EPnP) method, and the direct least squares (DLS) method. These methods, especially the P3P method, can be complemented by the Random Sample Consensus (RANSAC) algorithm, which is an iterative method for improving parameter estimation from a set of data that contains outliers.

[0070] In particular, the rotation matrix TIFF0007673091000026.tif9170 Translation vector TIFF0007673091000027.tif8170 and / or the overall scale s may be calculated by using a direct linear transformation and / or a Levenberg-Marquardt optimization.

[0071] Constructing a head pose estimation vector makes it possible to improve the accuracy of the location of the first viewpoint under operating conditions without real-world constraints, for example, without forcing the user to assume a given predetermined position with respect to the screen and / or recording device.

[0072] In another embodiment of the method according to the invention, the construction of the head pose estimation vector is performed using at least a 3D face model, which uses as input at least the first and second facial landmark locations.

[0073] In particular, the three-dimensional face model may use as input a location in the first image of each facial landmark of a set of n2 facial landmarks, the set of n2 facial landmarks comprising the first facial landmark and the second facial landmark. For example, n2 is between 2 and 80, particularly between 3 and 75. Moreover, n2 may be between 4 and 70, more particularly equal to either 6 or 68. For example, the set of n2 facial landmarks is a subset of the set of n1 facial landmarks, and in particular, the set of n1 facial landmarks coincides with the set of n2 facial landmarks.

[0074] The three-dimensional model is in particular an algorithm that makes it possible to calculate the three-dimensional coordinates of each element of the set of n2 facial landmarks by using as input the facial landmark locations, in particular the two-dimensional image coordinates, of the elements of the set of at least n2 facial landmarks. The three-dimensional facial landmark model may in particular be the Surrey facial model or the Candide facial model.

[0075] In particular, the 3D face model is based on the rotation matrix TIFF0007673091000028.tif8170 and translation vector TIFF0007673091000029.tif8170, and finally to compute the head pose estimation vector, for example by using equations (5) and (6).

[0076] For example, the location of the generic jth element of a set of n2 facial landmarks in the first image is given by coordinates (x f,j ,y f,j ), the 3D model is the 3D location of the generic elements in the 3D world reference frame. TIFF0007673091000030.tif10170. In this case, in particular, the rotation matrix TIFF0007673091000031.tif8170 Translation vector TIFF0007673091000032.tif8170 and / or the overall scale s is given by the equation, i.e., Calculated by solving TIFF0007673091000033.tif50170.

[0077] The use of a 3D model allows for the reliable construction of a head pose estimation vector from a relatively small amount of landmark locations, thereby reducing the processing burden of the method of the present invention.

[0078] An embodiment of the method according to the invention comprises: initiating acquisition of at least a second image; ● commencing locating a third facial landmark location of the first facial landmark in the second image; and initiating an estimation of a fourth facial landmark location of the first facial landmark in the first image, where the estimation of the fourth facial landmark location is performed using an optical flow equation and a third facial landmark location; Initiating detection of a fifth facial landmark location of the first facial landmark in the first image; The method includes the steps of:

[0079] Locating the first facial landmark location in the first image is based on the fourth facial landmark location and on the fifth facial landmark location.

[0080] Locating the first facial landmark location in the first image may include calculating a midpoint of a segment connecting the fourth facial landmark location and the fifth facial landmark location, in particular, the first facial landmark location substantially corresponds to a midpoint of a segment connecting the fourth facial landmark location and the fifth facial landmark location.

[0081] For example, the second image is a vector image or a second two-dimensional grid of pixels, for example a rectangular grid of pixels. Moreover, the second image may be encoded by at least a bitmap.

[0082] The entries of the second two-dimensional grid may be arranged in columns and rows and may be enumerated in ascending order, with each column and each row associated with a column number and a row number, respectively. In particular, the location of a pixel in the second image may be determined by the row number of the row to which the pixel belongs. TIFF0007673091000034.tif11170 and the column number of the column to which the pixel belongs The two-dimensional image coordinates of the pixel in the second image can be uniquely determined by the two-dimensional vector For example, the two-dimensional image coordinates of a pixel in the second image are the coordinates of the pixel in the image plane reference frame of the second image.

[0083] The first and second images are in particular images with the same face captured at different times, in particular the first and second images may be, for example, two frames of a video captured by a recording device.

[0084] In particular, the time t1 when the first image is captured and the time t2 when the second image is captured can be expressed as (t1-t2)=Δ t For example, Δ t is in the range between 0.02 seconds and 0.07 seconds, in particular in the range between 0.025 seconds and 0.06 seconds, and more particularly in the range between 0.03 seconds and 0.05 seconds. For example, Δ t can be equal to 0.04 seconds.

[0085] Moreover, if the recording device is operated to capture images at a frequency υ, then Δ t is the formula Δ t =N t / υ. Integer N t may be in the range between 1 and 5, in particular in the range between 1 and 3. More particularly, the integer N t may be equal to 1, in other words the first image immediately follows the second image. The frequency υ may be in the range between 15 Hz and 40 Hz, in particular in the range between 20 Hz and 35 Hz, and more particularly in the range between 25 Hz and 30 Hz. For example, the frequency υ may be equal to 27 Hz.

[0086] The two-dimensional reference frame in the plane of the second image and the two-dimensional reference frame in the plane of the first image may be defined in such a way that the first image and the second image form a three-dimensional grid of voxels, e.g., pixels comprising a first layer and a second layer, where the first layer and the second layer of pixels are a first rectangular grid and a second rectangular grid, respectively.

[0087] The location of a pixel in a voxel may be expressed by three-dimensional voxel coordinates (x,y,t), where the coordinates (x,y) are the two-dimensional coordinates of the pixel in the two-dimensional reference frame of the layer to which the pixel belongs. The time coordinate t specifies which layer the pixel belongs to. In particular, if t=t1 or t=t2, the pixel belongs to the first layer or the second layer, respectively.

[0088] The location of the facial landmark in the first image and / or in the second image may be represented by the location of the corresponding reference pixel in the voxel. In particular, the location of the facial landmark may be represented by the three-dimensional voxel coordinates (x, y, t) of the reference pixel. In particular, the time coordinate of the facial landmark location in the first image is t1, and the time coordinate of the facial landmark location in the second image is t2.

[0089] For example, the method for constructing a viewpoint on the screen is performed several times during an operation time interval of a device implementing the method of the present invention. In particular, the operation time interval is an interval during which the construction of a viewpoint is considered important, for example, to enable hands-free human-machine interaction. In this case, in particular, a second image is acquired at time t2 and is used to locate a third facial landmark location using the method of the present invention and, finally, to calculate a viewpoint on the screen. The information of the third facial landmark location can then be used to obtain the first landmark location in the first image using an optical flow equation.

[0090] In particular, the voxel coordinates (x4, y4, t1) of the fourth facial landmark location in the first image are given by the following equation: TIFF0007673091000037.tif17170, where (x3,y3,t2) are the voxel coordinates of the location of the third facial landmark in the second image. x,3 ,v y,3 ) represents the velocity of the third facial landmark location in the second image. The velocity vector can be obtained by using an optical flow equation, which is located at the voxel location (x, y, t) and has a velocity (v x ,v y ) is subject to the following constraints: I x (x,y,t)v x +I y (x,y,t)v y =-I t (x,y,t) (9) Meet the following.

[0091] I(x,y,t) is the intensity at the voxel location (x,y,t). I x , I y and I tare the derivatives of the intensity with respect to x, y and t, respectively. The derivatives are evaluated specifically at the voxel location (x,y,t). In particular, the optical flow equation at the location of the third facial landmark in the second image is given by: I x (x3,y3,t1)v x,3 +I y (x3,y3,t1)v y,3 =-I t (x3,y3,t1) (10) It is written as follows.

[0092] In particular, the optical flow equation allows predicting the change in location of the first facial landmark in the scene during the time interval from time t2 to time t1, such that the fourth facial landmark location is an estimate of the location of the first facial landmark in the first image based on the location of the first facial landmark in the second image.

[0093] Equations (9) and (10) can be expressed as the velocity vector (v x,3 ,v y,3 In particular, the method according to the present invention may include a step of starting to locate a plurality of pixel locations in the second image. In particular, the distance, e.g., the Euclidean distance, between each of the plurality of pixel locations in the second image and the location of the third facial landmark may be used to calculate an upper bound distance (d U ) or the upper limit distance (d U ) is equivalent to the method p,1 ,y p,1 ,t1), (x p,2 ,y p,2 ,t1),...,(x p,P ),y p,P ), t1) to locate P pixel locations, the velocity vector (v x,3 ,v y,3 ) is expressed by the following equation: It can be calculated by solving TIFF0007673091000038.tif33170.

[0094] The simultaneous equations (11) can be solved by using the least squares technique. Moreover, the following relationship: TIFF0007673091000039.tif43170 can be filled.

[0095] The upper distance limit may be expressed in units of pixels and may be in the range between 1 pixel and 10 pixels, in particular in the range between 2 pixels and 8 pixels, for example, the upper distance limit may be in the range between 3 pixels and 6 pixels, in particular may be equal to 3 pixels or 4 pixels.

[0096] The present invention may include a step of starting to locate a first pixel location in the second image. In particular, a distance, e.g., a Euclidean distance, between the first pixel location in the second image and a third facial landmark location is less than or equal to an upper limit distance. According to the Lucas-Kanade method, a velocity vector (v x,3 ,v y,3 ) can be calculated by using equation (11) with P=1, (x p,1 ,y p,1 , t1) are the voxel coordinates of the first pixel location.

[0097] For example, locating the third facial landmark location in the second image may be obtained using the first location algorithm. Moreover, detecting the fifth facial landmark location of the first facial landmark in the first image may be performed, for example, by locating said landmark location in the first image using the first location algorithm.

[0098] In this embodiment, the location of the first facial landmark in the first image is improved by consistently including information about the previous location of the first facial landmark, in other words about the location of said landmark in the second image, in this way the use of information gathered during the operation time interval of the device is at least partially optimized.

[0099] In another embodiment of the present invention, the first facial landmark location is equal to a weighted average between the fourth facial landmark location and the fifth facial landmark location.

[0100] In particular, the weighted average comprises a first weight w1 and a second weight w2, which respectively multiply the fourth and fifth facial landmark locations. The weighted average is expressed as follows: TIFF0007673091000040.tif13170, where (x1, y1, t1) and (x5, y5, t1) are the voxel coordinates of the first and fifth facial landmark locations in the first image, respectively. In particular, the first facial landmark location is a midpoint between the fourth and fifth facial landmark locations, whereby the coordinates of the first facial landmark location may be obtained by using equation (13) with w1=w2=½.

[0101] In particular, in this embodiment, the accuracy of locating the first facial landmark location in the first image may be improved if the values ​​of the weights w1 and w2 are based on the accuracy of locating the third facial landmark location and the fifth facial landmark location. For example, if the location of the third facial landmark location in the second image is more accurate than the location of the fifth facial landmark location in the first image, the first weight may be greater than the second weight. Alternatively, if the location of the third facial landmark location in the second image is less accurate than the location of the fifth facial landmark location in the first image, the first weight may be less than the second weight. The accuracy of locating the facial landmark location in the image may depend, among other things, on the pitch, yaw and roll of the face depicted in the image.

[0102] According to an embodiment of the present invention, locating the first facial landmark location in the first image is based on a landmark distance, the landmark distance being a distance between the third facial landmark location and the fourth facial landmark location.

[0103] In particular, the landmark distance d L is the Euclidean distance between the third and fourth facial landmark locations. The landmark distance may be expressed in units of pixels and / or as the following: It can be written as TIFF0007673091000041.tif13170.

[0104] In a further embodiment of the invention, the first weight is a monotonically decreasing function of the landmark distance and the second weight is a monotonically increasing function of the landmark distance.

[0105] According to this embodiment, the contribution of the fourth facial landmark location to the weighted sum becomes more marginal as the landmark distance increases. In this way, the accuracy of localizing the first facial landmark location in the first image is improved as the landmark distance decreases, since the accuracy of the fourth facial landmark location, which is an estimate of the location of the first facial landmark obtained by using the third facial landmark location, decreases.

[0106] According to the invention, a monotonically increasing function of a distance, e.g., of a landmark distance, is in particular a function that does not decrease as said distance increases. Alternatively, a monotonically decreasing function of a distance, e.g., of a landmark distance, is in particular a function that does not increase as the distance decreases.

[0107] For example, the first weight and the second weight may be: It is written as TIFF0007673091000042.tif19170, and min(d L,0 ,d L ) is d L,0 and d L d means the minimum value between L,0 may be expressed in units of pixels and / or may range between 1 pixel and 10 pixels. In particular, d L,0 is in the range between 3 and 7 pixels, more specifically, d L,0 is equal to 5 pixels.

[0108] An embodiment of the method of the present invention comprises the steps of: - initiating a localization of a second viewpoint on the screen, the localization of the second viewpoint being performed using at least the first line of sight vector; The method includes the steps of:

[0109] In this embodiment, the location of a first viewpoint on the screen is performed using a second viewpoint.

[0110] The second viewpoint is in particular the intersection between the screen and the first line of sight. The localization of the second viewpoint above the screen can be obtained by modeling the screen in terms of a surface and by constructing the second viewpoint as the intersection between said surface and the first line of sight.

[0111] In this case, in particular, the location of the first viewpoint on the screen may depend only indirectly on the first line of sight vector, and only due to the dependence of the location of the second viewpoint on the first line of sight vector.

[0112] In an embodiment of the invention, localization of the second viewpoint on the screen is performed using a calibration function, which depends on at least the location of the first calibration viewpoint and an estimate of the location of the first calibration viewpoint.

[0113] In particular, the calibration function φ is expressed as the 2D screen coordinates of the second viewpoint on the screen as follows: 2D screen coordinates of the first viewpoint on the screen from TIFF0007673091000043.tif12170 This allows TIFF0007673091000044.tif10170 to be calculated. TIFF0007673091000045.tif16170

[0114] For example, the method of the present invention comprises: Initiating acquisition of at least a first calibration image; Initiating locating a facial landmark location of a facial landmark in a first calibration image; Initiating locating a further facial landmark location of a further facial landmark in the first calibration image; and initiating a selection of a first ROI in a first calibration image, wherein the selection of the first ROI in the first calibration image is performed by using at least a facial landmark location in the first calibration image and a further facial landmark location in the first calibration image; - initiating construction of a first calibration gaze vector, where the construction of the first calibration gaze vector is performed using at least an ANN, where the ANN uses at least a first ROI of the first calibration image as an input; - starting to build an estimate of a location of a first calibration viewpoint on the screen, said building being performed using at least a first calibration line of sight vector; The method may include at least the step of performing the above.

[0115] The estimate of the location of the first calibration point is in particular the intersection between the screen and the calibration line. The calibration line is in particular a line that intersects with the eyeball center or iris center of the eye included in the first ROI of the first calibration image and is parallel to the first calibration gaze vector. The construction of the estimate of the location of the first calibration point on the screen can be obtained by modeling the screen in terms of a surface and constructing said estimate as the intersection between said surface and the calibration line.

[0116] The method of the present invention may further include a step of initiating prompting the user to look at a calibration point on the screen. In particular, the prompting may be performed before the acquisition of the first calibration image. Moreover, the method may also include a step of initiating construction of a first calibration head pose estimation vector, the construction of which may be performed by using at least the facial landmark locations in the first calibration image and further facial landmark locations. In this case, an estimate of the localization of the first calibration viewpoint on the screen may be based on the first calibration head pose estimation vector. The calibration function may also depend on the first calibration gaze vector and / or on the first calibration head pose estimation vector.

[0117] In particular, the calibration function depends on the location of each element of the set of n3 calibration viewpoints and on an estimate of said location of each element of the set of n3 calibration viewpoints, for example, n3 being between 2 and 20, in particular between 3 and 15. Moreover, n3 may be between 4 and 10, and more particularly is equal to 5.

[0118] For example, an estimate of the location of each element of the set of n calibration viewpoints may be obtained by using steps that result in the construction of an estimate of the location of a first calibration viewpoint. In particular, the estimate of the location of each element of the set of n calibration viewpoints is constructed using a corresponding calibration gaze vector and, optionally, a corresponding calibration head pose estimation vector. For example, a calibration function may depend on multiple calibration gaze vectors and / or on multiple calibration head pose estimation vectors.

[0119] The calibration allows constructing the first viewpoint by taking into account the actual situation in which the user is looking at the screen, thereby improving the accuracy of the location of the first viewpoint under real setup conditions, which may include for example the position of the screen relative to the screen, the fact that the user is wearing glasses or is affected by strabismus.

[0120] For example, the calibration function may comprise or consist of a radial basis function. In particular, the radial basis function may be a linear function or a polyharmonic spline with an exponent equal to 1. The radial basis function allows for an improvement in the accuracy of the location of the first viewpoint even when the calibration is based on a relatively low amount of calibration points. The first input data may comprise information for identifying a position of the second viewpoint on the screen. In particular, said information may comprise or consist of two-dimensional screen coordinates of the second viewpoint. Alternatively, or in conjunction with the above, the first input data may comprise information for characterizing the first gaze vector and / or the head pose estimation vector. The information for characterizing the three-dimensional vectors, e.g. the first gaze vector and the head pose estimation vector, may comprise or consist of three-dimensional coordinates of said vectors.

[0121] According to an embodiment of the present invention, the calibration function depends on the distance, for example the Euclidean distance, between at least the location of the first calibration viewpoint and an estimate of the location of the first calibration viewpoint.

[0122] The dependence on distance between the location of the first calibration viewpoint and the estimate of the location of the first calibration viewpoint enables improvement in the accuracy of the location of the first viewpoint even when the calibration is based on a relatively low amount of calibration points, for example less than 10, in particular less than 6 calibration points.

[0123] For example, the location on the screen of the generic jth element of a set of n3 calibration points is expressed as a 2D screen coordinate (a c,j ,b c,j ) and the estimate of the location of the generic element is expressed as a two-dimensional screen coordinate (a e,j ,b e,j ), the calibration function is: It can be written as TIFF0007673091000046.tif15170, TIFF0007673091000047.tif21170 is TIFF0007673091000048.tif13170 and TIFF0007673091000049.tif16170. In particular, The file is TIFF0007673091000050.tif18170.

[0124] n 3-dimensional matrix TIFF0007673091000051.tif9170 is a matrix TIFF0007673091000052.tif9170. In particular, the generic entry of the latter matrix TIFF0007673091000053.tif11170 is Given by TIFF0007673091000054.tif16170.

[0125] The calibration function defined in equation (15) results in a surprisingly accurate calibration.

[0126] According to a further embodiment of the invention, the localization of the first viewpoint on the screen is performed by means of a Kalman filter.

[0127] In particular, the Kalman filter is an algorithm that uses a series of estimates of the viewpoint on the screen to improve the accuracy of the first viewpoint localization. For example, the k-th iteration of the filter calculates the 2D screen coordinates (a k-1 ,b k-1 ) and velocity (v a,k-1 ,v b,k-1 ) and the 2D screen coordinates of the kth intermediate viewpoint The 2D screen coordinates of the kth estimated viewpoint (a k ,b k ) and velocity (v a,k ,v b,k ) can be calculated. TIFF0007673091000056.tif24170

[0128] Matrix K k teeth, The matrix below S k =R k +H k [F k P k-1|k-1 (F k ) T +Q k ](H k ) T (18) With respect to the following, namely: K k =[F k P k-1|k-1 (F k ) T +Q k ](H k ) T (S k ) -1 (17) It can be expressed as:

[0129] In particular, for each value of k, the matrix P k|k is the following, i.e. P k|k =[(14-K k H k )[F k P k-1|k-1 (F k ) T +Q k ](14-K k H k ) T +K k R k (K k ) T ] (19) where 14 is the 4×4 identity matrix. k and Q k are the covariance of the observation noise at the kth iteration and the covariance of the process noise at the kth iteration, respectively. k and H kare the state transition model matrix at the kth iteration and the observation model matrix at the kth iteration, respectively. In particular, the following relationship is satisfied: At least one of the following is true: TIFF0007673091000057.tif25170.

[0130] In particular, δ t is δ for the frequency υ described above t = 1υ. The matrix R k , Q k , F k and / or H k may be iterative independent. The following relations, i.e., At least one of the following may be true: TIFF0007673091000058.tif24170, where 0 is a null matrix. For example, the first and second viewpoints may be calculated by using a Kalman filter with one iteration, for example, by using equations (16)-(21) with substitution k→1. In this case, the first and second viewpoints are the first estimated viewpoint and the first intermediate viewpoint, respectively. In this case, in particular, the 0th estimated viewpoint is calculated by the two-dimensional screen coordinates (a0,b1) as specified in equation (21). o ) and has a velocity (v a,0 ,v b,o ) can be a screen point with

[0131] The first perspective is P 0|0and V0, iteratively calculated by using a Kalman filter with M iterations. In this case, in particular, the first viewpoint and the second viewpoint are the Mth estimated viewpoint and the Mth intermediate viewpoint, respectively. For example, M is a given integer. Alternatively, M may be the smallest number of iterations required to satisfy a stopping condition. For example, the stopping condition may comprise a requirement that the distance between the Mth estimated viewpoint and the Mth intermediate viewpoint is less than a first threshold, and / or a requirement that the distance between the Mth estimated viewpoint and the (M-1)th estimated viewpoint is less than a second threshold. In particular, the first threshold may be equal to the second threshold.

[0132] According to an embodiment of the method of the present invention, the localization of the first viewpoint on the screen is performed using a third viewpoint and a covariance matrix of process noise, the covariance matrix of process noise comprising a plurality of entries, the entries being monotonically increasing functions of the distance between the second viewpoint and the third viewpoint. In particular, each of the entries of the covariance matrix may be a monotonically increasing function of the distance between the second viewpoint and the third viewpoint.

[0133] For example, the distance between the first viewpoint and the third viewpoint may be expressed in units of pixels. In particular, the distance between the second viewpoint and the third viewpoint is the Euclidean distance between the viewpoints. In particular, if the first viewpoint is the Mth intermediate viewpoint, the third viewpoint may be the (M-1)th estimated viewpoint, and the distance d between the first viewpoint and the third viewpoint may be expressed in units of pixels. M,M-1 can be written as follows: TIFF0007673091000059.tif14170

[0134] The process noise covariance matrix is ​​the process noise covariance matrix Q k and the following relationship: At least one of the following is possible: TIFF0007673091000060.tif26170.

[0135] For example, q is the distance d k,k-1is a function of: It can be written as TIFF0007673091000061.tif27170, d k,k-1 can be obtained by using equation (23) with substitution M→k. In particular, q0 is 10 -35 From 10 -25 More specifically, in the range between 10 -32 From 10 -28 Moreover, q1 can range between 10 -25 From 10 -15 In particular, in the range between 10 -22 From 10 -18 For example, q2 can range between 10 -15 From 10 -5 More specifically, in the range between 10 -12 From 10 -8 For example, q3 is in the range between 10 -5 From 10 -1 In particular, in the range between 10 -3 In addition, q0, q1, q2, and / or q3 are each in the range of 10 -30 , 10 -20 , 10 -10 , and 10 -2 In addition, d0 may be in the range between 50 pixels and 220 pixels, more particularly in the range between 100 pixels and 200 pixels. In particular, d1 is in the range between 200 pixels and 400 pixels, more particularly in the range between 220 pixels and 300 pixels. In addition, d2 is in the range between 300 pixels and 600 pixels, more particularly in the range between 400 pixels and 550 pixels. Moreover, d0, d1 and / or d2 may be equal to 128 pixels, 256 pixels and 512 pixels, respectively.

[0136] In this embodiment, the numerical significance of the covariance matrix of the process noise is reduced when the distance between the intermediate viewpoint of the kth iteration and the estimated viewpoint of the (k-1)th iteration decreases. In this way, the Lukas-Kanade filter allows achieving a relatively high accuracy of the localization of the first viewpoint by using a relatively small number of iterations.

[0137] In a further embodiment of the method of the present invention, the localization of the first viewpoint on the screen is performed using a covariance matrix of a third viewpoint and an observation noise. The covariance matrix of the observation noise comprises a plurality of entries, the entries being monotonically increasing functions of the distance between the second viewpoint and the third viewpoint. For example, each of the entries of the covariance matrix may be a monotonically increasing function of the distance between the second viewpoint and the third viewpoint.

[0138] The covariance matrix of the observation noise is the covariance matrix of the observation noise at the kth iteration, R k and the following relationship: At least one of the following is possible: TIFF0007673091000062.tif14170.

[0139] For example, r is the distance d k,k-1 is a function of: It can be written as TIFF0007673091000063.tif15170.

[0140] In particular, r0 is 10 -35 From 10 -25 More specifically, in the range between 10 -32 From 10 -28 Moreover, r1 can range between 10 -5 From 10 -1 In particular, in the range between 10 -3 Furthermore, r0 and / or r1 may each range between 10 -30 and 10 -2Furthermore, d3 may be in the range between 50 pixels and 220 pixels, more particularly between 100 pixels and 200 pixels. In particular, d3 is equal to 128 pixels.

[0141] An embodiment of the method comprises: ● commencing locating a sixth facial landmark location of a sixth facial landmark in the first image; ● commencing locating a seventh facial landmark location of a seventh facial landmark in the first image; initiating a selection of a second region of interest in the first image, wherein the selection of the second region of interest is performed by using at least a sixth facial landmark location and a seventh facial landmark location; - commencing construction of a second gaze vector using at least an artificial neural network, the artificial neural network using at least a second region of interest as an input; The method includes the steps of:

[0142] In particular, according to this embodiment, locating the first viewpoint on the screen is performed using a first gaze vector and a second gaze vector. For example, the second ROI comprises at least the second eye or a part of the second eye. In particular, the second eye is different from the first eye. The second ROI may be encoded by a second bitmap. In this embodiment, the location of the first viewpoint is obtained by using information about the gaze vectors of both eyes. The accuracy of the location is thereby improved.

[0143] When constructing the second gaze vector, the ANN uses the second ROI as an input, where in particular the ANN input comprises information about the positions and colors of the pixels of the second ROI, and more particularly the ANN input may comprise or consist of a second bitmap.

[0144] The method according to the invention may comprise a step of initiating a localization of a fourth viewpoint and a fifth viewpoint, the localization of the fourth viewpoint and the fifth viewpoint being performed using a first line of sight vector and a second line of sight vector, respectively. In particular, the screen coordinates of the first viewpoint or the second viewpoint may be determined by the first line of sight vector and the second line of sight vector. TIFF0007673091000064.tif10170 is the screen coordinates of the fourth viewpoint, as follows: TIFF0007673091000065.tif10170 and the screen coordinates of the 5th viewpoint TIFF0007673091000066.tif10170 can be obtained by a weighted sum of TIFF0007673091000067.tif12170

[0145] For example, the first viewpoint or the second viewpoint may substantially coincide with a midpoint between the location of the fourth viewpoint and the location of the fifth viewpoint, i.e. TIFF0007673091000068.tif10170 can be obtained from the above equation by setting u1 = u2 = 1 / 2. The weights u1 and u2 may depend on the head pose estimation vector. In particular, the dependence of these weights is as follows, i.e. - u1 is less than u2 if the construction of the fourth viewpoint is less accurate than the construction of the fifth viewpoint due to the user's head pose; - u1 is greater than u2 if the reconstruction of the fourth viewpoint is more accurate than the reconstruction of the fifth viewpoint due to the user's head pose; It's like that.

[0146] For example, the construction of the fourth viewpoint is more accurate than the construction of the fifth viewpoint if the reconstruction of the first eye is more accurate than the reconstruction of the second eye, and is less accurate than the construction of the fifth viewpoint if the reconstruction of the first eye is less accurate than the reconstruction of the second eye, respectively. For example, the reconstruction of one eye of a user's face may be more accurate than the reconstruction of the other eye when the user's head is rotated in such a way that the latter eye is at least partially obscured in the first image, for example, by the user's nose.

[0147] The method of the present invention may further comprise the step of receiving a location of the first viewpoint on the screen and / or initiating a display of the location of the first viewpoint on the screen.

[0148] According to the present invention, the step of initiating an action, such as acquiring an image, locating a landmark or a viewpoint, selecting an area of ​​interest, constructing a vector, and / or estimating a facial landmark, may in particular be performed by performing said action. For example, the step of initiating acquisition of an image, e.g. a first image or a second image, may be performed by acquiring said image. Analogously, the step of initiating locating a facial landmark location in an image may be performed by locating said facial landmark location in said image. For example, the step of initiating construction of a vector, such as a first gaze vector or a head pose estimation vector, may be performed by constructing said vector.

[0149] Initiating an action, such as image acquisition, landmark or viewpoint location, ROI selection, vector construction, and / or facial landmark estimation, may in particular be performed by instructing a dedicated device to perform said action. For example, initiating acquisition of an image, e.g., a first image or a second image, may be performed by instructing a recording device to acquire said image. Analogously, initiating locating a facial landmark location in an image may be performed by instructing a dedicated computing device to locate said facial landmark location in said image.

[0150] According to the present invention, the step of initiating a first action may be performed together with the step of initiating one or more other actions. For example, the step of initiating locating a first landmark and the step of initiating locating a second landmark may be performed together by initiating locating a plurality of facial landmarks, the plurality including the first landmark and the second landmark.

[0151] Moreover, the step of initiating the selection of the first ROI may be performed together with the step of initiating the location of the first and second facial landmarks. For example, a computing device performing the present invention may instruct another computing device to locate the first and second facial landmarks. The latter device may be configured in such a way that after the latter device performs this location, the latter device performs the selection of the first ROI. In this case, according to the present invention, the step of initiating the location of the first and second facial landmarks also initiates the selection of the first ROI.

[0152] In one embodiment, all steps of the method according to the present invention can be implemented together. For example, a computing device performing the present invention can command another computing device to acquire a first image. After the latter device acquires said image, the latter device identifies the location of the first and second landmarks, selects a first ROI, constructs a first gaze vector, and then identifies the location of the first viewpoint on the screen. In this case, according to the present invention, the step of starting the acquisition of the first image also starts the location of the first and second facial landmark locations, the selection of the first ROI, the construction of the first gaze vector, and the location of the first viewpoint on the screen.

[0153] Throughout this specification, steps of the methods of the present invention are disclosed according to a given order. However, this given order does not necessarily reflect the chronological order in which the steps of the present invention are performed.

[0154] The invention also refers to a data processing system comprising at least a processor adapted to implement the method according to the invention.

[0155] Moreover, the invention refers to a computer program product comprising instructions which, when executed by a computing device, cause the computing device to perform the method according to the invention.

[0156] The invention also relates to a computer-readable storage medium comprising instructions which, when executed by a computing device, cause the computing device to perform the method according to the invention. The computer-readable medium is in particular non-transitory.

[0157] Exemplary embodiments of the invention are described below with reference to the accompanying drawings, in which the drawings and the corresponding detailed description merely serve to provide a better understanding of the invention and do not constitute any limitation of the scope of the invention as defined in the claims. [Brief description of the drawings]

[0158] [Figure 1] 1 is a schematic diagram of a first embodiment of a data processing system according to the present invention; [Diagram 2] FIG. 2 is a flow diagram of the operation of a first embodiment of the method according to the invention. [Figure 3a] 3 is a schematic representation of a first image obtained by implementing a first embodiment of the method according to the invention; [Figure 3b] 2 is a schematic representation of facial landmarks located by implementing a first embodiment of the method according to the invention; [Figure 3c] 3 is a schematic representation of a first ROI selected by implementing a first embodiment of the method according to the invention; [Figure 4a] 3 is a schematic representation of a heat map obtained by implementing a first embodiment of the method according to the invention; [Figure 4b] 3 is a schematic representation of a heat map obtained by implementing a first embodiment of the method according to the invention; [Diagram 5] FIG. 4 is a flow diagram of the operation of a second embodiment of the method according to the present invention. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0159] 1 is a schematic diagram of a first embodiment of a data processing system 100 according to the present invention. The data processing system 100 may be a computing device or a cluster of computing devices. The data processing system 100 comprises a processing element 110 and a storage means 120 in data communication with each other.

[0160] The processing element 110 may consist of or comprise a CPU and / or a GPU. Moreover, the processing element 110 comprises several modules 111-116 configured to perform steps of the method according to the invention. The first initiating module 111 is configured to initiate acquiring at least a first image. For example, the first initiating module 111 may be an acquiring module configured to acquire, e.g. capture, the first image.

[0161] The second initiation module 112 is configured to initiate locating a first facial landmark location of the first facial landmark in the first image. The second initiation module 112 may be a first locating module configured to locate a first facial landmark location of the first facial landmark in the first image. The third initiation module 113 is configured to initiate locating a second facial landmark location of the second facial landmark in the first image. In particular, the second initiation module 112 may be a second locating module configured to locate a second landmark location of the second facial landmark.

[0162] The third initiation module 113 and the second initiation module 112 may be the same initiation module and may be configured, for example, to initiate locating a plurality of facial landmark locations in the first image, the plurality including a first facial landmark location of the first facial landmark and a second facial landmark location of the second facial landmark.

[0163] The fourth initiation module 114 is configured to initiate a selection of a first ROI, in particular, a selection module configured to select the first ROI by using at least the first and second facial landmark locations. The fifth initiation module 115 is configured to initiate construction of a first gaze vector. For example, the fifth initiation module 115 is a construction module configured to construct the gaze vector using an ANN.

[0164] The sixth start module 116 is instead configured to start locating the first viewpoint on the screen. For example, the sixth start module 116 is a third locating module configured to locate the first viewpoint on the screen using at least the first line of sight vector.

[0165] The storage means 120 may comprise a volatile primary memory 121 and / or a non-volatile primary memory 122. The storage means 120 may further comprise a secondary memory 123 which may store an operating system and / or an ANN. Moreover, the secondary memory 123 may store a computer program product comprising instructions which, when executed by the processing element 110, cause the data processing system 100 to perform the method according to the present invention. The secondary memory 123 may store information about the first image and / or the first ROI.

[0166] The secondary memory 123, the primary memories 121, 122, and the processing element 110 need not be physically contained within the same housing, but may instead be spatially separated from one another. In particular, the secondary memory 123, the primary memories 121, 122, and the processing element 110 may be spatially separated from one another and may exchange data with one another via wired and / or wireless media (not shown).

[0167] Data processing system 100 may further include an input / output (I / O) interface 140 that enables the system 100 to communicate with input / output devices (eg, a display, a keyboard, a touch screen, a printer, a mouse, etc.).

[0168] The data processing system 100 may further comprise a network interface controller (NIC) 130 configured to connect the system 100 with a suitable network (not shown). In accordance with the present invention, the suitable network may be, for example, an intranet, the Internet, or a cellular network. For example, the NIC 130 may enable the data processing system 100 to exchange data with another computing device (not shown) that performs, for example, locating the facial landmarks, selecting the first ROI, constructing the first gaze vector, and / or locating the first gaze point.

[0169] In particular, the data processing system 100 comprises a recording device 160 configured to capture at least a first image. For example, the recording device may be a photo camera and / or a video camera. As shown in FIG. 1, the recording device 160 may be connected to the processing element 110 via the I / O interface 140. For example, the recording device 160 may be wirelessly connected to the I / O interface via the NIC 130. The data processing system 100 may comprise a display unit 150, comprising a screen 151, connected to the processing element 110 via the I / O interface 140. In particular, said unit 140 may be wirelessly connected to the I / O interface via the NIC 130. The recording device 160 and / or the display unit 150 may be intelligent devices with their own memory for storing relevant instructions and data for use with the I / O interface 140 or peripheral devices.

[0170] 2 is a flow diagram 200 of the operation of a first embodiment of the method according to the invention. In particular, the first embodiment of the method according to the invention may be implemented by a first computing device (not shown), which may be, for example, the data processing system 100 described above and depicted diagrammatically in FIG. 1. In step 210, the first computing device starts the acquisition of a first image 300, which is depicted diagrammatically in FIG. 3a. In particular, the first image 300 is acquired by using the recording device 160.

[0171] In steps 220 and 230, the first computing device starts locating a first facial landmark location 301 of the first facial landmark in the first image 300 and a second facial landmark location 302 of the second facial landmark in the first image 300. In particular, the above steps can be performed together by starting to locate in the first image 68 facial landmarks whose locations are represented in FIG. 3b by crossing dots. For example, the location of the facial landmarks is performed using a first location algorithm. In particular, the first facial landmark and the second facial landmark are the third eye landmark 301 and the fourth eye landmark 302, respectively. Moreover, the set of 68 facial landmarks comprises a fifth eye landmark 303, a sixth eye landmark 304, a seventh eye landmark 305, and an eighth eye landmark 306.

[0172] In step 240, the first computing device starts selecting a first ROI 310, which is diagrammatically represented in Fig. 3c. The selection of the first ROI 310 is performed by using the six eye landmarks 301-306 described above. In particular, the first ROI 310 consists of pixels having column numbers between integers C3 and C4 and row numbers between integers R3 and R4, where C3, C4, R3 and R4 satisfy equation (4).

[0173] In step 250, the first computing device starts constructing a first gaze vector. The construction of this gaze vector is performed by an ANN. In particular, the ANN output comprises eighteen heat maps 401-418, ranging from the first to the eighteenth. These heat maps are illustrated diagrammatically in Fig. 4a and Fig. 4b. More specifically, the heat maps 401 and 402 encode the pixel-wise confidence of the iris center and the eyeball center, respectively. Each of the heat maps 403-410 encodes the pixel-wise confidence of the location of one of the first to eighth eye boundary landmarks, in such a way that different eye boundary landmarks are associated with different heat maps. Moreover, each of the heat maps 411-418 encodes the pixel-wise confidence of the location of one of the first to eighth iris boundary landmarks, in such a way that different iris boundary landmarks are associated with different heat maps.

[0174] In the heatmaps 401-418, which are depicted diagrammatically in Fig. 4a and Fig. 4b, the regions with the highest per-pixel confidence values ​​are shaded or highlighted in black. Moreover, the per-pixel confidence values ​​are greater in the dark regions than in the shaded regions. The locations of the eye landmarks associated with each of the heatmaps 401-418 may correspond to the locations of points in the respective dark regions of the heatmaps 401-418. In particular, the locations of the iris center, the eyeball center, the eight eye boundary landmarks, and the eight iris boundary landmarks are obtained by processing the above-mentioned 18 heatmaps 401-418 using a soft-argmax layer.

[0175] The construction of the gaze vector is performed using the SVR algorithm, which in particular uses as input the eyeball center location, the iris center location, the eight iris boundary landmark locations, and the eight eye boundary landmark locations obtained by using the eighteen heat maps 401-418.

[0176] In step 260, the first computing device starts locating a first viewpoint on the screen. The first viewpoint is in particular an intersection between the screen 151 and the first line of sight. For example, the localization of the first viewpoint on the screen 151 can be obtained by modeling the screen in terms of a plane and constructing the first viewpoint as an intersection between the plane and the first line of sight.

[0177] 5 is a flow diagram 500 of the operation of a second embodiment of the method according to the invention. In particular, said embodiment may be implemented by a second computing device (not shown), which may be, for example, the data processing system 100 described above and depicted diagrammatically in FIG. 1. According to this embodiment, the first viewpoint is P 0|0 and is calculated iteratively by using a Kalman filter starting from V0.

[0178] The general iteration of the Kalman filter is represented by a counter m that is initialized to a value 0 in step 505. During the m-th iteration, the second computing device starts acquiring the m-th intermediate image (step 515), which may be acquired in particular by using the recording device 160. In particular, the first image and the m-th intermediate image comprise the same subject and are captured at different times. For example, the first image and the m-th intermediate image are, for example, two frames of a video captured by the recording device 160 of the computing device of the present invention.

[0179] In steps 520 and 525 of the m-th iteration, the second computing device begins locating facial landmark locations of the facial landmarks and further facial landmark locations of the further facial landmarks in the m-th intermediate image. The above steps may be performed together by beginning to locate the 68 facial landmarks using the first location algorithm.

[0180] In particular, the distribution of the 68 facial landmarks in the mth intermediate image is similar to the distribution of the 68 facial landmarks, which is illustrated diagrammatically in Fig. 3a. In particular, the set of 68 facial landmarks located in the mth intermediate image comprises 6 eye landmarks for the left eye and 6 eye landmarks for the right eye.

[0181] In step 530 of the m-th iteration, the second computing device starts to select the m-th intermediate ROI. In particular, the m-th intermediate ROI comprises an eye, for example, the left eye, and is selected by using six ocular landmarks of the left eye. In particular, the m-th intermediate ROI consists of pixels with column numbers between integers C3 and C4 and row numbers between integers R3 and R4. The integers C3, C4, R3 and R4 can be calculated by using Equation (4) and the coordinates in the m-th intermediate image of the six ocular landmarks of the left eye. The position of the m-th intermediate ROI in the m-th intermediate image is similar to the position of the first ROI, which is illustrated in FIG. 3b.

[0182] In step 535, the second computing device starts to construct the m-th intermediate gaze vector by using an ANN. The ANN of the second embodiment is equivalent to the ANN used by the first embodiment of the method according to the present invention. In particular, the ANN uses the pixels of the m-th intermediate ROI as input and provides 18 heat maps as output. The heat maps are similar to the heat maps depicted diagrammatically in Fig. 4a and Fig. 4b and are used to find the location in the m-th intermediate image of the iris center of the left eye, the eyeball center of the left eye, the eight eye boundary landmarks of the left eye, and the eight iris boundary landmarks of the left eye. For example, the landmarks can be obtained by processing the 18 heat maps using a soft-argmax layer.

[0183] The locations of the iris center of the left eye, the eyeball center of the left eye, the eight eye boundary landmarks of the left eye, and the eight iris boundary landmarks of the left eye are then used as inputs for the SVR algorithm to construct the mth intermediate gaze vector.

[0184] In step 540 of the m-th iteration, the second computing device starts to locate the m-th intermediate viewpoint on the screen. The viewpoint is specifically the intersection between the screen and the m-th intermediate line of sight. For example, the location of the first viewpoint on the screen can be obtained by modeling the screen with respect to a plane and constructing the first viewpoint as the intersection between the plane and the m-th intermediate line of sight. In particular, the m-th intermediate line of sight is the line that intersects with the eyeball center of the eye that is included in the m-th intermediate ROI and is parallel to the m-th intermediate line of sight vector.

[0185] In step 545 of the m-th iteration, the second computing device starts to calculate the m-th estimated viewpoint on the screen, which is performed in particular by using the m-th intermediate viewpoint, the (m-1)-th estimated viewpoint calculated during the (m-1)-th iteration, and equations (16)-(26) with substitution k→m.

[0186] In step 550 of the m-th iteration, the second computing device checks whether the m-th estimated viewpoint satisfies a stopping condition, in particular, the stopping condition comprises a requirement that the Euclidean distance between the m-th estimated viewpoint and the m-th intermediate viewpoint is less than a first threshold and / or a requirement that the distance between the m-th estimated viewpoint and the (m-1)-th estimated viewpoint is less than a second threshold.

[0187] If the stopping condition is not met, the second computing device increments the counter value by 1 (see step 510) and performs the (m+1)th iteration. If the stopping condition is met, in step 555, the second computing device starts locating the first viewpoint, which is performed by setting the viewpoint to be equal to the mth estimated viewpoint constructed in the mth iteration.

Claims

1. A method for locating a first viewpoint on a screen (151) by a computing device, the method comprising: Initiating the acquisition (210) of at least a first image (300); Initiating a location (220) of a first facial landmark location (301) of a first facial landmark in the first image (300); Initiating a location (230) of a second facial landmark location (302) of a second facial landmark in the first image (300); and initiating a selection (240) of a first region of interest (310) in said first image (300), said selection of said first region of interest (310) being performed by using at least said first facial landmark location (301) and said second facial landmark location (302); - initiating a construction (250) of a first gaze vector, said construction of said first gaze vector being performed using at least an artificial neural network, said artificial neural network using at least said first region of interest (310) as an input; Initiating a localization (250) of the first viewpoint on the screen (151), the localization of the first viewpoint being performed using at least the first line of sight vector, and the localization of the first viewpoint on the screen (151) being performed using a Kalman filter; The method includes at least the steps of:

2. 2. The method of claim 1, wherein the artificial neural network detects at least a first eye landmark location (403) of a first eye landmark and a second eye landmark location (404) of a second eye landmark in the first region of interest (310).

3. 3. The method of claim 2, wherein the construction of the gaze vector is performed using a support vector regression algorithm, the support vector regression algorithm using at least the first eye landmark location (403) and the second eye landmark location (404) as input.

4. 4. The method of claim 1, wherein the artificial neural network is an hourglass neural network.

5. Initiating construction of a head pose estimation vector, said construction of said head pose estimation vector being performed by using at least said first facial landmark location (301) and said second facial landmark location (302). The method further includes the step of: said localization of said first viewpoint on said screen (151) being based on said head pose estimation vector; 5. The method according to any one of claims 1 to 4.

6. 6. The method of claim 5, wherein the construction of the head pose estimation vector is performed using at least a 3D face model, the 3D face model using at least the first facial landmark location (301) and the second facial landmark location (302) as input.

7. initiating acquisition of at least a second image; and initiating locating a third facial landmark location of the first facial landmark in the second image; and initiating an estimation of a fourth facial landmark location of the first facial landmark in the first image (300), wherein the estimation of the fourth facial landmark location is performed using an optical flow equation and the third facial landmark location; initiating detection of a fifth facial landmark location of the first facial landmark in the first image (300); The method further includes the step of: the locating of the first facial landmark location (301) in the first image (300) is based on the fourth facial landmark location and the fifth facial landmark location; 7. The method according to any one of claims 1 to 6.

8. 8. The method of claim 7, wherein the locating of the first facial landmark location (301) in the first image (300) is based on a landmark distance, the landmark distance being a distance between the third facial landmark location and the fourth facial landmark location.

9. The method of claim 7 or 8, wherein the first facial landmark location (301) is equal to a weighted average between the fourth facial landmark location and the fifth facial landmark location.

10. - initiating a localization of a second viewpoint on said screen (151), said localization of said second viewpoint being performed using at least said first line of sight vector. The method further includes the step of: said locating said first viewpoint on said screen (151) is performed using said second viewpoint; 10. The method according to any one of claims 1 to 9.

11. 11. The method of claim 10, wherein the localization of the second viewpoint on the screen (151) is performed using a calibration function, the calibration function depending on at least a location of a calibration viewpoint and an estimate of the location of the calibration viewpoint.

12. 11. The method of claim 10, wherein the localization of the first viewpoint on the screen (151) is performed using a third viewpoint and a covariance matrix of process noise, the covariance matrix of the process noise comprising a plurality of entries, the entries being a monotonically increasing function of the distance between the first viewpoint and the third viewpoint.

13. A data processing system (100) comprising at least a processor (110) configured to perform the method of any one of claims 1 to 12.

14. A computer program comprising instructions which, when executed by a computing device, cause the computing device to perform a method according to any one of claims 1 to 12.

15. A computer readable storage medium comprising instructions which, when executed by a computing device, cause the computing device to perform the method of any one of claims 1 to 12.

Citation Information

Patent Citations

  • Method and device to determine trigger intent of user

    KR1020190085466A

  • Method and apparatus to determine trigger intent of user

    US20190212815A1

  • Image processing device, image processing method, image processing program, and recording medium

    WO2005006251A1