Line of sight estimation method, apparatus, device, and medium

By acquiring the calibration information and light source position of the target eye image, and combining it with the gaze estimation model, the problem of insufficient gaze estimation accuracy is solved, and a more accurate and reliable gaze direction estimation is achieved.

CN119851331BActive Publication Date: 2025-12-30BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311346260.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-10-17
Publication Date
2025-12-30
Estimated Expiration
2043-10-17

AI Technical Summary

Technical Problem

Existing line-of-sight estimation methods are insufficient in accuracy, which affects their effectiveness in different application scenarios.

Method used

By acquiring the calibration information, camera position, and light source position of the target eye image, and combining them with a pre-set gaze estimation model, the calibration information is introduced to improve the accuracy of gaze direction estimation. The gaze estimation model is used to solve the problem based on the acquired information, especially through a self-attention mechanism and pre-set relational constraints, to determine the gaze direction.

Benefits of technology

It improves the accuracy of gaze estimation, reduces the negative impact of camera state changes on estimation results, and reduces the cost and data acquisition required for training the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119851331B_ABST
    Figure CN119851331B_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure relate to a line-of-sight estimation method, device, equipment and medium, wherein the method comprises: obtaining calibration information corresponding to a target eye image to be estimated; wherein the calibration information comprises coordinate information corresponding to pixels of the target eye image in a device coordinate system, and the device coordinate system is determined based on a camera used to collect the target eye image; obtaining a camera position corresponding to the camera and a light source position corresponding to a target light source; and estimating a line-of-sight direction corresponding to the target eye image by using a preset line-of-sight estimation model based on the target eye image, the calibration information, the camera position and the light source position. The line-of-sight direction estimated by the embodiments of the present disclosure is more accurate and reliable.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of computer technology, and in particular to a line-of-sight estimation method, apparatus, device, and medium. Background Technology

[0002] In many fields such as gaming, healthcare, and intelligent control, it is necessary to identify the direction of human gaze in order to implement appropriate strategies based on the estimated gaze direction. However, the current accuracy of gaze direction estimation can negatively impact the application scenarios, thus necessitating an improvement in the accuracy of gaze estimation results. Summary of the Invention

[0003] To solve the above-mentioned technical problems, or at least partially solve them, this disclosure provides a line-of-sight estimation method, apparatus, device, and medium.

[0004] In a first aspect, embodiments of this disclosure provide a gaze estimation method, the method comprising: acquiring calibration information corresponding to a target eye image for which the gaze is to be estimated; wherein the calibration information includes coordinate information of pixels of the target eye image in a device coordinate system, the device coordinate system being determined based on a camera used to acquire the target eye image; acquiring a camera position corresponding to the camera and a light source position corresponding to a target light source; and estimating the gaze direction corresponding to the target eye image using a preset gaze estimation model based on the target eye image, the calibration information, the camera position, and the light source position.

[0005] Secondly, embodiments of this disclosure also provide a gaze estimation device, comprising: a calibration information acquisition module, configured to acquire calibration information corresponding to a target eye image for which the gaze is to be estimated; wherein the calibration information includes coordinate information of pixels of the target eye image in a device coordinate system, the device coordinate system being determined based on a camera used to acquire the target eye image; a position acquisition module, configured to acquire a camera position corresponding to the camera and a light source position corresponding to a target light source; and a gaze estimation module, configured to estimate the gaze direction corresponding to the target eye image using a preset gaze estimation model based on the target eye image, the calibration information, the camera position, and the light source position.

[0006] Thirdly, embodiments of this disclosure also provide an electronic device, the electronic device comprising: a processor; a memory for storing executable instructions of the processor; the processor being configured to read the executable instructions from the memory and execute the instructions to implement the line-of-sight estimation method as provided in embodiments of this disclosure.

[0007] Fourthly, embodiments of this disclosure also provide a computer-readable storage medium storing a computer program for performing the line-of-sight estimation method as provided in embodiments of this disclosure.

[0008] Fifthly, embodiments of this disclosure also provide a computer program product, including a computer program that, when executed by a processor, implements the line-of-sight estimation method as provided in embodiments of this disclosure.

[0009] The technical solution provided in this disclosure fully considers the influence of the camera on the imaging result of the eye image. Therefore, calibration information is introduced when estimating the gaze direction of the eye image, and information such as the light source position and camera position corresponding to the target eye image can be further obtained. The gaze direction is estimated based on the obtained target eye image and its corresponding calibration information, light source position, camera position, and other information using a gaze estimation model. The gaze direction estimated by the above calculation method is more accurate and reliable.

[0010] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description

[0011] The accompanying drawings, which are incorporated in and form a part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure.

[0012] To more clearly illustrate the technical solutions in the embodiments of this disclosure or the prior art, the accompanying drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0013] Figure 1 A schematic flowchart illustrating a line-of-sight estimation method provided in an embodiment of this disclosure;

[0014] Figure 2 A schematic diagram of positional relationships provided for an embodiment of this disclosure;

[0015] Figure 3 A schematic diagram illustrating the training of a gaze estimation model provided in an embodiment of this disclosure;

[0016] Figure 4 This is a schematic diagram of the structure of a line-of-sight estimation device provided in an embodiment of the present disclosure;

[0017] Figure 5This is a schematic diagram of the structure of an electronic device provided in an embodiment of this disclosure. Detailed Implementation

[0018] To better understand the above-mentioned objectives, features, and advantages of this disclosure, the solutions disclosed herein will be further described below. It should be noted that, unless otherwise specified, the embodiments and features described herein can be combined with each other.

[0019] Numerous specific details are set forth in the following description in order to provide a full understanding of this disclosure, but this disclosure may also be implemented in other ways different from those described herein; obviously, the embodiments in the specification are only some, and not all, of the embodiments of this disclosure.

[0020] Figure 1 This is a flowchart illustrating a gaze estimation method provided in an embodiment of the present disclosure. The method can be executed by a gaze estimation device, which can be implemented in software and / or hardware, and is generally integrated into an electronic device. Figure 1 As shown, the method mainly includes the following steps S102 to S106:

[0021] Step S102: Obtain the calibration information corresponding to the target eye image whose line of sight is to be estimated; wherein, the calibration information includes the coordinate information of the pixels of the target eye image in the device coordinate system, and the device coordinate system is determined based on the camera used to acquire the target eye image.

[0022] The target eye image can be the eyes of the target object (such as a user). For example, the target user can wear electronic devices such as VR glasses, smart glasses, eye trackers, or other electronic devices that require estimation of the user's eye gaze direction. The camera in this electronic device can capture images of the target user's eyes as the target eye image. Alternatively, the user can also directly capture the target user's eye image through a camera at a preset position without wearing any device. There are no restrictions on the method of capturing the target eye image. It should be noted that a prompt message can be sent to the user before capturing the eye image, and the user's eye image can only be captured after obtaining the user's authorization.

[0023] In this embodiment, taking into full account the influence of camera parameters on the imaging result of the eye image, this embodiment additionally obtains calibration information. The calibration information includes the coordinate information of the pixels of the target eye image in the camera's device coordinate system. In specific implementation, each pixel in the target eye image can be projected into the 3D space of the device coordinate system to obtain the X-axis and Y-axis coordinates of each pixel in the device coordinate system. The Z-axis coordinate of each pixel in the device coordinate system is set to a preset value, such as a uniform value. This embodiment does not limit the preset value corresponding to the Z-axis coordinate. By obtaining the calibration information in the above way, it is equivalent to projecting the 2D eye image into 3D space, which helps to further estimate the gaze direction corresponding to the eye image in 3D space.

[0024] Step S104: Obtain the camera position corresponding to the camera and the light source position corresponding to the target light source.

[0025] In practical applications, the number of target light sources can be one or more, including but not limited to those using light sources such as LEDs (Light Emitting Diodes). In this embodiment, in addition to configuring a camera, one or more light sources are configured to assist in estimating the eye's gaze direction based on the light spot formed on the eye. In practical applications, both the target light source and the camera can be integrated into electronic devices that require estimating the user's gaze direction, such as VR glasses or eye trackers, for example, integrated into the lens barrel of the electronic device. Alternatively, an external light source can be used directly as the target light source, without limitation.

[0026] Step S106: Based on the target eye image, calibration information, camera position, and light source position, the gaze direction corresponding to the target eye image is estimated using a preset gaze estimation model. Compared to the gaze estimation model trained directly using eye images labeled with gaze directions in related technologies, in this embodiment, the gaze direction can be estimated based on the target eye image, calibration information, camera position, and light source position using a gaze estimation model. Specifically, the model input of the gaze estimation model can be determined based on the target eye image, calibration information, camera position, and light source position, and then the gaze direction corresponding to the target eye image can be determined based on the model output. This embodiment introduces additional calibration information into the input of the gaze estimation model, which helps to improve the accuracy of gaze direction estimation. The main reason is that the inventors have found that the camera parameters of different devices or the same device worn at different times may be different, and the relative position between the camera and the eye may also be different. These can all be regarded as changes in camera state, which may have a certain impact on the imaging results, and thus affect the gaze direction represented by the eye image acquired by the camera. In practice, the eye image samples used to train gaze estimation models are generally insufficient to fully cover eye images captured under different camera conditions. As a result, the trained gaze estimation models struggle to predict the impact of changes in camera conditions on gaze direction. To address this, the inventors added calibration information as input to the gaze estimation model, directly incorporating it into the estimation process. This not only improves the accuracy of gaze estimation but also eliminates the need for costly acquisition of eye images under different camera conditions during model training.

[0027] This disclosure provides an implementation example of estimating the gaze direction corresponding to a target eye image using a preset gaze estimation model. Specifically, the preset gaze estimation model can be used to determine the corneal center position and pupil position of the eye in the target eye image, and the gaze direction corresponding to the target eye image can be estimated based on the corneal center position and pupil position. The gaze direction is the line connecting the corneal center position and the pupil position. In practical applications, the gaze estimation model can output the corneal center position and pupil position, thereby further estimating the gaze direction based on the corneal center position and pupil position; alternatively, the step of estimating the gaze direction based on the corneal center position and pupil position can be built into the gaze estimation model, and the gaze estimation model directly outputs the gaze direction.

[0028] For example, the corneal center position and pupil position obtained by the gaze estimation model described above can be 3D positions, and the corresponding gaze direction can be a 3D direction. Furthermore, 3D can be converted to 2D as needed, without limitation.

[0029] In some implementations, the gaze estimation model may perform the following steps A to B when determining the corneal center position and pupil position of the eye in the target eye image:

[0030] Step A involves merging the target eye image and calibration information to obtain merged information. For example, based on a preset channel dimension, the target eye image and calibration information are merged. This merging process includes, but is not limited to, stitching methods. Specifically, the target eye image and the coordinate information of each pixel in the target eye image (such as the X-axis and Y-axis coordinates in the device coordinate system) can be stitched together along the channel dimension to obtain merged information. This method can also be viewed as converting a two-dimensional image into a three-dimensional image. This three-dimensional image is encoded in coordinate form and represented by the merged information. The merged information presents information about the target eye image itself as well as spatial information mapping the target eye image to three-dimensional space, facilitating subsequent analysis and processing.

[0031] Step B: Based on the merged information, camera position, and light source position, determine the position of the light spot on the eye, the center position of the cornea, and the position of the pupil corresponding to the target light source; wherein, the position of the light spot and the center position of the cornea satisfy the preset relationship constraint conditions, and the pupil position is related to both the position of the light spot and the center position of the cornea.

[0032] The aforementioned preset relational constraints are the solution logic that the gaze estimation model must follow. This ensures that the relationship between the output spot position and the corneal center position conforms to the preset solution logic, thereby effectively guaranteeing the accuracy of the model's output. In some specific implementation examples, the preset relational constraints include: the angle between the target line and the angle bisector of the target angle is less than a preset angle threshold. Here, the target line is the line connecting the spot position and the corneal center position, and the target angle is determined based on the angle between the first and second lines. The first line is the line connecting the spot position and the light source position, and the second line is the line connecting the spot position and the camera position. For easier understanding, please refer to... Figure 2 The diagram illustrates a positional relationship, showing the eyeball, the camera position for capturing images of the eye, and the target light source. It should be noted that... Figure 2 The image only illustrates one target light source; in practical applications, there can be multiple target light sources. Figure 2In this model, the center point of the eyeball is considered as the corneal center C. The target light source M forms a light spot Q on the eyeball. The line connecting the camera position O and the light spot Q is abbreviated as L1, and the line connecting the target light source M and the light spot Q is labeled as L2. Theoretically, the line connecting the corneal center C and the light spot Q (labeled as L3) is the angle bisector of L1 and L2. In other words, the angle between the aforementioned target connection line (i.e., L3) and the angle bisector of the aforementioned target angle (i.e., the angle between L1 and L2) should theoretically be zero. However, considering calculation errors, a small range of fluctuations is allowed. Therefore, a small angle threshold can be set based on the allowed fluctuation range. An angle between the target connection line and the angle bisector of the target angle that is less than the preset angle threshold is considered to satisfy the preset relationship constraint. By setting the above preset relationship constraint, it is equivalent to introducing the solution logic of eye-related information into the gaze estimation process of the gaze estimation model, thereby further reducing the negative impact of camera state changes on the accuracy of gaze estimation results, as mentioned above.

[0033] In some specific implementation examples, step B above can be performed by referring to steps B1 to B3 as follows:

[0034] Step B1, based on the merged information, camera position, and light source position, determines the position of the light spot on the eye corresponding to the target light source and the center position of the cornea. In practical applications, feature extraction can be performed on the merged information, camera position, and light source position, and then the position of the light spot formed by the light source on the eye and the center position of the cornea can be determined based on the corresponding features. For ease of understanding, in some specific implementation examples, step B1 can be performed as follows: Steps B1.1 to B1.4

[0035] Step B1.1 involves extracting features from the merged information to obtain merged features. This embodiment of the invention does not limit the feature extraction algorithm or the feature extraction network used.

[0036] Step B1.2 involves performing a first encoding process on the camera position to obtain camera position features, and a second encoding process on the light source position to obtain light source position features. If there are multiple target light sources, each light source position is encoded separately to obtain multiple light source position features. The dimensions of the camera position features and light source position features obtained through the above encoding method are consistent with the dimensions of the aforementioned merged features.

[0037] Step B1.3: Based on the merged features, camera position features, and light source position features, the light source association features corresponding to the target light source are obtained. The light source association features, obtained by combining the merged features, camera position features, and light source position features, can present information related to the target light source. For example, the light source position features are multiplied with the merged features, and the light source association features are obtained based on the difference between the product result and the camera position features. Assuming there are N target light sources in total, the light source association feature `feature_i` corresponding to the i-th target light source can be expressed as:

[0038] feature_i=Ledi_PE*F-Camera_PE

[0039] Where Ledi_PE represents the light source position feature corresponding to the i-th target light source Led, F represents the merged feature, and Camera_PE represents the camera position feature.

[0040] Step B1.4: Obtain the corneal center features, and based on the corneal center features and the light source association features corresponding to the target light source, determine the position of the light spot on the eye corresponding to the target light source and the position of the corneal center of the eye. The corneal center features can be a preset feature vector, such as using a randomly initialized feature vector to represent the corneal center. The dimension of the corneal center features is consistent with the dimensions of the aforementioned merged features, camera position features, and light source position features.

[0041] To accurately determine the location of the light spot and the corneal center, embodiments of this disclosure can employ a self-attention mechanism based on corneal center features and light source association features corresponding to the target light source to determine the location of the light spot on the eye and the location of the corneal center of the eye. It is understood that the self-attention mechanism can effectively establish associations between different features. Given the corneal center features and light source association features, the self-attention mechanism can perform association analysis on different features, extracting effective information, thereby accurately and reliably determining the corneal center location and the light spot location. In practical applications, there can be multiple target light sources. First, a first self-attention process can be performed on the light source association features corresponding to multiple target light sources to obtain the light spot information corresponding to each target light source. Then, based on the light spot information, a second self-attention process is performed on the corneal center features and the light source association features corresponding to multiple target light sources to obtain the location of the light spot on the eye and the location of the corneal center of the eye. Specifically, the multiple light source association features (assuming N light source association features) can first be mutually treated with each other using a first attention process. For each target light source, by combining its own light source association features with the light source association features of other target light sources for self-attention processing, the spot information can be extracted more effectively. Based on this, corneal center features can be introduced, and combined with the aforementioned N light source association features, a total of N+1 features are subjected to pairwise self-attention processing. This further optimizes the spot information obtained from the first attention processing, determining more reliable N spot positions and a corneal center position.

[0042] Step B2: Based on the merged information, determine the pupil direction of the eye in the device coordinate system. Specifically, the merged features corresponding to the merged information can be obtained, and based on these features, the pupil direction of the eye in the device coordinate system can be predicted. The pupil direction can be a straight line passing through the pupil in three-dimensional space represented by the device coordinate system.

[0043] Step B3: Determine the pupil position based on the location of the light spot, the center of the cornea, and the pupil direction. Once this information is determined, the pupil position can be further calculated logically. In some specific implementation examples, steps B3.1 to B3.2 can be used to determine the pupil position:

[0044] Step B3.1: Determine the corneal radius corresponding to the eye based on the location of the light spot and the location of the corneal center. For example, the corneal radius R can be calculated based on the Eulerian distance between the light spot location and the corneal center location.

[0045] Step B3.2: Determine the pupil position based on the corneal center, corneal radius, and pupil direction. It can be understood that the corneal center and corneal radius can form a sphere; a straight line in the pupil direction intersects this sphere, and the intersection point can be considered the 3D position of the pupil.

[0046] The above-described method provided in this disclosure can be implemented using a gaze estimation model. This disclosure does not limit the structure of the gaze estimation model. To ensure that the gaze estimation model can output reliable results, based on the foregoing, this disclosure provides a training method for the gaze estimation model. That is, the gaze estimation model is obtained through the following steps 1 to 4:

[0047] Step 1: Obtain an eye sample image; wherein, the eye sample image carries the corresponding calibration information, camera position, light source position, and theoretical line of sight.

[0048] Step 2: Based on the eye sample image and the corresponding calibration information, camera position, and light source position, determine the target information corresponding to the eye sample image using the initial estimation model; the target information corresponding to the eye sample image includes the corneal center position and pupil position corresponding to the eye sample image.

[0049] Step 3: Based on the corneal center position and pupil position corresponding to the eye sample image, determine the estimated gaze direction corresponding to the eye sample image, and determine the first loss based on the difference between the estimated gaze direction and the theoretical gaze direction. For example, a preset loss function is used to determine the first loss based on the difference between the estimated gaze direction and the theoretical gaze direction. In practical applications, the estimated gaze direction can be a three-dimensional direction. In this case, the three-dimensional gaze direction can be converted into a two-dimensional gaze direction, which makes it easier to calculate the difference between the gaze directions and determine the first loss.

[0050] Step 4: Train an initial estimation model based on the first loss, and then obtain a gaze estimation model based on the trained initial estimation model. Adjust the parameters of the initial estimation model in the direction of reducing the first loss until the preset training termination condition is met to end the training, and obtain the trained initial estimation model, which is then used as the gaze estimation model.

[0051] To ensure the effectiveness of model training and the accuracy of the gaze estimation model's output, in some implementations, the target information corresponding to the eye sample image also includes the location of the light spot corresponding to the eye sample image; based on this, step 4 above can be performed with reference to steps 4.1 to 4.3 below:

[0052] Step 4.1: Obtain the correlation between the location of the light spot corresponding to the eye sample image and the location of the corneal center. Based on the difference between the correlation and the preset correlation, determine the second loss. In practical applications, if there are multiple target light sources, the corresponding second loss can be calculated for each target light source separately. Subsequently, multiple second losses can be fused using methods such as averaging.

[0053] In some implementations, the preset relationship may include: the line connecting the spot position and the corneal center position is the angle bisector of the target angle, and the target angle is determined based on the angle between a first line and a second line, where the first line connects the spot position and the light source position, and the second line connects the spot position and the camera position. Based on this, the difference between the correlation relationship and the preset relationship can be measured by the angle between the line connecting the spot position and the corneal center position corresponding to the eye sample image output by the initial estimation model and the angle bisector of the target angle; the smaller this angle, the smaller the difference between the correlation relationship and the preset relationship.

[0054] Based on the above implementation method, for ease of calculation, the line connecting the spot position and the corneal center position is theoretically the angle bisector of the target angle, which can be represented by vector operations. Specifically, a first unit vector is determined based on the line connecting the spot position and the light source position; a second unit vector is determined based on the line connecting the spot position and the camera position; the vector obtained by subtracting the second unit vector and the first unit vector can be regarded as the eye tangent vector; a third vector is determined based on the line connecting the spot position and the corneal center position, and the third vector is perpendicular to the eye tangent vector, and the product of the third vector and the eye tangent vector is zero. Based on this, the difference between the correlation relationship and the preset relationship can be measured by the difference between the product of the third vector calculated based on the camera position, the spot position output by the initial estimation model, and the corneal center position, and the eye switching vector, and zero.

[0055] All of the above methods can reasonably yield the second loss, but they are all illustrative examples. In practical applications, other methods can also be used to determine the second loss, which are not limited here.

[0056] Step 4.2 involves weighting the first loss and the second loss to obtain the total loss. The weights of the first loss and the second loss in this embodiment can be flexibly set according to requirements and are not limited herein.

[0057] Step 4.3: Train the initial estimation model based on the total loss. In practical applications, the training of the initial estimation model can be stopped when the total loss converges to within a preset threshold.

[0058] For ease of understanding, embodiments of this disclosure also provide Figure 3The diagram illustrates the training of a gaze estimation model, simply showing a camera and four target light sources (LED1 to LED4). Based on a positional encoding method, corresponding encoding vectors are obtained for each light source: camera position features, first light source position features, second light source position features, third light source position features, and fourth light source position features. Feature extraction is performed on the merged information obtained from the eye image and calibration information to obtain merged features. Based on the camera position features, the first to fourth light source position features, and the merged features, the associated features of the first, second, third, and fourth light sources can be obtained. Figure 3 The diagram simply illustrates the gaze estimation model, which includes a first neural network and a second neural network. Features associated with the first to fourth light sources are input into the first neural network, which outputs the positions of the first, second, third, and fourth light spots, as well as the corneal center position. These merged features are then input into the second neural network, which outputs the pupil direction. Based on the positions of the first to fourth light spots and the corneal center position, a solution loss (the aforementioned second loss) can be obtained; its calculation method is detailed in the previous sections. Similarly, based on the pupil direction and the corneal center position, a gaze direction loss (the aforementioned first loss) can be obtained. Specifically, the pupil position can be determined based on the pupil direction, and the estimated gaze direction can be obtained from the pupil position and the corneal center position. Combining this with the theoretical gaze direction allows the determination of the gaze direction loss. Finally, the model can be trained based on the solution loss and the gaze direction loss.

[0059] This disclosure describes embodiments of... Figure 3 The structures of the first and second neural networks are not limited. For example, the first neural network can be a Transformer network. The aforementioned step of performing a first self-attention processing on the light source association features corresponding to multiple target light sources to obtain the light spot information corresponding to each target light source can be implemented using an encoder in a Transformer network. The aforementioned step of performing a second self-attention processing on the corneal center features and the light source association features corresponding to multiple target light sources based on the light spot information to obtain the light spot position of the target light source on the eye and the corneal center position of the eye can be implemented using a decoder in a Transformer network. It should be noted that the above is only an illustrative example; other network structures can be used in practical applications, and no limitation is imposed here. Furthermore, Figure 3 The diagram does not fully illustrate all the network structures of the gaze estimation model. For example, the network used for position encoding can also be considered part of the gaze estimation model, but... Figure 3 The Chinese side did not indicate this.

[0060] In summary, compared with the gaze estimation model obtained by directly training an eye image labeled with the gaze direction in related technologies, the gaze estimation model obtained in this embodiment can perform gaze estimation by using logical solution, and the obtained gaze estimation result is more accurate and reliable.

[0061] Corresponding to the aforementioned line-of-sight estimation method, this disclosure provides a line-of-sight estimation device. Figure 4 This is a schematic diagram of a line-of-sight estimation device provided in an embodiment of the present disclosure. The device can be implemented by software and / or hardware, and is generally integrated into an electronic device, such as... Figure 4 As shown, the line-of-sight estimation device includes:

[0062] The calibration information acquisition module 402 is used to acquire calibration information corresponding to the target eye image whose line of sight is to be estimated; wherein, the calibration information includes the coordinate information of the pixels of the target eye image in the device coordinate system, and the device coordinate system is determined based on the camera used to acquire the target eye image;

[0063] The position acquisition module 404 is used to acquire the camera position corresponding to the camera and the light source position corresponding to the target light source;

[0064] The gaze estimation module 406 is used to estimate the gaze direction corresponding to the target eye image based on the target eye image, calibration information, camera position, and light source position using a preset gaze estimation model.

[0065] The apparatus provided in this disclosure fully considers the influence of the camera on the imaging result of the eye image. Therefore, calibration information is introduced when estimating the gaze direction of the eye image. Furthermore, information such as the light source position and camera position corresponding to the target eye image can be further obtained. The gaze direction is estimated based on the obtained target eye image and its corresponding calibration information, light source position, camera position, and other information using a gaze estimation model. The gaze direction estimated by the above calculation method is more accurate and reliable.

[0066] In some embodiments, the gaze estimation module 406 is specifically used to: determine the corneal center position and the pupil position of the eye in the target eye image using a preset gaze estimation model, and estimate the gaze direction corresponding to the target eye image based on the corneal center position and the pupil position.

[0067] In some embodiments, the gaze estimation module 406 is specifically used to: merge the target eye image and the calibration information to obtain merged information; based on the merged information, the camera position, and the light source position, determine the position of the light spot on the eye corresponding to the target light source, the position of the corneal center of the eye, and the position of the pupil of the eye; wherein the position of the light spot and the position of the corneal center satisfy a preset relationship constraint condition; and the pupil position is related to both the position of the light spot and the position of the corneal center.

[0068] In some implementations, the preset relationship constraint includes: the angle between the target line and the angle bisector of the target angle is less than a preset angle threshold, wherein the target line is the line connecting the spot position and the corneal center position, and the target angle is determined based on the angle between the first line and the second line, wherein the first line is the line connecting the spot position and the light source position, and the second line is the line connecting the spot position and the camera position.

[0069] In some embodiments, the gaze estimation module 406 is specifically used to: determine the position of the light spot on the eye corresponding to the target light source and the position of the corneal center of the eye based on the merged information, the camera position and the light source position; determine the pupil direction of the eye in the device coordinate system based on the merged information; and determine the pupil position of the eye according to the light spot position, the corneal center position and the pupil direction.

[0070] In some embodiments, the gaze estimation module 406 is specifically used for: extracting features from the merged information to obtain merged features; performing a first encoding process on the camera position to obtain camera position features, and performing a second encoding process on the light source position to obtain light source position features; obtaining light source association features corresponding to the target light source based on the merged features, the camera position features, and the light source position features; acquiring corneal center features, and determining the position of the light spot on the eye corresponding to the target light source and the corneal center position of the eye based on the corneal center features and the light source association features corresponding to the target light source.

[0071] In some embodiments, the gaze estimation module 406 is specifically used to: determine the position of the light spot on the eye corresponding to the target light source and the position of the corneal center of the eye based on the corneal center features and the light source association features corresponding to the target light source using a self-attention mechanism.

[0072] In some embodiments, there are multiple target light sources, and the position acquisition module 404 is specifically used to: perform a first self-attention processing on the light source association features corresponding to the multiple target light sources to obtain light spot information corresponding to each target light source; and perform a second self-attention processing on the corneal center feature and the light source association features corresponding to the multiple target light sources based on the light spot information to obtain the light spot position of the target light source on the eye and the corneal center position of the eye.

[0073] In some embodiments, the gaze estimation module 406 is specifically used to: determine the corneal radius corresponding to the eye based on the position of the light spot and the position of the corneal center; and determine the pupil position of the eye based on the corneal center position, the corneal radius, and the pupil direction.

[0074] In some embodiments, the apparatus further includes a model training module for obtaining the gaze estimation model through the following steps: acquiring the eye sample image; wherein the eye sample image carries calibration information samples, camera position samples, light source position samples, and theoretical gaze direction corresponding to the eye sample image; determining target information corresponding to the eye sample image based on the eye sample image and the calibration information, camera position, and light source position corresponding to the eye sample image using an initial estimation model; the target information corresponding to the eye sample image includes the corneal center position and pupil position corresponding to the eye sample image; determining the estimated gaze direction corresponding to the eye sample image based on the corneal center position and pupil position corresponding to the eye sample image, and determining a first loss based on the difference between the estimated gaze direction and the theoretical gaze direction; training the initial estimation model based on the first loss to obtain the gaze estimation model based on the trained initial estimation model.

[0075] In some implementations, the model training module is specifically used to: obtain the correlation between the spot position and the corneal center position corresponding to the eye sample image; determine a second loss based on the difference between the correlation and a preset relationship; perform weighted processing on the first loss and the second loss to obtain a total loss; and train the initial estimation model based on the total loss.

[0076] The line-of-sight estimation device provided in this disclosure can execute the line-of-sight estimation method provided in any embodiment of this disclosure, and has the corresponding functional modules and beneficial effects of executing the method.

[0077] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working process of the above-described device embodiments can be referred to the corresponding process in the method embodiments, and will not be repeated here.

[0078] This disclosure provides an electronic device, which includes: a storage device storing a computer program thereon; and a processing device for executing the computer program in the storage device to implement the steps of any method of this disclosure.

[0079] The following is for reference. Figure 5 The diagram illustrates a structural schematic of an electronic device 500 suitable for implementing embodiments of the present disclosure. The terminal devices in the embodiments of the present disclosure may include, but are not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. Figure 5 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments disclosed herein.

[0080] like Figure 5 As shown, the electronic device 500 may include a processing unit (e.g., a central processing unit, a graphics processor, etc.) 501, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 502 or a program loaded from a storage device 508 into a random access memory (RAM) 503. The RAM 503 also stores various programs and data required for the operation of the electronic device 500. The processing unit 501, ROM 502, and RAM 503 are interconnected via a bus 504. An input / output (I / O) interface 505 is also connected to the bus 504.

[0081] Typically, the following devices can be connected to I / O interface 505: input devices 506 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 507 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 508 including, for example, magnetic tapes, hard disks, etc.; and communication devices 509. Communication device 509 allows electronic device 500 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 5 An electronic device 500 with various devices is shown; however, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively.

[0082] In particular, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device 509, or installed from a storage device 508, or installed from a ROM 502. When the computer program is executed by the processing device 501, it performs the functions defined in the methods of embodiments of this disclosure.

[0083] In addition to the methods and devices described above, embodiments of this disclosure can also be computer program products, comprising computer program instructions that, when executed by a processor, cause the processor to perform the image processing methods provided in the embodiments of this disclosure. The computer program product can be written in any combination of one or more programming languages ​​to perform the operations of the embodiments of this disclosure. The programming languages ​​include object-oriented programming languages ​​such as Java and C++, as well as conventional procedural programming languages ​​such as C or similar languages. The program code can be executed entirely on a user's computing device, partially on a user's device, as a standalone software package, partially on a user's computing device and partially on a remote computing device, or entirely on a remote computing device or server.

[0084] Furthermore, embodiments of this disclosure may also be computer-readable storage media storing computer program instructions that, when executed by a processor, cause the processor to perform the line-of-sight estimation method provided in embodiments of this disclosure.

[0085] The computer-readable storage medium may be any combination of one or more readable media. A readable medium may be a readable signal medium or a readable storage medium. A readable storage medium may, for example, include, but is not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatuses, or devices, or any combination thereof. More specific examples of readable storage media (a non-exhaustive list) include: electrical connections having one or more wires, portable disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.

[0086] This disclosure also provides a computer program product, including a computer program / instructions that, when executed by a processor, implement the line-of-sight estimation method in this disclosure.

[0087] It is understood that before using the technical solutions disclosed in the various embodiments of this disclosure, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in this disclosure in an appropriate manner in accordance with relevant laws and regulations, and user authorization should be obtained.

[0088] For example, upon receiving a user's active request, a prompt message is sent to the user to explicitly inform them that the requested operation will require the acquisition and use of the user's personal information, such as collecting images of the user's eyes. This allows the user to autonomously choose whether to provide personal information to the software or hardware, such as the electronic device, application, server, or storage medium executing the operation of this disclosure, based on the prompt message. In this embodiment, the user's eye image is only acquired with the user's permission.

[0089] As an optional but non-limiting implementation, in response to a user's active request, sending a prompt message to the user can be done via a pop-up window, where the prompt message can be presented in text format. Furthermore, the pop-up window can also include a selection control allowing the user to choose "agree" or "disagree" to provide personal information to the electronic device.

[0090] It is understood that the above notification and user authorization process are merely illustrative and do not constitute a limitation on the implementation of this disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of this disclosure.

[0091] It should be noted that, in this document, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0092] The above description is merely a specific embodiment of this disclosure, enabling those skilled in the art to understand or implement it. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this disclosure. Therefore, this disclosure is not to be limited to the embodiments described herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A line of sight estimation method, characterized by, The method comprises: obtaining calibration information corresponding to a target eye image of a line of sight to be estimated; wherein the calibration information comprises coordinate information corresponding to pixels of the target eye image in a device coordinate system, and the device coordinate system is determined based on a camera used to collect the target eye image; obtaining a camera position corresponding to the camera and a light source position corresponding to a target light source; estimating a line of sight direction corresponding to the target eye image based on the target eye image, the calibration information, the camera position, and the light source position, and using a preset line of sight estimation model; The line of sight estimation model is obtained by the following steps: obtaining an eye sample image; wherein the eye sample image carries calibration information corresponding to the eye sample image, a camera position, a light source position, and a theoretical line of sight direction; determining target information corresponding to the eye sample image based on the eye sample image and the calibration information corresponding to the eye sample image, the camera position, and the light source position, and using an initial estimation model; the target information corresponding to the eye sample image includes a corneal center position and a pupil position of the eye sample image; determining an estimated line of sight direction corresponding to the eye sample image according to the corneal center position and the pupil position of the eye sample image, and determining a first loss based on the difference between the estimated line of sight direction and the theoretical line of sight direction; training the initial estimation model based on the first loss to obtain a line of sight estimation model based on the trained initial estimation model.

2. The method of claim 1, wherein, The use of the preset line of sight estimation model to estimate the line of sight direction corresponding to the target eye image comprises: determining a corneal center position of an eye in the target eye image and a pupil position of the eye based on the preset line of sight estimation model, and estimating a line of sight direction corresponding to the target eye image based on the corneal center position and the pupil position.

3. The method of claim 2, wherein, The determination of the corneal center position of the eye in the target eye image and the pupil position of the eye comprises: performing merging processing on the target eye image and the calibration information to obtain merged information; determining a light spot position of the target light source on the eye, a corneal center position of the eye, and a pupil position of the eye based on the merged information, the camera position, and the light source position; wherein the light spot position and the corneal center position satisfy a preset relationship constraint condition, and the pupil position is related to the light spot position and the corneal center position.

4. The method of claim 3, wherein, The preset relationship constraint condition comprises: an included angle between a target line and an angle bisector of a target angle is less than a preset angle threshold, wherein the target line is a line between the light spot position and the corneal center position, the target angle is determined based on an included angle between a first line and a second line, the first line is a line between the light spot position and the light source position, and the second line is a line between the light spot position and the camera position.

5. The method of claim 3, wherein, determining, based on the merged information, the camera position, and the light source position, a light spot position on the eye corresponding to the target light source and a corneal center position of the eye, comprises: determining, based on the merged information, the camera position, and the light source position, a light spot position on the eye corresponding to the target light source and a corneal center position of the eye, comprises: determining, based on the merged information, a pupil direction of the eye in the device coordinate system; determining, based on the light spot position, the corneal center position, and the pupil direction, a pupil position of the eye.

6. The method of claim 5, wherein, determining, based on the merged information, the camera position, and the light source position, a light spot position on the eye corresponding to the target light source and a corneal center position of the eye, comprises: performing feature extraction on the merged information to obtain merged features; performing first encoding processing on the camera position to obtain camera position features, and performing second encoding processing on the light source position to obtain light source position features; obtaining light source correlation features corresponding to the target light source according to the merged features, the camera position features, and the light source position features; obtaining a corneal center feature, and determining, based on the corneal center feature and the light source correlation features corresponding to the target light source, a light spot position on the eye corresponding to the target light source and a corneal center position of the eye.

7. The method of claim 6, wherein, determining, based on the corneal center feature and the light source correlation features corresponding to the target light source, the light spot position on the eye corresponding to the target light source and the corneal center position of the eye, comprises: determining, based on the corneal center feature and the light source correlation features corresponding to the target light source, the light spot position on the eye corresponding to the target light source and the corneal center position of the eye using a self-attention mechanism.

8. The method of claim 7, wherein, The number of target light sources is multiple, and determining, based on the corneal center feature and the light source correlation features corresponding to the target light source, the light spot position on the eye corresponding to the target light source and the corneal center position of the eye using a self-attention mechanism, comprises: performing first self-attention processing on the light source correlation features corresponding to multiple target light sources to obtain light spot information corresponding to each target light source; performing second self-attention processing on the corneal center feature and the light source correlation features corresponding to multiple target light sources based on the light spot information to obtain a light spot position on the eye corresponding to the target light source and a corneal center position of the eye.

9. The method of claim 5, wherein, determining, based on the light spot position, the corneal center position, and the pupil direction, a pupil position of the eye, comprises: determining, based on the light spot position and the corneal center position, a corneal radius corresponding to the eye; determining, based on the corneal center position, the corneal radius, and the pupil direction, a pupil position of the eye.

10. The method of claim 1, wherein, The target information corresponding to the eye sample image further includes a light spot position corresponding to the eye sample image; and the step of training the initial estimation model based on the first loss comprises: obtain a correlation between a light spot position corresponding to the eye sample image and a corneal center position, determine a second loss based on a difference between the correlation and a preset correlation; perform weighted processing on the first loss and the second loss to obtain a total loss; train the initial estimation model based on the total loss.

11. A line of sight estimation apparatus characterized by comprising: comprise: a calibration information acquisition module configured to obtain calibration information corresponding to a target eye image of a line of sight to be estimated, wherein the calibration information comprises coordinate information corresponding to pixels of the target eye image in a device coordinate system, and the device coordinate system is determined based on a camera used to collect the target eye image; a position acquisition module configured to obtain a camera position corresponding to the camera and a light source position corresponding to a target light source; a line of sight estimation module configured to estimate a line of sight direction corresponding to the target eye image based on the target eye image, the calibration information, the camera position, and the light source position, using a preset line of sight estimation model; a model training module configured to obtain the line of sight estimation model by the following steps: obtaining an eye sample image, wherein the eye sample image carries calibration information samples corresponding to the eye sample image, camera position samples, light source position samples, and a theoretical line of sight direction; determining target information corresponding to the eye sample image based on the eye sample image and the calibration information, the camera position, and the light source position corresponding to the eye sample image, using an initial estimation model; the target information corresponding to the eye sample image comprises a corneal center position and a pupil position corresponding to the eye sample image; determining an estimated line of sight direction corresponding to the eye sample image according to the corneal center position and the pupil position corresponding to the eye sample image, and determining a first loss based on a difference between the estimated line of sight direction and the theoretical line of sight direction; training the initial estimation model based on the first loss to obtain a line of sight estimation model based on the trained initial estimation model.

12. An electronic device, comprising: The electronic device comprises: a storage device having a computer program stored thereon; a processing device configured to execute the computer program in the storage device to implement the steps of the line of sight estimation method of any one of claims 1-10.

13. A computer-readable storage medium, characterized in that, The storage medium stores a computer program, and the computer program is used to execute the line of sight estimation method of any one of the preceding claims 1-10.

14. A computer program product, characterised in that, The computer program comprises a computer program, and the computer program is used to implement the line of sight estimation method of any one of claims 1-10 when executed by a processor.

Citation Information

Patent Citations

  • Pupil center-corneal reflection (PCCR) based sight line evaluation method in sight line tracking system

    CN102125422A

  • Sight line estimation method and apparatus

    CN107358217A