Method and apparatus for determining gaze information, and eye movement tracking device
The method and device enhance gaze information accuracy by assessing eye image reliability and combining gaze information from both eyes, optimizing camera angles to address non-ideal capture issues in conventional eye tracking.
Patent Information
- Application Number
- JP2025534364
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-12-14
- Filing Date
- 2023-12-13
- Publication Date
- 2025-12-05
AI Technical Summary
Conventional eye tracking technologies face challenges in accurately determining gaze information due to compact device layouts that require cameras to capture eye images at large angles, leading to reduced accuracy and difficulty in applying conventional methods.
A method and device that determine gaze information by acquiring two eye images, assessing their reliability based on target feature points, and using a gaze estimation model to combine gaze information from both eyes, even if one eye's reliability is low, while optimizing the shooting angle of the image acquisition device.
Improves the accuracy and adaptability of gaze information determination by combining model predictions with reliability assessments, addressing non-ideal shooting angles and binocular constraints.
Smart Images

Figure 2025539569000001_ABST
Abstract
Description
[Technical Field]
[0001] [CROSS-REFERENCE TO RELATED APPLICATIONS] This application claims priority to a Chinese patent application submitted to the China Patent Office on December 14, 2022, bearing application number 202211610186.3 and entitled "Method and apparatus for determining gaze information and eye movement tracking device," the entire contents of which are incorporated herein by reference.
[0002] The present application relates to the field of eye tracking technology, and more particularly to a method, apparatus and eye movement tracking device for determining gaze information. [Background technology]
[0003] Eye tracking technology involves illuminating a user's eyes with a set of infrared lights, collecting images of the eyes with a camera, and analyzing the image features in the human eye area to detect key features of the human eye and calculate the gaze direction or gaze point location of the human eye. Current eye tracking technology requires reliable image data to be collected from both the left and right eyes, and then using a gaze estimation algorithm to determine the gaze direction or gaze point location based on the collected image data. However, many eye tracking wearable devices have a compact layout, requiring the camera to capture eye images at a relatively large angle, which causes the gaze direction of the captured human eye to be away from the camera. This makes it difficult to apply conventional eye tracking methods and reduces the accuracy of determining eye gaze information. How to improve the accuracy of determining gaze information is an urgent issue. Summary of the Invention [Problem to be solved by the invention]
[0004] In view of the above problems, the embodiments of the present application propose a method, an apparatus, and an eye movement tracking device for determining gaze information to improve the above problems. [Means for solving the problem]
[0005] According to one aspect of an embodiment of the present application, there is provided a method for determining gaze information for use in an eye movement tracking device, the method including: acquiring two eye images collected by the eye movement tracking device; determining target feature points for each of the two eye images; and determining a reliability corresponding to each of the two eye images based on the target feature points for each of the two eye images; and determining a reliability corresponding to each of the two eye images when the reliability corresponding to a first eye image of the two eye images is smaller than a predetermined reliability and the reliability corresponding to a second eye image of the two eye images is equal to or greater than the predetermined reliability. The method includes determining first gaze information of the second eyeball image based on the target feature points of the second eyeball image and determining the first gaze information as first gaze information of the first eyeball image; inputting the two eyeball images into a gaze estimation model respectively to obtain second gaze information of each of the two eyeball images output by the gaze estimation model; and determining target gaze information corresponding to the two eyeball images based on the first gaze information of the first eyeball image, the first gaze information of the second eyeball image, the second gaze information of the first eyeball image, and the second gaze information of the second eyeball image.
[0006] According to one aspect of an embodiment of the present application, there is provided a gaze information determining device for use in an eye movement tracking device, the device including: an image acquisition module for acquiring two eye images collected by the eye movement tracking device; a target feature point determination module for determining target feature points for each of the two eye images and determining a reliability corresponding to each of the two eye images based on the target feature points for each of the two eye images; and a target feature point determination module for determining a reliability corresponding to each of the two eye images when a reliability corresponding to a first eye image of the two eye images is smaller than a predetermined reliability and a reliability corresponding to a second eye image of the two eye images is equal to or greater than the predetermined reliability. and a second gaze information determination module for inputting the two eyeball images into a gaze estimation model and obtaining second gaze information of each of the two eyeball images output by the gaze estimation model. The eyeball image processing system further includes a target gaze information determination module for determining target gaze information corresponding to the two eyeball images based on the first gaze information of the first eyeball image, the first gaze information of the second eyeball image, the second gaze information of the first eyeball image, and the second gaze information of the second eyeball image.
[0007] According to one aspect of an embodiment of the present application, there is provided an eye movement tracking device, the eye movement tracking device including an image collection device and a device body, the image collection device being installed at a target position of the device body, wherein the target position is a position located on the buccal and / or nasal side of the target subject when the eye movement tracking device is attached to the target subject. [Effects of the Invention]
[0008] In the solution of the present application, target feature points are determined in two previously acquired eye images, and reliability corresponding to each of the two eye images is determined based on the target feature points. If it is determined that the reliability corresponding to one of the two eye images is equal to or greater than a predetermined reliability and the reliability corresponding to the other eye image is lower than the predetermined reliability, first gaze information of the eye image whose corresponding reliability is higher than the predetermined reliability is calculated, and this first gaze information is also used as the first gaze information of the other eye image. Second gaze information is determined by inputting the two eye images into a gaze estimation model, and target gaze information can then be determined based on the first gaze information corresponding to the two eye images and the second gaze information corresponding to the two eye images.
[0009] This application determines whether to combine model prediction methods to determine target gaze information based on the reliability of two eye images, thereby improving the accuracy of target gaze information and the adaptability and robustness of the eye movement tracking device. Furthermore, by optimizing the shooting angle of the image acquisition device of the eye movement tracking device, the problem of non-ideal shooting angles of both cameras can be avoided, providing high-quality image data for gaze information determination. At the same time, validity judgments can be made based on the image data of a single eye, and then combined with model predictions to determine target gaze information, solving the binocular constraint problem in eye movement tracking algorithms.
[0010] It is to be understood that both the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the invention. [Brief explanation of the drawings]
[0011] The drawings herein are incorporated into the specification and constitute a part of this specification, show embodiments that are applicable to the present application, and are used together with the specification to explain the principles of the present application. It is obvious that the drawings in the following description are merely some embodiments of the present application, and those skilled in the art can obtain other drawings based on these drawings without any creative effort. [Figure 1] 1 is a schematic diagram of an eye movement tracking device according to an embodiment of the present application; [Figure 2] 1 is a schematic diagram of an eye movement tracking device according to another embodiment of the present application; [Figure 3] 1 is an exploded view of an eye movement tracking device according to an embodiment of the present application. [Figure 4] 1 is a flowchart of a method for determining gaze information according to an embodiment of the present application; [Figure 5] 1 is a flowchart of specific steps of step 230 according to an embodiment of the present application. [Figure 6] 2 is a schematic diagram of step 240 according to one embodiment of the present application. [Figure 7] 4 is a flowchart of a method for determining gaze information according to another embodiment of the present application. [Figure 8] 3 is a flowchart of specific steps of step 330 according to an embodiment of the present application. [Figure 9] 4 is a flowchart of a method for determining gaze information according to another embodiment of the present application. [Figure 10] 1 is a flowchart of a method for determining gaze information according to a further embodiment of the present application; [Figure 11] 10 is a flowchart of specific steps of step 560 according to an embodiment of the present application. [Figure 12] FIG. 1 is a block diagram of an apparatus for determining gaze information according to an embodiment of the present application; [Figure 13] 1 shows a structural schematic diagram of a computer system suitable for implementing an eye movement tracking device according to an embodiment of the present application. DETAILED DESCRIPTION OF THE INVENTION
[0012] Exemplary embodiments will now be more fully described with reference to the drawings. However, exemplary embodiments may be embodied in various forms and should not be construed as being limited to the examples set forth herein. On the contrary, the provision of these embodiments will make this application more thorough and complete and will fully convey the concept of exemplary embodiments to those skilled in the art.
[0013] It should be noted that the described features, structures, or characteristics may be combined in any suitable manner in one or more embodiments. In the following description, numerous specific details are provided to provide a thorough understanding of the embodiments of the present application. However, those skilled in the art will recognize that the technical solutions of the present application can be practiced without one or more of the specific details, or that other methods, components, devices, steps, etc. can be employed. In other cases, well-known methods, devices, implementations, or operations are not shown or described in detail to avoid obscuring aspects of the present application.
[0014] FIG. 1 is a schematic diagram of an eye movement tracking device according to one embodiment of the present application. As shown in FIG. 1, the eye movement tracking device 100 includes an image collecting device 110 and a device main body 120. The image collecting device 110 is installed at a target position on the device main body 120. Optionally, the target position is a position located on the cheek side of the target subject when the eye movement tracking device 100 is attached to the target subject, i.e., the position shown in FIG. 1. In some other embodiments, the target position may be a position located on the nasal side of the target subject when the eye movement tracking device 100 is attached to the target subject. As shown in FIG. 2, the image collecting device 110 is installed at a target position on the device main body 120, and the target position is a position located on the nasal side of the target subject when the eye movement tracking device 100 is attached to the target subject. Here, the target subject is the subject wearing the eye movement tracking device 100. 3, the eye movement tracking device 100 optionally further includes a lampshade 130, a filter 140, a light source circuit board 150, an optical mechanism 160, and a motherboard 170. The motherboard 170 is located in the device body 120, which has two recess structures, and the lampshade 130, the filter 140, the light source circuit board 150, and the optical mechanism 160 are all installed in the recess structures. The optical mechanism 160 is installed in the innermost layer of the recess structure, the light source circuit board is installed on the optical mechanism 160, the image collecting device 110 is located between the optical mechanism 160 and the light source circuit board 150, and the image collecting device 110 is disposed on the target position, the filter 140 is installed in front of the light source circuit board 150 at the same position as the image collecting device 110, and the lampshade 130 is installed in the outermost layer of the recess structure.
[0015] Optionally, the eye movement tracking device 100 may collect eye images at a target tilt angle when collecting images of the target subject's eye using the image collection device 110. Here, the target tilt angle is the angle between the axial direction of the image collection device 110 (the direction in which the center point of the image collection device and the optical center of the lens of the image collection device 110 are located) and the direction in which the optical axis is located, and the target tilt angle is greater than 50°. In other embodiments, the target tilt angle may be other angles.
[0016] Referring to FIG. 4, FIG. 4 illustrates a method for determining gaze information according to an embodiment of the present application. In a specific embodiment, the method for determining gaze information may be used in the gaze information determining device 600 shown in FIG. 12 and the eye movement tracking device 100 (FIG. 1 or FIG. 2) in which the gaze information determining device 600 is configured. The following describes a specific flow of this embodiment. It should be understood that this method may be performed by an eye movement tracking device, which includes an image collecting device and a device main body, and the image collecting device is installed at a target position of the device main body, where the target position is located at a buccal and / or nasal contact position of the target object when the eye movement tracking device is worn on the target object. The following describes the flow shown in FIG. 4 in detail, and the method for determining gaze information may specifically include the following steps:
[0017] Step 210: Obtain two eye images collected by the eye tracking device.
[0018] In one method, when an infrared light source is illuminated, the pupil area in the collected eye image appears black, which is called the dark pupil effect. At the same time, when infrared light is illuminated, the infrared light source is reflected by the cornea to form a highlighted light spot, also known as a Purkinje image. The dark pupil effect enhances the contrast between the pupil and the light spot, and the pupil center and the light spot position change with changes in the gaze direction of the eye. Therefore, eye images collected under an infrared light source can be used for eye movement tracking. In one method, an infrared light source can be attached to the eye movement tracking device, and the infrared light source can be turned on when collecting eye images to collect eye images under the infrared light source.
[0019] In one method, after obtaining two eye images collected by an eye movement tracking device, noise removal processing is required for the two eye images to improve the accuracy of determining gaze information.
[0020] Step 220: determining target feature points for each of the two eye images; and determining a reliability corresponding to each of the two eye images based on the target feature points for each of the two eye images.
[0021] In one approach, the target feature point is a light spot in the two eye images that satisfies a predetermined condition. Here, satisfying the predetermined condition may mean that the light spots in the two eye images have points that fit an ellipse equation corresponding to the pupil of the spherical image. Alternatively, the target feature point may be determined by an ellipse fitting method. Here, the ellipse fitting algorithm involves finding an ellipse for a given set of sample points and fitting it as close as possible to these sample points. That is, all light spots in the images are fitted with an ellipse equation as a model, and the light spot that contains as many light spots as possible in an ellipse equation, i.e., that satisfies this ellipse equation, is the target feature point.
[0022] As one method, the number of target feature points and the number of all light spots in the two eye images are respectively statistically calculated, and the ratio of the number of target feature points corresponding to each of the two eye images to the number of all light spots is calculated, and this ratio may be used as the reliability of each corresponding eye image.
[0023] Step 230: If the reliability corresponding to a first eyeball image among the two eyeball images is smaller than a predetermined reliability, and the reliability corresponding to a second eyeball image among the two eyeball images is equal to or greater than the predetermined reliability, first gaze information of the second eyeball image is determined based on the target feature point of the second eyeball image, and the first gaze information is determined as the first gaze information of the first eyeball image.
[0024] In one embodiment, the first gaze information may be the gaze direction and gaze point of the corresponding eye image. Optionally, in step 220, an environmental image collection device may be correspondingly installed to collect a scenario image seen by the human eye. After determining the target feature points using an ellipse fitting algorithm, an ellipse equation corresponding to the second eye image may be determined based on the determined target feature points. The ellipse parameter area, ellipse parameters, and position information of the target feature points may be determined. The pupil center may then be determined. A homography mapping method may be used to perform mapping transformation based on the coordinate of the pupil center in the second eye image and the scenario image to determine a mapping relationship between the pupil center of the second eye image and the scenario image. When the homography mapping method is used, there is a one-to-one correspondence between the eye image coordinate system and the scenario image coordinate system. Furthermore, the gaze direction or gaze point of the second eye image may be determined based on the mapping relationship between the pupil center of the second eye image and the scenario image.
[0025] Optionally, the eye tracking device needs to be calibrated when in use to determine the mapping relationship between the pupil center of the eye image and the scenario image. Optionally, the eye tracking device can be calibrated using a nine-point calibration method. Here, the nine-point calibration method involves using a laser transmitter fixed to the eye tracking device to sequentially emit nine laser beams in different directions and project them onto a screen or object in front of the user's eyes, and having the user constantly gaze at the nine light spots that appear in front of them. When the user gazes at each light spot, the correspondence between the coordinates of the pupil center in the eye image and the coordinates of the corresponding light spot in the scenario image is determined, thereby establishing a mapping transformation matrix between the two images, thereby achieving the purpose of calibration.
[0026] In one method, when collecting eyeball images, it is not possible to guarantee that the reliability of each of the two collected eyeball images can be equal to or greater than a predetermined reliability. In current technology, if the reliability of one eyeball image is lower than a predetermined reliability, the eyeball image cannot be used, and gaze information cannot be determined, or the determined gaze information will be relatively large in discrepancy with the actual gaze information of the eyeball. To avoid this situation, if it is determined that the reliability corresponding to one of the two eyeball images is lower than a predetermined reliability, the first gaze information determined by the other eyeball image based on the assumption that the lines of sight of both eyes are parallel may also be determined as the first gaze information of the eyeball image whose reliability is lower than the predetermined reliability.
[0027] In another method, if it is determined that the reliability corresponding to each of the two eyeball images is greater than a preset reliability, the gaze information of the two eyeball images is determined directly based on the mapping relationship between the pupil centers of the two eyeball images and the scenario image, and then a fusion process is performed based on the gaze information of the two eyeball images to further determine target gaze information.
[0028] Alternatively, if it is determined that the reliability of each of the two eye images is lower than a predetermined reliability, the two eye images need to be reacquired. This situation may indicate a problem with the eye movement tracking device, which may require adjustment of the eye movement tracking device or the user may need to adjust the position of the eye movement tracking device.
[0029] One method is to explain why the coordinate values of the pupil center determined by the ellipse algorithm fitting are not located in the central region of the image. This is because the current gaze point or gaze point is located at a relatively extreme position on the plane. In such cases, the determination reliability is generally low. In such cases, the corneal presentation area is incomplete, causing problems such as the pupil center being misaligned, preventing the ellipse fitting algorithm from fitting the complete corneal area and resulting in a large number of outliers. However, this does not mean that a good eye feature point cannot be found. Therefore, the confidence threshold can be proportionally attenuated, and optionally, a good attenuation coefficient can be obtained by using a two-dimensional Gaussian distribution and adjusting the covariance matrix. where the covariance matrix is JPEG2025539569000002.jpg17152, where d is 2, μ is the mean value, X is the set of coordinate values of the pupil center determined by the ellipse fitting algorithm, and x1 to x n are the abscissas corresponding to the multiple pupil centers determined by the ellipse algorithm fitting, and y1 to y n are the ordinates corresponding to the pupil centers determined by the ellipse algorithm fitting.
[0030] Alternatively, different users may have different confidence thresholds, and some users may have significantly more outliers than others. In such cases, using a fixed confidence level would result in a significantly worse user experience for that user due to model prediction. In such cases, the confidence threshold is adjusted by collecting and analyzing the results of the first n pictures, comparing the results determined based on the algorithm with the results determined based on the model, and calculating an average confidence level based on the algorithm.
[0031] In addition, timing information has a guiding significance for adjusting the reliability threshold. Generally, the change between the previous and next frames of the eye image collected by the eye movement tracking device is generally not too large, and if a frame in a stable sequence is very bad (i.e., the difference between the previous and next frames is very large), the collected picture may be incorrect, and in this case the threshold can be reduced or the current frame can be discarded.
[0032] In some embodiments, as shown in FIG. 5, step 230 includes:
[0033] Step 231: determining a pupil center of the second eye image based on the target feature point of the second eye image.
[0034] In one method, the second eye image may be converted into a grayscale image first, an estimated pupil center is randomly selected in the pupil region of the grayscale image of the second eye image, a difference value is calculated based on the grayscale value of the estimated pupil center and the grayscale values corresponding to all light spots in the image, light spots whose difference value is smaller than a difference value threshold are set as pupil contour points, an ellipse fitting algorithm is used based on the determined pupil contour points to determine parameters for setting the estimated pupil center as the ellipse center, and the number of light spots of the second eye image that form a subset of this ellipse is determined, an estimated pupil center is randomly selected in the pupil region, the calculation is repeated, and based on the determined number of light spots of the second eye image that form a subset of the corresponding ellipse, the ellipse that includes the largest number of light spots is set as the ellipse corresponding to the pupil contour, light spots in the subset that corresponds to this ellipse are set as target feature points, and the pupil center of the second eye image is determined based on the target feature points. Here, the second eye image may be an eye image collected by the left eye or the right eye under illumination with an infrared light source.
[0035] In one method, since the light spot illuminated by the infrared light may hit the pupil boundary and cause occlusion, the highlight points in the grayscale image corresponding to the second eye image are first taken as the light spot reflected from the cornea, and the structure of the light spot is modeled using a multivariate Gaussian distribution, and finally a radial interpolation algorithm is used to remove the light spot that hits the pupil boundary.
[0036] Step 232: determining an iris region in the grayscale image of the second eye image based on the pupil center.
[0037] Based on the structural relationship of the human eye, the direction of the line connecting the three-dimensional center of the pupil and the three-dimensional center of the iris is the gaze direction of the human eye, so the gaze direction or the position of the contemplation point of the human eye can be determined based on the two-dimensional center of the pupil and the two-dimensional center of the iris in the eyeball image.
[0038] In one method, an iris region in the grayscale image of the second eye image can be determined by performing iris identification on the grayscale image of the second eye image, and the iris center of the eye of the second eye image can be determined in the iris region. Optionally, because the iris, pupil, and sclera (white of the eye) have different effects depending on the magnitude of the grayscale value in the grayscale image, iris identification can be performed primarily on the grayscale image of the second eye image. Optionally, the iris region in the grayscale image of the second eye image can be identified using a circular difference algorithm.
[0039] Alternatively, a histogram equalization algorithm may be used to non-linearly stretch the grayscale values of each pixel in the grayscale image of the second eye image to enhance the iris region, so that the iris region has a relatively high contrast in the grayscale image of the second eye image.
[0040] Step 233: determining the maximum grayscale value in the grayscale image of the second eye image of the pupil of the second eye image.
[0041] In one method, the maximum grayscale value of the pupil may be determined by comparing the grayscale value of the pupil in the second eye image with the grayscale value in the grayscale image of the second eye image, where the pixel value is represented by one byte (8b) after quantization. For example, when grayscale values having a continuous variation of black, gray, and white are quantized into 256 grayscale levels, the grayscale value ranges from 0 to 255, representing a range of brightness from dark to light, and a color in the corresponding grayscale image from black to white.
[0042] Step 234: determining a reference iris edge point of the second eye image based on the maximum grayscale value and the iris region.
[0043] In one approach, it may be determined that since the pupil is within the iris region, the grayscale value corresponding to the iris region should be between the maximum grayscale value of the pupil and 255. max If , the grayscale value of the iris region is (T max 255), and then calculates an intermediate value of the grayscale values, sets this intermediate value as a starting threshold, performs iterative calculations based on the starting threshold to determine a target threshold for iris segmentation, performs iris segmentation on the grayscale image of the second eyeball image based on the target threshold, and finally identifies the eye image after iris segmentation to determine a reference iris edge point of the second eyeball image.
[0044] where: A starting threshold may be determined based on JPEG2025539569000003.jpg1886. Then, based on the starting threshold, a first average value of pixel points in the grayscale image whose grayscale values are greater than the starting threshold and a second average value of pixel points whose grayscale values are less than the starting threshold are determined. A third average value is determined based on the second average value and the first average value, which is the average of the first and second average values. Finally, a difference value between the third average value and the starting threshold is calculated. If the difference value is not zero, the third average value is set as a new starting threshold. The previous steps may be repeated until the difference value becomes zero or the number of iterations reaches a threshold value, thereby obtaining a threshold value for iris segmentation. The iris may then be segmented based on this threshold value, and reference iris edge points in the second eye image may be determined. Note that the greater the number of iterations, the greater the impact on reliability. Therefore, if the number of iterations is greater than a predetermined number, a reliability coefficient must be generated and multiplied by this reliability coefficient when calculating the reliability.
[0045] Optionally, after segmenting the iris of the second eyeball image, a light spot on the iris boundary is determined based on the light spot and the segmented iris area, and this light spot is set as a reference iris edge point.
[0046] Step 235: Determine the iris center of the second eye image based on the reference iris edge points using an ellipse fitting method.
[0047] In one approach, the method of step 231 is used to determine the ellipsoid parameters of the iris boundary, where the ellipse center in the ellipse parameters is the iris center.
[0048] Step 236: Determine the first gaze information of the second eye image based on the pupil center and the iris center, and determine the first gaze information as first gaze information of the first eye image.
[0049] In one method, the two-dimensional coordinates of the pupil center and the iris center in the second eye image are mapped to three-dimensional coordinates, and the first gaze information of the second eye image is determined by connecting the pupil center and the iris center in the three-dimensional coordinate system.
[0050] In one method, after determining the first gaze information, the relationship between the pupil center and cornea center in the second eye image and the first gaze information is learned, and then the eye corresponding to the second eye image is continuously tracked according to this relationship, thereby realizing eye movement tracking.
[0051] Still referring to FIG. 4, in step 240, the two eye images are input into a gaze estimation model respectively, and second gaze information of the two eye images is output by the gaze estimation model respectively.
[0052] In one embodiment, the gaze estimation model is used to determine the gaze of an eyeball corresponding to the eye image based on eye feature information in the eye image, optionally including, but not limited to, the corneal center, pupil center, light spot, etc. of the eyeball corresponding to the eye image.
[0053] Optionally, the gaze estimation model may include a feature extraction layer and a gaze estimation layer, where the feature extraction layer is used to extract feature information from two eye images, and the gaze estimation layer performs discrimination and calculation based on the feature information extracted by the feature extraction layer to obtain second gaze information corresponding to each of the two eye images. Optionally, the gaze estimation layer may be a deconvolution layer. As shown in FIG. 6, the gaze estimation model may include a feature extraction layer a and a deconvolution layer b, where the eye images are used as input information for the gaze estimation model, where the feature extraction layer a first extracts feature information from the eye images, and then the deconvolution layer b determines a two-dimensional corneal center, a two-dimensional light spot, and a two-dimensional pupil center based on the extracted feature information. The two-dimensional corneal center, the two-dimensional light spot, and the two-dimensional pupil center are then three-dimensionally mapped to obtain a three-dimensional corneal center and a three-dimensional pupil center. The optical axis of the eye is determined based on the three-dimensional corneal center and the three-dimensional pupil center, and the second gaze information is determined based on the optical axis.
[0054] Optionally, when the deconvolution layer performs identification and calculation based on the feature information extracted by the feature extraction layer, the deconvolution layer identifies the two-dimensional coordinate information of the feature based on the feature information previously extracted by the feature extraction layer, and converts the two-dimensional coordinate information of the feature into three-dimensional coordinate information. Since the determined features include the pupil center and the corneal center, etc., the optical axis information of the eyeball image can be determined. Furthermore, since a Kappa angle exists between the optical axis and the visual axis, the visual axis information of the eyeball image can be determined based on this Kappa angle, and second gaze information of the eyeball image can be further obtained.
[0055] In another method, before determining the second gaze information, the first eye image and the first gaze information of the first eye image, and the second eye image and the first gaze information of the second eye image can be used as training data for the gaze estimation model, and the gaze estimation model can be trained in real time using the training data to adjust the model parameters in the gaze estimation model, so that the gaze estimation model can be applied to different target objects, and gaze estimation can be performed by the gaze estimation model for the eyes of different target objects.
[0056] In another method, after determining the second gaze information, the gaze estimation model can be trained based on the second gaze information and the corresponding eye images to learn the eye characteristics of the user of the corresponding eye movement tracking device, which can facilitate subsequent continuous eye tracking and improve the rate at which the estimation model can determine the second gaze information.
[0057] Step 250: Determine target gaze information corresponding to the two eye images based on the first gaze information of the first eye image, the first gaze information of the second eye image, the second gaze information of the first eye image, and the second gaze information of the second eye image.
[0058] In one method, the first reference gaze information of the eyes corresponding to the two eye images can be first determined by fusing the first gaze information of the first eye image and the first gaze information of the second eye image, or alternatively, the first reference gaze information of the eyes corresponding to the two eye images can be determined by calculating an average value of the first gaze information of the first eye image and the first gaze information of the second eye image. It can be understood that a similar method can be used to fusing the second reference gaze information of the eyes corresponding to the two eye images, and finally, the target gaze information can be determined based on the first reference gaze information and the second reference gaze information.
[0059] In another method, from a physiological viewpoint, everyone has a dominant eye, which may be the left eye or the right eye. The brain preferentially receives what is seen through the dominant eye. That is, there is a certain disparity between the position information of an object observed through the dominant eye and the position information of an object observed through the secondary eye, and the position information of an object observed through the dominant eye is closer to the actual position information of the object. To reduce the increase in disparity caused by the defects of the secondary eye, a weighting factor corresponding to each eye image is obtained when determining the first reference gaze information and the second reference gaze information. Here, the weighting factor corresponding to each eye image is preset by the user according to his or her dominant eye and secondary eye. Optionally, the weighting factor of the dominant eye is greater than the weighting factor of the secondary eye.
[0060] In an embodiment of the present application, target feature points are determined in two previously acquired eyeball images, and reliability corresponding to each of the two eyeball images is determined based on the respective target feature points. If it is determined that the reliability corresponding to one of the two eyeball images is equal to or greater than a predetermined reliability and the reliability corresponding to the other eyeball image is lower than the predetermined reliability, first gaze information of the eyeball image whose corresponding reliability is higher than the predetermined reliability is calculated, and this first gaze information is also used as the first gaze information of the other eyeball image. Second gaze information is determined by inputting the two eyeball images into a gaze estimation model, and target gaze information can be determined based on the first gaze information corresponding to the two eyeball images and the second gaze information corresponding to the two eyeball images.
[0061] This application determines whether to combine model prediction methods to determine target gaze information based on the reliability of two eye images, thereby improving the accuracy of target gaze information and the adaptability and robustness of the eye movement tracking device. Furthermore, by optimizing the shooting angle of the image acquisition device of the eye movement tracking device, the problem of non-ideal shooting angles of both cameras can be avoided, providing high-quality image data for gaze information determination. At the same time, validity judgments can be made based on the image data of a single eye, and then combined with model predictions to determine target gaze information, solving the binocular constraint problem in eye movement tracking algorithms.
[0062] Referring to Figure 7, Figure 7 shows a method for determining gaze information according to an embodiment of the present application. The following describes the flow shown in Figure 7 in detail, and the method for determining gaze information may specifically include the following steps:
[0063] Step 310: Obtain two eye images collected by the eye movement tracking device.
[0064] Step 320: grayscale processing is performed on the two eyeball images respectively to obtain grayscale images corresponding to the two eyeball images.
[0065] One way is that in an RGB image, when R=G=B, the color represents a grayscale color, where the R=G=B value is called the grayscale value. Therefore, each pixel in a grayscale image only requires one byte to store the grayscale value (also called intensity value or brightness value). The grayscale range is 0 to 255, with 255 representing the brightest (pure white) and 0 representing the darkest (pure black). To obtain a grayscale image of the two eye images, the image grayscale processing may be performed using the maximum value method, the average value method, or the weighted average method. Optionally, the maximum value method directly takes the largest value among the three R, B, and G components (0 being the minimum value and 255 being the maximum value), and uses this value as the value of the other components, as expressed as R=G=B=max(R, G, B). The average value method calculates the average value of the three R, B, and G components, and uses this average value as the value of the other components, as expressed as R=G=B=(R+G+B) / 3. The weighted average method weights the three components according to their importance and other indicators. For example, since the human eye is most sensitive to green and least sensitive to blue, the weighted average of the three RGB components is calculated as GRAY=R*0.299+G*0.587+B*0.114, which can obtain a more reasonable grayscale image.
[0066] Step 330: determine light spots in grayscale images corresponding to each of the two eye images; and determine target feature points from the light spots included in the grayscale images corresponding to each of the two eye images, where the target feature points are spots among the light spots that meet a condition.
[0067] In one method, in the grayscale image, when infrared light is irradiated, the cornea reflects and forms a highlight light spot. When grayscaling is performed, the brightness of the light spot is greater than the brightness of other areas, i.e., the grayscale value of the position where the light spot is located is greater than the grayscale values of other pixel points and is close to 255. Therefore, the point with the grayscale value close to 255 can be determined to further determine the light spot in the two eyeball images.
[0068] As one method, for a grayscale image corresponding to each eye image, an estimated pupil center may be randomly selected in the pupil region of the grayscale image, and a difference value may be calculated based on the grayscale value of this estimated pupil center and the grayscale values corresponding to all light spots in the corresponding grayscale image, and the light spots whose difference value is smaller than a difference value threshold may be set as pupil contour points. An ellipse fitting algorithm may be used based on the determined pupil contour points to determine parameters for an ellipse center that is the estimated pupil center, and the number of light spots corresponding to the two eye images that form a subset of this ellipse may be determined. An estimated pupil center may be randomly selected in the pupil region, and this calculation may be repeated. Based on the determined number of light spots corresponding to the corresponding eye image that form a subset of the corresponding ellipse, the ellipse that includes the largest number of light spots may be set as the ellipse corresponding to the pupil contour, and the light spots in the subset corresponding to this ellipse may be set as target feature points.
[0069] In some embodiments, for the grayscale images corresponding to each of the two eye images, as shown in FIG. 8, step 330 includes:
[0070] Step 331: determine a reference pupil center in the grayscale image.
[0071] In one method, a point in the area where the pupil is located may be randomly selected as the reference pupil center. Alternatively, after determining the pupil contour reference point at the current reference pupil center, a reference pupil center at another position may be randomly selected and repeated multiple times to determine as many contour reference points as possible, thereby increasing the accuracy of the determined target feature point.
[0072] Step 332: Calculate grayscale difference values between the reference pupil center and each of the light spots in the grayscale image.
[0073] In one method, after grayscaling processing is performed on the eyeball image, each pixel point has its own grayscale value, and an elliptical light spot corresponding to the pupil boundary with each reference pupil center as a circular point may be further determined by calculating the difference value between the grayscale value of each light spot and the grayscale value of the reference pupil center.
[0074] Step 333: determine the light spot whose gray scale difference value is greater than the difference value threshold as the contour reference point.
[0075] In one method, if the grayscale difference value between the reference pupil center and any one of the light spots is greater than a difference threshold, it may be determined that the light spot satisfies the requirement for the pupil contour reference point, where the difference threshold may be set according to actual needs and is not specifically limited herein.
[0076] Step 334: Determine the target feature points from the contour reference points using an ellipse fitting method.
[0077] In one method, since not all contour reference points are necessarily pupil contour points, in order to determine the point in the eyeball image that is closest to the pupil boundary contour, an ellipse fitting algorithm is used to determine the point among the contour reference points that is closest to the pupil boundary contour, and the determined point that is closest to the pupil boundary contour may be called the target feature point. Here, a random sampling consensus algorithm (RANSAC) is used to fit the ellipse equation that determines the pupil boundary multiple times, and the number of contour reference points on each ellipse equation is counted, and the ellipse equation with the largest number is determined as the ellipse equation for the pupil boundary, and the contour reference point on this ellipse equation is used as the target feature point.
[0078] Still referring to FIG. 7 , in step 340, a confidence level corresponding to each of the two eye images is determined based on the number of the light spots and the number of the target feature points included in the grayscale images corresponding to each of the two eye images.
[0079] In one method, the ratio of the target feature points to the light spots in two eye images may be calculated, and the ratio may be used as the reliability of the corresponding eye image. For example, if the number of target feature points in a first eye image is n and the number of light spots is m, the reliability corresponding to the first eye image is n / m.
[0080] Step 350: if the reliability corresponding to a first eyeball image among the two eyeball images is smaller than a predetermined reliability and the reliability corresponding to a second eyeball image among the two eyeball images is equal to or greater than the predetermined reliability, determine first gaze information of the second eyeball image based on the target feature point of the second eyeball image, and determine the first gaze information as the first gaze information of the first eyeball image.
[0081] Step 360: inputting the two eye images into a gaze estimation model respectively, and obtaining second gaze information for each of the two eye images output by the gaze estimation model.
[0082] Step 370: Determine target gaze information corresponding to the two eye images based on the first gaze information of the first eye image, the first gaze information of the second eye image, the second gaze information of the first eye image, and the second gaze information of the second eye image.
[0083] Here, for a detailed description of step 310 and steps 350 to 370, please refer to step 210 and steps 230 to 250, and no further description will be given here.
[0084] In this embodiment, grayscale images of the two eyeball images are obtained by performing grayscale conversion processing on the two eyeball images, and then a boundary contour feature point of the pupil is determined at the light spot based on the difference value between the grayscale value of the light spot and the grayscale value of the pupil center in the grayscale image. An ellipse fitting algorithm is then performed based on the contour feature point to determine a target feature point that satisfies the conditions, thereby improving the accuracy of determining the target feature point.
[0085] Referring to Figure 9, Figure 9 illustrates a method for determining gaze information according to an embodiment of the present application. The following describes the flow shown in Figure 9 in detail, where the gaze estimation model includes a feature extraction layer and a gaze estimation layer, and the method for determining gaze information may specifically include the following steps:
[0086] Step 410: Obtain two eye images collected by the eye movement tracking device.
[0087] Step 420: determine target feature points for each of the two eye images; and determine a reliability corresponding to each of the two eye images based on the target feature points for each of the two eye images.
[0088] Step 430: if the reliability corresponding to a first eye image among the two eye images is smaller than a predetermined reliability and the reliability corresponding to a second eye image among the two eye images is equal to or greater than the predetermined reliability, determine first gaze information of the second eye image based on the target feature point of the second eye image, and determine the first gaze information as the first gaze information of the first eye image.
[0089] Here, for a specific description of steps 410 to 430, please refer to steps 210 to 230, and no further description will be given here.
[0090] Step 440: inputting the two eye images into the feature extraction layer respectively to obtain eye feature information corresponding to the two eye images respectively.
[0091] In one embodiment, the eye feature information includes feature information of the cornea center of the eye image, feature information of the pupil center of the eye image, feature information of the light spot in the eye image, etc. Optionally, the eye feature information extracted by the feature extraction layer may be two-dimensional feature information in the eye image.
[0092] In one embodiment, the feature extraction layer may include a convolutional neural network, which may optionally be a neural network configured with RepVGG operators. Optionally, the feature extraction layer may include multiple neural network blocks. For example, the feature extraction layer may include five neural network blocks, each with a neural network configuration as shown in the table below.
[0093] Table 1 Neural network configuration for each block JPEG2025539569000004.jpg57167
[0094] Here, both the 3x3 convolution and the Relu layer can be standard network layers. Using the network structure in Table 1 as the feature extraction layer has the following advantages: 1. The convolutional network itself is constructed using a combination of 3x3 convolution and Relu activation layer operators, which effectively increases computational density; 2. The single-channel network architecture has high parallelism and saves video memory under the same computational load; 3. The single-channel architecture has good flexibility, allowing the network width to be changed using quantization means. It is compatible with most chips and is friendly to edge devices.
[0095] In step 450, a corneal center corresponding to each of the two eye images and a pupil center corresponding to each of the two eye images are determined based on the eye feature information corresponding to each of the two eye images.
[0096] In one method, a deconvolution layer may be set in the gaze estimation model, and this deconvolution layer may be used as the gaze estimation layer, so that the corneal centers corresponding to each of the two eyeball images and the pupil centers corresponding to each of the two eyeball images can be determined by this deconvolution layer.
[0097] Optionally, deconvolution is a special type of convolution, which generally expands the size of the input image by padding zeros into the image matrix that has been previously scaled by a certain ratio. Optionally, the network structure of this deconvolution layer is as shown in Table 2.
[0098] Table 2. Network structure of the deconvolution layer JPEG2025539569000005.jpg60168
[0099] Here, FC is a fully connected network, and its deconvolution layer has three branched neural network structures, which are used to determine the two-dimensional corneal center of the eyeball image based on corneal feature information, determine the two-dimensional light spot of the eyeball image based on light spot (reflection point) feature information, and determine the two-dimensional pupil center of the eyeball image based on pupil feature information. Here, the branched neural network structure for determining the two-dimensional corneal center has a two-layer neural network, including a 3*3 deconvolution network and a fully connected neural network. The branched neural network structure for determining the two-dimensional light spot has a three-layer neural network, including two 3*3 deconvolution networks with different output sizes and one fully connected neural network, where the number of two-dimensional light spots output by the fully connected neural network is six. The branched neural network structure for determining the two-dimensional pupil center has a three-layer neural network, including two 3*3 deconvolution networks with different output sizes and one fully connected neural network.
[0100] Step 460: inputting the corneal centers corresponding to the two eyeball images and the pupil centers corresponding to the two eyeball images into the gaze estimation layer, respectively, to obtain second gaze information for each of the two eyeball images.
[0101] In one embodiment, the gaze estimation layer may determine the two-dimensional corneal center and the two-dimensional light spot previously determined in step 450 as the three-dimensional corneal center, convert the two-dimensional pupil center to the three-dimensional pupil center, and determine the optical axis of the eyeball based on the three-dimensional corneal center and the three-dimensional pupil center. Then, based on the kappa angle between the optical axis and the visual axis, determine the visual axis according to the optical axis. The gaze estimation layer may then estimate gaze information based on the visual axis to obtain second gaze information for each of the two eyeball images. Alternatively, the kappa angle is generally between 0 and 5 degrees (circumferential degrees), and different people have different kappa angles. Alternatively, the kappa angle may be preset by the user when using the eye movement tracking device. Alternatively, the eye movement tracking device may be constructed using a fully connected network, which may include multiple fully connected network layers, or other neural networks, but these are not limited thereto.
[0102] Step 470: Determine target gaze information corresponding to the two eye images based on the first gaze information of the first eye image, the first gaze information of the second eye image, the second gaze information of the first eye image, and the second gaze information of the second eye image.
[0103] Here, for the specific step description of step 470, please refer to step 250, and no further description will be given here.
[0104] In this embodiment, if the reliability of one of the two eyeball images is lower than a reliability threshold, it can be combined with a gaze estimation model to determine second gaze information of the two eyeball images, and further determine target gaze information based on the first gaze information and the second gaze information. The gaze estimation model includes a feature extraction layer and a gaze estimation layer, where the feature extraction layer can extract eye feature information of the two eyeball images, and the gaze estimation layer can determine the second gaze information based on the eye feature information of the two eyeball images extracted by the feature extraction layer, thereby improving the efficiency of determining the second gaze information and saving the calculation resources of the eye movement tracking device.
[0105] Referring to Figure 10, Figure 10 illustrates a method for determining gaze information according to an embodiment of the present application. The following describes the flow shown in Figure 10 in detail, and the method for determining gaze information may specifically include the following steps:
[0106] Step 510: Obtain two eye images collected by the eye movement tracking device.
[0107] Step 520: determining target feature points for each of the two eye images; and determining a reliability corresponding to each of the two eye images based on the target feature points for each of the two eye images.
[0108] Step 530: if the reliability corresponding to a first eyeball image among the two eyeball images is smaller than a predetermined reliability and the reliability corresponding to a second eyeball image among the two eyeball images is equal to or greater than the predetermined reliability, determine first gaze information of the second eyeball image based on the target feature point of the second eyeball image, and determine the first gaze information as the first gaze information of the first eyeball image.
[0109] Step 540: inputting the two eye images into a gaze estimation model respectively, and obtaining second gaze information for each of the two eye images output by the gaze estimation model.
[0110] Here, for a specific description of steps 510 to 540, please refer to steps 210 to 240, and no further description will be given here.
[0111] Step 550: Determine first reference gaze information based on the first gaze information of the first eye image and the first gaze information of the second eye image, and determine second reference gaze information based on the second gaze information of the first eye image and the second gaze information of the second eye image.
[0112] The first reference gaze information is gaze information obtained by fusing the first gaze information of the first eyeball image with the first gaze information of the second eyeball image, and similarly, the second reference gaze information is gaze information obtained by fusing the second gaze information of the first eyeball image with the second gaze information of the second eyeball image.
[0113] In one method, a weight of the influence of the eyeball corresponding to the first eyeball image on the gaze information and a weight of the influence of the eyeball corresponding to the second eyeball image on the gaze information may be preset, and a weighted calculation may be performed on the first gaze information of the first eyeball image and the first gaze information of the second eyeball image based on the weights to obtain first reference gaze information. Similarly, a weighted calculation may be performed on the second gaze information of the first eyeball image and the second gaze information of the second eyeball image based on the weights corresponding to each eyeball to obtain second reference gaze information. Optionally, different weights may be set for the calculation of the first reference gaze information and the second reference gaze information, or the same weight may be set. Optionally, the weighted calculation may be a weighted average calculation.
[0114] Step 560: Obtain a first weight of the first reference gaze information and a second weight of the second reference gaze information.
[0115] In one embodiment, the first weight and the second weight are weights for determining target gaze information. Because the accuracy of the second reference gaze information corresponding to the first reference gaze information determined by the algorithm and the second reference gaze information determined by the model is different, they need to be combined to determine the target gaze information, so a first weight for the first reference gaze information and a second weight for the second reference gaze information need to be obtained.
[0116] In some embodiments, as shown in FIG. 11, step 560 includes:
[0117] Step 561: determining a method for determining the target gaze information;
[0118] As a method, determining the target gaze information may refer to two different methods. One method is to determine the target gaze information by combining gaze information determined by an algorithm (rule) and gaze information determined by a model. Another method is to use first reference gaze information determined based on an algorithm as the target gaze information when a single output is selected. Optionally, when a single output is selected, second reference gaze information may be used as the target gaze information.
[0119] Step 562: if the determining manner is the first manner, determine the first reference gaze information as target gaze information;
[0120] In one mode, the determination mode is set by the user when using the eye movement tracking device. Optionally, a selection switch may be set on the eye movement tracking device, and the user can use the selection switch to determine the determination mode of the target gaze information. When the user selects the first mode, flag information of the first mode is generated correspondingly, and when determining the target gaze information, the first reference gaze information is used as the target gaze information based on the flag information.
[0121] Step 563: if the determination manner is the second manner, obtain a first weight of the first reference gaze information and a second weight of the second reference gaze information;
[0122] In one method, when a user can select the second method as the method for determining target gaze information by using a selection switch on the eye movement tracking device, flag information for the second method is generated accordingly, and when determining target gaze information, it is determined based on this flag information that the target gaze information needs to be determined by combining the gaze information determined based on the algorithm and the model, so that a first weight for the first reference gaze information and a second gaze weight for the second gaze information can be obtained based on this flag information.
[0123] Continuing to refer to FIG. 10, in step 570, a weighting calculation is performed based on the first weight, the first reference gaze information, the second weight, and the second reference gaze information to determine target gaze information corresponding to the two eye images.
[0124] In one mode, when the determination mode is the second determination mode, the target gaze information is determined by performing a weighting calculation on the first reference gaze information and the second reference gaze information based on the first weight and the second weight. Optionally, the first weight and the second weight may be preset different weights or the same weight. Optionally, the first weight and the second weight may be normalized to obtain a dynamic weight.
[0125] In this embodiment, first reference gaze information is determined based on the first gaze information of the first eye image and the first gaze information of the second eye image that have been determined previously, and second reference gaze information is determined based on the second gaze information of the first eye image and the second gaze information of the second eye image that have been determined, and then, depending on the method of determining target gaze information, it is determined whether to actually use the first reference gaze information as target gaze information, or to perform weighting calculation based on the first reference gaze information and the second reference gaze information to determine target gaze information, thereby improving the user's experience and improving the efficiency of determining target gaze information when the user considers the computing resources of the eye movement tracking device.
[0126] The following introduces an embodiment of the device of the present application, which can be used to perform the method in the above embodiment of the present application. For details not disclosed in the device embodiment of the present application, please refer to the above embodiment of the method of the present application.
[0127] FIG. 12 is a block diagram of an apparatus for determining gaze information according to an embodiment of the present application, which is used in an eye movement tracking device. As shown in FIG. 12 , the apparatus for determining gaze information 600 includes: an image acquisition module 610, a target feature point determination module 620, a first gaze information determination module 630, a second gaze information determination module 640, and a target gaze information determination module 650.
[0128] The image acquisition module is used to acquire two eye images collected by the eye movement tracking device, the target feature point determination module is used to determine target feature points of each of the two eye images and determine a reliability corresponding to each of the two eye images based on the target feature points of each of the two eye images, and the first gaze information determination module is used to determine the target feature points of the second eye image when the reliability corresponding to a first eye image of the two eye images is smaller than a predetermined reliability and the reliability corresponding to a second eye image of the two eye images is equal to or greater than the predetermined reliability. the second gaze information determination module is used to input the two eyeball images into a gaze estimation model respectively and obtain second gaze information for each of the two eyeball images output by the gaze estimation model; and the target gaze information determination module is used to determine target gaze information corresponding to the two eyeball images based on the first gaze information of the first eyeball image, the first gaze information of the second eyeball image, the second gaze information of the first eyeball image, and the second gaze information of the second eyeball image.
[0129] In some embodiments, the target feature point determination module includes: a grayscale processing submodule for performing grayscale processing on each of the two eyeball images to obtain grayscale images corresponding to each of the two eyeball images; a target feature point determination submodule for determining light spots in the grayscale images corresponding to each of the two eyeball images and determining target feature points from the light spots included in the grayscale images corresponding to each of the two eyeball images, where the target feature points are spots among the light spots that meet a condition; and a reliability determination submodule for determining reliability corresponding to each of the two eyeball images based on the number of the light spots and the number of the target feature points included in the grayscale images corresponding to each of the two eyeball images.
[0130] In some embodiments, for a grayscale image corresponding to each of the two eye images, the target feature point determination sub-module includes: a reference pupil center determination unit for determining a reference pupil center in the grayscale image; a grayscale difference value calculation unit for calculating a grayscale difference value between the reference pupil center and each of the light spots in the grayscale image; a contour reference point determination unit for determining a light spot whose grayscale difference value is greater than a difference value threshold as a contour reference point; and a target feature point determination unit for determining the target feature point from the contour reference point using an ellipse fitting method.
[0131] In some embodiments, the first gaze information determination module includes a pupil center determination submodule for determining a pupil center of the second eye image based on the target feature point of the second eye image; an iris region determination submodule for determining an iris region in a grayscale image of the second eye image based on the pupil center; a maximum grayscale value determination submodule for determining a maximum grayscale value in the grayscale image of the second eye image of the pupil of the second eye image; a reference iris edge point determination submodule for determining a reference iris edge point of the second eye image based on the maximum grayscale value and the iris region; an iris center determination submodule for determining an iris center of the second eye image based on the reference iris edge point using an ellipse fitting method; and a first gaze information determination submodule for determining the first gaze information of the second eye image based on the pupil center and the iris center and determining the first gaze information as first gaze information of the first eye image.
[0132] In some embodiments, the second gaze information determination module includes an eye feature information extraction submodule for inputting each of the two eye images into the feature extraction layer and obtaining eye feature information corresponding to each of the two eye images; a determination submodule for determining a corneal center corresponding to each of the two eye images and a pupil center corresponding to each of the two eye images based on the eye feature information corresponding to each of the two eye images; and a second gaze information determination submodule for inputting the corneal center corresponding to each of the two eye images and the pupil center corresponding to each of the two eye images into the gaze estimation layer and obtaining second gaze information for each of the two eye images.
[0133] In some embodiments, the target gaze information determination module includes a reference gaze information determination submodule for determining first reference gaze information based on first gaze information of the first eye image and first gaze information of the second eye image, and for determining second reference gaze information based on second gaze information of the first eye image and second gaze information of the second eye image; a weight acquisition submodule for acquiring a first weight of the first reference gaze information and a second weight of the second reference gaze information; and a target gaze information determination submodule for performing weighting calculations based on the first weight, first reference gaze information, the second weight, and the second reference gaze information, and determining target gaze information corresponding to the two eye images.
[0134] In some embodiments, the weight acquisition submodule includes: a determination method determination unit for determining a determination method of the target gaze information; a target gaze information determination unit for determining the first reference gaze information as target gaze information when the determination method is a first method; and a weight acquisition unit for obtaining a first weight of the first reference gaze information and a second weight of the second reference gaze information when the determination method is a second method.
[0135] According to one aspect of the embodiments of the present application, there is further provided an eye movement tracking device, as shown in FIG. 13, the eye movement tracking device 100 includes a processor 180 and one or more memories 190, the one or more memories 190 are used to store program instructions to be executed by the processor 180, and the processor 180 performs the above-mentioned gaze information determining method when executing the program instructions.
[0136] Furthermore, processor 180 may include one or more processing cores. Processor 180 runs or executes instructions, programs, code sets, or instruction sets stored in memory 190, as well as data stored in memory 190. Alternatively, processor 180 may be implemented using at least one hardware form of a digital signal processing (DSP), a field-programmable gate array (FPGA), or a programmable logic array (PLA). Processor 180 may also incorporate one or more combinations of a central processing unit (CPU), a graphics processing unit (GPU), and a modem. Here, the CPU is primarily responsible for processing the operating system, user interface, and application programs, the GPU is responsible for rendering and drawing display content, and the modem is used for wireless communication. It should be understood that the modem may not be integrated into the processor, but may be implemented independently by a single communications chip.
[0137] According to one aspect of the present application, the present application further provides a computer-readable storage medium, which may be included in the electronic device described in the above embodiments, or may exist independently of the eye movement tracking device, the computer-readable storage medium carrying computer-readable instructions that, when executed by a processor, implement the method of any one of the above embodiments.
[0138] It should be noted that the computer-readable medium in the embodiments of the present application may be a computer-readable signal medium, a computer-readable storage medium, or any combination of the above. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of the computer-readable storage medium may include, but are not limited to, an electrical connection having one or more conductors, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM), a flash memory, an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical memory device, a magnetic memory device, or any suitable combination of the above. In the present application, a computer-readable storage medium may be any tangible medium that contains or stores a program, which may be used by or in combination with an instruction execution system, apparatus, or device. However, in this application, a computer-readable signal medium may include a propagated data signal, either in baseband or as part of a carrier, bearing computer-readable program code. Such propagated data signals may take various forms, including, but not limited to, electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium may be any computer-readable medium other than a computer-readable storage medium, which can transmit, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium may be transmitted by any suitable medium, including, but not limited to, wireless, wired, etc., or any suitable combination of the above.
[0139] It should be noted that although the above detailed description refers to multiple modules or units of an apparatus for performing operations, such division is not mandatory. In fact, according to embodiments of the present application, the features and functions of two or more of the modules or units described above may be embodied in one module or unit. Conversely, the features and functions of one module or unit described above may be further divided to be embodied by multiple modules or units.
[0140] Other embodiments of the present application will be readily apparent to those skilled in the art after considering the specification and practicing the embodiments disclosed herein. This application is intended to cover any variations, uses, or adaptations of the present application, including common general knowledge or commonly used technical means in the art that are not disclosed herein, in accordance with the general principles of the present application.
[0141] It is to be understood that the present application is not limited to the exact construction already described above and shown in the drawings, and that various modifications and changes can be made without departing from the scope thereof, which is limited only by the appended claims.
Claims
1. 1. A method for determining gaze information for use in an eye movement tracking device, comprising: acquiring two eye images collected by the eye tracking device; determining target feature points for each of the two eye images; and determining a reliability corresponding to each of the two eye images based on the target feature points for each of the two eye images; When a reliability corresponding to a first eyeball image of the two eyeball images is smaller than a predetermined reliability and a reliability corresponding to a second eyeball image of the two eyeball images is equal to or greater than the predetermined reliability, determining first gaze information of the second eyeball image based on the target feature point of the second eyeball image, and determining the first gaze information as first gaze information of the first eyeball image; inputting the two eyeball images into a gaze estimation model, respectively, and obtaining second gaze information for each of the two eyeball images output by the gaze estimation model; A method for determining gaze information, comprising determining target gaze information corresponding to the two eyeball images based on first gaze information of the first eyeball image, first gaze information of the second eyeball image, second gaze information of the first eyeball image, and second gaze information of the second eyeball image.
2. determining a target feature point for each of the two eye images and determining a reliability corresponding to each of the two eye images based on the target feature point for each of the two eye images, performing a grayscale process on each of the two eyeball images to obtain grayscale images corresponding to each of the two eyeball images; determining light spots in grayscale images corresponding to the two eye images, and determining target feature points from the light spots included in the grayscale images corresponding to the two eye images, wherein the target feature points are spots among the light spots that meet a condition; and determining a reliability corresponding to each of the two eye images based on the number of the light spots and the number of the target feature points included in a grayscale image corresponding to each of the two eye images.
3. For grayscale images corresponding to each of the two eye images, determining light spots in the grayscale images corresponding to each of the two eye images and determining target feature points from the light spots included in the grayscale images corresponding to each of the two eye images includes: determining a reference pupil center in the grayscale image; calculating a grayscale difference value between the reference pupil center and each of the light spots in the grayscale image; determining a light spot having a grayscale difference value greater than a difference value threshold as a contour reference point; and determining the target feature points from the contour reference points using an ellipse fitting method.
4. Determining first gaze information of the second eyeball image based on the target feature point of the second eyeball image and determining the first gaze information as first gaze information of the first eyeball image includes: determining a pupil center of the second eye image based on the target feature point of the second eye image; determining an iris region in a grayscale image of the second eye image based on the pupil center; determining a maximum grayscale value in a grayscale image of the second eye image of a pupil of the second eye image; determining a reference iris edge point of the second eye image based on the maximum grayscale value and the iris region; determining an iris center of the second eye image based on the reference iris edge points using an ellipse fitting method; 2. The method of claim 1, further comprising: determining the first gaze information of the second eye image based on the pupil center and the iris center; and determining the first gaze information as the first gaze information of the first eye image.
5. The gaze estimation model includes a feature extraction layer and a gaze estimation layer, and inputting the two eyeball images into the gaze estimation model and obtaining second gaze information of each of the two eyeball images output by the gaze estimation model includes: inputting the two eyeball images to the feature extraction layer, respectively, to obtain eye feature information corresponding to the two eyeball images; determining a corneal center corresponding to each of the two eyeball images and a pupil center corresponding to each of the two eyeball images based on eye feature information corresponding to each of the two eyeball images; 2. The method of claim 1, further comprising inputting a corneal center corresponding to each of the two eyeball images and a pupil center corresponding to each of the two eyeball images into the gaze estimation layer, respectively, to obtain second gaze information for each of the two eyeball images.
6. determining target gaze information corresponding to the two eye images based on first gaze information of the first eye image, first gaze information of the second eye image, second gaze information of the first eye image, and second gaze information of the second eye image; determining first reference gaze information based on first gaze information of the first eye image and first gaze information of the second eye image, and determining second reference gaze information based on second gaze information of the first eye image and second gaze information of the second eye image; Obtaining a first weight of the first reference gaze information and a second weight of the second reference gaze information; 2. The method of claim 1, further comprising: performing a weighting calculation based on the first weight, first reference gaze information, the second weight, and the second reference gaze information to determine target gaze information corresponding to the two eye images.
7. obtaining a first weight of the first reference gaze information and a second weight of the second reference gaze information; determining a method for determining the target gaze information; When the determining manner is a first manner, determining the first reference gaze information as target gaze information; 7. The method of claim 6, further comprising: if the determination method is the second method, obtaining a first weight of the first reference gaze information and a second weight of the second reference gaze information.
8. 1. An apparatus for determining gaze information for use in an eye movement tracking device, comprising: an image capture module for capturing two eye images collected by the eye movement tracking device; a target feature point determination module for determining target feature points for each of the two eye images and determining a reliability corresponding to each of the two eye images based on the target feature points for each of the two eye images; a first gaze information determination module for determining first gaze information of the second eyeball image based on the target feature point of the second eyeball image when a reliability corresponding to a first eyeball image of the two eyeball images is smaller than a predetermined reliability and a reliability corresponding to a second eyeball image of the two eyeball images is equal to or greater than the predetermined reliability, and determining the first gaze information as the first gaze information of the first eyeball image; a second gaze information determination module for inputting the two eyeball images into a gaze estimation model respectively and obtaining second gaze information for each of the two eyeball images output by the gaze estimation model; and a target gaze information determination module for determining target gaze information corresponding to the two eyeball images based on first gaze information of the first eyeball image, first gaze information of the second eyeball image, second gaze information of the first eyeball image, and second gaze information of the second eyeball image.
9. 1. An eye movement tracking device comprising: an image collection device and a device main body, the image collection device being installed at a target position of the device main body, wherein the target position is a position located on the cheek side and / or nose side of the target subject when the eye movement tracking device is attached to the target subject.
10. 10. The eye movement tracking device of claim 9, wherein the image capture device performs eye image capture at a target tilt angle, wherein the target tilt angle is greater than 50 degrees.