Method and device with slip detection and correction function
Obtaining eye images through imaging sensors and calculating pupil or cornea limbal contours solves the problem of user gaze tracing, achieving accurate tracking in virtual reality, augmented reality and autonomous driving, and improving interaction and monitoring effects.
Patent Information
- Application Number
- CN202210445768.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2017-06-30
- Filing Date
- 2018-06-23
- Publication Date
- 2025-08-08
- Estimated Expiration
- 2038-06-23
AI Technical Summary
The prior art is difficult to quickly and intuitively track user concerns, especially in virtual reality, augmented reality and autonomous driving environments, affecting the effectiveness and security of user interaction with computers.
By using an imaging sensor to acquire images of the human eye, extract the contour of the pupil or limbus, determine the positional relationship between the center of the eye and the imaging sensor, and combine head movement to calculate the relative motion and orientation relationship of the imaging sensor relative to the center of the eye to achieve eye tracking.
It realizes accurate tracking of user's sight, supports interaction in virtual reality and augmented reality, improves driver monitoring in autonomous driving environments, and enhances the sensitivity and reliability of user interaction with computers.
Smart Images

Figure CN114742863B_ABST
Abstract
Description
[0001] This invention patent application is a divisional application of the Chinese invention patent application with international application number PCT / US2018 / 039185, Chinese application number 201880044298.2, international application date June 23, 2018, and invention name “Wearable eye tracking system with slip detection and correction function”.
[0002] Cross-reference to related applications
[0003] The disclosure of U.S. Provisional Application No. 62 / 527,562 is incorporated herein by reference in its entirety. Technical Field
[0004] The present disclosure relates to a method and apparatus with slip detection and correction capabilities. Background Art
[0005] Human-computer interaction (HCI), or more generally, human-computer interaction, focuses on the design and use of computer technology and the interface between users and computers. HCI relies on sensitive, intuitive, and accurate measurement of human input movements. Mouse, keyboard, and touch screen are traditional input devices that require manual operation by the user. Some input devices, such as Microsoft The user's body or hand gestures can be tracked without any physical contact.In this disclosure, the term "user" and the term "person" can be used interchangeably.
[0006] Recent advances in virtual reality (VR) technology have brought VR goggles to the consumer market. VR goggles create an immersive three-dimensional (3D) experience for users. Users can look around the virtual world by turning their heads, just as they would in the real world.
[0007] Augmented reality (AR) is another rapidly developing field. A key difference between AR and VR is that AR operates in real-world scenes in real time, while VR operates solely within computer-created or recorded scenes. In both VR and AR, understanding where the user is looking and what actions they intend to take with respect to desired objects is extremely valuable. Effective and reliable eye tracking will enable widespread application in this context.
[0008] Today, autonomous vehicles are also on the cutting edge. There are situations where a car in autonomous mode might require the driver's attention due to updated road or traffic conditions or a change in driving style. Therefore, it is useful to continuously monitor where the driver is looking.
[0009] Machine learning and artificial intelligence (AI) likely work in a cycle of learning, modeling, and prediction. Quickly and intuitively tracking user attention for data collection and validation can play an important role in this cycle. Summary of the Invention
[0010] This article discloses a method, including: using an imaging sensor to obtain a first image and a second image of a human eye; and determining a positional relationship between an eyeball center and the imaging sensor based on the first image and the second image.
[0011] According to an embodiment, the method further comprises determining a relative movement of the imaging sensor with respect to the center of the eyeball based on the positional relationship.
[0012] According to an embodiment, the first image and the second image are different.
[0013] According to an embodiment, the method of determining the positional relationship includes extracting a contour of the pupil or the limbus from the first image and the second image, and determining the positional relationship based on the contour.
[0014] According to an embodiment, the first image and the second image are acquired during intentional or random movement of the eye.
[0015] According to an embodiment, the method further comprises determining an orientation relationship between the imaging sensor and the head of the person.
[0016] According to an embodiment, the method for determining the orientation relationship includes: obtaining a first set of gaze vectors when a person's head rotates around a first axis and the eyes remain at a first gaze point; obtaining a second set of gaze vectors when the person's head rotates around a second axis and the eyes remain at a second gaze point, or obtaining a third gaze vector when the eyes remain at a third gaze point; determining the orientation relationship based on the first set of gaze vectors and based on the second set of gaze vectors or the third gaze vectors.
[0017] According to an embodiment, the imaging sensor comprises more than one camera.
[0018] According to an embodiment, the imaging sensor is configured to acquire images of more than one eye of a person.
[0019] Disclosed herein is an apparatus comprising: an imaging sensor configured to acquire a first image and a second image of a human eye; and a computing device comprising a processor configured to determine a positional relationship between an eyeball center of the eye and the imaging sensor based on the first image and the second image.
[0020] According to an embodiment, the processor may be further configured to determine a relative movement of the imaging sensor with respect to the center of the eyeball based on the positional relationship.
[0021] According to an embodiment, the first image and the second image are different.
[0022] According to an embodiment, the processor may be further configured to extract a contour of the pupil or limbus from the first image and the second image, and determine the positional relationship based on the contour.
[0023] According to an embodiment, the processor may be further configured to determine the positional relationship by determining the position of the center of the eyeball in the two-dimensional image plane of the imaging sensor based on the contour.
[0024] According to an embodiment, the imaging sensor is configured to acquire the first image and the second image during intentional or random movement of the eye.
[0025] According to an embodiment, the processor may be configured to determine a positional relationship between the imaging sensor and the person's head.
[0026] According to an embodiment, the processor can also be configured to determine the orientation relationship by the following method: when a person's head rotates around a first axis and the eyes remain at a first gaze point, obtaining a first set of gaze vectors; when the person's head rotates around a second axis and the eyes remain at a second gaze point, obtaining a second set of gaze vectors, or obtaining a third gaze vector when the eyes remain at a third gaze point; based on the first set of gaze vectors, and based on the second set of gaze vectors or the third gaze vectors, determine the orientation relationship.
[0027] According to an embodiment, the imaging sensor comprises more than one camera.
[0028] According to an embodiment, the imaging sensor is configured to acquire images of more than one eye of a person.
[0029] This article discloses a method, including: when a person's head rotates about a first axis and the eyes remain at a first gaze point, using an imaging sensor to acquire a first set of images of the person's eyes; when the person's head rotates about a second axis and the eyes remain at a second gaze point, using the imaging sensor to acquire a second set of images of the person's eyes, or when the eyes remain at a third gaze point, using the imaging sensor to acquire a third image; and based on the first set of images and based on the second set of images or the third image, determining an orientation relationship between the imaging sensor and the person's head.
[0030] According to an embodiment, the method also includes: obtaining a first group of gaze vectors based on the first group of images; obtaining a second group of gaze vectors based on the second group of images, or obtaining a third gaze vector based on the third image; wherein, the orientation relationship is determined based on the first group of gaze vectors and based on the second group of gaze vectors or the third gaze vectors.
[0031] According to an embodiment, the first gaze point and the second gaze point are different; or wherein the first gaze point and the third gaze point are different.
[0032] According to an embodiment, the imaging sensor comprises more than one camera.
[0033] According to an embodiment, the imaging sensor is configured to acquire images of more than one eye of a person.
[0034] Disclosed herein is an apparatus comprising an imaging sensor and a computing device including a processor. The imaging sensor may be configured to capture a first set of images of a person's eyes when the person's head rotates about a first axis and the eyes remain at a first gaze point. The imaging sensor may also be configured to capture a second set of images of the person's eyes when the person's head rotates about a second axis and the eyes remain at a second gaze point, or to capture a third set of images when the eyes remain at a third gaze point. The computing device includes a processor configured to determine an orientation relationship between the imaging sensor and the person's head based on the first set of images and based on either the second set of images or the third set of images.
[0035] According to an embodiment, the processor is also configured to: obtain a first group of gaze vectors based on the first group of images; obtain a second group of gaze vectors based on the second group of images, or obtain a third gaze vector based on the third image; and determine the orientation relationship based on the first group of gaze vectors and based on the second group of gaze vectors or the third gaze vectors.
[0036] According to an embodiment, the first gaze point and the second gaze point are different; or wherein the first gaze point and the third gaze point are different.
[0037] According to an embodiment, the imaging sensor comprises more than one camera.
[0038] According to an embodiment, the imaging sensor is configured to acquire images of more than one eye of a person.
[0039] This article discloses a method, comprising: acquiring a first image of a human eye using an imaging sensor; extracting a contour of a pupil or corneal limbus from the first image; determining a position of an eyeball center of the eye in a two-dimensional image plane of the imaging sensor; and determining a positional relationship between the eyeball center of the eye and the imaging sensor based on the contour and the position.
[0040] According to an embodiment, the step of determining the position comprises extracting the outline of the pupil or the limbus from two images of the human eye acquired using an imaging sensor.
[0041] According to an embodiment, the first image is one of said two images.
[0042] Disclosed herein is an apparatus comprising: an imaging sensor configured to acquire an image of a human eye; and a computing device comprising a processor configured to extract a contour of a pupil or a limbus from at least one image to determine a position of an eyeball center of the eye in a two-dimensional image plane of the imaging sensor, and to determine a positional relationship between the eyeball center of the eye and the imaging sensor based on the contour and the position.
[0043] According to an embodiment, the processor may be further configured to determine the position of the eyeball center of the eye in the two-dimensional image plane of the imaging sensor by extracting the outline of the pupil or the corneal limbus from the at least two images. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] Figure 1 A two-dimensional (2D) coordinate system is shown.
[0045] Figure 2 A three-dimensional coordinate system is shown.
[0046] Figure 3 A straight line in two-dimensional space is shown.
[0047] Figure 4 An ellipse in two dimensions.
[0048] Figure 5 It shows that the eyeball center point image on the two-dimensional image plane of the camera can be calculated based on at least two pupil images at different positions.
[0049] Figure 6 Three three-dimensional coordinate systems are shown: the head coordinate system; the camera coordinate system; and the eye coordinate system.
[0050] Figure 7 The head coordinate system is shown.
[0051] Figure 8 The cameras are shown fixed to a frame so that the relative position and orientation of the cameras do not change.
[0052] Figure 9 The first step of a calibration process is shown.
[0053] Figure 10 The second step of the calibration process is shown.
[0054] Figure 11A A flow chart of a method according to an embodiment is schematically shown.
[0055] Figure 11B Shown Figure 11A A flowchart of the optional procedures in the flowchart.
[0056] Figure 12A flow chart of a method according to another embodiment is schematically shown. DETAILED DESCRIPTION
[0057] In the following detailed description, examples with extensive detail are provided to provide a thorough understanding of the relevant teachings. However, it will be apparent to those skilled in the art that the present teachings can be practiced without these details. In other cases, well-known methods, procedures, and components have been described at a relatively high level without detailed description to avoid unnecessarily obscuring aspects of the present teachings.
[0058] Appendix Section A1.1 defines the three-dimensional coordinate system using the right-hand rule. Appendix Section A1.2 defines the two-dimensional coordinate system in the camera image frame.
[0059] The mathematical utility functions used in this disclosure are listed in Section A2, Section A3, and Section A4 of the Appendix. Quaternions, vectors, and matrix mathematics are discussed. Quaternions are widely used in the present invention. Functions using quaternions can also be expressed using matrices, Euler angles, or other suitable mathematical expressions.
[0060] The present invention relates to three three-dimensional coordinate systems, such as Figure 6 As shown: the head coordinate system Xh-Yh-Zh-Oh, abbreviated as CS-H; the camera coordinate system Xc-Yc-Zc-Oc, abbreviated as CS-C; and the eye coordinate system Xe-Ye-Ze-Oe, abbreviated as CS-E. The origin of CS-E is at the center of the eyeball. The units of CS-C and CS-E are the radius of the eyeball. The order of inclusion of these coordinate systems is: CS-H > CS-C > CS-E.
[0061] The head coordinate system CS-H is as follows Figure 7 As shown in Figure 1, the X-axis points from the user's left ear to their right ear; the Y-axis points from the base of the chin to the top of the head; and the Z-axis points from the tip of the nose to the back of the head. Therefore, the X-axis is aligned with the left-right horizontal direction, the Y-axis is aligned with the vertical direction, and the Z-axis is aligned with the front-back horizontal direction. The rotational directions of the axes are defined in Section A1.1 of the Appendix.
[0062] The eye coordinate system CS-E is fixed to one of the user's eyes. When the eye is in its primary position (e.g., looking straight ahead), the X-axis, Y-axis, and Z-axis of CS-E point in the same directions as the X-axis, Y-axis, and Z-axis of CS-H, respectively.
[0063] This disclosure uses abbreviations of the form "ABC": A stands for type; B stands for specific; and C stands for context. For example, the abbreviation "qch" represents the position of camera "c" in head coordinate system "h" using quaternion "q." See Section A2 of the Appendix.
[0064] The line of focus is the line passing through the center of the eyeball and the center of the pupil.
[0065] The gaze vector is a unit vector pointing from the origin of the CS-E to the negative Z-axis of the CS-E. The gaze vector indicates the direction of the gaze line.
[0066] The gaze point is the point on which the user's gaze falls.
[0067] Gaze is a situation where the gaze vector or gaze point is restricted to a small area for a period of time.
[0068] The following functions can be achieved by using an imaging sensor (e.g., a camera) that faces the user's eyes at a fixed position and orientation relative to the user's head. In the present invention, the terms "imaging sensor" and "camera" are used interchangeably. "Imaging sensor" is not limited to static cameras.
[0069] From two or more images of the user's eye, the position of the eye center in the camera coordinate system CS-C can be calculated using an ellipse extracted from the outline of the pupil or limbus. There are no specific requirements for the user's actions to obtain these images. The user's eye movements can be either active or random.
[0070] Knowing the position of the eye center in the CS-C coordinate system, the gaze vector in the CS-C coordinate system can be calculated. The eye tracking process can be started and the gaze vector in the CS-C coordinate system can be continuously calculated.
[0071] During eye tracking, by continuously calculating and updating the position of the eye center in the CS-C coordinate system, camera slip caused by changes in the camera's relative position to the user's eye center can be detected and corrected.
[0072] During camera-head calibration, the camera's orientation in the user's head coordinate system (CS-H) is calculated. Step 1: The user selects a first gaze point from a distance and fixates on it. Without changing the first gaze point, the user rotates their head around the X-axis of the head coordinate system. This generates a first set of gaze vectors associated with this rotation. This generates a vector vhcx in the CS-C coordinate system, aligned with the X-axis of the CS-H coordinate system. Step 2: The user selects a second gaze point from a distance. Without changing the second gaze point, the user rotates their head around the Y-axis of the head coordinate system. This generates a second set of gaze vectors associated with this rotation. The first and second gaze points can be the same or different. This generates a vector vhcy in the CS-C coordinate system, aligned with the Y-axis of the CS-H coordinate system. Step 3: The user looks straight ahead, selects a gaze point, and records the gaze vector at this location. In the CS-C coordinate system, a vector vhcz, aligned with the Z-axis of the CS-H coordinate system, can be considered to point in the negative direction of this gaze vector. If any two of the X, Y, and Z axes of the coordinate system are known, the third axis can be calculated. Therefore, using the vector obtained in any of the three steps above, the direction of the camera relative to the user's head can be calculated.
[0073] The gaze vector of the eye relative to the user's head can be calculated using the gaze vector in CS-C and the orientation of the camera relative to the user's head, that is, the orientation of CS-C relative to CS-H.
[0074] Multiple configurations based on the above description are also possible: Configuration A uses one camera facing the user's eyes; Configuration B uses multiple cameras facing the user's eyes; Configuration C uses one camera facing both eyes of the user; Configuration D can be a combination of Configuration A, Configuration B and Configuration C.
[0075] It is possible to: use a camera with configuration A to get the gaze vector of one eye; use a single camera with configuration B to get the gaze vector of one eye; use one camera with configuration A per eye to get the gaze vector of both eyes; use more than one camera with configuration B per eye to get the gaze vector of both eyes; use one camera with configuration C to get the gaze vector of both eyes; and use configuration D to find the gaze vector of both eyes.
[0076] Various hardware components can be used to implement the functions and processes in this disclosure. One hardware component is a camera. A camera can measure the brightness and color of light. In this disclosure, the term "camera coordinate system" or "CS-C" is used interchangeably with the term "imaging sensor coordinate system" or "CS-I". A camera can capture color images or grayscale images. A camera can also capture infrared images or non-infrared (such as visible light) images. Parameters of a camera include its physical size, resolution, and the focal length of its lens. A two-dimensional camera image frame coordinate system is defined in Appendix Section A1.2.
[0077] One hardware component is a headgear that secures the camera to the user's head. Depending on the application, this headgear can be a pair of eyeglass frames, a headband, or a helmet. Other headgear forms are also possible.
[0078] A hardware component can be a computer, such as an embedded system or a desktop system.
[0079] like Figure 8 As shown, the camera can be fixed on a rigid frame so that its position and orientation relative to the user's head remain unchanged. In eye tracking mode, the gaze vector of the user's eyes can be calculated.
[0080] The position of the eye center relative to the camera is calculated as follows.
[0081] Before calculating the gaze vector, the location of the eye center ("vec") in the CS-C must be determined. To do this, the camera captures a series of images of the eye. For each image, the pupil outline is extracted. The pupil outline can be visualized as an ellipse. Using two or more ellipses representing the pupil outline in the image sequence, vec can be calculated. The detailed steps are described in Sections C1 and C2 of the algorithm section.
[0082] While two ellipses derived from two eye images are sufficient to calculate VEC, in practice, using ellipses from more eye images can yield more accurate and robust results. Specifically, by combining a series of eye images, more VEC solutions can be calculated. These solutions may vary due to measurement errors and system noise. However, these solutions can be further processed using various methods to achieve even higher accuracy and robustness. Examples of these methods include averaging, least squares regression, and Kalman filtering.
[0083] Ideally, the position and shape of the pupil should vary sufficiently from one eye image to another. To obtain images of pupils in different positions, the user can either gaze at a fixed position and turn their head in different directions, or keep their head in a fixed position and turn their eyes to look in different directions, or a combination of the two.
[0084] The ellipse obtained from the pupil outline in the eye image should not be circular. Specifically, the ratio of the length of its semi-major axis a to the semi-minor axis b (see the Appendix) should be greater than a threshold, that is, the following condition should be met: a / b > threshold.
[0085] Alternatively, the limbus can be used in place of the pupil to extract ellipses from eye images. Ellipses extracted from the limbus in two or more eye images can be used to calculate the eye center vec in the CS-C coordinate system using the same method as ellipses extracted from the pupil in two or more images. Ellipses extracted from both the pupil and limbus can also be used to improve flexibility and robustness.
[0086] The following describes how to calculate the gaze vector in the camera coordinate system. After obtaining the eye center vec, the eye tracking process begins. The gaze vector vgc in CS-C can be calculated from the captured eye image, more specifically, from the center of the pupil pc detected in the eye image.
[0087] Assumptions:
[0088] vec = (xc, yc, zc) is the coordinate of the origin of CS-E in CS-C;
[0089] pc = (x, y) is the center of the pupil in the camera image plane;
[0090] vgc = (xg, yg, zg) is the gaze vector in CS-C;
[0091] vgc can be calculated from vec and pc. Detailed information can be found in C3 of the algorithm section.
[0092] The following describes methods for detecting and correcting camera slip. In CS-C, the eye center position (vec) can be calculated in a separate stage or in parallel with the eye tracking process. If the eye center position (vec) is calculated continuously, possible camera slip due to changes in the camera position relative to the eye center can be detected by comparing the current value of vec with the existing value. During eye tracking, camera slip relative to the eye center can be effectively detected. vec can be a parameter that is updated in real time during eye tracking.
[0093] The following describes a method for calibrating the orientation of the camera relative to the user's head.
[0094] During camera-head calibration, the camera's orientation in the head coordinate system CS-H is calculated. The user is instructed to select a gaze point from a distance and gaze at it. First, without changing their gaze point, the user rotates their head around the X-axis of the head coordinate system, obtaining a first set of gaze vectors associated with this rotation. This yields a vector vhcx aligned with the X-axis of CS-H in CS-C. Second, the user can maintain their first gaze point or select a new one from a distance. Without changing their gaze point, they rotate their head around the Y-axis of the head coordinate system, obtaining a second set of gaze vectors associated with this rotation. This yields a vector vhcy aligned with the Y-axis of CS-H in CS-C. Third, the user can look straight ahead and select a gaze point, recording the gaze vector at that location. Therefore, in CS-C, the vector vhcz aligned with the Z-axis of CS-H can be considered to point in the negative direction of this gaze vector. Given any two of the X, Y, and Z axes of a coordinate system, the third can be calculated. Therefore, using the vectors obtained in any of the three steps above, the direction of the camera relative to the user's head can be calculated.
[0095] During the calibration process in one embodiment, as Figure 9 and Figure 10 As shown. Figure 9 In [1], the user selects a distant fixation point. While gazing at this fixation point, the user moves their head up and down. The camera captures images of the eye that show the location of the pupil center relative to different head positions. Figure 10 In the example, the user can maintain the first gaze point or select a new distant gaze point. Without changing the gaze point, the user can turn their head left and right. The camera captures images of the eyes, showing the position of the pupils in relation to different head positions.
[0096] In another calibration process, the user looks straight ahead and selects a distant gaze point. The gaze vector at this position is recorded. Then the Z-axis direction of CS-H in the CS-C coordinate system can be regarded as pointing in the negative direction of this gaze vector. Then, the user can choose to perform Figure 9 The first step and Figure 10 The second step in .
[0097] Assume that qch is the orientation quaternion of the camera coordinate system in the head coordinate system, that is, the orientation of CS-C in CS-H. The details of calculating qch can be found in Section C4 of the algorithm section.
[0098] More than two gaze vectors can be used to detect the X or Y axis of CS-H in CS-C, as described in algorithm sections C4.2 and C4.3. Multiple gaze vector pairs can generate multiple solutions, which can be obtained using methods such as averaging, least squares regression, and Kalman filtering. Similarly, multiple vectors can be used to detect the Z axis of CS-H in CS-C, as described in algorithm section C4.8.
[0099] The following describes how to calculate the gaze vector in the head coordinate system.
[0100] Given the gaze vector vgc in CS-C pointing from the origin of CS-E to the target the user is looking at, and the CS-C direction quaternion qch in CS-H, the gaze vector vgh in CS-H can be calculated as vgh = qvq_trans(qch, vgc). See Appendix Section A2.3.6.
[0101] Figure 11A A flow chart of a method according to an embodiment is schematically shown. In step 1210, a first image and a second image of a person's eye are acquired using an imaging sensor. The imaging sensor may include one or more cameras. The imaging sensor may acquire images of one or both eyes of a person. The first image and the second image may be different. The first image and the second image may be acquired while the eyes are moving intentionally or randomly. In step 1220, a positional relationship between the center of the eye and the imaging sensor may be determined based on the first image and the second image, for example, using a processor in a computing device. For example, the outline of the pupil or the limbus may be extracted from the first image and the second image (for example, using a processor), and the positional relationship may be determined based on the outline (for example, using a processor). In optional step 1230, the relative movement of the imaging sensor relative to the center of the eye may be determined based on the positional relationship, for example, using a processor. In optional step 1240, the orientation relationship between the imaging sensor and the person's head may be determined, for example, using a processor.
[0102] Figure 11B A flowchart for optional step 1240 is shown. In step 1241, a first set of gaze vectors is obtained (e.g., using a processor) when the person's head is rotated about a first axis and the eyes are maintained at a first gaze point. In step 1242, a second set of gaze vectors is obtained (e.g., using a processor) when the person's head is rotated about a second axis and the eyes are maintained at the second gaze point; or a third set of gaze vectors is obtained (e.g., using a processor) when the eyes are maintained at a third gaze point. In step 1243, an orientation relationship is determined based on the first set of gaze vectors and based on either the second set of gaze vectors or the third set of gaze vectors (e.g., using a processor).
[0103] Figure 12A flow chart of a method according to another embodiment is schematically shown. In step 1310, while a person's head is rotated about a first axis and the eyes are maintained at a first gaze point, a first set of images of the person's eyes is acquired using an imaging sensor. The imaging sensor may include one or more cameras. The imaging sensor may acquire images of one or both eyes of the person. In optional step 1311, a first set of gaze vectors is acquired based on the first set of images (e.g., using a processor of a computing device). In step 1320, while the person's head is rotated about a second axis and the eyes are maintained at a second gaze point, a second set of images of the eyes is acquired using the imaging sensor; or while the eyes are maintained at a third gaze point, a third image is acquired using the imaging sensor. In optional process 1321, a second set of gaze vectors is obtained based on the second set of images (e.g., using a processor); or a third gaze vector is obtained based on the third image (e.g., using a processor). At step 1340, a positional relationship between the imaging sensor and the person's head is determined (e.g., using a processor of a computing device) based on the first set of images and based on the second set of images or the third set of images (e.g., based on the first set of gaze vectors, the second set of gaze vectors, and the third set of gaze vectors obtained from the first set of images, the second set of images, and the third set of images, respectively). The first gaze point and the second gaze point may be different, or the first gaze point and the third gaze point may be different.
[0104] algorithm
[0105] C1. Calculate the position of the eyeball center image in the camera's two-dimensional image plane
[0106] Given an image of an eye, assume that the pupil outline can be extracted. The shape of the pupil outline can be considered as an ellipse. (See Section A4 of the Appendix)
[0107] An ellipse in two-dimensional space can be represented by its vertices and minor axis vertices, as Figure 4 shown.
[0108] Although the center of the eye cannot be seen from the camera, Figure 5 As shown, the point pe of the eyeball center image in the camera's two-dimensional image plane can be calculated based on at least two pupil images at different positions. Assume:
[0109] E0 is the ellipse extracted from the eye image based on the pupil at time t0;
[0110] E1 is the ellipse extracted from the eye image based on the pupil at time t1.
[0111] Assumptions:
[0112] L0 is a straight line passing through the two minor axis vertices (co-vertices) of E0;
[0113] L1 is a straight line passing through the two minor axis vertices of E1.
[0114] The method for solving the linear equations of L0 and L1 is in A3.1 of the Appendix.
[0115] The intersection point pe of L0 and L1 can be calculated (see Appendix Section A3.2).
[0116] C2. Calculate the position of the eye center in CS-C
[0117] Assume that the ellipse E represents the outline of the pupil in the captured eye image, and assume that E is defined by its vertex and minor axis vertex, such as Figure 4 As shown;
[0118] p0 and p1 are two points at the ends of two vertices, the order is not important;
[0119] p2 and p3 are the endpoints of the two minor axis vertices, the order is not important;
[0120] pc is the center point of the ellipse, which can be calculated from p0, p1, p2, and p3 (see Appendix Section A4.2);
[0121] pe is the eye center image obtained in the C1 part of the algorithm.
[0122] The above six points can be converted into corresponding three-dimensional unit vectors on the two-dimensional image plane of the camera, pointing from the CS-C origin to the image point on the two-dimensional image plane of the camera, as described in Appendix A.1.3.
[0123] v0=v_frm_p(p0)
[0124] v1=v_frm_p(p1)
[0125] v2=v_frm_p(p2)
[0126] v3=v_frm_p(p3)
[0127] vc=v_frm_p(pc)
[0128] ve=v_frm_p(pe)
[0129] cos_long=v_dot(v0,v1)
[0130] cos_short = v_dot(v2, v3)
[0131] tan_long_h=sqrt((1-cos_long) / (1+cos_long))
[0132] tan_short_h=sqrt((1-cos_short) / (1+cos_short))
[0133] cos_alfa=v_dot(vc,ve)
[0134] sin_beta=tan_short_h / tan_long_h
[0135] sin_r=sqrt(1-cos_alfa*cos_alfa)
[0136] sin_d=sqrt(1-sin_beta*sin_beta)
[0137] dist = sin_d / sin_r
[0138] Known:
[0139] ve, the unit vector from the origin CS-C to the center of the eyeball,
[0140] dist, the distance from the origin of CS-C to the center of the eyeball,
[0141] vec, the position of the eye center in CS-C is: vec = (x, y, z)
[0142] in,
[0143] x=xe*dist
[0144] y=ye*dist
[0145] z=ze*dist
[0146] v=(xe,ye,ze)
[0147] C3. Calculate the gaze vector in the camera coordinate system
[0148] Assume that the position of the eye center has been obtained using the method described in parts C1 and C2 of the algorithm, where
[0149] vec = (xc, yc, zc) is the coordinate of the origin of CS-E in CS-C
[0150] pc = (x, y) is the center of the pupil in the camera coordinate system (see Appendix Section A1.2)
[0151] vgc = (xg, yg, zg) is the gaze vector from the origin of CS-E to the gaze point in CS-C. The calculation method of Vgc is:
[0152] h = -DEFOX(x) (see Appendix Section A1.3)
[0153] v=-DEFOY(y)
[0154] a=h*h+v*v+1
[0155] b=2*(((a-1)*zc-h*xc-v*yc))
[0156] c=(xc-h*zc)*(xc-h*zc)+(yc-v*zc)*(yc-v*zc)-1
[0157] p=b*b-4*a*c
[0158] k=sqrt(p)
[0159] z1=(-b+k) / (2*a)
[0160] z2=(-bk) / (2*a)
[0161]
[00109] z1 and z2 are both candidate solutions for zv. z1 is selected when z2 points in the direction of the camera. Therefore, we have:
[0162] zg=z1
[0163] xg=h*(zc+zg)-xc
[0164] yg=v*(zc+zg)-yc
[0165] C4. Get the camera's position relative to the user's head
[0166] Assuming the eye center position is known, calculate the gaze vector vgc in CS-C. The goal is to obtain the orientation quaternion qch of CS-C in CS-H.
[0167] During a C4.1 calibration process, the user selects a distant gaze point. While maintaining this gaze point, the user first rotates their head around the X-axis of the CS-H, obtaining a first gaze vector v0 and a second gaze vector v1 associated with this axis. The user can then maintain this first gaze point or select a new distant gaze point and maintain it. Without changing their gaze point, the user rotates their head around the Y-axis of the CS-H, obtaining a first gaze vector v2 and a second gaze vector v3 associated with this axis. v0, v1, v2, and v3 are all vectors in the CS-C.
[0168] C4.2 calculates the vector vhcx in CS-C that is aligned with the X axis of CS-H:
[0169] v10=v_crs(v1, v0)
[0170] vhcx=v_uni(v10)
[0171] C4.3 calculates the vector vhcy in CS-C that is aligned with the Y axis of CS-H:
[0172] v32=v_crs(v3, v2)
[0173] vhcy=v_uni(v32)
[0174] C4.4 Knowing vhcx and vhcx, we can calculate the 3x3 matrix mhc from CS-H to CS-C:
[0175] vx=vhcx
[0176] vz = v_crs(vhcx, vhcy)
[0177] vy=v_crs(vz,vx)
[0178] mhc=m_frm_v(vx,vy,vz)
[0179] C4.5 Knowing mhc, calculate the quaternion qch from CS-C to CS-H:
[0180] qhc=q_frm_m(mhc)
[0181] q_cnj(qhc)
[0182] C4.6: It's worth noting that the user can also rotate their head around the Y-axis of CS-H to obtain the first gaze vector v2 and the second gaze vector v3 along that axis. Then, rotate around the X-axis of CS-H to obtain the first gaze vector v0 and the second gaze vector v1 along that axis. This gives us v0, v1, v2, and v3 in CS-C. The rest of the steps are identical.
[0183] C4.7: It's worth noting that, as described in C4.2 and C4.3, more than two vectors can be used to calculate vhcx or vhcy. Taking vhcx as an example, if multiple pairs of gaze vectors are used, there are multiple solutions, such as vhcx0, vhcx1, vhcx2, and so on. Using well-known methods such as averaging, least squares regression, or Kalman filtering, the final result vhcx can be calculated based on these solutions. Similarly, vhcy can be calculated in the same way.
[0184] C4.8 In another calibration process, the user looks straight ahead and selects a distant gaze point. The gaze vector at this position is recorded. The direction of the Z axis of CS-H in CS-C can then be considered to point in the negative direction of this gaze vector. In CH-C, a vector vhcz aligned with the Z axis of CS-H can be obtained. The orientation of CS-H in CH-C can be calculated from any two vectors among vhcx, vhcy, and vhcz. vhcx and / or vhcy can be obtained (see algorithm section C4.1). The process of calculating qch is similar to the description in algorithm sections C4.2 to C4.5.
[0185] appendix
[0186] The above algorithm uses the mathematical tools listed in the appendix.
[0187] A1 coordinate system
[0188] A1.1 As Figure 2 As shown in Figure 1, a three-dimensional coordinate system has three axes: X, Y, and Z. The order of the axes and the positive rotation direction can be determined by the right-hand rule.
[0189] Any two axes can form a plane. Therefore, in a three-dimensional coordinate system, there are three planes defined as P-XY, P-YX, and P-ZX.
[0190] A1.2 As Figure 1 As shown, there are two axes, X and Y, in the two-dimensional coordinate system of the camera image frame.
[0191] A1.3 Convert points in the two-dimensional camera image frame coordinate system to the three-dimensional camera coordinate system.
[0192] A three-dimensional camera coordinate system CS-C has the x-axis pointing to the right, the y-axis pointing to the top, and the z-axis pointing away from the lens.
[0193] The two-dimensional image plane can be considered as:
[0194] Parallel to the XY plane of CS-C;
[0195] The origin is in the upper left corner;
[0196] The center of its image is located at (0, 0, 1) of CS-C;
[0197] Its X-axis is parallel to the X-axis of CS-C and points in the same direction;
[0198] Its Y-axis is parallel to the Y-axis of CS-C and points in the opposite direction;
[0199] The units are different from those of CS-C. Specifically, FOCAL LEN is the focal length of the camera, measured in pixels.
[0200] Calculate the unit vector vu from the origin of CS-C to point p on the two-dimensional plane of the camera image:
[0201] vu=v_frm_p(p)
[0202] in,
[0203] p=(x,y)
[0204] vu=(vx,vy,vz)
[0205] in,
[0206] vu=v_uni(v)
[0207] v=(h,v,−1.0)
[0208] in,
[0209] h=DEFOX(x)=(x-x_center) / FOCAL_LEN
[0210] v=DEFOY(y)=(y_center-y) / FOCAL_LEN
[0211] Among them, (x_center, y_center) is the center coordinate of the camera's two-dimensional image frame.
[0212] A1.4 A point in a three-dimensional coordinate system can be represented by a three-dimensional vector v = (x, y, z), which points from the origin of the coordinate system to the position of the point.
[0213] A2 Quaternions, 3D Vectors, 3x3 Matrices and 2D Vector Math
[0214] A quaternion consists of 4 elements
[0215] q=(w,x,y,z)
[0216] A2.1.2 Unit quaternion:
[0217] q=q_idt(q)=(1,0,0,0)
[0218] A2.1.3 Conjugation of quaternions:
[0219] q_cnj(q)=(w, -x, -y, -z)
[0220] A2.1.4 Length of quaternion:
[0221] q_len(q)=sqrt(w*w+x*x+y*y+z*z)
[0222] sqrt() is the square root of a floating point number
[0223] A2.1.5 The length of a unit quaternion is 1
[0224] Normalized quaternion q:
[0225] u=q_uni(q)
[0226] in,
[0227] q=(w,x,y,z)
[0228] u=(uw,ux,uy,uz)
[0229] uw=x / len
[0230] ux=x / len
[0231] uy=y / len
[0232] uz=z / len
[0233] len=q_len(q)
[0234] A2.1.62 Products of Quaternions p and q
[0235] t=q_prd2(q,p)=q*p
[0236] in,
[0237] q=(qw,qx,qy,qz)
[0238] p=(pw,px,py,pz)
[0239] t=(tw,tx,ty,tz)
[0240] and,
[0241] tw=(qw*pw-qx*px-qy*py-qz*pz)
[0242] tx=(qw*px+qx*pw+qy*pz-qz*py)
[0243] ty=(qw*py-qx*pz+qy*pw+qz*px)
[0244] tz=(qw*pz+qx*py-qy*px+qz*pw)
[0245] Because quaternions can be used to represent rotational transformations, if q2=q_prd2(q1,q0) is the product of two quaternions, then using q2 as an orientation transformation is equivalent to first applying q0 and then applying q1.
[0246] A2.1.7 Product of three quaternions:
[0247] q=q_prd3(q1, q2, q3)=q_prd2(q1, q_prd2(q2, q3))
[0248] A2.1.8 Product of four quaternions:
[0249] q=q_prd4(q1, q2, q3, q4)=q_prd2(q1, q_prd3(q2, q3, q4))
[0250] A2.2.13-dimensional vector has 3 elements:
[0251] v = (x, y, z)
[0252] A2.2.2 Length of three-dimensional vector:
[0253] v_len(v)=sqrt(x*x+y*y+z*z)
[0254] A2.2.3 The length of a unit three-dimensional vector is 1
[0255] Normalize a 3D vector:
[0256] u=v_uni(v)
[0257] in,
[0258] v = (x, y, z)
[0259] u=(ux,uy,uz)
[0260] ux=x / len
[0261] uy=y / len
[0262] uz=z / len
[0263] len=v_len(v)
[0264] A2.2.4 A unit quaternion can be interpreted as a combination of a rotation vector and an angle around that vector:
[0265] q=(w,x,y,z)
[0266] v = (vx, vy, vz) is the rotation vector
[0267] theta is the rotation angle
[0268] in,
[0269] w=cos(theta / 2)
[0270] x=vx*sin(theta / 2)
[0271] y=vy*sin(theta / 2)
[0272] z=vz*sin(theta / 2)
[0273] A2.2.5 The dot product of two 3D vectors va and vb:
[0274] d=v_dot(va, vb)=va.vb=ax*bx+ay*by+az*bz
[0275] in,
[0276] va=(ax,ay,az)
[0277] vb=(bx,by,bz)
[0278] Vector dot products are of great significance.
[0279] Assume theta is the angle between va and vb:
[0280] Then cos(theta)=v_dot(va,vb)
[0281] A2.2.6 The cross product of two three-dimensional vectors va and vb:
[0282] vc = v_crs(va, vb) = va x vb
[0283] in,
[0284] va=(ax,ay,az)
[0285] vb=(bx,by,bz)
[0286] vc=(cx,cy,cz)
[0287] cx=ay*bz-az*by
[0288] cy=az*bx-ax*bz
[0289] cz=ax*by-ay*bx
[0290] A2.3.1 3×3 Matrix
[0291]
[0292] A2.3.2 Identity 3×3 Matrix
[0293]
[0294] A2.3.3 Matrix Subtraction
[0295]
[0296]
[0297]
[0298] A2.3.4 Matrix-Vector Multiplication
[0299] vd=mv_prd(m,v)=m*vs
[0300]
[0301] vs = (x, y, z)
[0302] vd=(dx,dy,dz)
[0303] in,
[0304] dx=Xx*x+Yx*y+Zx*z
[0305] dy=Xy*x+Yy*y+Zy*z
[0306] dz=Xz*x+Yz*y+Zz*zA2.3.5 Quaternion Matrix
[0307] m=m_frm_q(q)
[0308] q=(qw,qx,qy,qz)
[0309] Where m is a 3×3 matrix
[0310]
[0311] and
[0312] Xx=1.0f-2.0f*qy*qy-2.0f*qz*qz
[0313] Xy=2.0f*qx*qy+2.0f*qw*qz
[0314] Xz=2.0f*qx*qz-2.0f*qw*qy
[0315] Yx=2.0f*qx*qy-2.0f*qw*qz
[0316] Yy=1.0f-2.0f*qx*qx-2.0f*qz*qz
[0317] Yz=2.0f*qy*qz+2.0f*qw*qx
[0318] Zx=2.0f*qx*qz+2.0f*qw*qy
[0319] Zy=2.0f*qy*qz-2.0f*qw*qx
[0320] Zz=1.0f-2.0f*qx*qx-2.0f*qy*qy
[0321] A2.3.6 Converting a 3D vector v using a quaternion q
[0322] vd=qvq_trans(q, vs)=mv_prd(m, vs)
[0323] in,
[0324] q is the quaternion, vs is the original three-dimensional vector
[0325] vd is the transformed three-dimensional vector
[0326] m is a 3×3 matrix
[0327] m=m_frm_q(q)
[0328] A2.3.7 Matrix for Rotating the x-Axis
[0329] m=m_frm_x_axis_sc(s,c)
[0330] in,
[0331]
[0332] s=sin(theta)
[0333] c=cos(theta)
[0334] and,
[0335] Xx=1.0
[0336] Yx=0.0
[0337] Zx=0.0
[0338] Xy=0.0
[0339] Yv=c
[0340] Zy=-s
[0341] Xz=0.0
[0342] Yz=s
[0343] Zz=c
[0344] A2.3.8 Matrix for rotating the y-axis
[0345] m=m_frm_y_axis_sc(s,c)
[0346] in,
[0347]
[0348] s=sin(theta)
[0349] c=cos(theta)
[0350] and,
[0351] Xx=c
[0352] Yx=0.0
[0353] Zx=s
[0354] Xy=0.0
[0355] Yy=1.0
[0356] Zy=0.0
[0357] Xz=-s
[0358] Yz=0.0
[0359] Zz=c
[0360] A2.3.9 Quaternions of matrices
[0361] q=q_frm_m(m)
[0362] in,
[0363] q=(w,x,y,z)
[0364]
[0365] and,
[0366]
[0367] A2.3.9 Vector matrices
[0368] m=m_frm_v(vx,vy,vz)
[0369] in,
[0370]
[0371] vx=(Xx,Xy,Xz)
[0372] vy=(Yx,Yy,Yz)
[0373] vz=(Zx,Zy,Zz)
[0374] A2.4.1 A point in two-dimensional space is a two-dimensional vector consisting of two elements:
[0375] p=(x,y)
[0376] A2.4.2 The distance between two-dimensional points pa and pb:
[0377] d=p_dist(pa, pb)=sqrt((xa-xb)*(xa-xb)+(ya-yb)*(ya-yb))
[0378] in
[0379] pa=(xa,ya)
[0380] pb=(xb,yb)
[0381] A2.4.3 The length of a two-dimensional vector;
[0382] p_len(p)=sqrt(x*x+y*y)
[0383] A2.4.4 The length of a unit two-dimensional vector is 1
[0384] Normalized two-dimensional vector p:
[0385] u=p_uni(p)
[0386] in,
[0387] p=(x,y)
[0388] u=(ux,uy)
[0389] ux=x / len
[0390] uy=y / len
[0391] len=p_len(v)
[0392] A2.4.5 The dot product of two-dimensional vectors pa and pb:
[0393] d=p_dot(pa,pb)=xa*xb+ya*yb
[0394] in,
[0395] pa=(xa,ya)
[0396] pb=(xb,yb)
[0397] The dot product of vectors is of great significance.
[0398] Assume that the angle between vectors pa and pb is theta
[0399] So
[0400] cos(theta)=p_dot(pa,pb)
[0401] A3. Lines in Two-Dimensional Space
[0402] A3.1 Two-point form of a line in two-dimensional space
[0403] like Figure 3 As shown, a line L in two-dimensional space can be represented by the two points p0 and p1 it passes through:
[0404] The linear equation of L is:
[0405] L:y=m*x+b
[0406] in,
[0407] m=(y1-y0) / (x1-x0)
[0408] b=y0-x0*(y1-y0) / (x1-x0)
[0409] p0=(x0,y0)
[0410] p1=(x1,y1)
[0411] A3.22 The intersection point p of the lines L0 and L1
[0412] Given the linear equations of L0 and L1:
[0413] L0:y=m0*x+b0
[0414] L1:y=m1*x+b1
[0415] Their intersection point can be calculated using the following formula:
[0416] p=(x,y)
[0417] in,
[0418] x=(b1-m1) / (m0-b0)
[0419] y=(m0*b1-b0*m1) / (m0-b0b)
[0420] Note: If a == b, then L0 and L1 are parallel.
[0421] A4. Ellipses in Two-Dimensional Space
[0422] A4.1 As Figure 4As shown, the ellipse E can be represented by any three of the following two pairs of vertices:
[0423] Vertex: p0, p1
[0424] Minor axis vertices: p2, p3
[0425] A4.2 The center of the ellipse, pc, can be calculated using the following formula:
[0426] pc=(x,y)
[0427] x=(x0+x1) / 2or(x2+x3) / 2
[0428] y=(y0+y1) / 2or(y2+y3) / 2
[0429] in,
[0430] p0=(x0,y0)
[0431] p1=(x1,y1)
[0432] p2=(x2,y2)
[0433] p3=(x3,y3)
[0434] A4.3 Semi-major axis a and semi-minor axis b:
[0435] a = q_dist(p0, p1)
[0436] b = q_dist(p2, p3)
[0437] Although the present invention discloses various aspects and embodiments, other aspects and embodiments will be apparent to those skilled in the art. The various aspects and embodiments disclosed in this disclosure are for illustration only and are not intended to be limiting, with the true scope and spirit of the invention being indicated by the appended claims. Those skilled in the art will recognize that various modifications and / or enhancements may be made to the present teachings.
[0438] While the foregoing has described what is believed to constitute the present teachings and / or other examples, it will be appreciated that various modifications may be made thereto, and that the subject matter disclosed herein may be implemented in a variety of forms and examples, and that the teachings may be applied in numerous applications, only some of which are described herein. It is intended that the appended claims claim any and all applications, modifications, and variations that fall within the true scope of the present teachings.
Claims
1. A method for determining a positional relationship between an eye center and an imaging sensor, wherein the imaging sensor faces a user's eye at a fixed position and orientation relative to the user's head, the method comprising: Acquire a first image and a second image of the user's eyes using an imaging sensor; extracting a first ellipse of the outline of the pupil or the limbus from the first image, and extracting a second ellipse of the outline of the pupil or the limbus from the second image; Determine an intersection position of a first straight line passing through two minor axis vertices of the first ellipse and a second straight line passing through two minor axis vertices of the second ellipse, wherein the intersection position is a position of an eyeball center in a two-dimensional image plane of the imaging sensor; A positional relationship between a center of the eye and an imaging sensor is determined based on the contour and the position.
2. An apparatus for determining a positional relationship between an eye center and an imaging sensor, wherein the imaging sensor faces a user's eye at a fixed position and orientation relative to the user's head, the apparatus comprising: an imaging sensor configured to acquire a first image and a second image of a user's eye; A computing device includes a processor configured to extract a first ellipse of the outline of the pupil or the corneal limbus from a first image, and extract a second ellipse of the outline of the pupil or the corneal limbus from a second image, to determine an intersection position of a first straight line passing through two minor axis vertices of the first ellipse and a second straight line passing through two minor axis vertices of the second ellipse, wherein the intersection position is the position of the eyeball center of the eye in a two-dimensional image plane of an imaging sensor, and to determine a positional relationship between the eyeball center and the imaging sensor based on the outline and the position.
Citation Information
Patent Citations
Information processing method and electronic equipment
CN106325510A
Image-based head position tracking method and system
US20130083976A1