Line-of-sight positioning method, head-mounted display device, computer device, and storage medium

By determining the coordinate transformation relationship between the image and the display screen in the eye-tracking system, the pupil center is transformed into the relative light spot position, which solves the accuracy problem of the eye-tracking system when sliding, realizes efficient and accurate eye-tracking positioning, and improves the stability and convenience of the system.

CN114967904BActive Publication Date: 2025-10-28BEIJING BOE OPTOELECTRONCIS TECH CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202110188536.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-02-19
Publication Date
2025-10-28
Estimated Expiration
2041-02-19

AI Technical Summary

Technical Problem

When the user's eye slides relative to the device, the pupil center is used as the model input, which reduces the accuracy of the gaze point calculation and affects the user experience. In addition, multiple calibrations are required before each use, which is inefficient and inconvenient.

Method used

By determining the transformation relationship between the image coordinate system and the light source coordinate system, as well as between the light source coordinate system and the display screen coordinate system, the line of sight is located by using the position of the pupil center relative to the light spot, thus avoiding the influence of relative sliding and eliminating the calibration process.

Benefits of technology

It achieves efficient and accurate gaze positioning, improves the accuracy, efficiency, stability and convenience of the gaze tracking system, avoids the decrease in accuracy caused by relative sliding, and requires no calibration.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114967904B_ABST
    Figure CN114967904B_ABST
Patent Text Reader

Abstract

The embodiments of the present application disclose a gaze positioning method, a head-mounted display device, a computer device, and a storage medium. In one embodiment, the gaze positioning method is used for a head-mounted display device including a display screen, a camera, and a light source. The method includes: determining a first transformation relationship and a second transformation relationship, wherein the first transformation relationship is the transformation relationship between an image coordinate system and a light source coordinate system, and the second transformation relationship is the transformation relationship between the light source coordinate system and a display screen coordinate system; calculating the coordinates of the pupil center in the image coordinate system based on an eye image of the user's eye captured by the camera; and obtaining the coordinates of the pupil center in the display screen coordinate system as the gaze point position based on the coordinates of the pupil center in the image coordinate system, the first transformation relationship, and the second transformation relationship. This embodiment can accurately and efficiently achieve gaze positioning without the need for calibration, thereby improving the accuracy, efficiency, stability, and convenience of the gaze tracking system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of display technology. More specifically, it relates to a gaze positioning method, a head-mounted display device, a computer device, and a storage medium. Background Technology

[0002] With the development of VR (Virtual Reality) technology, eye-tracking technology, or gaze tracking technology, has gained attention for its applications in VR interaction, foveated rendering, and other areas.

[0003] Currently, gaze tracking systems typically use multinomial regression models or 3D geometric models to calculate gaze points, with the pupil center in the eye image as the input. Before using the gaze tracking system, multi-point calibration, such as 5-point or 9-point calibration, is performed on the user to determine the model parameters suitable for that user. This approach has several drawbacks. First, it requires at least one multi-point calibration before each use, which is inefficient and inconvenient. Second, in actual use, when the user's eyes slide relative to the initial calibration state (e.g., relative sliding between the user's eyes and the VR headset), the gaze point error increases because the pupil center is an absolute value. If the pupil center is still used as the model input, the calculated gaze point will drift significantly, resulting in a severe decrease in gaze point calculation accuracy and negatively impacting the user experience. Summary of the Invention

[0004] The purpose of this application is to provide a gaze positioning method, a head-mounted display device, a computer device, and a storage medium to solve at least one of the problems existing in the prior art.

[0005] To achieve the above objectives, this application adopts the following technical solution:

[0006] The first aspect of this application provides a gaze positioning method for a head-mounted display device, the head-mounted display device including a display screen, a camera, and a light source, the method comprising:

[0007] Determine the first transformation relationship and the second transformation relationship. The first transformation relationship is the transformation relationship between the image coordinate system and the light source coordinate system, and the second transformation relationship is the transformation relationship between the light source coordinate system and the display screen coordinate system.

[0008] Calculate the coordinates of the pupil center in the image coordinate system based on the eye image of the user captured by the camera;

[0009] Based on the coordinates of the pupil center in the image coordinate system, the first transformation relationship, and the second transformation relationship, the coordinates of the pupil center, which serves as the fixation point, in the display screen coordinate system are obtained.

[0010] The gaze localization method provided in the first aspect of this application transforms the pupil center's position into a relative value with respect to the light spot formed by the light source on the user's eye by performing two coordinate system transformations. This relative value is only related to changes in the gaze point and is independent of relative slippage (such as slippage between the user's head and the VR headset). Therefore, it avoids the decrease in gaze localization accuracy caused by relative slippage. Furthermore, since it uses the position of the relative light spot at the pupil center for gaze localization or gaze point calculation, calibration is unnecessary. In summary, the gaze localization method provided in the first aspect of this application can achieve accurate and efficient gaze localization without requiring calibration, thus improving the accuracy, efficiency, stability, and convenience of the gaze tracking system.

[0011] Optionally, determining the first transformation relationship includes:

[0012] Multiple light spots are formed in the user's eye using multiple light sources, and the camera is used to capture an eye image of the user's eye containing the multiple light spots;

[0013] The coordinates of each light spot in the user's eye image containing the multiple light spots are calculated in the image coordinate system, and the coordinates of each light source in the light source coordinate system are determined according to the setting position of the multiple light sources.

[0014] Based on the coordinates of each light spot in the image coordinate system and the corresponding coordinates of the light source in the light source coordinate system, the first transformation matrix between the image coordinate system and the light source coordinate system is calculated to determine the first transformation relationship.

[0015] This optional method can accurately and efficiently determine the transformation relationship between the image coordinate system and the light source coordinate system.

[0016] Optionally, the first transformation matrix includes linear transformation coefficients and translation transformation coefficients.

[0017] Optionally, determining the second transformation relationship includes:

[0018] The coordinates of each light source in the light source coordinate system are determined based on the setting positions of the multiple light sources, and the coordinates of each light source in the display screen coordinate system are determined based on the relative positional relationship between the setting positions of the multiple light sources and the setting position of the display screen.

[0019] Based on the coordinates of each light source in the light source coordinate system and the corresponding coordinates of the light source in the display screen coordinate system, a second transformation matrix between the light source coordinate system and the display screen coordinate system is calculated to determine the second transformation relationship.

[0020] This optional method can accurately and efficiently determine the transformation relationship between the light source coordinate system and the display screen coordinate system.

[0021] Optionally, the second transformation matrix includes linear transformation coefficients and translation transformation coefficients.

[0022] A second aspect of this application provides a head-mounted display device, comprising:

[0023] Multiple light sources, cameras, displays, and processors;

[0024] The camera is used to capture images of the user's eyes under the illumination of the light source;

[0025] The processor is configured to determine a first transformation relationship and a second transformation relationship, wherein the first transformation relationship is a transformation relationship between the image coordinate system and the light source coordinate system, and the second transformation relationship is a transformation relationship between the light source coordinate system and the display screen coordinate system; calculate the coordinates of the pupil center in the image coordinate system based on the eye image of the user's eye; and obtain the coordinates of the pupil center in the display screen coordinate system, which serves as the fixation point, based on the coordinates of the pupil center in the image coordinate system, the first transformation relationship, and the second transformation relationship.

[0026] The head-mounted display device provided in the second aspect of this application transforms the pupil center position into a relative value with respect to the light spot formed by the light source on the user's eye by performing two coordinate system transformations. This relative value is only related to changes in the gaze point and is independent of relative slippage (such as slippage between the user's head and the VR head-mounted display device). Therefore, it avoids the decrease in gaze positioning accuracy caused by relative slippage. Furthermore, since it uses the position of the relative light spot at the pupil center for gaze positioning or gaze point calculation, calibration is not required. In summary, the head-mounted display device provided in the second aspect of this application can achieve accurate and efficient gaze positioning without calibration, improving the accuracy, efficiency, stability, and convenience of the gaze tracking system.

[0027] Optionally,

[0028] The multiple light sources are also used to form multiple light spots in the user's eyes;

[0029] The camera is also used to capture eye images of the user's eye containing the multiple light spots;

[0030] The processor is configured to determine the first transformation relationship by: calculating the coordinates of each light spot in the eye image containing the plurality of light spots in the image coordinate system, and determining the coordinates of each light source in the light source coordinate system according to the setting positions of the plurality of light sources; and calculating the first transformation matrix between the image coordinate system and the light source coordinate system according to the coordinates of each light spot in the image coordinate system and the coordinates of the corresponding light source in the light source coordinate system, so as to determine the first transformation relationship.

[0031] This optional method can accurately and efficiently determine the transformation relationship between the image coordinate system and the light source coordinate system.

[0032] Optionally, in a plane parallel to the display surface of the screen, the plurality of light sources are evenly distributed circumferentially around a preset position on the plane.

[0033] Optionally, the processor is configured to determine the second transformation relationship by: determining the coordinates of each light source in the light source coordinate system based on the setting positions of the multiple light sources, and determining the coordinates of each light source in the display screen coordinate system based on the relative positional relationship between the setting positions of the multiple light sources and the setting position of the display screen; and calculating a second transformation matrix between the light source coordinate system and the display screen coordinate system based on the coordinates of each light source in the light source coordinate system and the corresponding coordinates of the light source in the display screen coordinate system, so as to determine the second transformation relationship.

[0034] This optional method can accurately and efficiently determine the transformation relationship between the light source coordinate system and the display screen coordinate system.

[0035] Optionally, the head-mounted display device is a virtual reality head-mounted display device.

[0036] A third aspect of this application provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the line-of-sight positioning method provided in the first aspect of this application.

[0037] The fourth aspect of this application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the line-of-sight positioning method provided in the first aspect of this application.

[0038] The beneficial effects of this application are as follows:

[0039] The technical solution described in this application transforms the pupil center's position into a relative value with respect to the light spot formed by the light source on the user's eye by performing two coordinate system transformations. This relative value is only related to changes in the gaze point and is independent of relative slippage (such as slippage between the user's head and the VR headset). Therefore, it avoids the decrease in gaze positioning accuracy caused by relative slippage. Furthermore, by using the position of the gaze point or relative light spot based on the pupil center, calibration is unnecessary. In summary, the technical solution described in this application can achieve accurate and efficient gaze positioning without calibration, improving the accuracy, efficiency, stability, and convenience of the gaze tracking system. Attached Figure Description

[0040] The specific embodiments of this application will be described in further detail below with reference to the accompanying drawings.

[0041] Figure 1 A flowchart illustrating the line-of-sight positioning method provided in an embodiment of this application is shown.

[0042] Figure 2 A schematic diagram of the structure of a VR head-mounted display device is shown.

[0043] Figure 3 A schematic diagram showing the correspondence between the image coordinate system and the infrared light source coordinate system.

[0044] Figure 4 A schematic diagram showing the correspondence between the infrared light source coordinate system and the display screen coordinate system.

[0045] Figure 5 A schematic diagram of the computer system structure is shown. Detailed Implementation

[0046] To more clearly illustrate this application, the following description, in conjunction with embodiments and accompanying drawings, further clarifies the application. Similar components in the drawings are indicated by the same reference numerals. Those skilled in the art should understand that the specific description below is illustrative rather than restrictive and should not be construed as limiting the scope of protection of this application.

[0047] Existing gaze tracking solutions for head-mounted displays, such as VR headsets, typically employ multinomial regression models or 3D geometric models to calculate gaze points, with the pupil center in the eye image as the model input. Before using the gaze tracking system, multi-point calibration, such as 5-point or 9-point calibration, is performed on the user to determine the model parameters suitable for that user. This approach has several drawbacks. First, it requires at least one multi-point calibration before each use, resulting in low efficiency and ease of use. Second, in actual use, when the user's eyes slide relative to the initial calibration state (e.g., relative sliding between the user's eyes and the VR headset), the gaze point error increases due to the absolute value of the pupil center. If the pupil center is still used as the model input in this situation, the calculated gaze point will drift significantly, severely reducing the accuracy of gaze point calculation and impacting the user experience.

[0048] Therefore, taking a VR head-mounted display device including a display screen, an infrared camera, and an infrared light source as an example, one embodiment of this application provides a gaze positioning method, such as... Figure 1 As shown, the method includes the following steps:

[0049] S110. Determine the first transformation relationship and the second transformation relationship. The first transformation relationship is the transformation relationship between the image coordinate system and the infrared light source coordinate system, and the second transformation relationship is the transformation relationship between the infrared light source coordinate system and the display screen coordinate system.

[0050] In a specific example, such as Figure 2 As shown, the VR head-mounted display device involved in this embodiment includes, for example, a housing and a VR display screen (…). Figure 2 (Not shown in the image) Two infrared cameras, two lenses, and multiple infrared light sources. The two lenses, for example, are Fresnel lenses, positioned directly in front of the left and right eyes of the user wearing the VR headset. Multiple infrared light sources are arranged in a plane parallel to the display surface of the screen. On one hand, this provides uniform illumination to the eyes, facilitating a clearer image of the eyes and thus helping to separate the pupil from the iris region. On the other hand, it forms multiple light spots in the user's eye area, obtaining a light spot image and providing a positional reference for gaze point calculation. The number of infrared light sources can be 6, 8, 10, 12, etc., for example... Figure 2 The two lenses shown are symmetrically arranged horizontally with the center of the display screen as the midpoint. The VR head-mounted display device includes twelve infrared light sources, six circumferentially distributed outside the left lens and six circumferentially distributed outside the right lens. Figure 2 The infrared LEDs 1 to 6 are arranged circumferentially around the center of the right lens to ensure both effective supplementary lighting and accurate gaze positioning. Two infrared cameras are positioned below the two lenses to capture images of the user's left and right eyes, including the pupils and light spots, under infrared illumination when the user is wearing the VR headset. Alternatively, only one infrared camera can be used to capture images of the left and right eyes for subsequent gaze positioning. It should be noted that the infrared light source is chosen in this embodiment because it does not interfere with the user's viewing of the display screen. However, it is understood that other types of light sources and corresponding cameras can also be used in this VR headset.

[0051] In the subsequent examples, only the coordinate system transformation relationship and pupil center position mapping of the right eye images captured by the infrared camera on the right under the illumination of infrared LED1 to infrared LED6 are explained.

[0052] In one possible implementation, determining the first transformation relationship (the transformation relationship between the image coordinate system and the infrared light source coordinate system) includes:

[0053] Multiple infrared light sources are used to form multiple light spots in the user's eye, and the infrared camera is used to capture an eye image of the user's eye containing the multiple light spots;

[0054] The coordinates of each light spot in the image coordinate system of the user's eye image containing the multiple light spots are calculated, and the coordinates of each infrared light source in the infrared light source coordinate system are determined according to the setting position of the multiple infrared light sources.

[0055] Based on the coordinates of each light spot in the image coordinate system and the corresponding coordinates of the infrared light source in the infrared light source coordinate system, the first transformation matrix between the image coordinate system and the infrared light source coordinate system is calculated to determine the first transformation relationship.

[0056] In this way, the transformation relationship between the image coordinate system and the infrared light source coordinate system can be determined accurately and efficiently, so that in subsequent steps, the coordinates of the pupil center in the image coordinate system can be accurately mapped to the coordinates of the pupil center in the infrared light source coordinate system.

[0057] In one possible implementation, the first transformation matrix includes linear transformation coefficients and translation transformation coefficients.

[0058] Continuing with the previous example, such as Figure 3 As shown, the coordinates of the centers of infrared LEDs 1 to 6 in the infrared light source coordinate system are [(x1, y1), ..., (x6, y6)], which can be obtained based on the preset installation positions of the six infrared light sources. See [link to relevant documentation]. Figure 3 Furthermore, in this embodiment, the infrared light source coordinate system refers to the Cartesian coordinate system within the plane containing the infrared light source. Infrared LEDs 1 to 6 will form six light spots in the user's right eye. These light spots (infrared spots) have high grayscale values ​​after imaging on the human cornea. Therefore, an adaptive threshold algorithm can be used to segment the light spots in the eye image of the user's eye containing the six light spots captured by the infrared camera. Then, connected regions are obtained to obtain the connected regions of the six light spots. By finding the centroids of these connected regions, the coordinates of the centers of the six light spots formed by infrared LEDs 1 to 6 in the image coordinate system are obtained, which are [(x1', y1'), ..., (x6', y6')]. See [link to relevant documentation]. Figure 3 To further clarify, in this embodiment, the image coordinate system refers to the Cartesian coordinate system of the eye image.

[0059] The relationship between the coordinates of the centers of infrared LEDs 1 to 6 in the infrared light source coordinate system [(x1, y1), ..., (x6, y6)] and the coordinates of the centers of the six light spots in the image coordinate system [(x1', y1'), ..., (x6', y6')] can be expressed as follows: [(x1, y1), ..., (x6, y6)] are obtained by linear transformation and translation transformation of [(x1', y1'), ..., (x6', y6')]. For example:

[0060] x i=m11*x′ i +n12*y′ i +m13

[0061] y i =m21*x′ i +m22*y′ i +m23

[0062] Represented in matrix form as follows:

[0063]

[0064] Among them, (x i ,y i (x′) is the coordinate of the center of the infrared LED in the infrared light source coordinate system. i ,y′ i Let be the coordinates of the center of the corresponding light spot in the image coordinate system. Then, matrix A, as the first transformation matrix between the image coordinate system and the infrared light source coordinate system, represents the first transformation relationship (the transformation relationship between the image coordinate system and the infrared light source coordinate system). Among them, m11, m12, m21, and m22 are the linear transformation coefficients of matrix A, and m13 and m23 are the translation transformation coefficients of matrix A.

[0065] It should be noted that the first transformation relationship can be determined when a user first uses the VR headset and stored in the VR headset's memory, or it can be determined by testers before the VR headset leaves the factory. When subsequent steps require mapping the pupil center position, the VR headset's processor can directly read it from the memory. Since the relative positions of the infrared camera and the infrared light source remain unchanged, the transformation relationship between the light spot and the image coordinate system and the infrared light source coordinate system remains constant during the user's use of the VR headset, regardless of relative sliding (sliding between the user's head and the VR headset) or a change of user. Therefore, it is not necessary to redetermine the first transformation relationship.

[0066] In one possible implementation, determining the second transformation relationship (the transformation relationship between the infrared light source coordinate system and the display screen coordinate system) includes:

[0067] The coordinates of each infrared light source in the infrared light source coordinate system are determined based on the setting positions of the multiple infrared light sources, and the coordinates of each infrared light source in the display screen coordinate system are determined based on the relative positional relationship between the setting positions of the multiple infrared light sources and the setting position of the display screen.

[0068] Based on the coordinates of each infrared light source in the infrared light source coordinate system and the corresponding coordinates of the infrared light source in the display screen coordinate system, a second transformation matrix between the infrared light source coordinate system and the display screen coordinate system is calculated to determine the second transformation relationship.

[0069] In this way, the transformation relationship between the infrared light source coordinate system and the display screen coordinate system can be determined accurately and efficiently, thereby enabling the precise mapping of the pupil center position from the infrared light source coordinate system to the display screen coordinate system in subsequent steps.

[0070] In one possible implementation, the second transformation matrix includes linear transformation coefficients and translation transformation coefficients.

[0071] Continuing with the previous example, such as Figure 4 As shown, the coordinates of the centers of infrared LEDs 1 to 6 in the infrared light source coordinate system are [(x1, y1), ..., (x6, y6)], which can be obtained based on the preset installation positions of the six infrared light sources. The coordinates of the centers of infrared LEDs 1 to 6 in the display screen coordinate system are [(x1, y1), ..., (x6, y6)], which can be obtained based on the relative positional relationship between the preset installation positions of the six infrared light sources and the preset installation position of the display screen. See [link to relevant documentation]. Figure 4 To further clarify, in this embodiment, the display screen coordinate system refers to the rectangular coordinate system within the display surface of the display screen.

[0072] The relationship between the coordinates [(X1, Y1), ..., (X6, Y6)] of the centers of infrared LEDs 1 to 6 in the display screen coordinate system and the coordinates [(x1, y1), ..., (x6, y6)] of the centers of infrared LEDs 1 to 6 in the infrared light source coordinate system can be expressed as follows: [(X1, Y1), ..., (X6, Y6)] are obtained by linear transformation and translation transformation of [(x1, y1), ..., (x6, y6)]. For example:

[0073] X i =n11*x i +n12*y i +n13

[0074] Y i =n21*x i +n22*y i +n23

[0075] Represented in matrix form as follows:

[0076]

[0077] Among them, (x i ,y i(X) is the coordinate of the center of the infrared LED in the infrared light source coordinate system. i ,Y i Let be the coordinates of the center of the infrared LED in the display screen coordinate system. Then, matrix B, as the second transformation matrix between the infrared light source coordinate system and the display screen coordinate system, represents the second transformation relationship (the transformation relationship between the infrared light source coordinate system and the display screen coordinate system). Among them, n11, n12, n21, and n22 are the linear transformation coefficients of matrix B, and n13 and n23 are the translation transformation coefficients of matrix B.

[0078] It should be noted that the second transformation relationship can be predetermined, for example, determined before the VR headset leaves the factory and stored in the VR headset's memory. When subsequent steps require mapping the pupil center position, the VR headset's processor can directly read it from the memory. Since the relative position of the display screen and the infrared light source remains unchanged, the transformation relationship between the infrared light source coordinate system and the display screen coordinate system remains constant during the user's use of the VR headset, regardless of relative sliding (sliding between the user's head and the VR headset) or a change of user. Therefore, it is not necessary to redetermine the second transformation relationship.

[0079] S120. Calculate the coordinates of the pupil center in the image coordinate system based on the eye image of the user captured by the infrared camera.

[0080] In a specific example, an adaptive threshold segmentation algorithm can be used to separate the pupil region from the eye image captured by an infrared camera. Then, the Canny edge detection algorithm is used to detect the pupil edge. The detected edge points are used to perform ellipse fitting, and the fitted pupil center value in the image coordinate system is obtained as (p′). x ,p′ y ).

[0081] S130. Based on the coordinates of the pupil center in the image coordinate system, the first transformation relationship, and the second transformation relationship, obtain the coordinates of the pupil center, which serves as the fixation point, in the display screen coordinate system.

[0082] Continuing with the previous example, we can first determine the value of the pupil center in the image coordinate system as (p′) based on the first transformation relationship. x ,p′ y Mapped to coordinates in the infrared light source coordinate system (p) x ,p y ):

[0083]

[0084] Then, according to the second transformation relationship, the coordinates (p) of the pupil center in the infrared light source coordinate system are determined. x,p y The mapping is the coordinates of the pupil center in the display screen coordinate system (P). x ,P y ):

[0085]

[0086] At this point, the position of the gaze point on the display screen is obtained, and the gaze localization is completed.

[0087] It should be noted that steps S120 and S130 are performed in real time during the line-of-sight positioning process.

[0088] In summary, the gaze positioning method provided in this embodiment achieves gaze positioning by performing two coordinate system transformations on the pupil center coordinates: mapping the pupil center position from the image coordinate system to the infrared light source coordinate system, and then mapping it to the display screen coordinate system. This yields the coordinates of the pupil center, which serves as the gaze point, in the display screen coordinate system, thus realizing gaze positioning. It is evident that the gaze positioning method provided in this embodiment transforms the pupil center position into a relative value with respect to the light spot formed by the infrared light source on the user's eye. This relative value is only related to changes in the gaze point and is independent of relative slippage (such as slippage between the user's head and the VR headset). Therefore, it avoids the decrease in gaze positioning accuracy caused by relative slippage. Furthermore, since the method uses the position of the pupil center relative to the light spot for gaze positioning or gaze point calculation, calibration is unnecessary. Since the relative positions of the infrared camera and the infrared light source remain unchanged (this relationship does not change even when the user's head slides between the VR headset or when a different user switches VR headsets), the transformation relationship of the light spot between the image coordinate system and the infrared light source coordinate system is constant. Similarly, since the relative positions of the display screen and the infrared light source remain unchanged, the transformation relationship of the light spot between the infrared light source coordinate system and the display screen coordinate system is also constant. Therefore, by calculating the coordinates of the pupil center in the image coordinate system based on the real-time captured images of the user's eyes, the gaze localization result can be obtained accurately and efficiently based on the two transformation relationships. In summary, the gaze localization method provided in this embodiment can achieve gaze localization accurately and efficiently without calibration, improving the accuracy, efficiency, stability, and convenience of the gaze tracking system.

[0089] The above embodiments are only used as examples of VR head-mounted display devices to illustrate the gaze positioning method provided in this application. Those skilled in the art will understand that, in addition to VR head-mounted display devices, the gaze positioning method provided in this application can also be applied to other head-mounted display devices such as AR head-mounted display devices.

[0090] Another embodiment of this application provides a head-mounted display device, including: multiple infrared light sources, an infrared camera, a display screen, and a processor;

[0091] The infrared camera is used to capture images of the user's eyes under the illumination of the infrared light source.

[0092] The processor is configured to determine a first transformation relationship and a second transformation relationship, wherein the first transformation relationship is a transformation relationship between the image coordinate system and the infrared light source coordinate system, and the second transformation relationship is a transformation relationship between the infrared light source coordinate system and the display screen coordinate system; calculate the coordinates of the pupil center in the image coordinate system based on the eye image of the user's eye; and obtain the coordinates of the pupil center in the display screen coordinate system as the fixation point position based on the coordinates of the pupil center in the image coordinate system, the first transformation relationship, and the second transformation relationship.

[0093] In one possible implementation,

[0094] The multiple infrared light sources are also used to form multiple light spots in the user's eyes;

[0095] The infrared camera is also used to acquire eye images of the user's eye containing the multiple light spots;

[0096] The processor is configured to determine the first transformation relationship by: calculating the coordinates of each light spot in the eye image containing the plurality of light spots in the image coordinate system, and determining the coordinates of each infrared light source in the infrared light source coordinate system according to the setting positions of the plurality of infrared light sources; and calculating the first transformation matrix between the image coordinate system and the infrared light source coordinate system according to the coordinates of each light spot in the image coordinate system and the coordinates of the corresponding infrared light source in the infrared light source coordinate system, so as to determine the first transformation relationship.

[0097] In one possible implementation, the plurality of infrared light sources are evenly distributed circumferentially around a predetermined position on a plane parallel to the display surface of the screen. For example, in a VR head-mounted display device, the plurality of infrared light sources are evenly distributed circumferentially around the lens of the VR head-mounted display device, with the lens center as the center.

[0098] In one possible implementation, the processor, for determining the second transformation relationship, includes: determining the coordinates of each infrared light source in the infrared light source coordinate system based on the placement positions of the plurality of infrared light sources, and determining the coordinates of each infrared light source in the display screen coordinate system based on the relative positional relationship between the placement positions of the plurality of infrared light sources and the placement position of the display screen; and calculating a second transformation matrix between the infrared light source coordinate system and the display screen coordinate system based on the coordinates of each infrared light source in the infrared light source coordinate system and the corresponding coordinates of the infrared light source in the display screen coordinate system, so as to determine the second transformation relationship.

[0099] In one possible implementation, the head-mounted display device is a VR head-mounted display device.

[0100] It should be noted that the principle and working process of the head-mounted display device provided in this embodiment are similar to the above-described gaze positioning method. Taking the head-mounted display device as a VR head-mounted display device as an example, its structure is similar to that of the VR head-mounted display device exemplified in the above embodiment. The relevant parts can be referred to the above description, and will not be repeated here.

[0101] like Figure 5 As shown, a computer system suitable for executing the line-of-sight positioning method provided in the above embodiments includes a central processing module (CPU), which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) or a program loaded from a storage portion into a random access memory (RAM). Various programs and data required for the operation of the computer system are also stored in the RAM. The CPU, ROM, and RAM are connected to each other via a bus. An input / output (I / O) interface is also connected to the bus.

[0102] The following components are connected to the I / O interface: input sections including keyboards, mice, etc.; output sections including liquid crystal displays (LCDs) and speakers, etc.; storage sections including hard disks, etc.; and communication sections including network interface cards such as LAN cards and modems, etc. The communication sections perform communication processing via networks such as the Internet. Drives are also connected to the I / O interface as needed. Removable media, such as disks, optical disks, magneto-optical disks, semiconductor memories, etc., are installed on the drives as needed so that computer programs read from them can be installed into the storage sections as required.

[0103] Specifically, according to this embodiment, the process described in the flowchart above can be implemented as a computer software program. For example, this embodiment includes a computer program product comprising a computer program tangibly embodied on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network via a communication component, and / or installed from a removable medium.

[0104] The flowcharts and schematic diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of the system, method, and computer program product of this embodiment. In this regard, each block in the flowchart or schematic diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than those shown in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the schematic diagram and / or flowchart, and combinations of blocks in the schematic diagram and / or flowchart, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0105] On the other hand, this embodiment also provides a non-volatile computer storage medium. This non-volatile computer storage medium can be the non-volatile computer storage medium included in the device described in the above embodiments, or it can be a separate non-volatile computer storage medium not assembled into the terminal. The non-volatile computer storage medium stores one or more programs. When the one or more programs are executed by a device, the device causes the device to:

[0106] Determine the first transformation relationship and the second transformation relationship. The first transformation relationship is the transformation relationship between the image coordinate system and the infrared light source coordinate system, and the second transformation relationship is the transformation relationship between the infrared light source coordinate system and the display screen coordinate system.

[0107] Calculate the coordinates of the pupil center in the image coordinate system based on the eye image of the user captured by the infrared camera;

[0108] Based on the coordinates of the pupil center in the image coordinate system, the first transformation relationship, and the second transformation relationship, the coordinates of the pupil center, which serves as the fixation point, in the display screen coordinate system are obtained.

[0109] In the description of this application, it should be noted that the terms "upper," "lower," etc., indicating the orientation or positional relationship are based on the orientation or positional relationship shown in the accompanying drawings, and are only for the convenience of describing this application and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of this application. Unless otherwise expressly specified and limited, the terms "installed," "connected," and "linked" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; they can refer to the internal communication between two elements. For those skilled in the art, the specific meaning of the above terms in this application can be understood according to the specific circumstances.

[0110] It should also be noted that, in the description of this application, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0111] Obviously, the above embodiments of this application are merely examples for clearly illustrating this application, and are not intended to limit the implementation of this application. For those skilled in the art, other variations or modifications can be made based on the above description. It is impossible to exhaustively list all implementation methods here. Any obvious variations or modifications derived from the technical solutions of this application are still within the protection scope of this application.

Claims

1. A gaze positioning method for a head-mounted display device, the head-mounted display device comprising a display screen, a camera, and a light source, characterized in that, The relative positions of the camera and the light source are fixed, and the relative positions of the light source and the display screen are fixed. The method includes: Determine the first transformation relationship and the second transformation relationship. The first transformation relationship is the transformation relationship between the image coordinate system and the light source coordinate system, and the second transformation relationship is the transformation relationship between the light source coordinate system and the display screen coordinate system. Calculate the coordinates of the pupil center in the image coordinate system based on the eye image of the user captured by the camera; Based on the coordinates of the pupil center in the image coordinate system, the first transformation relationship, and the second transformation relationship, the coordinates of the pupil center, which serves as the fixation point, in the display screen coordinate system are obtained. Determining the first transformation relationship includes: Multiple light spots are formed in the user's eye using multiple light sources, and the camera is used to capture an eye image of the user's eye containing the multiple light spots; The coordinates of each light spot in the user's eye image containing the multiple light spots are calculated in the image coordinate system, and the coordinates of each light source in the light source coordinate system are determined according to the setting position of the multiple light sources. Based on the coordinates of each light spot in the image coordinate system and the corresponding light source in the light source coordinate system, a first transformation matrix between the image coordinate system and the light source coordinate system is calculated to determine the first transformation relationship. The first transformation matrix includes linear transformation coefficients and translation transformation coefficients. The relationship between the image coordinate system, the light source coordinate system, and the first transformation matrix is ​​expressed as follows: in, This represents the coordinates of the center of the light source in the light source coordinate system. The coordinates of the center of the corresponding light spot in the image coordinate system are represented by A, which represents the first transformation matrix. m11, m12, m21, and m22 are the linear transformation coefficients of the first transformation matrix A, and m13 and m23 are the translation transformation coefficients of the first transformation matrix A. Determining the second transformation relationship includes: The coordinates of each light source in the light source coordinate system are determined based on the setting positions of the multiple light sources, and the coordinates of each light source in the display screen coordinate system are determined based on the relative positional relationship between the setting positions of the multiple light sources and the setting position of the display screen. Based on the coordinates of each light source in the light source coordinate system and the corresponding coordinates of the light source in the display screen coordinate system, a second transformation matrix between the light source coordinate system and the display screen coordinate system is calculated to determine the second transformation relationship.

2. The method according to claim 1, characterized in that, The second transformation matrix includes linear transformation coefficients and translation transformation coefficients.

3. A head-mounted display device, characterized in that, include: The system includes multiple light sources, a camera, a display screen, and a processor, wherein the relative positions of the camera and the light source are fixed, and the relative positions of the light source and the display screen are fixed. The camera is used to capture images of the user's eyes under the illumination of the light source; The processor is configured to determine a first transformation relationship and a second transformation relationship, wherein the first transformation relationship is a transformation relationship between the image coordinate system and the light source coordinate system, and the second transformation relationship is a transformation relationship between the light source coordinate system and the display screen coordinate system; calculate the coordinates of the pupil center in the image coordinate system based on the eye image of the user's eye; and obtain the coordinates of the pupil center in the display screen coordinate system, which serves as the fixation point, based on the coordinates of the pupil center in the image coordinate system, the first transformation relationship, and the second transformation relationship. The multiple light sources are also used to form multiple light spots in the user's eyes; The camera is also used to capture eye images of the user's eye containing the multiple light spots; The processor is configured to determine the first transformation relationship by: calculating the coordinates of each light spot in the eye image containing the plurality of light spots in the image coordinate system, and determining the coordinates of each light source in the light source coordinate system according to the setting positions of the plurality of light sources; and calculating a first transformation matrix between the image coordinate system and the light source coordinate system according to the coordinates of each light spot in the image coordinate system and the coordinates of the corresponding light source in the light source coordinate system, so as to determine the first transformation relationship, wherein the first transformation matrix includes linear transformation coefficients and translation transformation coefficients. The relationship between the image coordinate system, the light source coordinate system, and the first transformation matrix is ​​expressed as follows: in, This represents the coordinates of the center of the light source in the light source coordinate system. The coordinates of the center of the corresponding light spot in the image coordinate system are represented by A, which represents the first transformation matrix. m11, m12, m21, and m22 are the linear transformation coefficients of the first transformation matrix A, and m13 and m23 are the translation transformation coefficients of the first transformation matrix A. The processor is configured to determine the second transformation relationship by: determining the coordinates of each light source in the light source coordinate system based on the setting positions of the multiple light sources, and determining the coordinates of each light source in the display screen coordinate system based on the relative positional relationship between the setting positions of the multiple light sources and the setting position of the display screen; and calculating the second transformation matrix between the light source coordinate system and the display screen coordinate system based on the coordinates of each light source in the light source coordinate system and the corresponding coordinates of the light source in the display screen coordinate system, so as to determine the second transformation relationship.

4. The head-mounted display device according to claim 3, characterized in that, In a plane parallel to the display surface of the screen, the plurality of light sources are evenly distributed circumferentially around a preset position on the plane.

5. The head-mounted display device according to claim 3, characterized in that, The head-mounted display device is a virtual reality head-mounted display device.

6. A computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the method as described in any one of claims 1-2.

7. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the method as described in any one of claims 1-2.

Citation Information

Patent Citations

  • Equipment control method and related equipment

    CN111061372A

  • System and method for eye gaze tracking using corneal image mapping

    US20030123027A1