A gaze tracking method and system based on 3D pose and 6DoF localization

By combining infrared spot images and inertial information in multi-view geometric calculations, the problem of helmet positioning loss when the helmet leaves the recognition area is solved, realizing the accuracy of real-time helmet positioning and eye tracking, and improving the user experience and tracking accuracy.

CN119440244BActive Publication Date: 2025-10-31SHENZHEN HUAHOM TETHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411208673.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-08-30
Publication Date
2025-10-31
Estimated Expiration
2044-08-30

AI Technical Summary

Technical Problem

When the user moves too much and moves out of the recognition area, the eye-tracking system based on 3D posture and 6DoF positioning will lose recognition, resulting in operation interruption and decreased tracking accuracy.

Method used

By combining infrared spot images acquired by multiple first cameras, the coordinate information of the helmet relative to the cabin is calculated through multi-view geometry principles, and inertial information is used for compensation. A fuzzy learning model is used to predict position errors, ensuring real-time positioning and eye tracking of the helmet.

Benefits of technology

When the helmet leaves the recognition area, the combination of inertial information and optical tracking ensures the helmet's real-time positioning and tracking accuracy, improving the user experience and tracking precision.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119440244B_ABST
    Figure CN119440244B_ABST
Patent Text Reader

Abstract

This invention relates to the field of VR recognition technology, specifically disclosing a gaze tracking method and system based on 3D posture and 6DoF positioning. The method includes: acquiring light spot images of infrared light emitted by infrared lamps at multiple fixed positions within the cabin using multiple first cameras, and calculating the first coordinate information of the helmet relative to the cabin; acquiring the helmet's 6DoF inertial information in real time; when the first coordinate information of the helmet cannot be obtained based on the light spot images, combining the most recently acquired first coordinate information with the corresponding 6DoF inertial information to calculate the latest second coordinate information of the helmet relative to the cabin; using a second camera inside the helmet to photograph the user's eyes, and calculating the third coordinate information of the user's gaze relative to the helmet; and calculating the fourth coordinate information of the user's gaze relative to the cabin. This invention can solve the problem of loss of 3D posture recognition caused by the user leaving the recognition area.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of VR recognition technology, and in particular to a gaze tracking method and system based on 3D pose and 6DoF positioning. Background Technology

[0002] Eye tracking is the process of identifying what someone is looking at and how they are looking, and it has been widely applied in various fields such as human-computer interaction, virtual reality, assisted driving, human factors analysis, and psychological research. From the physiological structure of the eye, humans primarily acquire visual data through the fovea, which provides only about 1–2 degrees of visual field. Although this area occupies only a tiny portion of the visual field, the information recorded through it contains 50% of the effective visual information transmitted to the brain via the optic nerve. Therefore, the human visual and attentional systems work around a primary goal: to focus the optical image of the object of interest onto the fovea. This is the most fundamental and primary reason for eye movement.

[0003] Based on their usage scenarios, eye-tracking devices can be mainly divided into two types: screen-based eye trackers and wearable eye trackers. Screen-based eye trackers are placed at a certain distance from the user to track the user's eye movements. Wearable eye trackers integrate the eye-tracking system and scene camera into a lightweight frame, such as glasses or a helmet, to collect the user's eye movements in a real-world environment. Users can move freely, maximizing the ecological validity of the experiment.

[0004] Therefore, combining wearable eye trackers with gaming pods can create many interesting experiences. The pod can accurately recognize the user's 3D posture, while the wearable eye tracker can recognize the user's gaze; combining the two allows for more virtual operations. However, if the user's movements are too large and they move out of the recognition area, 3D posture recognition will be lost. Summary of the Invention

[0005] In view of the above technical problems, the present invention provides a gaze tracking method and system based on 3D pose and 6DoF positioning to solve the problem of loss of 3D pose recognition caused by the user leaving the recognition area.

[0006] Other features and advantages of this disclosure will become apparent from the following detailed description, or may be learned in part from practice of this disclosure.

[0007] According to one aspect of the present invention, a gaze tracking method based on 3D pose and 6DoF localization is disclosed, the method comprising:

[0008] Multiple first cameras acquire light spot images of infrared light emitted by infrared lamps at multiple fixed positions inside the cabin. The first cameras are mounted on the head-mounted display helmet. The position of the infrared lamp is calculated based on the light spot images. Based on the position of the infrared lamp and the multi-view geometry principle, the first coordinate information of the helmet relative to the cabin is calculated. The first coordinate information includes three-dimensional coordinates and attitude. The first camera is an infrared camera.

[0009] The 6DoF inertial information of the helmet is acquired in real time. When the first coordinate information of the helmet cannot be obtained based on the spot image, the most recently acquired first coordinate information is combined with the corresponding 6DoF inertial information to calculate the latest second coordinate information of the helmet relative to the cabin.

[0010] Based on eye tracking, the user's eyes are captured by a second camera inside the helmet, and the third coordinate information of the user's eye gaze relative to the helmet is calculated.

[0011] By associating the third coordinate information with the first coordinate information or the third coordinate information with the second coordinate information, a fourth coordinate information of the user's line of sight relative to the cabin is obtained.

[0012] Furthermore, the process of reverse-calculating the first coordinate information of the helmet specifically includes:

[0013] Based on image processing algorithms, the pixel coordinates of the light spot are extracted from the light spot image;

[0014] The pixel coordinates captured by multiple first cameras are matched to identify the projection position of the same infrared light in different first cameras;

[0015] Based on the multi-view geometry principle, the three-dimensional coordinates of the infrared light relative to the first camera are calculated by triangulation using the pixel coordinates of the matched light spot and the relative position of the first camera.

[0016] The first coordinate information is calculated based on the three-dimensional coordinates of the infrared light relative to the first camera and the three-dimensional coordinates relative to the cabin, as well as the different positions of the first camera on the helmet.

[0017] Furthermore, after calculating the latest second coordinate information of the helmet relative to the cabin, the method further includes:

[0018] The positional error between the first coordinate information and the second coordinate information is calculated in real time based on the first coordinate information of the helmet obtained from the light spot image;

[0019] Based on the fuzzy prediction algorithm, the 6DoF inertial information is used as input and the position error is used as output to train and construct a fuzzy learning model.

[0020] When the first coordinate information of the helmet cannot be obtained based on the light spot image, the fuzzy learning model is used to predict the position error.

[0021] Based on the 6DoF inertial information and the first coordinate information, the second coordinate information is derived, and the predicted position error is compensated into the second coordinate information.

[0022] Furthermore, the training and construction of the fuzzy learning model includes:

[0023] Define the structure of the fuzzy learning model, which includes n fuzzy subsystems and m enhancement nodes;

[0024] Multiple fuzzy rules are defined for each of the fuzzy subsystems, and the fuzzy rules are used to describe the relationship between the 6DoF inertial information and the position error;

[0025] A Gaussian membership function is defined for each fuzzy set, and its center and width are determined based on the K-means clustering algorithm;

[0026] Calculate the activation intensity of each of the fuzzy rules as the product of the Gaussian membership functions of the 6DoF inertial information;

[0027] Based on the activation intensity and the parameters of the fuzzy subsystem, the output function of each fuzzy subsystem is constructed.

[0028] The 6DoF inertial information is organized into a training dataset, represented by a matrix, and the parameters of the fuzzy subsystem are trained and adjusted to obtain the fuzzy learning model.

[0029] Furthermore, calculating the user's eye gaze relative to the helmet's third coordinate information specifically includes:

[0030] Based on at least two second cameras continuously capturing pupil images of the same eye of the user, the bounding box of the pupil is calculated based on a deep learning algorithm to detect the center of the pupil;

[0031] Based on the principle of multi-view geometry and the intrinsic and extrinsic parameters of different cameras, different second cameras are projected onto the corresponding pupil center in three-dimensional space to obtain rays. The intersection or closest point of different rays is taken as the three-dimensional position of the pupil center. Multiple three-dimensional positions of the pupil center are fitted and combined with the empirical constant of the eyeball radius to obtain the user's eyeball model.

[0032] The user's pupil center is continuously tracked, and the vector line from the eyeball model to the pupil center is taken as the user's line of sight. Based on the coordinate relationship between the second camera and the first camera, the third coordinate information of the line of sight relative to the first camera is obtained.

[0033] Furthermore, among the multiple second cameras installed inside the helmet, at least one infrared camera and one RGB camera are included.

[0034] Furthermore, after establishing the eyeball model, at least one infrared camera is used to continuously track the center of the user's pupil, and at least one RGB camera is used to periodically acquire the center of the user's pupil to verify the accuracy of the infrared camera.

[0035] According to a second aspect of this disclosure, a gaze tracking device based on 3D posture and 6DoF positioning is provided, comprising: an optical tracking module for acquiring light spot images of infrared light emitted by infrared lamps at multiple fixed positions inside the cabin through multiple first cameras, wherein the first cameras are mounted on a head-mounted display helmet; calculating the position of the infrared lamps based on the light spot images; and calculating the first coordinate information of the helmet relative to the cabin based on the position of the infrared lamps and multi-view geometry principles, wherein the first coordinate information includes three-dimensional coordinates and posture; and wherein the first cameras are infrared cameras.

[0036] An inertial tracking module is used to acquire the 6DoF inertial information of the helmet in real time. When the first coordinate information of the helmet cannot be obtained based on the spot image, the most recently acquired first coordinate information is combined with the corresponding 6DoF inertial information to calculate the latest second coordinate information of the helmet relative to the cabin.

[0037] An eye-tracking module is used to capture the user's eyes using a second camera inside the helmet based on eye tracking, and calculate the third coordinate information of the user's eye gaze relative to the helmet.

[0038] The calculation module is used to associate the third coordinate information with the first coordinate information or the third coordinate information with the second coordinate information to obtain the fourth coordinate information of the user's eye line of sight relative to the cabin.

[0039] The technical solution disclosed herein has the following beneficial effects:

[0040] By combining the helmet's 6DoF inertial information with optical tracking, the helmet's current position and attitude can be deduced from the 6DoF inertial information even when the helmet cannot be identified by the first camera. This ensures that the helmet tracking process is not interrupted, improving the user experience and tracking accuracy. Attached Figure Description

[0041] Figure 1 This is a flowchart of a gaze tracking method based on 3D pose and 6DoF localization in an embodiment of this specification;

[0042] Figure 2 This is a structural block diagram of a gaze tracking device based on 3D pose and 6DoF positioning, as described in an embodiment of this specification. Detailed Implementation

[0043] Example embodiments will now be described more fully with reference to the accompanying drawings. However, example embodiments can be implemented in many forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided to make this disclosure more comprehensive and complete, and to fully convey the concept of the example embodiments to those skilled in the art. The described features, structures, or characteristics may be combined in any suitable manner in one or more embodiments. In the following description, numerous specific details are provided to give a full understanding of embodiments of this disclosure. However, those skilled in the art will recognize that the technical solutions of this disclosure may be practiced with one or more of the specific details omitted, or other methods, components, apparatus, steps, etc., may be employed. In other instances, well-known technical solutions are not shown or described in detail to avoid obscuring various aspects of this disclosure.

[0044] Furthermore, the accompanying drawings are merely illustrative of this disclosure. The same reference numerals in the drawings denote the same or similar parts, and therefore repeated descriptions of them will be omitted. Some block diagrams shown in the drawings are functional entities and do not necessarily correspond to physically or logically independent entities. These functional entities may be implemented in software, in one or more hardware modules or integrated circuits, or in different network and / or processor devices and / or microcontroller devices.

[0045] like Figure 1 As shown in the embodiments of this specification, a gaze tracking method based on 3D pose and 6DoF positioning is provided. The execution subject of this method can be a computer, server, wearable smart device, etc. Specifically, the method may include the following steps S101~S104:

[0046] In step S101, multiple first cameras acquire images of infrared light spots emitted by infrared lights at multiple fixed positions inside the cabin. The first cameras are mounted on the head-mounted display helmet. The position of the infrared light is calculated based on the image of the light spot. Based on the position of the infrared light and the principle of multi-view geometry, the first coordinate information of the helmet relative to the cabin is calculated. The first coordinate information includes three-dimensional coordinates and attitude. The first camera is an infrared camera.

[0047] Specifically, the process of back-calculating the first coordinate information of the helmet includes: extracting the pixel coordinates of the light spot from the light spot image based on an image processing algorithm; matching the pixel coordinates captured by multiple first cameras to identify the projection positions of the same infrared light in different first cameras; calculating the three-dimensional coordinates of the infrared light relative to the first camera by performing triangulation based on the multi-view geometry principle, using the pixel coordinates of the matched light spot and the relative positions of the first cameras; and calculating the first coordinate information based on the three-dimensional coordinates of the infrared light relative to the first camera and relative to the cabin, as well as the positions of the different first cameras on the helmet.

[0048] The first camera is an infrared camera, and the image processing algorithm can include threshold segmentation, morphological operations, and edge detection. Threshold segmentation can be performed by setting a grayscale threshold to binarize the image, making the light spot white and the background black. Morphological operations can be opening and closing operations to remove noise from the image and retain the light spot area. Edge detection can be performed on the processed image to find the center position of the light spot and record its pixel coordinates. In images captured by multiple first cameras (infrared cameras), the projection position of the light spot image is different from different viewpoints. To calculate the three-dimensional coordinates of the infrared lamp in space, it is necessary to identify the corresponding light spot of the same infrared lamp in different camera images. The specific process includes feature matching, that is, using the shape, size, and brightness features of the light spot to match the same light spot in different camera images, which can be represented as:

[0049] ;

[0050]

[0051] Multi-view geometry principles include epipolar geometric constraints, which further verify the accuracy of matching through epipolar geometric relationships. For two cameras, the corresponding point of any given point must be on the epipolar line of the other camera, expressed as:

[0052] ;

[0053] in, are homogeneous coordinates on the image plane, and F is the fundamental matrix.

[0054] Triangulation is a method that uses the pixel coordinates of a matched light spot in different cameras to calculate the three-dimensional coordinates of an infrared light relative to each camera using multi-view geometry. In other words, it is a method that inversely calculates the three-dimensional coordinates of an object from the two-dimensional image plane projection from the perspectives of two or more cameras.

[0055] In step S102, the 6DoF inertial information of the helmet is acquired in real time. When the first coordinate information of the helmet cannot be obtained based on the spot image, the most recently acquired first coordinate information is combined with the corresponding 6DoF inertial information to calculate the latest second coordinate information of the helmet relative to the cabin.

[0056] In helmet positioning and tracking systems, to ensure the continuity and accuracy of the user's helmet position and attitude (i.e., 6DoF—six degrees of freedom, including position and attitude angles in three-dimensional space) information under different environmental conditions, the system relies not only on infrared spot data captured by a camera but also on real-time data from an inertial measurement unit (IMU). Specifically, the IMU typically integrates an accelerometer, gyroscope, and sometimes a magnetometer, providing real-time information on the helmet's acceleration, angular velocity, and orientation changes. This data can be integrated and filtered to derive the helmet's instantaneous position and attitude. Accelerometer: measures the helmet's linear acceleration in space. Gyroscope: measures the helmet's angular velocity around three axes. Magnetometer: provides orientation reference, typically used to compensate for errors caused by gyroscope drift. Due to ambient light interference, poor infrared light reflection, or other external factors, the camera may sometimes fail to reliably capture infrared spot images. In such cases, the system cannot calculate the helmet's initial coordinates (i.e., precise spatial positioning data based on the spot) from the spot data. In this scenario, the system reverts to using the most recently obtained first coordinate information calculated from the optical spot (i.e., the accurate position and attitude of the helmet in the last successful capture), and combines this with the currently acquired 6DoF inertial information to infer the helmet's current position and attitude (second coordinate information). The last first coordinate information is then fused with the acceleration and angular velocity data provided by the IMU, and algorithms such as a Kalman filter or an extended Kalman filter (EKF) are used to predict the helmet's current position and attitude. IMU data provides high-frequency attitude and position updates, thus compensating for the impact of missing optical data in a short time. A Kalman filter can be applied; it is a recursive algorithm that predicts the current state based on the system model and corrects prediction errors based on sensor data (IMU information). Specifically: In the prediction phase, the next position and attitude of the helmet are predicted based on real-time IMU data; in the update phase, once the optical spot data is recovered, the system uses the latest optical information to correct the helmet position and attitude predicted by the IMU. Finally, the latest helmet position and attitude are calculated: The second coordinate information is calculated by fusing and processing IMU data with the most recent first coordinate information to estimate the helmet's latest position and attitude; this result is the second coordinate information. Position estimation is achieved by double integration of the IMU data to derive the helmet's displacement information. Attitude estimation is performed by deriving the helmet's rotation angle change from the IMU's gyroscope data and combining it with previous attitude information to calculate the current attitude. Error compensation is also performed; to reduce the inherent drift and accumulated errors of the inertial sensor, the system may employ error correction algorithms, such as real-time correction via communication with the ground station or vision-based error feedback.

[0057] In step S103, based on eye tracking, the user's eyes are captured by the second camera inside the helmet, and the third coordinate information of the user's eye gaze relative to the helmet is calculated.

[0058] Specifically, calculating the third coordinate information of the user's eye gaze relative to the helmet includes: based on pupil images of the same eye continuously captured by at least two second cameras, calculating the bounding box of the pupil based on a deep learning algorithm to detect the pupil center; based on multi-view geometry principles and the intrinsic and extrinsic parameters of different cameras, projecting different second cameras to their corresponding pupil centers into three-dimensional space to obtain rays, using the intersection or closest point of different rays as the three-dimensional position of the pupil center, fitting multiple three-dimensional positions of the pupil center, and combining with the empirical constant of the eyeball radius to obtain the user's eyeball model; continuously tracking the user's pupil center, using the vector line from the eyeball model to the pupil center as the user's eye gaze, and obtaining the third coordinate information of the eye gaze relative to the first camera based on the coordinate relationship between the second camera and the first camera.

[0059] The second camera includes at least one infrared camera and one RGB camera. After the eye model is established, at least one infrared camera is used to continuously track the center of the user's pupil, and at least one RGB camera is used to periodically acquire the center of the user's pupil to verify the accuracy of the infrared camera.

[0060] In step S104, the third coordinate information is associated with the first coordinate information or the third coordinate information is associated with the second coordinate information to obtain the fourth coordinate information of the user's eye line of sight relative to the cabin.

[0061] In one embodiment, after calculating the latest second coordinate information of the helmet relative to the cabin, the method further includes: calculating the position error between the first coordinate information and the second coordinate information based on the first coordinate information of the helmet obtained from the spot image in real time; training and constructing a fuzzy learning model based on a fuzzy prediction algorithm, using the 6DoF inertial information as input and the position error as output; predicting the position error using the fuzzy learning model when the first coordinate information of the helmet cannot be obtained based on the spot image; deriving the second coordinate information based on the 6DoF inertial information and the first coordinate information, and compensating the predicted position error into the second coordinate information.

[0062] The process of training and constructing the fuzzy learning model includes: defining the structure of the fuzzy learning model, which comprises n fuzzy subsystems and m augmentation nodes; defining multiple fuzzy rules for each fuzzy subsystem, which describe the relationship between the 6DoF inertial information and the position error; defining a Gaussian membership function for each fuzzy set and determining its center and width based on the K-means clustering algorithm; calculating the activation intensity of each fuzzy rule as the product of the Gaussian membership functions of the 6DoF inertial information; constructing the output function of each fuzzy subsystem based on the activation intensity and the parameters of the fuzzy subsystem; organizing the 6DoF inertial information into a training dataset, representing it as a matrix, and training and adjusting the parameters of the fuzzy subsystems to obtain the fuzzy learning model.

[0063] Based on the same line of thought, such as Figure 2 As shown, an exemplary embodiment of this disclosure also provides a gaze tracking system based on 3D pose and 6DoF positioning, including:

[0064] The optical tracking module 201 is used to acquire light spot images of infrared light emitted by infrared lamps at multiple fixed positions inside the cabin through multiple first cameras. The first cameras are set on the head-mounted display helmet. The position of the infrared lamp is calculated based on the light spot image. Based on the position of the infrared lamp and the multi-view geometry principle, the first coordinate information of the helmet relative to the cabin is calculated. The first coordinate information includes three-dimensional coordinates and attitude. The first camera is an infrared camera.

[0065] The inertial tracking module 202 is used to acquire the 6DoF inertial information of the helmet in real time. When the first coordinate information of the helmet cannot be obtained based on the spot image, the most recently acquired first coordinate information is combined with the corresponding 6DoF inertial information to calculate the latest second coordinate information of the helmet relative to the cabin.

[0066] The eye-tracking module 203 is used to capture the user's eyes using a second camera inside the helmet based on eye tracking, and calculate the third coordinate information of the user's eye gaze relative to the helmet.

[0067] The calculation module 204 is used to associate the third coordinate information with the first coordinate information or the third coordinate information with the second coordinate information to obtain the fourth coordinate information of the user's eye line of sight relative to the cabin.

[0068] The system combines the helmet's 6DoF inertial information with optical tracking, enabling the helmet to deduce its current position and attitude even when it cannot be identified by the first camera. This ensures that the helmet tracking process is uninterrupted, improving the user experience and tracking accuracy.

[0069] From the above description of the embodiments, those skilled in the art will readily understand that the exemplary embodiments described herein can be implemented by software or by combining software with necessary hardware. Therefore, the technical solutions according to the embodiments of this disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, USB flash drive, external hard drive, etc.) or on a network, including several instructions to cause a computing device (such as a personal computer, server, terminal device, or network device, etc.) to execute the method according to the exemplary embodiments of this disclosure.

[0070] Furthermore, the above figures are merely illustrative representations of the processes included in the methods according to exemplary embodiments of this disclosure, and are not intended to be limiting. It is readily understood that the processes shown in the above figures do not indicate or limit the temporal order of these processes. Additionally, it is readily understood that these processes may be executed synchronously or asynchronously, for example, in multiple modules.

[0071] Other embodiments of this disclosure will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not disclosed herein. The specification and embodiments are to be considered exemplary only, and the true scope and spirit of this disclosure are indicated by the claims.

[0072] It should be understood that this disclosure is not limited to the precise structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this disclosure is limited only by the appended claims.

Claims

1. A gaze tracking method based on 3D pose and 6DoF localization, characterized in that, The method includes: Multiple first cameras acquire light spot images of infrared light emitted by infrared lamps at multiple fixed positions inside the cabin. The first cameras are mounted on the head-mounted display helmet. The position of the infrared lamp is calculated based on the light spot images. Based on the position of the infrared lamp and the multi-view geometry principle, the first coordinate information of the helmet relative to the cabin is calculated. The first coordinate information includes three-dimensional coordinates and attitude. The first camera is an infrared camera. The system acquires the 6DoF inertial information of the helmet in real time. When the first coordinate information of the helmet cannot be obtained based on the spot image, the most recently acquired first coordinate information is combined with the corresponding 6DoF inertial information to calculate the latest second coordinate information of the helmet relative to the cabin. After calculating the latest second coordinate information of the helmet relative to the cabin, the system further includes: calculating the position error between the first coordinate information and the second coordinate information based on the first coordinate information of the helmet acquired from the spot image in real time; training and constructing a fuzzy learning model based on a fuzzy prediction algorithm, using the 6DoF inertial information as input and the position error as output; predicting the position error using the fuzzy learning model when the first coordinate information of the helmet cannot be obtained based on the spot image; deriving the second coordinate information based on the 6DoF inertial information and the first coordinate information, and compensating the predicted position error into the second coordinate information. Based on eye tracking, the user's eyes are captured by a second camera inside the helmet, and the third coordinate information of the user's eye gaze relative to the helmet is calculated. By associating the third coordinate information with the second coordinate information, a fourth coordinate information is obtained, which represents the user's line of sight relative to the cabin.

2. The gaze tracking method based on 3D pose and 6DoF localization according to claim 1, characterized in that, The process of back-calculating the first coordinate information of the helmet specifically includes: Based on image processing algorithms, the pixel coordinates of the light spot are extracted from the light spot image; The pixel coordinates captured by multiple first cameras are matched to identify the projection position of the same infrared light in different first cameras; Based on the multi-view geometry principle, the three-dimensional coordinates of the infrared light relative to the first camera are calculated by triangulation using the pixel coordinates of the matched light spot and the relative position of the first camera. The first coordinate information is calculated based on the three-dimensional coordinates of the infrared light relative to the first camera and the three-dimensional coordinates relative to the cabin, as well as the different positions of the first camera on the helmet.

3. The gaze tracking method based on 3D pose and 6DoF localization according to claim 1, characterized in that, The training and construction of the fuzzy learning model includes: Define the structure of the fuzzy learning model, which includes n fuzzy subsystems and m enhancement nodes; Multiple fuzzy rules are defined for each of the fuzzy subsystems, and the fuzzy rules are used to describe the relationship between the 6DoF inertial information and the position error; A Gaussian membership function is defined for each fuzzy set, and its center and width are determined based on the K-means clustering algorithm; Calculate the activation intensity of each of the fuzzy rules as the product of the Gaussian membership functions of the 6DoF inertial information; Based on the activation intensity and the parameters of the fuzzy subsystem, the output function of each fuzzy subsystem is constructed. The 6DoF inertial information is organized into a training dataset, represented by a matrix, and the parameters of the fuzzy subsystem are trained and adjusted to obtain the fuzzy learning model.

4. The gaze tracking method based on 3D pose and 6DoF localization according to claim 1, characterized in that, Calculating the user's line of sight relative to the helmet's third coordinate information specifically includes: Based on at least two second cameras continuously capturing pupil images of the same eye of the user, the bounding box of the pupil is calculated based on a deep learning algorithm to detect the center of the pupil; Based on the principle of multi-view geometry and the intrinsic and extrinsic parameters of different cameras, the lines connecting different second cameras and the corresponding pupil centers are projected into three-dimensional space to obtain rays. The intersection or closest point of different rays is taken as the three-dimensional position of the pupil center. Multiple three-dimensional positions of the pupil center are fitted and combined with the empirical constant of the eyeball radius to obtain the user's eyeball model. The user's pupil center is continuously tracked, and the vector line from the eyeball model to the pupil center is taken as the user's line of sight. Based on the coordinate relationship between the second camera and the first camera, the third coordinate information of the line of sight relative to the first camera is obtained.

5. The gaze tracking method based on 3D pose and 6DoF localization according to claim 4, characterized in that, The helmet contains a plurality of second cameras, including at least one infrared camera and one RGB camera.

6. The gaze tracking method based on 3D pose and 6DoF localization according to claim 5, characterized in that, After establishing the eye model, at least one infrared camera is used to continuously track the center of the user's pupil, and at least one RGB camera is used to periodically acquire the center of the user's pupil to verify the accuracy of the infrared camera.

7. A gaze tracking system based on 3D pose and 6DoF positioning, characterized in that, include: An optical tracking module is used to acquire light spot images of infrared light emitted by infrared lamps at multiple fixed positions inside the cabin through multiple first cameras. The first cameras are set on the head-mounted display helmet. The position of the infrared lamp is calculated based on the light spot image. Based on the position of the infrared lamp and the multi-view geometry principle, the first coordinate information of the helmet relative to the cabin is calculated. The first coordinate information includes three-dimensional coordinates and attitude. The first camera is an infrared camera. An inertial tracking module is used to acquire the 6DoF inertial information of the helmet in real time. When the first coordinate information of the helmet cannot be obtained based on the spot image, the most recently acquired first coordinate information is combined with the corresponding 6DoF inertial information to calculate the latest second coordinate information of the helmet relative to the cabin. After calculating the latest second coordinate information of the helmet relative to the cabin, the method further includes: calculating the position error between the first coordinate information and the second coordinate information based on the first coordinate information of the helmet obtained from the spot image in real time; training and constructing a fuzzy learning model based on a fuzzy prediction algorithm, using the 6DoF inertial information as input and the position error as output; predicting the position error using the fuzzy learning model when the first coordinate information of the helmet cannot be obtained based on the spot image; deriving the second coordinate information based on the 6DoF inertial information and the first coordinate information, and compensating the predicted position error into the second coordinate information; An eye-tracking module is used to capture the user's eyes using a second camera inside the helmet based on eye tracking, and calculate the third coordinate information of the user's eye gaze relative to the helmet. The calculation module is used to associate the third coordinate information with the second coordinate information to obtain the fourth coordinate information of the user's eye line of sight relative to the cabin.

Citation Information

Patent Citations

  • VR head-mounted all-in-one machine

    CN112416125A

  • Using 6DOF pose information to align images from separated cameras

    US11049277B1