Head-mounted visual line tracking device and slip compensation method
Patent Information
- Application Number
- CN202610858601.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-15
- Publication Date
- 2026-09-04
AI Technical Summary
即使头戴装置整体佩戴状态未发生明显位移,装置相对于佩戴者头部仍可能产生难以直接察觉的微小滑移或姿态偏差
[0045]This invention achieves a high degree of lightweighting and integration by systematically designing the head-mounted acquisition unit, connection unit, and edge computing unit. This ensures functionality such as scene image acquisition, binocular image acquisition, inertial information acquisition, data processing, and real-time display. This design facilitates independent deployment in pilot training, flight training, simulation operation training, and experimental environments, reducing reliance on external equipment and complex installation environments.
Smart Images

Figure CN122698733A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to intelligent sensing devices and multimodal information fusion, and more particularly to a head-mounted eye-tracking device and a slip compensation method. Background Technology
[0002] Currently, head-mounted eye-tracking technology has reached a relatively mature product form. Existing solutions typically include a scene image acquisition module, left and right eye image acquisition modules, a head-mounted support structure, and a matching data processing module. Some systems also integrate an inertial measurement module to output acceleration, angular velocity, or attitude-related information.
[0003] However, for applications such as cockpits, flight cabins, simulation training cabins, and other environments that emphasize prolonged wear, dynamic movements, vibration interference, and real-time on-site observation, there is still room for further optimization of existing head-mounted eye-tracking systems.
[0004] I. Existing head-mounted eye-tracking devices typically require simultaneous acquisition of left and right eye images and scene images, and in some applications, simultaneous acquisition of inertial information is also necessary. The consistency of these multi-source data in terms of time reference directly affects gaze point estimation, eye movement event recognition, and the results of multi-source information fusion processing. Especially under the conditions of lightweight and miniaturized system design, how to achieve stable synchronous acquisition and unified management of multiple image data and inertial data within a limited volume remains a significant technical problem affecting system performance.
[0005] Second, during piloting, flying, and high-intensity operational training, wearers often experience rapid head rotation, continuous vibration, sweating, changes in facial expressions, and cable traction. Even if the overall wearing position of the head-mounted device does not change significantly, the device may still experience slight slippage or posture deviation relative to the wearer's head that is difficult to perceive directly. Such changes alter the relative geometric relationship between the eye image acquisition module, the scene image acquisition module, and the head, causing a shift in the gaze mapping relationship established based on the initial calibration. This leads to a drift in the gaze point, affecting the stability and reliability of training evaluation and experimental measurement results.
[0006] Third, for application scenarios such as training, on-site assessment, and experimental debugging, in addition to recording line-of-sight data, operators often need to view the scene, eye images, and system operating status locally in real time. How to integrate multi-channel acquisition, synchronous processing, status display, and portable deployment into a single design without significantly increasing system size and wearing burden is also a problem that needs to be solved in practical engineering applications. Summary of the Invention
[0007] This invention addresses the shortcomings of existing technologies by providing a head-mounted eye-tracking device and a slip compensation method. While maintaining the system's lightweight design and wearability, it achieves effective synchronization of eye images, scene images, and inertial information, and enhances the device's adaptability to micro-slippage under high-dynamic usage conditions, thereby improving the stability and practicality of eye-tracking results.
[0008] To achieve the above objectives, the present invention adopts the following technical solution: a head-mounted eye-tracking device, comprising a head-mounted acquisition unit, a connection unit, and an edge computing unit;
[0009] The head-mounted acquisition unit is used to acquire images of the wearer's left and right eyes, images of the scene in front, and inertial data related to head movement;
[0010] The connection unit connects the head-mounted acquisition unit and the edge computing unit to enable data and power transmission.
[0011] The edge computing unit includes an embedded computing platform, which establishes a unified time reference and constructs a unified synchronization data group for the front scene image, left and right eye images and inertial data based on a unified system clock;
[0012] The embedded computing platform is also used to form an initial gaze mapping model in the initial calibration stage, calculate the gaze landing point based on the mapping model in the running stage, and integrate the eye stability reference feature offset residual, inertial data and scene image motion consistency residual to perform wear micro-slip detection and hierarchical compensation processing, thereby compensating and updating the gaze landing point or mapping parameters.
[0013] Furthermore, the head-mounted acquisition unit is mounted on a lightweight head-mounted frame with a lensless open structure and an overall weight not exceeding 90g. The head-mounted acquisition unit includes a scene image acquisition module, a left-eye image acquisition module, a right-eye image acquisition module, and an inertial measurement module. The inertial measurement module is fixedly installed on the mounting base of the scene image acquisition module or in a rigid area of the head-mounted frame.
[0014] Furthermore, the sampling frame rates of the left-eye image acquisition module, the right-eye image acquisition module, and the inertial measurement module are all higher than the sampling frame rate of the scene image acquisition module; the left-eye image acquisition module and the right-eye image acquisition module are each equipped with a near-infrared illumination module to provide auxiliary illumination to the eye area.
[0015] Furthermore, when the edge computing unit constructs a unified synchronization data group, for each frame of scene image, it selects the image frame with the timestamp closest to that frame time from the left and right eye image sequences as the matching frame, and matches the inertial data at the corresponding time.
[0016] A head-mounted gaze tracking and slip compensation method, based on the gaze mapping model, includes the following steps:
[0017] S1. In the initial stable wearing state, the head-mounted acquisition unit collects multi-source synchronous data, performs initial calibration, determines the reference mapping parameters according to the preset eye feature to scene coordinate mapping model, and forms an initial mapping model based on the reference mapping parameters.
[0018] S2. During operation, continuously receive front scene images, left and right eye images, and inertial data aligned with a unified time reference;
[0019] S3. Extract eye stabilization reference features based on the received left and right eye images, and calculate the eye stabilization reference feature offset residual based on the change of the eye stabilization reference features relative to the initial stable wearing state or the short-term sliding window reference state.
[0020] The eye stability reference features include one or more of the following: inner corner of the eye, outer corner of the eye, short line of eyelid margin, and the relative geometric relationship between eyelid margin and corner of the eye;
[0021] Based on the received inertial data, the theoretical motion trend corresponding to the scene image is predicted, and the scene motion consistency residual is calculated in combination with the actual motion characteristics of the scene image. A comprehensive slip judgment index is constructed based on the eye stability reference feature offset residual, scene motion consistency residual, line of sight landing point error under known reference target and data quality status. When the comprehensive slip judgment index is greater than the set threshold or continuously meets the preset frame number condition, it is determined that the device has micro-slipped relative to the wearer's head.
[0022] S4. After determining that micro-slip exists, the edge computing unit performs dynamic compensation on the line-of-sight landing point, performs online compensation update on the initial mapping model parameters, or triggers a recalibration process based on the magnitude or credibility of the comprehensive slip determination index.
[0023] Further, in step S1, the initial calibration specifically includes:
[0024] The wearer is instructed to sequentially focus on multiple preset calibration reference points in the scene image;
[0025] At each preset calibration reference point, the corresponding left and right eye images are acquired and the eye feature vectors are extracted;
[0026] The eye feature vector and the corresponding scene image coordinates are used to form a calibration sample pair;
[0027] The preset mapping function from eye feature vectors to scene image coordinates is invoked, and the baseline mapping parameters are determined based on minimizing the loss function, thereby constructing an initial mapping model.
[0028] Further, in step S3, the construction of the comprehensive slip judgment index specifically involves: performing row normalization on the eye stability reference feature offset residual, scene motion consistency residual, and prediction error or mapping residual of the current gaze point, and then weighting and summing them according to preset weights to calculate the comprehensive slip judgment index.
[0029] The eye-stabilizing reference feature offset residual is used to characterize the offset, rotation, or local affine change of the head-stabilizing reference features in the eye image relative to the initial wearing reference state during operation. It is the main basis for determining whether the head-mounted device has micro-slipped relative to the wearer's head. The scene motion consistency residual is used to characterize the consistency between the scene motion trend predicted by inertial data and the actual motion features of the scene image. It serves as an auxiliary credibility constraint for slip determination and compensation update. The prediction error or mapping residual of the current gaze point only participates in the calculation of the comprehensive slip determination index when the current gaze target is a known reference target, a preset calibration point, a fixed target on the system interface, or a known target point in the training task.
[0030] The conditions under which the device determines a slight slippage trend include: the comprehensive slippage determination index is greater than the first threshold and continues for more than a set number of frames;
[0031] Alternatively, the residual offset of the eye-stabilized reference feature is greater than the second threshold and continues for more than a certain number of frames;
[0032] Alternatively, the rate of change of the eye-stabilized reference feature offset residual over a unit of time is greater than a set rate of change threshold.
[0033] Alternatively, when a known reference target exists, the line-of-sight mapping residual continuously increases within a short sliding window, and the eye-stabilized reference feature offset residual increases synchronously.
[0034] Furthermore, the online compensation update and recalibration process in step S4 includes one or more of the following methods:
[0035] Method 1: When the slip is less than a set threshold, the local geometric correction amount in the eye image coordinate system is estimated based on the offset residual of the eye stable reference feature, and the eye features are corrected online. The corrected eye features are then input into the initial mapping model or the current mapping model to achieve dynamic compensation under small slip.
[0036] Method 2: When the degree of slippage is in a moderate range and there is a known reference target, the mapping model parameters are locally updated online based on the mapping residual between the position of the known reference target in the scene image coordinate system and the current line of sight, combined with the effective observation data within the short-term sliding window.
[0037] Method 3: When the degree of slippage continues to increase or the residual of the line of sight landing point is still higher than the set threshold after online compensation by Method 1 and Method 2, a prompt message is output and the recalibration process is triggered to guide the wearer to look at a small number of preset calibration points (the number of which is less than that required by regular calibration) to locally update the model parameters or re-establish the mapping model.
[0038] Furthermore, Method 1 specifically includes:
[0039] Extract stable reference features of the eye from the current eye image and match them with the positional relationship of the reference features in the initial calibration stage to estimate the translation, rotation, or local affine transformation parameters of the current eye image relative to the initial wearing reference state; based on the transformation parameters, perform reverse correction on the pupil center or eye feature vector of the current frame to obtain the corrected eye features; input the corrected eye feature vector into the initial mapping model or the current mapping model to calculate the compensated gaze point.
[0040] Furthermore, method 2 specifically includes:
[0041] When the system detects that the wearer is gazing at a known reference target, it obtains the pixel coordinates of the known reference target in the scene image coordinate system and calculates the mapping residual between these pixel coordinates and the pixel coordinates of the gaze point output by the current mapping model. A compensation loss function containing the mapping residual term and parameter constraint term is constructed, and the mapping model parameters are locally updated based on minimizing the compensation loss function. The parameter constraint term limits the variation of the updated mapping model parameters relative to the initial mapping model parameters, preventing excessive drift in the mapping model due to short-term errors, blinking, reflections, abnormal scene motion, or incorrect gaze judgment. The updated mapping model parameters are used as the current online mapping model parameters for subsequent gaze point calculations.
[0042] Furthermore, method 3 specifically includes:
[0043] When the residual offset of the stable reference feature of the eye is continuously greater than the set threshold in multiple consecutive synchronous data sets, or when the residual of the gaze landing point after compensation by method 1 and method 2 is still continuously greater than the set threshold, the edge computing unit outputs a recalibration prompt message and presents a small number of preset calibration points in the scene image or display interface; after the wearer looks at the small number of preset calibration points in sequence, the system collects the corresponding eye feature vector and the scene coordinates of the calibration points, and performs local updates on the mapping model parameters based on the collected samples, or re-establishes the gaze mapping model in the current wearing state.
[0044] Compared with the prior art, the present invention has the following advantages.
[0045] This invention achieves a high degree of lightweighting and integration by systematically designing the head-mounted acquisition unit, connection unit, and edge computing unit. This ensures functionality such as scene image acquisition, binocular image acquisition, inertial information acquisition, data processing, and real-time display. This design facilitates independent deployment in pilot training, flight training, simulation operation training, and experimental environments, reducing reliance on external equipment and complex installation environments.
[0046] This invention sets up a unified timing management module to manage and synchronize scene images, left and right eye images, and inertial data using a unified time reference. This reduces the errors caused by timing inconsistencies in gaze mapping, event recognition, and multi-source fusion analysis, thereby improving the consistency and reliability of gaze tracking results.
[0047] This invention uses the eye-stabilized reference feature offset residual as the main criterion for micro-slippage of the head-mounted device relative to the wearer's head. It combines inertial data and scene image motion consistency residuals to constrain the credibility of the compensation update process. Based on this, it performs hierarchical processing according to the degree of slippage. While ensuring the accuracy of eye tracking, it reduces the number of complete recalibrations, reduces the interruption time in training, assessment or experimentation, and improves the accuracy of slippage detection results and the stability of the compensation process.
[0048] The edge computing unit of this invention has local processing and display capabilities, and can directly display scene images, eye images, line-of-sight superposition results and system status information, which makes it convenient for instructors, experimenters or operators to view the equipment operating status and adjust parameters in real time on site, thereby improving the ease of use of the device in training, teaching and experimental debugging scenarios.
[0049] This invention adopts a modular system architecture, with clear division of labor among the head-mounted acquisition, connection and transmission, edge computing and display modules, which facilitates subsequent maintenance, replacement and expansion; at the same time, multiple modules work together under a unified timing and processing framework, which is conducive to forming an engineering system suitable for practical application deployment. Attached Figure Description
[0050] The present invention will be further described below with reference to the accompanying drawings and specific embodiments. The scope of protection of the present invention is not limited to the following description.
[0051] Figure 1 This is a schematic diagram of the overall structure of the device in the embodiment.
[0052] Figure 2 This is a schematic diagram of the structure of the head-mounted acquisition unit in an embodiment.
[0053] Figure 3 This is a schematic diagram of the edge computing unit in an embodiment.
[0054] Figure 4This is a roadmap for multi-source data synchronization technology in an example.
[0055] Figure 5 This is a roadmap for micro-slip detection and online compensation technology in an example. Detailed Implementation
[0056] The embodiments of this application will be further described in detail below with reference to the accompanying drawings and examples. The detailed description of the following embodiments and the accompanying drawings are used to illustrate the principles of this application by way of example, but should not be used to limit the scope of this application. This application can be implemented in many different forms and is not limited to the specific embodiments disclosed herein, but includes all technical solutions falling within the scope of the claims.
[0057] These embodiments are provided to make the application thorough and complete, and to fully express the scope of the application to those skilled in the art. It should be noted that, unless otherwise specifically stated, the relative arrangement of components and steps, material composition, numerical expressions, and values illustrated in these embodiments should be interpreted as merely exemplary and not as limiting.
[0058] It should be noted that, in the description of this application, unless otherwise stated, "a plurality of" means two or more; the terms "upper," "lower," "left," "right," "inner," and "outer," etc., indicating orientation or positional relationship, are only for the convenience of describing this application and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of this application. When the absolute position of the described object changes, the relative positional relationship may also change accordingly.
[0059] Furthermore, the terms "first," "second," and similar terms used in this application do not indicate any order, quantity, or importance, but are merely used to distinguish different parts. "Vertical" is not strictly vertical, but within the permissible margin of error. "Parallel" is not strictly parallel, but within the permissible margin of error. Terms such as "including" or "contains" mean that the element preceding the word encompasses the element listed after it, and do not exclude the possibility of encompassing other elements as well.
[0060] It should also be noted that, in the description of this application, unless otherwise expressly specified and limited, the terms "installation," "connection," and "joining" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a direct connection or an indirect connection through an intermediate medium. Those skilled in the art can understand the specific meaning of the above terms in this application depending on the specific circumstances. When a specific device is described as being located between a first device and a second device, an intermediary device may or may not be present between the specific device and the first or second device.
[0061] All terms used in this application have the same meaning as understood by one of ordinary skill in the art to which this application pertains, unless otherwise specifically defined. It should also be understood that terms defined in general dictionaries should be interpreted as having meanings consistent with their meanings in the context of the relevant art, and not as idealized or highly formalized, unless expressly defined herein.
[0062] Techniques, methods, and equipment known to those skilled in the art may not be discussed in detail, but where appropriate, they should be considered part of the specification.
[0063] like Figure 1-5 As shown, the embodiment provides a head-mounted gaze tracking device and a slip compensation method, wherein the overall structure of the device is as follows:
[0064] The head-mounted eye-tracking device includes a head-mounted acquisition unit 1, a connection unit 2, and an edge computing unit 3. The head-mounted acquisition unit 1 is worn on the user's head and is used to acquire scene images, left and right eye images, and inertial data; the connection unit 2 is used to complete data and power transmission; and the edge computing unit 3 is used to complete multi-source data synchronization, image processing, attitude calculation, slip compensation, and result display.
[0065] See Figure 1-3 The head-mounted acquisition unit is mounted on a lightweight head-mounted frame 4. The head-mounted frame preferably adopts an open structure, including a nose pad structure 5, a temple structure 6, and multiple module mounting parts.
[0066] The open-frame acquisition unit 4 is preferably made of nylon, carbon fiber reinforced plastic, or a combination thereof, and can be manufactured using 3D printing, injection molding, or machining. The overall frame is preferably a lensless open-frame structure, and the weight of the entire headband is preferably no more than 90g to reduce the burden of prolonged wear. The nose pad 5 is preferably made of flexible silicone, thermoplastic elastomer, or memory material to form a bonding layer, improving wearing comfort and anti-slip properties. The temples 6 may be equipped with an elastic clamping structure, a replaceable ear hook structure, or a headband auxiliary fixing structure to improve stability during intense action scenarios. The circuit board fixing part 7 is located in the center of the frame and / or at the base of the temples for mounting the camera interface board, IMU module, and connecting devices.
[0067] The scene image acquisition module 8 is located at the center of the front of the head-mounted frame, with its optical axis direction basically consistent with the wearer's forward viewing direction, and is used to acquire first-person perspective scene images. The scene image acquisition module 8 preferably uses a color camera with a resolution of 1280×720, 1920×1080 or higher, and a frame rate of 30fps-60fps; its lens is preferably a wide-angle lens, and can be corrected in real time through distortion correction parameters.
[0068] The left-eye image acquisition module 9 and the right-eye image acquisition module 10 are respectively installed on the lower side of the corresponding positions of the left and right eyes. The left-eye image acquisition module 9 and the right-eye image acquisition module 10 preferably employ near-infrared cameras, with a preferred resolution of 320×240, 640×480, or higher, and a preferred frame rate of 60fps-240fps. Preferably, the eye image acquisition modules are also equipped with near-infrared illumination modules to enhance the pupil boundary and reflection features in the eye images, thereby improving the stability of subsequent eye movement feature extraction.
[0069] In this embodiment, the head-mounted acquisition unit is also equipped with an inertial measurement unit (IMU) module 11. The IMU module 11 is preferably a six-axis or nine-axis inertial measurement module, fixedly installed in a rigid area of the head-mounted acquisition unit to reduce errors introduced by loose installation. Preferably, the IMU module 11 is installed adjacent to the scene image acquisition module 8, so that its output more accurately represents the attitude changes of the head-mounted subject. Since the position of the IMU relative to the head-mounted frame is known after assembly, the relative relationship between the IMU coordinate system and the head-mounted reference coordinate system can be pre-established during system initialization or calibration, thereby providing a coordinate basis for subsequent attitude calculation, slip recognition, and online compensation.
[0070] As an optional implementation, connection unit 2 connects the acquisition unit to the edge computing unit via a flexible cable and integrates a multi-channel data synchronization transmission module. This unit supports parallel transmission of image data from the scene camera, left and right eye cameras, and IMU data via a single cable.
[0071] To simplify wiring and reduce weight, the connection unit adopts a design that integrates multiple cables into a single flexible cable, reducing the number of cables and wiring complexity while improving ease of wear. The outer layer of the cable is protected by a bend-resistant braided mesh, enhancing tensile and abrasion resistance and preventing breakage or damage during long-term use.
[0072] The connection unit offers cable options of varying lengths to meet diverse usage scenarios. For example, short cables are suitable for pilots' personal portable training, while long cables can be used by rear-seat instructors to view the edge computing unit's processing results in real time, or for real-time data monitoring in larger spaces such as simulators / flight cabins.
[0073] To ensure data stability and acquisition accuracy, a lightweight vibration damping module is designed inside the connection unit, which can effectively absorb the mechanical impact caused by the movement of the acquisition unit or external vibration, and reduce the impact of camera shake on image quality.
[0074] Meanwhile, the connection unit adopts a multi-layer electromagnetic shielding structure to isolate electromagnetic interference (EMI) from the acquisition unit and the external environment, prevent data transmission errors or signal noise, and improve the stability and reliability of the entire eye-tracking system.
[0075] Through the above design, the connection unit not only achieves efficient data and power transmission, but also takes into account mechanical stability, anti-interference ability and wearing comfort, meeting the needs of high-intensity training and multi-scenario experiments.
[0076] As an optional implementation, the edge computing unit 3 includes an embedded computing platform 12, a display module 13, and a power supply module 14. The embedded computing platform 12 is preferably a Raspberry Pi 5, Jetson Orin Nano, or other equivalent platform; the display module 13 is preferably an LCD or OLED touch screen; the power supply module 14 can be a rechargeable lithium battery pack, preferably with a battery life of not less than 4 hours, more preferably not less than 8 hours.
[0077] The embedded computing platform 12 receives scene images, left and right eye images, and IMU data sent by the head-mounted acquisition unit, and performs processing steps such as time synchronization, image preprocessing, eye movement feature extraction, gaze calculation, slip detection, online compensation, and result output. The platform has built-in high-speed flash memory and SD card interfaces for storing raw and processed experimental data, and provides USB, Ethernet, and GPIO interfaces to enable extended connections to external sensors or servers and remote control functions.
[0078] Display module 13 is connected to the embedded platform via an HDMI interface to display scene images, eye images, gaze point overlay results, and system operating status in real time, providing operators with an intuitive interface and preview function. Users can set, start, or stop data acquisition, and preview real-time data and adjust system parameters through the touch operating system.
[0079] The power supply module 14 employs a voltage regulation design to ensure stable operation of the embedded platform, camera, and LED lighting system. It also incorporates overvoltage, overcurrent, overtemperature, and short-circuit protection circuits to enhance system safety. The power supply module is removable for easy maintenance or replacement and supports expansion to improve battery life. The power supply module 14 powers the head-mounted acquisition unit and edge computing unit, enabling the entire device to operate independently of an external fixed host.
[0080] Based on the hardware collaboration of the aforementioned head-mounted acquisition unit, connection unit, and edge computing unit, the system achieves lightweight and portable deployment while possessing the ability to acquire multimodal data streams in parallel. However, due to differences in the sampling rates and transmission delays of each sensor, unified timing management of the acquired multi-source data is necessary to ensure the coordinated accuracy of subsequent line-of-sight calculation and slip detection. The specific synchronization mechanism for multi-source data in this embodiment is described in detail below.
[0081] I. Multi-source data synchronization:
[0082] In this invention, the multi-source data includes: left-eye image, right-eye image, scene image, and IMU inertial data. Since the sampling frequency, transmission path, processing delay, and caching mechanism of the above data may differ, a unified synchronization mechanism is needed to ensure the temporal consistency of the data used in line-of-sight calculation.
[0083] II. Image Data Synchronization:
[0084] In this embodiment, the left-eye and right-eye cameras preferably acquire eye images at the same frame rate, such as 60fps, while the scene camera preferably acquires scene images at a lower frame rate, such as 30fps. Since scene images are primarily used to characterize the environment observed in front of the wearer, while eye images are used to extract pupil center and eye movement features, the scene camera and eye camera can operate at different frame rates to balance system bandwidth, processing load, and gaze tracking accuracy requirements. To maintain the temporal correspondence between multi-source data under different frame rate conditions, this embodiment uses timestamp marking under a unified time base and a cross-frame rate matching method to achieve image synchronization.
[0085] Specifically, a unified system clock is provided by the synchronization management module in the front-end control circuit or edge computing unit, and the corresponding acquisition timestamp is recorded when each frame of scene image, left-eye image, and right-eye image is received. Let the timestamp of the k-th frame in the scene image sequence be... The timestamp of the m-th frame in the left eye image sequence is The timestamp of the nth frame in the right eye image sequence is Because the frame rate of the left and right eye cameras is higher than that of the scene camera, two or more frames of eye images are usually captured between two adjacent scene images. Therefore, when constructing a synchronous data set based on scene images, the edge computing unit performs edge computing on each frame of scene image... Select the images with the closest timestamps from the left-eye and right-eye image sequences, respectively. The matching relationship between the image frames used as matching frames can be expressed as follows:
[0086]
[0087]
[0088] Thus, a synchronized image group corresponding to the k-th frame of the scene image is constructed:
[0089]
[0090] When satisfied
[0091]
[0092] At that time, the above three frames are regarded as the same synchronization group, where, A preset time threshold is used to match the eye image. Preferably, the time threshold can be set to half of the single frame period of the eye camera or preset according to the system's allowable error range.
[0093] In another embodiment, two eye images falling within the scene frame time window can also be fused, for example, by taking the closest frame, averaging the eye features of the two frames, or using interpolation to estimate the eye feature values corresponding to the scene frame time, in order to further reduce the timing error caused by cross-frame rate sampling.
[0094] For example, suppose the timestamps of two adjacent left-eye images are respectively and And satisfy
[0095]
[0096] Then, the left-eye feature quantity corresponding to the scene frame time can be estimated based on linear interpolation. :
[0097]
[0098] in, The image of the left eye at the timestamp The extracted eye feature vector is obtained from the image at point 1. The same logic applies to the right eye image. This allows for further temporal alignment of eye features with scene image time, thereby improving the temporal consistency of subsequent gaze mapping. For applications with high real-time requirements, nearest neighbor matching is preferred; for applications with high accuracy requirements and sufficient computational resources, interpolation alignment is preferred.
[0099] Through the unified timestamp synchronization mechanism at different frame rates, this embodiment does not require the scene camera and the left and right eye cameras to work at the same frame rate. It can maintain a stable temporal correspondence between scene images and eye images under conditions of low system bandwidth and low edge computing load, providing a reliable data foundation for subsequent gaze estimation, event recognition and multi-source fusion processing.
[0100] III. IMU Data Synchronization:
[0101] In this embodiment, the IMU module is used to output inertial measurement information of the head-mounted device during use. The inertial measurement information preferably includes three-axis angular velocity data and three-axis acceleration data, and in some embodiments, it may further include attitude information calculated by inertial analysis. Since the sampling frequency of the IMU is typically higher than the frame rate of the scene camera and the eye camera, for example, 100Hz, 200Hz, 400Hz, or higher, the IMU usually generates multiple sets of inertial sampling data within any image frame time interval. To ensure that the inertial information corresponds to the scene image and the eye image under a unified time reference, this embodiment also uses a unified timestamp and time-aligned synchronization method for the IMU data.
[0102] Specifically, when the IMU module outputs each set of inertial measurement values, the corresponding timestamp is recorded by the synchronization management module in the front-end acquisition and control circuit or the edge computing unit. Let the IMU at time... Output raw measurement values:
[0103]
[0104] in, Indicates time Three-axis angular velocity, Indicates time Triaxial acceleration.
[0105] Since gaze estimation is typically based on scene image time or synchronized image group time, therefore at any scene frame time... For or synchronize data group time At this point, the edge computing unit selects the corresponding inertial information from the IMU data sequence. If a timestamp satisfies:
[0106]
[0107] The inertial state at that moment can then be estimated by interpolation using two adjacent IMU sampling points, and its expression can be written as:
[0108]
[0109] in, Indicates time The corresponding interpolated inertial measurement values. Using this method, even when the IMU sampling frequency is higher than the image sampling frequency, time-corresponding inertial information can be provided for each frame of scene image or each group of synchronized images.
[0110] In another implementation, to improve the pose representation capability under short-term dynamic actions, instead of directly using a single interpolation result, multiple IMU sample values can be integrated or statistically processed within the image frame time window. For example, the timestamps of two adjacent scene images can be used. and By constructing a time window, angular velocity and acceleration can be averaged, integrated, or weighted within that time window to obtain short-term motion state quantities corresponding to that scene frame. Let there be a total of [number missing] within this window. With IMU sampling points, the average angular velocity of the window can be expressed as:
[0111]
[0112] The window-average acceleration can be expressed as:
[0113]
[0114] When the head rotates rapidly or there is instantaneous vibration interference, the above window statistics can better characterize the overall motion trend of the head-mounted device within the time period corresponding to the scene frame, which is beneficial for subsequent attitude calculation, slip recognition and online compensation.
[0115] Through the above-described IMU data synchronization method, this embodiment can achieve unified time alignment between inertial measurement information and scene images and eye images without requiring the IMU sampling frequency to be consistent with the image frame rate, thus providing a reliable data foundation for subsequent head pose analysis, multi-source fusion, and slip compensation.
[0116] IV. Construction of Unified Synchronization Data Packets:
[0117] After completing the timestamp marking and alignment of scene images, left and right eye images, and IMU data, the edge computing unit further constructs a unified synchronization data set, which serves as the standard input for gaze tracking calculation, slip detection, and online compensation. Since in this embodiment, the left and right eye cameras preferably operate at 60fps, the scene camera preferably operates at 30fps, and the IMU module preferably operates at 100Hz, the unified synchronization data set is preferably constructed using the scene image frame time as the reference time. That is, the timestamp corresponding to each frame of the scene image is used as the anchor point of the current processing cycle, and the data that best matches this anchor point time is extracted from the left and right eye image sequences and the IMU sequence.
[0118] Specifically, for the k-th frame scene image in the scene image sequence Its timestamp is The edge computing unit selects the images with the closest timestamps from the left-eye and right-eye image sequences, respectively. The image frame is denoted as... and And obtain the time according to the aforementioned IMU data synchronization method. Corresponding inertial measurement value Therefore, the k-th unified synchronization data group is constructed:
[0119]
[0120] In a preferred embodiment, if the time difference between the left and right eye images and the scene image is less than a preset image synchronization threshold, and there are valid sampling points or reliable interpolations in the IMU data near that moment, then the synchronized data set is considered valid and input into the subsequent processing module. If a data source is missing or the time deviation exceeds a preset threshold, the following processing strategy can be selected according to the system operating mode: In real-time preview mode, the current data set is allowed to participate in the display output in a degraded manner; in high-precision analysis mode, the data set is discarded or marked as a low-confidence sample and does not participate in compensation parameter updates and precise statistical analysis.
[0121] To further reduce timing errors caused by different sampling frequencies, this embodiment also allows interpolation alignment of eye image features rather than the original image frames. Specifically, the edge computing unit first extracts eye features from two left-eye images and two right-eye images at temporally adjacent scene frame times, and then estimates the eye feature values corresponding to the scene image times using interpolation. Let the left-eye image feature be... The right eye image features are Then at scene frame time The aligned binocular features can be obtained:
[0122]
[0123] Building upon this, the unified synchronization data set can also be further represented as a fused data set containing scene images, aligned binocular features, and inertial information:
[0124]
[0125] In this embodiment, the edge computing unit performs subsequent processing based on the constructed unified synchronous data set. First, it extracts the pupil center, eyeball contour, or other eye movement features based on the left and right eye images or binocular features, and calculates the point where the gaze falls in the scene image. Simultaneously, it calculates the motion state of the head-mounted device based on the inertial information at the corresponding time, and further combines the scene image motion information and eye configuration change information to determine whether there is a micro-slippage of the head-mounted device relative to the wearer's head. Since all the above inputs are aligned under a unified time reference, the impact of temporal mismatch on gaze estimation and slippage recognition can be effectively reduced.
[0126] Through the above-mentioned unified and synchronized data group construction method, this embodiment can achieve unified organization and fusion processing of multi-source data even when the sampling frequencies, data formats and transmission delays of scene cameras, left and right eye cameras and IMU modules are different, so as to provide technical support for the stable operation of the head-mounted eye tracking device of the present invention in high dynamic application scenarios.
[0127] V. Synchronization Anomaly Handling:
[0128] In actual operation, occasional frame loss, latency jitter, or temporary buffer congestion may occur. To address this, this embodiment implements the following exception handling strategies:
[0129] (1) When any camera frame is lost, the strategy of preserving the previous valid frame, local interpolation, or skipping the current frame is adopted.
[0130] (2) When the IMU data missing time is less than the threshold, nearest neighbor retention or linear extrapolation is used;
[0131] (3) When the time deviation between the image and the IMU exceeds the synchronization threshold, the current data packet is marked as low confidence and will not participate in the slip compensation parameter update. It is only used for preview display or downgraded output.
[0132] Through the above synchronous design, unified time-series management of multi-source data on the edge computing platform can be achieved by using commercially available general-purpose cameras and general-purpose IMU modules.
[0133] VI. Micro-slip detection and compensation mechanism:
[0134] This embodiment achieves micro-slip detection and hierarchical compensation of the head-mounted device relative to the wearer's head through the collaborative processing of eye image information, scene image information, and inertial data. Specifically, stable reference features in the eye image are used as the primary criterion for micro-slip of the head-mounted device relative to the wearer's head, while inertial data and scene image motion information are used to evaluate the reliability of the current multi-source data fusion and assist in compensation updates.
[0135] 1. Initial calibration and establishment of wearing reference state:
[0136] Upon first use, the system undergoes an initial calibration process. During calibration, the user sequentially fixates on multiple known scene reference points. The edge computing unit establishes an initial mapping model from eye features to scene coordinates based on eye features such as pupil center and pupil ellipse parameters in the left and right eye images, and the coordinates of the reference points in the scene image. Let the left and right eyes be at time... The feature vectors are as follows:
[0137]
[0138] These can then be combined to form a binocular feature vector:
[0139]
[0140] During calibration, the wearer is instructed to sequentially view multiple known target points in the scene. Their coordinates in the scene image pixel coordinate system are:
[0141]
[0142] Collect a set of sample pairs:
[0143]
[0144] A polynomial mapping is used to establish a mapping function from eye features to scene points:
[0145]
[0146] in, This represents the point where the gaze lands, as output by the initial gaze mapping model. For the preset mapping function, These are the initial mapping model parameters. The preset mapping function can be a multinomial mapping function, a linear regression model, a radial basis function model, a neural network model, or other mapping models from eye features to scene image coordinates.
[0147] The initial mapping model parameters are obtained by minimizing the following loss function:
[0148]
[0149] Simultaneously, the system records the positional relationship of stable reference features in the left and right eye images under initial calibration, serving as the initial wearing reference state for subsequent micro-slip detection. Under stable wearing conditions, certain structural relationships in the eye images should remain relatively stable for a short period. These features, unlike eye-movement features such as the pupil center and pupil ellipse direction which change with the gaze direction, can more directly reflect the relative positional changes of the eye image acquisition module relative to the wearer's head.
[0150] Therefore, one or more of the following features can be extracted:
[0151] (1) The positions of the inner and outer canthi points in the eye diagram;
[0152] (2) The position and direction of the short line at the edge of the eyelid or the local contour of the eyelid;
[0153] (3) The relative geometric relationship between the eyelid edge and the corner of the eye.
[0154] Let the ocular stability reference features recorded during the initial calibration be:
[0155]
[0156] in, This represents the pixel coordinates of the i-th stable reference feature point in the eye image coordinate system under the initial wearing reference state, and N is the number of valid reference feature points participating in the recording.
[0157] 2. IMU Initialization and Attitude Reference Establishment:
[0158] An inertial measurement unit (IMU) is fixedly mounted in a rigid area of the head-mounted device, and its output is used to describe the overall motion state of the head-mounted device. In this embodiment, the inertial data is mainly used to assist in evaluating the reliability of scene image motion estimation and to provide motion interpretation information for multi-source fusion determination.
[0159] Before system startup or initial calibration, the edge computing unit performs initialization processing on the inertial measurement module. Initialization processing includes stationary state sampling, zero-bias estimation, initial attitude recording, and output frequency setting. The data output frequency of the inertial measurement module can be set from 1Hz to 200Hz. In a preferred embodiment, the data output frequency is set to 100Hz or 200Hz, which is higher than the sampling frame rate of the scene image acquisition module and the eye image acquisition module.
[0160] Let the gyroscope output at the j-th inertial sampling time during the initialization phase be... The accelerometer output is Then, the gyroscope zero bias and accelerometer zero bias can be estimated based on multiple sets of static sampled values during the initialization phase:
[0161]
[0162]
[0163] in, This represents the number of inertial data sets collected during the initialization phase. To achieve zero bias in the gyroscope, This provides zero bias for the accelerometer. During operation, it can perform zero bias compensation, filtering, and attitude calculation on inertial measurement values.
[0164] Let the attitude output by the inertial measurement module in the initial stable wearing state be: The current output posture is Then, the change in posture at the current moment relative to the initial wearing state can be expressed as:
[0165]
[0166] When using rotation matrices or quaternions to represent posture, the motion trend of the head-mounted device can also be represented by posture increments relative to the initial posture. These posture changes are used for subsequent calculation of scene motion consistency residuals.
[0167] 3. Calculation of residual offset of eye-stabilized reference feature:
[0168] During the operation phase, the edge computing unit extracts the current eye-stabilized reference features in each valid synchronized data set.
[0169] Let the ocular stability reference features recorded during the initial calibration be: At present The extracted eye-stabilizing reference features are as follows:
[0170]
[0171] in, This represents the pixel coordinates of the feature point corresponding to the i-th initial eye-stabilized reference feature at the current time in the eye image coordinate system.
[0172] The ocular stability reference feature offset residual is used to measure the degree of difference between the current ocular stability reference feature and the baseline ocular stability reference feature under the initial wearing reference state. The residual can be expressed as:
[0173]
[0174] in, This represents the distance function or residual function.
[0175] In one implementation, the eye-stabilizing reference feature offset residual can be calculated based on the weighted average distance of the matched feature points:
[0176]
[0177] in, The weight of the i-th stable reference feature point in the eye is used to characterize the detection confidence or matching reliability of that feature point. When there is blinking, occlusion, strong reflection, or false detection of features, the weight of the corresponding feature point can be reduced, or low-confidence feature points can be removed.
[0178] To reduce the impact of single-frame detection noise on slip determination, a short-time sliding window can be used to smooth the residual offset of the ocular stability reference feature. The short-time sliding window refers to a time window consisting of several consecutive valid synchronized data sets selected backward from the current synchronized data set; its duration is 0.2s to 2.0s, or includes 5 to 60 consecutive valid synchronized data sets, preferably 0.5s to 1.0s, or includes 10 to 30 consecutive valid synchronized data sets. The historical features within the short-time sliding window are mainly used to smooth the residual, confirm whether the offset persists, or remove abnormal frames; they do not replace the baseline ocular stability reference features recorded under the initial wearing reference state.
[0179] 4. Calculation of residual for scene motion consistency:
[0180] Let adjacent scene image frames and Edge computing units can use optical flow, feature point matching, global image registration, or other image motion estimation methods to obtain motion features of the actual scene.
[0181]
[0182] The actual scene motion features can represent the overall translation, rotation, scale change, local optical flow statistics, or feature point matching motion trends in the scene image.
[0183] Based on the attitude change output by the inertial measurement module It can predict the theoretical motion trend that the scene image should present under the current motion state of the head-mounted device:
[0184]
[0185] in, This represents a function that predicts the motion trend of a scene image based on changes in inertial attitude.
[0186] By comparing the actual scene motion characteristics with the theoretical motion trends, the scene motion consistency residuals are obtained:
[0187]
[0188] When the scene motion consistency residual is small, it indicates that the inertial data can explain the scene image motion well, and the multi-source fusion of the current synchronized data group has high credibility. When the scene motion consistency residual is greater than the set credibility threshold, it is determined that the credibility of the current scene motion estimation or inertial data interpretation is reduced. This type of synchronized data group can be marked as a low credibility data group and will not participate in the compensation parameter update and online update of the mapping model.
[0189] 5. Line-of-sight mapping residual
[0190] When a known reference target exists, the line-of-sight mapping residual is calculated. The known reference target includes a preset calibration point, a fixed target on the system interface, a known target point in the training task, a visual marker point, or other reference targets whose locations are known.
[0191] When the system detects that the wearer is looking at a known reference target, let the coordinates of that known reference target in the scene image coordinate system be... The current mapping model outputs the gaze point as follows: Then the residual of the line-of-sight mapping is:
[0192]
[0193] When there is no known reference target at the current moment, or when it cannot be confirmed that the wearer is looking at the reference target, the gaze point mapping residual is not included in the comprehensive slip judgment or its corresponding weight is set to zero, so as to avoid unsupervised erroneous updates in the absence of real gaze point constraints.
[0194] 6. Comprehensive judgment of microslip and classification of slip degree:
[0195] The edge computing unit comprehensively evaluates the eye-stabilized reference feature offset residual, scene motion consistency residual, and gaze-point mapping residual. Among them, the eye-stabilized reference feature offset residual is the main criterion for micro-slip detection; the scene motion consistency residual is a constraint on the credibility of fused data; and the gaze-point mapping residual is only used as an auxiliary criterion when a known reference target exists.
[0196] In one implementation, a comprehensive slip determination index can be constructed:
[0197]
[0198] in, , and These represent the normalized eye-stabilized reference feature offset residual, scene motion consistency residual, and gaze point mapping residual, respectively. , , These are the weighting coefficients. Since the residual offset of the eye-stabilized reference feature is the primary basis for determining the microslippage of the head-mounted device relative to the wearer's head, therefore... Preferred > and When there is no known reference target, It can be zero; when the scene motion consistency residual exceeds the confidence threshold, the current synchronized data group will not participate in the compensation parameter update.
[0199] A head-mounted device can be determined to have a micro-slippage tendency when one of the following conditions is met:
[0200] (1) Comprehensive slip judgment index Greater than the first threshold and lasting for more than M frames;
[0201] (2) Eye-stabilized reference feature offset residual Greater than the second threshold and lasting for more than M frames;
[0202] (3) Eye-stabilized reference feature offset residual The rate of change per unit time is greater than a set rate of change threshold;
[0203] (4) When a known reference target exists, the line-of-sight mapping residual The residual shift of the eye-stabilized reference feature continues to increase within a short sliding window. Increase synchronously.
[0204] Where M is the preset number of consecutive frames, which can be set to 5 to 30 frames, preferably 10 to 20 frames.
[0205] After determining the existence of a slight slip trend, the edge computing unit further classifies the degree of slip into small slip, medium slip, and large slip based on the magnitude, duration, rate of change of the residual of the eye-stabilized reference feature offset, and the residual of the gaze point after compensation. When the residual of the eye-stabilized reference feature offset is small and the change of the gaze point is stable, it is determined to be small slip; when the residual of the eye-stabilized reference feature offset is in the medium range, or when there is a known reference target and the mapping residual can be reduced through local updates, it is determined to be medium slip; when the residual of the eye-stabilized reference feature offset is consistently large, or when the residual is still consistently higher than the set threshold after online compensation, it is determined to be large slip.
[0206] 7. Graded slip compensation
[0207] When a micro-slip is detected, the system does not immediately interrupt the acquisition, but performs graded slip compensation according to the degree of slip.
[0208] (1) Small slip compensation
[0209] For small slips, eye-motion feature coordinate correction based on eye-stabilizing reference features is preferred. The system estimates the eye coordinate correction transformation for slip compensation based on the matching relationship between the current and initial eye-stabilizing reference features.
[0210]
[0211] in, This represents a two-dimensional geometric transformation from the current eye image coordinate state to the initial calibration coordinate state. This describes the process of estimating a two-dimensional geometric transformation based on the matching relationship between the current stable eye reference features and the initial stable eye reference features. The two-dimensional geometric transformation can be a translation transformation, a rigid transformation, a similarity transformation, or a local affine transformation.
[0212] Taking affine transformation as an example, the eye coordinate correction transformation satisfies:
[0213]
[0214] in, It is a two-dimensional linear transformation matrix. This is a two-dimensional translation vector. The transformation parameters can be obtained by minimizing the error between the transformed current eye-stabilized reference features and the initial eye-stabilized reference features:
[0215]
[0216] After obtaining the eye coordinate correction transformation, this transformation is applied to the pupil center, major and minor axes of the pupil ellipse, pupil ellipse direction, or other eye movement features extracted at the current moment for gaze mapping. Let the eye movement feature vector at the current moment be... The eye-tracking feature vector corrected to the initial calibrated coordinate state is:
[0217]
[0218] Subsequently, the corrected eye-tracking feature vector is input into the initial gaze mapping model to obtain the gaze point after slip compensation:
[0219]
[0220] in, This is the compensated gaze point. This method does not change the initial gaze mapping model parameters and can quickly restore the consistency between the current eye movement features and the initial calibration coordinate state under small slip conditions.
[0221] (2) Medium slip compensation
[0222] For moderate slippage, dynamic correction of the line-of-sight landing point or local update of mapping parameters can be performed on the basis of small slippage compensation.
[0223] In one implementation, the pixel coordinate compensation offset can be estimated based on the eye-stabilized reference feature offset trend, the gaze point offset trend, or the known reference target error within a short-term sliding window. The short-term sliding window refers to a time window consisting of several consecutively selected effective synchronized data groups, ending with the current synchronized data group. The duration of the short-term sliding window is 0.2s to 2.0s, or includes 5 to 60 consecutive effective synchronized data groups; preferably, the duration of the short-term sliding window is 0.5s to 1.0s, or includes 10 to 30 consecutive effective synchronized data groups. Let the gaze point obtained after eye-tracking feature coordinate correction be... The pixel coordinate compensation offset is The further corrected line of sight is:
[0224]
[0225] In another implementation, when a known reference target, preset calibration point, fixed target on the system interface, or known target point in the training task exists, the mapping model parameters can be locally updated using effective observation data within a short-term sliding window. In this case, a compensation loss function containing mapping residual terms and parameter constraint terms can be constructed:
[0226]
[0227] in, This represents the set of valid synchronized data groups within a short-term sliding window. This represents the coordinates of a known reference target in the scene image coordinate system. This is the regularization coefficient. The parameter constraint term is used to limit the magnitude of change of the updated model parameters relative to the initial model parameters, preventing excessive drift of model parameters due to blinking, reflections, incorrect gaze judgments, or abnormal scene motion.
[0228] The locally updated mapping model parameters can be expressed as:
[0229]
[0230] The updated viewpoint can be represented as:
[0231]
[0232] During parameter updates, you can set the maximum update magnitude, learning rate, forgetting factor, residual descent condition, or low-confidence data removal rules to ensure that the mapping model only performs local compensation updates.
[0233] (3) Large slip compensation
[0234] When the residual offset of the eye-stabilized reference feature remains large, or when the residual of the line of sight landing point remains higher than the set threshold after small and medium slip compensation, the system determines that the current slip degree has exceeded the range that online compensation can stably handle, and triggers a rapid recalibration process.
[0235] The rapid recalibration process includes: the edge computing unit outputting a slip alarm or recalibration prompt; guiding the wearer to gaze at a small number of preset calibration points; simultaneously collecting the eye movement feature vectors and corresponding scene coordinates when the wearer gazes at these few preset calibration points; and locally updating the mapping model parameters based on the newly collected samples, or re-establishing the gaze mapping model in the current wearing state.
[0236] The number of preset calibration points is less than the number required for conventional initial calibration. Compared to complete recalibration, rapid recalibration requires less time, reducing interruptions during training, evaluation, or experimentation.
[0237] After completing dynamic compensation, local updates, or rapid recalibration, the system outputs the current gaze point result and uses the compensated model state or gaze point correction state for subsequent synchronous data group processing. If the eye stability reference feature offset residual, gaze point mapping residual, or comprehensive slip judgment index recovers to the set range after compensation, the system maintains the current compensation state and continues to run; if the residual continues to increase after compensation, the slip level is increased and a higher-level compensation or recalibration process is initiated.
[0238] The embodiments of this application have now been described in detail. To avoid obscuring the concept of this application, some details known in the art have not been described. Those skilled in the art can fully understand how to implement the technical solutions disclosed herein based on the above description.
[0239] While specific embodiments of this application have been described in detail by way of examples, those skilled in the art should understand that the above examples are for illustrative purposes only and are not intended to limit the scope of this application. Those skilled in the art should understand that modifications can be made to the above embodiments or equivalent substitutions can be made to some technical features without departing from the scope and spirit of this application. In particular, as long as there is no structural conflict, the various technical features mentioned in the embodiments can be combined in any manner.
Claims
1. A head-mounted eye-tracking device, characterized in that, It includes a head-mounted acquisition unit (1), a connection unit (2), and an edge computing unit (3); The head-mounted acquisition unit (1) is used to acquire images of the wearer's left and right eyes, images of the scene in front, and inertial data related to head movement; The connection unit (2) is connected between the head-mounted acquisition unit (1) and the edge computing unit (3) to realize data and power transmission; The edge computing unit (3) includes an embedded computing platform (12), which establishes a unified time reference and constructs a unified synchronization data group for the front scene image, left and right eye images and inertial data based on a unified system clock; The embedded computing platform (12) is also used to form an initial gaze mapping model in the initial calibration stage, calculate the gaze landing point based on the mapping model in the running stage, and integrate the eye stability reference feature offset residual, inertial data and scene image motion consistency residual to perform wear micro-slip detection and hierarchical compensation processing, thereby compensating and updating the gaze landing point or mapping parameters.
2. The head-mounted eye-tracking device according to claim 1, characterized in that, The head-mounted acquisition unit (1) is mounted on a lightweight head-mounted frame (4) with an open structure without lenses and an overall weight not exceeding 90g. The head-mounted acquisition unit (1) includes a scene image acquisition module (8), a left eye image acquisition module (9), a right eye image acquisition module (10), and an inertial measurement module (11). The inertial measurement module (11) is fixedly installed on the mounting base of the scene image acquisition module (8) or in the rigid area of the head-mounted frame (4).
3. The head-mounted eye-tracking device according to claim 2, characterized in that, The sampling frame rates of the left eye image acquisition module (9), the right eye image acquisition module (10), and the inertial measurement module (11) are all higher than the sampling frame rate of the scene image acquisition module (8); the left eye image acquisition module (9) and the right eye image acquisition module (10) are respectively equipped with near-infrared illumination modules to provide auxiliary illumination to the eye area.
4. The head-mounted eye-tracking device according to claim 1, characterized in that, When the edge computing unit (3) constructs a unified synchronization data group, for each frame of scene image, it selects the image frame with the timestamp closest to the time of the frame from the left and right eye image sequences as the matching frame, and matches the inertial data at the corresponding time.
5. A head-mounted gaze tracking and slip compensation method based on the device described in claim 1, characterized in that, Based on the aforementioned gaze mapping model, the following steps are included: S1. In the initial stable wearing state, the head-mounted acquisition unit collects multi-source synchronous data, performs initial calibration, determines the reference mapping parameters according to the preset eye feature to scene coordinate mapping model, and forms an initial mapping model based on the reference mapping parameters. S2. During operation, continuously receive front scene images, left and right eye images, and inertial data aligned with a unified time reference; S3. Extract eye stabilization reference features based on the received left and right eye images, and calculate the eye stabilization reference feature offset residual based on the change of the eye stabilization reference features relative to the initial stable wearing state or the short-term sliding window reference state. The eye stability reference features include one or more of the following: inner corner of the eye, outer corner of the eye, short line of eyelid margin, and the relative geometric relationship between eyelid margin and corner of the eye; Based on the received inertial data, the theoretical motion trend corresponding to the scene image is predicted, and the scene motion consistency residual is calculated in combination with the actual motion characteristics of the scene image. A comprehensive slip judgment index is constructed based on the eye stability reference feature offset residual, scene motion consistency residual, line of sight landing point error under known reference target and data quality status. When the comprehensive slip judgment index is greater than the set threshold or continuously meets the preset frame number condition, it is determined that the device has micro-slipped relative to the wearer's head. S4. After determining that there is micro-slip, the edge computing unit (3) performs dynamic compensation on the line of sight landing point, performs online compensation update on the parameters of the initial mapping model, or triggers the recalibration process according to the magnitude or credibility of the comprehensive slip judgment index.
6. The head-mounted gaze tracking slip compensation method according to claim 5, characterized in that, In step S1, the initial calibration specifically includes: The wearer is instructed to sequentially focus on multiple preset calibration reference points in the scene image; At each preset calibration reference point, the corresponding left and right eye images are acquired and the eye feature vectors are extracted; The eye feature vector and the corresponding scene image coordinates are used to form a calibration sample pair; The preset mapping function from eye feature vectors to scene image coordinates is invoked, and the baseline mapping parameters are determined based on minimizing the loss function, thereby constructing an initial mapping model.
7. The head-mounted gaze tracking and slip compensation method according to claim 5, characterized in that, In step S3, the construction of the comprehensive slip judgment index specifically involves: performing row normalization on the eye stability reference feature offset residual, scene motion consistency residual, and prediction error or mapping residual of the current gaze point, and then weighting and summing them according to preset weights to calculate the comprehensive slip judgment index. The eye-stabilizing reference feature offset residual is used to characterize the offset, rotation, or local affine change of the head-stabilizing reference features in the eye image relative to the initial wearing reference state during operation. It is the main basis for determining whether the head-mounted device has micro-slipped relative to the wearer's head. The scene motion consistency residual is used to characterize the consistency between the scene motion trend predicted by inertial data and the actual motion features of the scene image. It serves as an auxiliary credibility constraint for slip determination and compensation update. The prediction error or mapping residual of the current gaze point only participates in the calculation of the comprehensive slip determination index when the current gaze target is a known reference target, a preset calibration point, a fixed target on the system interface, or a known target point in the training task. The conditions under which the device determines a slight slippage trend include: the comprehensive slippage determination index is greater than the first threshold and continues for more than a set number of frames; Alternatively, the residual offset of the eye-stabilized reference feature is greater than the second threshold and continues for more than a certain number of frames; Alternatively, the rate of change of the eye-stabilized reference feature offset residual over a unit of time is greater than a set rate of change threshold. Alternatively, when a known reference target exists, the line-of-sight mapping residual continuously increases within a short sliding window, and the eye-stabilized reference feature offset residual increases synchronously.
8. The head-mounted gaze tracking slip compensation method according to claim 5, characterized in that, The online compensation update and recalibration process in step S4 includes one or more of the following methods: Method 1: When the slip is less than a set threshold, the local geometric correction amount in the eye image coordinate system is estimated based on the offset residual of the eye stable reference feature, and the eye features are corrected online. The corrected eye features are then input into the initial mapping model or the current mapping model to achieve dynamic compensation under small slip. Method 2: When the degree of slippage is in a moderate range and there is a known reference target, the mapping model parameters are locally updated online based on the mapping residual between the position of the known reference target in the scene image coordinate system and the current line of sight, combined with the effective observation data within the short-term sliding window. Method 3: When the degree of slippage continues to increase or the residual of the line of sight landing point is still higher than the set threshold after online compensation by Method 1 and Method 2, a prompt message is output and the recalibration process is triggered to guide the wearer to look at a small number of preset calibration points (the number of which is less than that required by regular calibration) to locally update the model parameters or re-establish the mapping model.
9. The head-mounted gaze tracking slip compensation method according to claim 8, characterized in that, Method 1 specifically includes: Extract stable reference features of the eye from the current eye image and match them with the positional relationship of the reference features in the initial calibration stage to estimate the translation, rotation, or local affine transformation parameters of the current eye image relative to the initial wearing reference state; based on the transformation parameters, perform reverse correction on the pupil center or eye feature vector of the current frame to obtain the corrected eye features; input the corrected eye feature vector into the initial mapping model or the current mapping model to calculate the compensated gaze point.
10. The head-mounted gaze tracking slip compensation method according to claim 8, characterized in that, Method 2 specifically includes: When the system detects that the wearer is gazing at a known reference target, it obtains the pixel coordinates of the known reference target in the scene image coordinate system and calculates the mapping residual between these pixel coordinates and the pixel coordinates of the gaze point output by the current mapping model. A compensation loss function containing the mapping residual term and parameter constraint term is constructed, and the mapping model parameters are locally updated based on minimizing the compensation loss function. The parameter constraint term limits the variation of the updated mapping model parameters relative to the initial mapping model parameters, preventing excessive drift in the mapping model due to short-term errors, blinking, reflections, abnormal scene motion, or incorrect gaze judgment. The updated mapping model parameters are used as the current online mapping model parameters for subsequent gaze point calculations. Method 3 specifically includes: When the residual offset of the stable reference feature of the eye is continuously greater than the set threshold in multiple consecutive synchronous data sets, or when the residual of the gaze landing point after compensation by method 1 and method 2 is still continuously greater than the set threshold, the edge computing unit outputs a recalibration prompt message and presents a small number of preset calibration points in the scene image or display interface; after the wearer looks at the small number of preset calibration points in sequence, the system collects the corresponding eye feature vector and the scene coordinates of the calibration points, and performs local updates on the mapping model parameters based on the collected samples, or re-establishes the gaze mapping model in the current wearing state.