Eye tracking method, control unit and eye tracking device
By acquiring structured light image data of the ocular surface, the geometric contribution of the cornea and the optical contribution of the tear film are decoupled and separated, enabling real-time optical compensation for dynamic changes in the tear film. This allows for the reconstruction of three-dimensional point cloud data, solving the problem of low accuracy in estimating the gaze direction and improving the accuracy and robustness of eye tracking.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- HUAQIN TECH CO LTD
- Filing Date
- 2026-03-13
- Publication Date
- 2026-08-04
AI Technical Summary
In existing eye-tracking technologies, the accuracy of gaze direction estimation is low, making it difficult for eye-tracking accuracy to meet the requirements of high-precision applications.
By acquiring structured light image data of the target eyeball surface reflection, extracting the wrapping phase distribution, and iteratively determining the tear film thickness distribution at the current moment based on the coupled surface model of the corneal basic shape and tear film thickness distribution, three-dimensional point cloud data is reconstructed, the viewing direction is determined, and real-time optical compensation for the dynamic changes of the tear film is achieved.
It significantly improves the accuracy and robustness of eye tracking, meeting the stringent requirements of sub-degree gaze estimation in clinical diagnosis and highly immersive virtual reality applications.
Smart Images

Figure CN121838247B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of eye-tracking technology, specifically to an eye-tracking method, a control unit, and an eye-tracking device. Background Technology
[0002] Eye-tracking technology, as one of the core means of human-computer interaction, has been widely applied in fields such as medical auxiliary diagnosis, virtual reality (VR), augmented reality (AR), autonomous driving, and the metaverse. Precise eye-tracking information not only enhances the naturalness and immersion of interaction but also provides objective quantitative evidence for visual function assessment, neurological disease diagnosis, and cognitive state analysis. Therefore, high-precision and robust eye-tracking methods have significant scientific research value and promising industrial application prospects.
[0003] In related technologies, the following two methods are typically used for eye tracking:
[0004] The first method is the pupil-corneal reflection method. For example, an infrared light source is used to form a reflective spot (such as a Pulcim spot) on the corneal surface. Combined with a high-speed camera to acquire eye images, the relative positional relationship between the pupil center and each reflective spot is extracted. Then, using a preset fixed-parameter eye geometry model, a mapping relationship between two-dimensional feature vectors and the gaze direction is established, and finally the coordinates of the user's gaze point or gaze angle are calculated.
[0005] The second method is a model fitting method based on the iris contour. For example, the iris image is obtained from the eye image, the center point of the iris image is calculated, and the center point is used as the two-dimensional coordinates of the pupil. Then, an individualized eyeball model is constructed based on the eye image, and the three-dimensional center and radius of the eyeball model are estimated to obtain the coordinates of the origin of the gaze and the radius of the eyeball. Finally, the gaze direction is determined by geometric projection relationship based on the two-dimensional coordinates of the pupil, the coordinates of the origin of the gaze, and the radius of the eyeball.
[0006] However, all of the above methods for eye tracking suffer from low accuracy in estimating the gaze direction, making it difficult to meet the requirements of high-precision applications. Summary of the Invention
[0007] This application provides an eye-tracking method, a control unit, and an eye-tracking device to solve the problem that the accuracy of gaze direction estimation is low in related technologies, which makes it difficult to meet the requirements of high-precision applications.
[0008] In a first aspect, this application provides an eye-tracking method, comprising:
[0009] Acquire structured light image data of the reflection from the surface of the target eyeball;
[0010] The wrapping phase distribution is extracted from the structured light image data, and the instantaneous normal vector of each point on the surface of the target eyeball is determined based on the wrapping phase distribution. The wrapping phase distribution represents the phase distribution of reflected light after being modulated by the basic shape of the cornea and the tear film thickness distribution.
[0011] Based on a pre-constructed coupled surface model containing the basic shape of the cornea and the tear film thickness distribution, the tear film thickness distribution at the current moment is determined iteratively according to the instantaneous normal vector and the initial value of the tear film thickness distribution of the target eyeball.
[0012] Based on the current tear film thickness distribution and corneal basic shape, reconstruct three-dimensional point cloud data representing the actual physical shape of the target eyeball surface;
[0013] Based on 3D point cloud data, the direction of the target eye's gaze is determined.
[0014] In one possible implementation, the coupled surface model satisfies the following form:
[0015]
[0016] in, This refers to the actual corneal surface height, including the tear film. For the basic shape of the cornea, Tear film thickness distribution; This is the residual term.
[0017] In one possible implementation, based on a pre-constructed coupled surface model containing the basic shape of the cornea and the tear film thickness distribution, the tear film thickness distribution at the current moment is iteratively determined according to the instantaneous normal vector and the initial value of the tear film thickness distribution of the target eyeball, including:
[0018] Based on the coupled surface model, determine the model normal vector corresponding to the current estimated tear film thickness distribution;
[0019] Construct an objective function that includes a difference term between the instantaneous normal vector and the model-based normal vector, as well as a physical rationality constraint term for the tear film thickness distribution. Adjust the relative importance of the difference term and the physical rationality constraint term by weighting coefficients.
[0020] The tear film thickness distribution is iteratively updated with the objective function minimized until the convergence condition is met, so as to obtain the tear film thickness distribution at the current time.
[0021] In one possible implementation, the three-dimensional point cloud data includes corneal region point clouds and scleral region point clouds. Based on the three-dimensional point cloud data, determining the gaze direction of the target eyeball includes:
[0022] Spherical fitting is performed on the point cloud of the corneal region to obtain the corneal center of the target eyeball;
[0023] Spherical fitting is performed on the point cloud of the scleral region to obtain the scleral center of the target eyeball;
[0024] Determine the direction of the target eye's line of sight based on the center of the cornea and the center of the sclera.
[0025] In one possible implementation, determining the line of sight direction of the target eye based on the corneal center and the scleral center includes:
[0026] Determine the direction of the optical axis of the target eyeball based on the corneal center and scleral center;
[0027] Based on the obtained Kappa angle parameters, the optical axis direction is corrected to obtain the line of sight direction of the target eye.
[0028] In one possible implementation, based on the current tear film thickness distribution and the basic corneal shape, three-dimensional point cloud data characterizing the actual physical shape of the target eyeball surface is reconstructed, including:
[0029] The tear film thickness distribution at the current moment is dynamically compensated and updated based on the change in the reflected light intensity of the target eyeball at the current moment relative to the reference reflected light intensity, and the tear film thickness distribution at the previous moment.
[0030] Based on the compensated and updated tear film thickness distribution and corneal basic shape, the current three-dimensional point cloud data is reconstructed.
[0031] In one possible implementation, extracting the wrapper phase distribution from structured light image data includes:
[0032] The phase deflection algorithm is used to demodulate the phase of the structured light image data and extract the wrapping phase distribution corresponding to at least two spatial directions.
[0033] In one possible implementation, the structured light image data is obtained by projecting a composite structured light pattern onto the surface of the target eyeball, the composite structured light pattern containing at least two sinusoidal fringe components with different spatial frequencies.
[0034] Secondly, this application provides an eye-tracking device, comprising:
[0035] The acquisition module is used to acquire structured light image data reflected from the surface of the target eyeball;
[0036] The processing module is used to extract the wrapping phase distribution from the structured light image data, and based on the wrapping phase distribution, determine the instantaneous normal vector of each point on the surface of the target eyeball. The wrapping phase distribution represents the reflected light phase distribution after being modulated by the basic shape of the cornea and the tear film thickness distribution.
[0037] The determination module is used to iteratively determine the tear film thickness distribution at the current moment based on a pre-built coupled surface model containing the basic shape of the cornea and the tear film thickness distribution, according to the instantaneous normal vector and the initial value of the tear film thickness distribution of the target eyeball.
[0038] The reconstruction module is used to reconstruct three-dimensional point cloud data representing the actual physical shape of the target eyeball surface based on the tear film thickness distribution and corneal basic shape at the current moment.
[0039] The determination module is also used to determine the gaze direction of the target eye based on 3D point cloud data.
[0040] In one possible implementation, the coupled surface model satisfies the following form:
[0041]
[0042] in, This refers to the actual corneal surface height, including the tear film. For the basic shape of the cornea, Tear film thickness distribution; This is the residual term.
[0043] In one possible implementation, the determining module is specifically used to: determine the model normal vector corresponding to the current estimated tear film thickness distribution based on the coupled surface model; construct an objective function, which includes a difference term between the instantaneous normal vector and the model normal vector, as well as a physical rationality constraint term for the tear film thickness distribution, and adjust the relative importance of the difference term and the physical rationality constraint term through weighting coefficients; and iteratively update the tear film thickness distribution with the goal of minimizing the objective function until the convergence condition is met, so as to obtain the tear film thickness distribution at the current moment.
[0044] In one possible implementation, the three-dimensional point cloud data includes corneal region point clouds and scleral region point clouds. The determination module is further configured to: perform spherical fitting on the corneal region point cloud to obtain the corneal center of the target eye; perform spherical fitting on the scleral region point cloud to obtain the scleral center of the target eye; and determine the gaze direction of the target eye based on the corneal center and the scleral center.
[0045] In one possible implementation, the determining module is further configured to: determine the optical axis direction of the target eyeball based on the corneal center and the scleral center; and correct the optical axis direction based on the obtained Kappa angle parameters to obtain the line of sight direction of the target eyeball.
[0046] In one possible implementation, the reconstruction module is specifically used to: dynamically compensate and update the tear film thickness distribution at the current moment based on the change in the reflected light intensity of the target eyeball at the current moment relative to the reference reflected light intensity and the tear film thickness distribution at the previous moment; and reconstruct the three-dimensional point cloud data at the current moment based on the compensated and updated tear film thickness distribution and the basic shape of the cornea.
[0047] In one possible implementation, the processing module is specifically used to: employ a phase deflection algorithm to perform phase demodulation on the structured light image data and extract the wrapping phase distribution corresponding to at least two spatial directions.
[0048] In one possible implementation, the structured light image data is obtained by projecting a composite structured light pattern onto the surface of the target eyeball, the composite structured light pattern containing at least two sinusoidal fringe components with different spatial frequencies.
[0049] Thirdly, this application provides a control unit, including: a memory and a processor;
[0050] The memory stores instructions that the computer executes;
[0051] The processor executes computer execution instructions stored in memory, causing the processor to perform the first aspect and / or various possible implementations of the first aspect as described above.
[0052] Fourthly, this application provides a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, are used to implement the first aspect and / or various possible embodiments of the first aspect.
[0053] Fifthly, this application provides a computer program product, including a computer program that, when executed by a processor, implements the first aspect and / or various possible implementations of the first aspect.
[0054] Sixthly, this application provides an eye-tracking system, comprising:
[0055] The projection unit is used to project incident light carrying a structured light pattern onto the surface of the target eyeball;
[0056] The image acquisition unit is used to acquire structured light image data formed after reflection from the surface of the target eyeball.
[0057] The control unit is electrically connected to the projection unit and the image acquisition unit, respectively, and is used to execute the first aspect and / or various possible implementations of the first aspect.
[0058] In a seventh aspect, this application provides an eye-tracking device that integrates an eye-tracking system as described in the sixth aspect, or an eye-tracking control unit as described in the third aspect.
[0059] The eye-tracking method, control unit, and eye-tracking device provided in this application include: acquiring structured light image data of reflections from the surface of a target eyeball; extracting the wrapping phase distribution from the structured light image data, and determining the instantaneous normal vector of each point on the surface of the target eyeball based on the wrapping phase distribution, wherein the wrapping phase distribution characterizes the reflected light phase distribution modulated by the corneal basic shape and tear film thickness distribution; iteratively determining the tear film thickness distribution at the current moment based on a pre-constructed coupled surface model containing the corneal basic shape and tear film thickness distribution, according to the instantaneous normal vector and the initial value of the tear film thickness distribution of the target eyeball; reconstructing three-dimensional point cloud data characterizing the actual physical shape of the surface of the target eyeball based on the tear film thickness distribution and the corneal basic shape at the current moment; and determining the gaze direction of the target eyeball based on the three-dimensional point cloud data. This application acquires structured light image data of the target eye's surface reflection and extracts the encapsulated phase distribution modulated by the corneal basic shape and tear film thickness distribution. This allows for the determination of the instantaneous normal vector at each point on the eye's surface, accurately capturing the impact of tear film dynamics on optical measurements. Based on this, a pre-constructed coupled surface model incorporating the corneal basic shape and tear film thickness distribution is used. The instantaneous normal vector and initial value of the tear film thickness distribution are combined to iteratively solve for the current tear film thickness distribution. This effectively decouples the corneal geometric contribution from the tear film optical contribution, eliminating the systematic errors caused by neglecting the tear film or using a fixed geometric model in traditional methods. Furthermore, by iteratively updating the current tear film thickness distribution, real-time monitoring of tear film dynamics is achieved. Optical compensation effectively eliminates the interference of tear film interference and thickness fluctuations on gaze direction estimation. Based on the decoupled tear film thickness distribution and corneal basic shape, the actual physical shape of the target eyeball surface is reconstructed, achieving accurate modeling and compensation for individual differences in the tear film. By reconstructing three-dimensional point cloud data of the entire ocular surface, the rich geometric information of the ocular surface is fully utilized, significantly improving information density and surface reconstruction accuracy, and providing more complete geometric constraints for gaze direction estimation. Finally, the gaze direction is determined based on three-dimensional point cloud data including tear film dynamic compensation. Under the premise of fully considering individual corneal differences and tear film dynamic changes, the accuracy and robustness of eye tracking are significantly improved, which can meet the stringent requirements of sub-degree gaze estimation in clinical diagnosis and highly immersive virtual reality applications. Attached Figure Description
[0060] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.
[0061] Figure 1 Flowchart of the eye-tracking method provided in the embodiments of this application Figure 1 ;
[0062] Figure 2 Flowchart of the eye-tracking method provided in the embodiments of this application Figure 2 ;
[0063] Figure 3 This is a schematic diagram of the structure of the eye-tracking device provided in the embodiments of this application;
[0064] Figure 4 This is a schematic diagram of the structure of the eye-tracking system provided in the embodiments of this application;
[0065] Figure 5 This is a schematic diagram of the structure of the control unit provided in the embodiments of this application;
[0066] Figure 6 This is a schematic diagram of the structure of the eye-tracking device provided in the embodiments of this application.
[0067] The accompanying drawings illustrate specific embodiments of this application, which will be described in more detail below. These drawings and descriptions are not intended to limit the scope of the concept in any way, but rather to illustrate the concept of this application to those skilled in the art through reference to particular embodiments. Detailed Implementation
[0068] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this application as detailed in the appended claims.
[0069] The terms “first,” “second,” etc., used in the specification and claims of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented, for example, in orders other than those illustrated or described herein. Furthermore, the terms “comprising” and “having,” and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, products, or apparatus.
[0070] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, use and processing of the relevant data must comply with relevant laws, regulations and standards, and corresponding operation entry points are provided for users to choose to authorize or refuse.
[0071] In related technologies, the first approach relies on sparse reflected light spots (typically only 2-6 light sources corresponding to approximately 8-12 effective sampling points) to characterize the complex corneal surface. This sampling density is far below the requirements of the Nyquist sampling theorem, failing to fully sample the rich geometric information of the eye surface (unable to completely acquire corneal curvature and aspherical characteristics). This results in the inability to break through the theoretical upper limit of sub-degree (<0.5°) gaze direction estimation. Furthermore, the use of a fixed-parameter eye geometry model ignores significant differences in corneal curvature radius (typically 7.2-8.5 mm), aspherical coefficient (typically -0.35 to -0.05), and tear film thickness (typically 3-40 μm) between individuals, potentially leading to large (e.g., exceeding 0.5°) accuracy drift in different subjects using the same eye-tracking device. The second approach simplifies the eyeball into a standard spherical model, ignoring the differences in the bispherical structure of the cornea and sclera, as well as the aspherical characteristics of the cornea. Furthermore, iris center detection is susceptible to changes in illumination, eyelid occlusion, and pupil dilation, leading to ambiguity in the two-dimensional to three-dimensional projection inversion. In addition, this approach also ignores the dynamic optical effects of the tear film, failing to consider the refraction and interference effects of the tear film as an optical thin film on the iris image. Uneven tear film thickness distribution and its changes over time (e.g., tear film flow velocity decreasing from 4.2 mm / s to 0.8 mm / s within 9 seconds after blinking) can cause phase shifts and intensity fluctuations in the reflected light spot. These optical artifacts are misinterpreted by the system as eye movements, introducing an additional jitter error of 0.2°~0.4°, further reducing the stability of gaze estimation.
[0072] In summary, the relevant technologies are limited by issues such as sparse sampling, model simplification, and poor adaptability to individual differences, making it difficult to achieve sub-degree level (e.g., less than 0.5°) eye-tracking accuracy, and thus unable to meet the stringent requirements of applications such as clinical diagnosis, human-computer interaction, and high-end virtual reality.
[0073] To address the aforementioned issues, this application provides an eye-tracking scheme. It acquires structured light image data reflecting across the entire ocular surface using structured light illumination, extracts the encapsulated phase distribution modulated by the corneal basic shape and tear film thickness distribution, and decouples and separates the corneal geometric contribution from the tear film optical contribution based on this phase distribution. By constructing a coupled surface model incorporating the corneal basic shape and tear film thickness distribution, iteratively solving for the tear film thickness distribution at the current moment, real-time optical compensation for dynamic changes in the tear film is achieved. Furthermore, high-density three-dimensional point cloud data of the entire ocular surface is reconstructed, fully utilizing the rich geometric information of the eyeball surface. Based on fully considering individual corneal differences and real-time tear film dynamic compensation, high-precision and robust gaze direction estimation is achieved.
[0074] First, the application scenarios of this application will be introduced.
[0075] This application applies to scenarios with extremely high requirements for eye-tracking accuracy and stability, such as VR, AR, medical diagnosis (e.g., eye disease screening, neurodegenerative disease assessment), and autonomous driving (driver attention monitoring). In VR and AR, real-time tracking with sub-degree accuracy (e.g., less than 0.5°) is needed to achieve natural interaction; in medical diagnosis, long-term stable tracking is required to capture minute physiological abnormalities; in autonomous driving, low-latency (e.g., less than 100ms) attention monitoring is required to ensure safety. Related technologies are limited by individual physiological differences (e.g., tear film thickness fluctuations, corneal curvature changes) and algorithm efficiency issues, making it difficult to meet the requirements of these scenarios.
[0076] The technical solution of this application and how the technical solution of this application solves the above-mentioned technical problems are described in detail below with specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments. The embodiments of this application will now be described with reference to the accompanying drawings.
[0077] Figure 1 Flowchart of the eye-tracking method provided in the embodiments of this application Figure 1 This eye-tracking method can be applied to the control unit in an eye-tracking system, which is electrically connected to the projection unit and image acquisition unit in the eye-tracking system.
[0078] like Figure 1 As shown, this eye-tracking method includes the following steps:
[0079] S101. Acquire structured light image data of the reflection from the surface of the target eyeball.
[0080] Structured light image data is reflected image data that is projected onto the surface of the target eyeball using a structured light pattern (such as a sinusoidal stripe pattern, where the stripe period can be preset according to the eyeball size and imaging field of view, typically 2-5 line pairs per millimeter to ensure recognizable stripe deformation in areas of corneal curvature variation) and simultaneously acquired by an image acquisition unit (such as a near-infrared camera). This structured light image data contains information on the light intensity distribution modulated by the eyeball surface.
[0081] For example, based on the system calibration parameters and target measurement requirements, a composite sinusoidal fringe pattern for projection is generated, which can be represented in the display coordinate system as follows: ,in The frequency in the horizontal direction, The frequencies are in the vertical direction, and the unit is cycles per pixel. This is the background intensity, used to ensure that the pattern grayscale is always positive. To modulate the intensity, control the stripe contrast; These are the pixel coordinates in the horizontal direction of the display. The pixel coordinates are in the vertical direction of the display. The composite sinusoidal fringe pattern is characterized by superimposing sinusoidal fringes in both the horizontal and vertical directions to form a grid-like interference pattern. This composite design allows phase information in both directions to be encoded simultaneously in a single projection. The generated composite sinusoidal fringe pattern is displayed on a high-resolution display and projected onto the target eye surface via a projection unit. Simultaneously, a synchronous trigger signal controls two high-speed cameras to simultaneously acquire distorted images reflected from the target eye surface. During acquisition, the left and right cameras are strictly synchronized to ensure that images captured at the same moment reflect the same eye state.
[0082] The acquired raw image data can be further preprocessed to obtain structured light image data. The preprocessing includes distortion correction, grayscale normalization, and region of interest extraction, thereby extracting the effective region containing the cornea and part of the sclera from the normalized image as input data for subsequent phase calculation, so as to reduce the amount of computation and improve processing efficiency.
[0083] S102. Extract the wrapping phase distribution from the structured light image data, and determine the instantaneous normal vector of each point on the surface of the target eyeball based on the wrapping phase distribution. The wrapping phase distribution represents the phase distribution of reflected light after being modulated by the basic shape of the cornea and the thickness distribution of the tear film.
[0084] In this embodiment, the basic corneal shape is a corneal geometric model obtained through individualized measurements, including parameters such as the corneal radius of curvature and aspheric coefficient, characterizing the ideal geometric shape of the corneal surface without tear film coverage. The tear film thickness distribution represents the change in the thickness of the tear film on the corneal surface at the current moment. It is a spatial distribution function that dynamically changes over time, and its value range is typically between 3 and 40 μm.
[0085] The encapsulation phase distribution is the result of the reflected light phase being modulated by the corneal basal shape and tear film thickness distribution, and its value range is limited to the interval (-π,π] or [0,2π). It should be understood that when a structured light pattern is projected onto the target eye surface, the phase of the reflected light is modulated by two factors: the geometric optical path determined by the corneal basal shape and the additional optical path determined by the tear film thickness. These two factors superimpose to form the coupled information in the encapsulation phase. The former is determined by corneal curvature and aspherical characteristics and is relatively stable; the latter is proportional to the tear film thickness and changes dynamically over time.
[0086] For example, in some embodiments, extracting the wrapping phase distribution from structured light image data includes: using a phase deflection algorithm to demodulate the phase of the structured light image data and extracting the wrapping phase distribution corresponding to at least two spatial directions.
[0087] The spatial direction includes both horizontal and vertical directions. For example, the extraction of the wrapper phase distribution can be achieved through the following two methods:
[0088] In one implementation, a frequency-domain phase demodulation method based on Fourier transform is employed. Utilizing the spectral distribution characteristics of the composite fringe image in the frequency domain, phase information in the horizontal and vertical directions is separated through frequency-domain filtering, and then the wrapping phase value of each pixel in the structured light image data is calculated. Specifically, a two-dimensional Fourier transform is performed on the structured light image data to convert it from the spatial domain to the frequency domain, obtaining a complex spectral distribution; the real and imaginary parts are extracted, for example, the wrapping phase distribution in the horizontal direction is calculated using the following formula: The vertical phase distribution of the package is calculated using the following formula: , ( () represents the camera pixel coordinates. This is a horizontal window function used to extract the horizontal frequency components from the Fourier spectrum. is a vertical window function used to extract the vertical frequency components in the Fourier spectrum. Im[·] represents taking the imaginary part of the complex spectrum distribution, and Re[·] represents taking the real part of the complex spectrum distribution.
[0089] Another implementation employs a spatial domain phase demodulation method based on Gabor filters. Leveraging the excellent time-frequency localization characteristics of Gabor filters, local phase information of each pixel is directly extracted in the spatial domain. Orthogonal phase distributions are then obtained through multi-directional filtering and phase decoupling. Specifically, a multi-directional Gabor filter bank (such as a two-dimensional Gabor filter with eight equally spaced directions) is constructed to perform convolution filtering on the structured light image data, obtaining the complex response distribution in each direction. , In the formula, This represents a two-dimensional convolution operation. For structured light image data, This is a Gabor filter. Further, the local phase of each pixel is extracted from the filtered response, for example, for each pixel... and each direction Local phase value Then, through direction-selective optimization and phase decoupling, the wrap-around phase distributions in the horizontal, vertical, and diagonal directions are obtained. For example, the direction with the largest response amplitude among the eight directions is selected as the dominant fringe direction at that point, and the corresponding phase value is selected as the optimal local phase at that point.
[0090]
[0091] In the formula, Furthermore, based on the constructed 3×8-dimensional decoupling matrix M_decomp (a constant coefficient matrix satisfying least squares constraints, pre-determined through offline calibration or theoretical derivation), the local phases in eight directions are mapped onto orthogonal bases in the horizontal, vertical, and diagonal directions:
[0092]
[0093] In the formula, The phase distribution is a horizontal wrap-around pattern. The phase distribution is a vertical wrap-around pattern; The phase distribution is a diagonal wrap-around pattern.
[0094] This embodiment employs a phase deflection algorithm, requiring only a single frame of structured light image to simultaneously extract the packaged phase distribution corresponding to at least two spatial directions. Compared to the traditional multi-step phase-shifting deflectometry (PMD), which requires acquiring 8-12 frames of images and introduces a system latency of over 100 milliseconds, this embodiment compresses the data acquisition to a single frame, significantly reducing the time overhead of image acquisition and transmission, and compressing the system latency to below the millisecond level, thus meeting the real-time eye-tracking requirements of 120Hz or even higher frame rates.
[0095] It should be understood that the instantaneous normal vector includes the unit normal direction vector of each point on the surface of the target eyeball at the current moment, which can be calculated from the geometric relationship of reflected light rays, and comprehensively reflects the surface orientation under the combined effect of the basic shape of the cornea and the tear film thickness distribution.
[0096] Correspondingly, based on the envelope phase distribution, the instantaneous normal vector of each point on the surface of the target eyeball can be determined in the following way:
[0097] For example, based on the extracted wrapper phase distribution, a precise correspondence between camera pixels and display pixels is established. Specifically, for each camera pixel coordinate ( According to its horizontal wrapping phase, and vertical wrap phase Calculate the corresponding display coordinates ( ), ( )= ,in The stripe period.
[0098] Then, calculate the direction vector of the reflected ray for each corresponding point. and the direction vector of the incident ray , Let be the three-dimensional coordinates of a point on the surface of the eyeball to be solved. Based on the geometric constraint that the angle of incidence equals the angle of reflection in the law of reflection, the surface normal vector is solved using the direction vectors of the incident and reflected rays: This allows us to obtain the instantaneous normal vectors of each point on the surface of the target eyeball.
[0099] S103. Based on a pre-constructed coupled surface model containing the basic shape of the cornea and the tear film thickness distribution, the tear film thickness distribution at the current moment is determined iteratively according to the instantaneous normal vector and the initial value of the tear film thickness distribution of the target eyeball.
[0100] Among them, the coupled surface model is a mathematical expression that couples the basic shape of the cornea with the tear film thickness distribution, and is used to describe the complete corneal surface morphology including the tear film.
[0101] It should be understood that the instantaneous normal vector contains comprehensive information about the corneal basal shape and tear film thickness distribution. Since the corneal basal shape is relatively stable and can be measured in advance, while the tear film thickness distribution changes dynamically over time, it is necessary to decouple the two using mathematical methods. This step employs an iterative optimization strategy, using the measured instantaneous normal vector as a constraint and the tear film thickness distribution as the variable to be solved. By continuously adjusting the estimated tear film thickness, the theoretical normal vector calculated based on the complete eyeball surface model (i.e., the coupled surface model) gradually approximates the measured normal vector, ultimately obtaining a tear film thickness distribution consistent with optical measurements and conforming to physical laws.
[0102] In step S103, before starting iterative optimization, a reasonable initial estimate of the tear film thickness distribution needs to be set. Considering the individual variability in tear film thickness, the initial value of the tear film thickness distribution of the target eye can be calibrated at the beginning of eye tracking. For example, the eye parameters of the target eye (individualized) can be measured using a preset measurement method. These eye parameters include: corneal radius of curvature, scleral radius of curvature, corneal asphericity, basic corneal shape, and the initial value of the tear film thickness distribution. The calibration of the initial value of the tear film thickness distribution can be achieved in the following way:
[0103] In one implementation, the initial value of the tear film thickness distribution is calibrated based on an optical reflection model that considers the tear film interference effect. This optical reflection model is expressed as:
[0104]
[0105] in, The direct reflection component of the corneal surface can be calculated from the measured individualized corneal baseline shape; The scattering component can be determined by empirical formulas obtained from system calibration or by a pre-calibrated scattering model; This refers to the tear film interference reflection component; To extract the actual measured total reflected light intensity from structured light image data.
[0106] According to the expression for the tear film interference reflection component:
[0107]
[0108] in, The air-tear film interface reflectance (calculated from a preset refractive index); The reflectance coefficient at the tear film-corneal interface (calculated from a preset refractive index); The incident light wavelength (850nm or 940nm); The change in the incident angle as calibrated by the system; The refractive index of the tear film can be set to 1.336. Given the incident light intensity, the initial tear film thickness distribution can be obtained by numerically solving the above equation. The estimated value. Optionally, since the above equation is a nonlinear equation and has multivaluedness, it can be solved by the following method: performing a grid search within the physiological range of tear film thickness (3-40 μm) to find the value that makes the theoretical calculation consistent with the measured value. The closest thickness value is used as the initial value of the tear film thickness; or the initial value of the tear film thickness is uniquely determined by using a combination of equations based on multi-wavelength or multi-angle measurement data.
[0109] Another implementation employs a tear film thickness estimation algorithm based on multi-frequency phase measurement. This leverages the difference in sensitivity to tear film thickness at different spatial frequencies to decouple corneal and tear film contributions. To meet the requirements of multi-frequency phase measurement, in some embodiments, structured light image data is obtained by projecting a composite structured light pattern onto the target eye surface. This composite structured light pattern contains sinusoidal fringe components at at least two different spatial frequencies. For example, the composite structured light pattern is a composite pattern containing a first spatial frequency (e.g., below 5 cycles / mm, used to acquire background phase information) and a second spatial frequency (e.g., above 15 cycles / mm, used to acquire tear film attenuation information). Correspondingly, the wrap-around phase distribution of the two different spatial frequencies, including the low-frequency phase response, can be extracted from the structured light image data. (corresponding to f < 5 cycles / mm) and high-frequency phase response (Corresponding to f > 15 cycles / mm). According to the principle of multi-frequency phase measurement, the two satisfy the following relationship:
[0110]
[0111]
[0112] in, The background phase is contributed by both the basic shape of the cornea and systematic errors. The coherence length is obtained through system calibration.
[0113] Solving the two equations above simultaneously, the background phase is eliminated. , get about The equation:
[0114] .
[0115] Furthermore, the fixed-point iteration method is used to solve the equation, with the iteration format as follows:
[0116]
[0117] In the formula, The iteration count is [number], and iterations continue until the difference between two consecutive iterations is less than a preset threshold (e.g., 0.1 μm). Convergence is then achieved. This is the initial value of the tear film thickness at the current pixel. The above iterative solution is performed pixel-by-pixel across the entire measurement area to obtain the initial value of the tear film thickness distribution across the entire field.
[0118] S104. Based on the tear film thickness distribution and corneal basic shape at the current moment, reconstruct three-dimensional point cloud data representing the actual physical shape of the target eyeball surface.
[0119] The three-dimensional point cloud data includes a discrete set of three-dimensional points that represent the actual physical shape of the eyeball surface, obtained through reconstruction of the eyeball surface. It may include point cloud data of the corneal region and the sclera region, and is used for subsequent calculation of the gaze direction.
[0120] This step aims to fuse the current tear film thickness distribution obtained after decoupling with the pre-measured corneal basal shape to reconstruct high-density three-dimensional point cloud data containing tear film dynamics, providing an accurate geometric basis for subsequent gaze direction estimation.
[0121] For example, the tear film thickness distribution at the current moment is superimposed point-by-point on the basic shape of the cornea. Since both are defined in the same two-dimensional coordinate system (usually with the corneal vertex as the origin and the xy plane as the corneal basal plane), pixel-level addition can be performed directly. Through this superposition operation, the complete corneal surface height value containing the tear film at each sampling point can be obtained. The obtained complete corneal surface after superposition is discretely sampled to generate three-dimensional point cloud data, with each sampling point corresponding to a three-dimensional coordinate ( ),in and For planar coordinates, This represents the surface height value.
[0122] If the sampling density matches the resolution of the original image, it can reach the megapixel level. Such high-density sampling ensures that the rich geometric information of the eye surface is fully preserved, providing sufficient geometric constraints for subsequent gaze direction estimation.
[0123] Optionally, the generated 3D point cloud data is organized and output in a structured manner, including: each point contains 3D coordinates ( (and its corresponding attribute information, such as confidence level, region identifier (cornea or sclera), normal vector, etc.)
[0124] Optionally, to further improve the quality and usability of 3D point cloud data, the following enhancement processes can be combined: outlier removal (including identifying and removing outliers caused by noise, reflection, or occlusion based on local consistency checks, for example, calculating the average distance between each point and its neighboring points, and marking them as outliers and removing them if the deviation exceeds 3 times the standard deviation), hole filling (including filling holes for local data loss caused by eyelid occlusion or weak reflection signals using interpolation methods, commonly including bilinear interpolation, radial basis function interpolation, or partial differential equation-based repair algorithms), and smoothing filtering (including performing appropriate smoothing on the point cloud while preserving geometric features to suppress measurement noise, for example, using Gaussian filtering, bilateral filtering, or anisotropic diffusion filtering, and adaptively adjusting the smoothing intensity according to the local curvature).
[0125] S105. Based on 3D point cloud data, determine the direction of the target eye's gaze.
[0126] It should be understood that the direction of the eye's gaze can be inferred from the geometric features of the eye's surface. Three-dimensional point cloud data fully records the spatial coordinates of the cornea and surrounding sclera, containing geometric attributes such as the eye's position, orientation, and surface curvature. By analyzing the spatial distribution characteristics of the point cloud data and fitting a parametric model characterizing the eye's geometry, the direction of the eye's axis of symmetry can be determined. Combined with the physiological correspondence between the gaze direction and the eye's geometric axis, the current gaze direction can be obtained.
[0127] For example, a model fitting method is used to estimate key parameters characterizing the geometry of the eyeball. For instance, the eyeball is approximated as a sphere, and the best-fitting center position and radius are fitted by minimizing the sum of squared distances from the point cloud to the sphere's surface. This center approximately corresponds to the eyeball's rotation center. Differential geometric features such as local curvature and normal vectors at each point in the point cloud data are calculated to analyze the curvature distribution pattern of the eyeball surface and identify the geometric boundary between the corneal and scleral regions. Based on the fitted center position and the surface features of the point cloud, the direction of the eyeball's geometric axis of symmetry is determined. This axis of symmetry reflects the principal axis direction of the eyeball as an approximate rotating body and has a definite geometric relationship with the gaze direction. Combining the prior correspondence between the gaze direction and the eyeball's geometric axis, the gaze direction at the current moment is calculated.
[0128] This application embodiment acquires structured light image data reflected from the target eyeball surface and extracts the wrapping phase distribution modulated by the corneal basic shape and tear film thickness distribution. This allows for the determination of the instantaneous normal vector at each point on the eyeball surface, accurately capturing the impact of dynamic tear film changes on optical measurements. Furthermore, based on a pre-constructed coupled surface model incorporating the corneal basic shape and tear film thickness distribution, the current tear film thickness distribution is iteratively solved using the instantaneous normal vector and initial tear film thickness distribution value. This effectively decouples the corneal geometric contribution from the tear film optical contribution, eliminating the systematic errors caused by ignoring the tear film or using a fixed geometric model in traditional methods. Simultaneously, by iteratively updating the current tear film thickness distribution, the dynamic changes of the tear film are accurately assessed. Real-time optical compensation effectively eliminates the interference of tear film interference and thickness fluctuations on gaze direction estimation. Based on the decoupled tear film thickness distribution and corneal basic shape, the actual physical shape of the target eyeball surface is reconstructed, achieving accurate modeling and compensation for individual differences in the tear film. By reconstructing three-dimensional point cloud data of the entire ocular surface, the rich geometric information of the ocular surface is fully utilized, significantly improving information density and surface reconstruction accuracy, and providing more complete geometric constraints for gaze direction estimation. Finally, the gaze direction is determined based on three-dimensional point cloud data including tear film dynamic compensation. Under the premise of fully considering individual corneal differences and tear film dynamic changes, the accuracy and robustness of eye tracking are significantly improved, which can meet the stringent requirements of sub-degree gaze estimation in clinical diagnosis and highly immersive virtual reality applications.
[0129] Optionally, considering that although single-frame measurement can obtain the gaze direction at the current moment, it may be affected by factors such as measurement noise, instantaneous fluctuations of the tear film and micro-movements of the eyeball, and random jitter may exist between adjacent frames, in order to further improve the stability and accuracy of gaze estimation, Kalman filtering is introduced to fuse and optimize the multi-frame measurement results. Through multi-frame fusion optimization using Kalman filtering, random jitter in the eye movement trajectory is effectively smoothed.
[0130] In some embodiments, the coupled surface model satisfies the following form:
[0131]
[0132] in, The actual corneal surface height, including the tear film, is expressed in millimeters (mm). It represents the height value of each point on the corneal surface along the optical axis in a coordinate system with the corneal vertex as the origin. The basic shape of the cornea is expressed in millimeters (mm), representing the ideal geometric shape of the corneal surface without tear film coverage. The tear film thickness distribution is expressed in micrometers (μm), representing the variation in the thickness of the tear film layer covering the corneal surface. This parameter is a dynamic variable to be solved, which varies with time and spatial location, and the normal physiological range is 3-40 μm. The residual term, measured in millimeters (mm), represents tiny deviations that the model cannot explain. These mainly include measurement noise, minor irregularities on the corneal surface, and model approximation errors. In actual optimization, it is usually assumed that the residual term follows a Gaussian distribution with a mean of zero.
[0133] The coupled surface model provided in this application decomposes the actual corneal surface into three parts: the basic shape of the cornea, the tear film thickness distribution, and residual terms, providing a precise mathematical framework for surface reconstruction and tear film compensation in eye tracking. This coupled surface model preserves the individualized geometric features of the cornea while fully considering the dynamic changes of the tear film, making it one of the keys to achieving high-precision gaze estimation.
[0134] In some embodiments, based on a pre-constructed coupled surface model containing the basic shape of the cornea and the tear film thickness distribution, the tear film thickness distribution at the current moment is iteratively determined according to the instantaneous normal vector and the initial value of the tear film thickness distribution of the target eyeball, including the following steps:
[0135] S1031. Based on the coupled surface model, determine the model normal vector corresponding to the current estimated value of tear film thickness distribution.
[0136] For example, based on the current estimated tear film thickness distribution and combined with a pre-measured basic corneal shape, the theoretical surface normal vector corresponding to the current estimated tear film thickness distribution is calculated using an ocular surface geometry model (i.e., a coupled surface model). Specifically, firstly, the complete ocular surface morphology is reconstructed based on the basic corneal shape and the current estimated tear film thickness distribution; then, differential geometric analysis is performed on this ocular surface morphology to calculate the unit normal direction at each point, obtaining the model normal vector distribution. This model normal vector reflects the geometric orientation that the ocular surface should have under the current estimated tear film thickness.
[0137] S1032. Construct an objective function, which includes a difference term between the instantaneous normal vector and the model-based normal vector, as well as a physical rationality constraint term for the tear film thickness distribution. Adjust the relative importance of the difference term and the physical rationality constraint term by weighting coefficients.
[0138] For example, an objective function is constructed to evaluate the consistency between the measured values (the instantaneous normal vector obtained in S102) and the model values (the model normal vector obtained in S1031). The objective function is expressed as:
[0139]
[0140] in, It is the instantaneous normal vector; This is the model normal vector corresponding to the current estimated tear film thickness distribution. This is a physical constraint term for the tear film thickness distribution, used to ensure that the solved tear film thickness distribution conforms to the physical characteristics of the tear film as a continuous liquid film. For example, it is at least one of the following: smoothness constraint, curvature constraint, total variation constraint, or physiological range constraint; These are weighting coefficients used to adjust the data fitting term. The relative importance of the physical rationality constraint.
[0141] S1033. With minimizing the objective function as the optimization objective, iteratively update the tear film thickness distribution until the convergence condition is met, so as to obtain the tear film thickness distribution at the current moment.
[0142] For example, with minimizing the above objective function as the optimization objective, a numerical optimization algorithm is used to iteratively update the estimated tear film thickness distribution. Each iteration includes: calculating the objective function value corresponding to the current estimated tear film thickness distribution; determining the update direction and step size based on the gradient information of the objective function with respect to the tear film thickness distribution; updating the tear film thickness distribution according to certain rules; and determining whether the preset convergence conditions are met. The optimization algorithm includes gradient descent, conjugate gradient, Gauss-Newton, or Levenberg-Marquardt algorithms, etc., and an appropriate algorithm can be selected based on the problem size and computational resources. Convergence conditions may include: the change in tear film thickness distribution between two consecutive iterations being less than the corresponding preset threshold (e.g., 0.1 μm); the decrease in the objective function value being less than the corresponding preset threshold; the number of iterations reaching a preset maximum value; or the root mean square error between the model normal vector and the measurement normal vector being less than the corresponding preset threshold. When the convergence conditions are met, the iteration terminates, and the current estimated tear film thickness distribution is output as the final result; otherwise, the iteration returns to step S1031 to continue.
[0143] In this embodiment, by constructing an objective function that includes data fitting terms and physical rationality constraints, and dynamically adjusting the balance between the two using weighting coefficients, it can effectively suppress noise interference and ensure the physical continuity of tear film thickness distribution while ensuring a high degree of consistency between the reconstructed surface and the measured optical data. This mechanism, through independent frame-by-frame iterative solving, accurately separates the coupling influence of fixed corneal geometry and dynamic tear film changes on the normal vector, achieving high-precision and robust real-time tracking of tear film thickness. This provides accurate and physically consistent dynamic tear film compensation data for subsequent gaze direction estimation.
[0144] The eyeball anatomically comprises the cornea and sclera, which have different radii of curvature but together form an approximate bispherical geometry. The corneal region has a steeper curvature, reflecting the optical surface of the anterior part of the eyeball; the scleral region has a gentler curvature, forming the main outer shell of the eyeball. By segmenting the three-dimensional point cloud data into regions and fitting the spherical parameters of the corneal and scleral regions respectively, two key centers representing the geometric features of the eyeball can be obtained. The spatial relationship between these two centers directly reflects the pose and orientation of the eyeball. Therefore, in some embodiments, the three-dimensional point cloud data includes corneal and scleral point clouds. Based on the three-dimensional point cloud data, the gaze direction of the target eyeball is determined, including the following steps:
[0145] S1051. Perform spherical fitting on the point cloud of the corneal region to obtain the corneal center of the target eyeball.
[0146] For example, the RANSAC algorithm is used to perform spherical fitting on the corneal region point cloud. The optimal spherical parameters are found through multiple random samplings and interior point counting. Specifically, a minimum subset is randomly selected from the corneal region point cloud to determine candidate spherical parameters; spherical fitting requires at least four non-coplanar points. Based on these four randomly selected points, the spherical equation is solved to obtain the candidate sphere centers. and candidate radius The solution method can be the least squares method or geometric analysis method, which obtains the spherical parameters by solving a system of linear equations; for each point in the corneal region point cloud. Calculate its distance to the candidate sphere. ,like If the value is less than a preset threshold (e.g., 0.05-0.1 mm), the point is determined to be an interior point; otherwise, it is an exterior point. The number of interior points reflects the degree of agreement between the candidate spherical parameter and the point cloud data. The above random sampling, spherical parameter calculation, and interior point determination process is repeated for a preset number of times M (e.g., 500-1000 times). In each iteration, the candidate spherical parameter with the most interior points is recorded as the optimal candidate.
[0147] After iterating through the RANSAC algorithm, the spherical parameters with the largest number of interior points are obtained. ) and its corresponding interior point set This inner point set is a high-quality point cloud that is consistent with the geometry of the cornea, eliminating outliers caused by tear film abnormalities, local reflections, or measurement noise.
[0148] To further improve the fitting accuracy, based on the obtained optimal interior point set... A weighted optimization method is used for refined spherical fitting. For example, the objective function is:
[0149]
[0150] In the formula, interior point set The first in Corneal point cloud coordinates; For point The surface normal vector at the point ensures that the point cloud not only lies on the sphere, but also that the normal direction points to the center of the sphere, increasing the geometric consistency of the fit. Let the center of the cornea be the target. Let be the radius of curvature of the cornea to be determined; This is a robust loss function used to further suppress the impact of residual outliers.
[0151] This optimization problem can be solved using nonlinear least squares algorithms, such as the Levenberg-Marquardt algorithm or the Gauss-Newton method. Algorithm iterative updates. and This continues until the objective function converges or the maximum number of iterations is reached.
[0152] S1052. Perform spherical fitting on the point cloud of the scleral region to obtain the scleral center of the target eyeball.
[0153] The fitting method is similar to that for the corneal region, and will not be described in detail here.
[0154] S1053. Determine the direction of vision of the target eye based on the corneal center and scleral center.
[0155] For example, the direction of the line connecting the centers of the cornea and sclera is calculated based on the corneal center and the scleral center. This direction reflects the geometrical symmetry axis of the eye and is the basis for determining the direction of vision. Specifically, the direction is calculated from the scleral center... Pointing towards the center of the cornea The vector, Then, the vector is normalized to obtain the unit direction vector. This unit direction vector represents the geometric optical axis of the eye, characterizing the eye as the axis of symmetry of the optical system. After obtaining the geometric optical axis, it can be used as a preliminary estimate of the line of sight or as input for subsequent correction calculations. This optical axis fully utilizes the geometric constraints provided by the fitting of the cornea and sclera, exhibiting higher accuracy and robustness compared to methods based solely on single-region fitting.
[0156] In this embodiment, by fitting the spherical parameters of the cornea and sclera in separate regions, the anatomical structure of the eyeball is fully considered. While obtaining the corneal and scleral centers with clear physiological significance, high-precision localization of the dual centers is achieved using the rich geometric constraints provided by high-density point clouds. Real-time fitting based on individual point cloud data can adapt to the geometric differences in the eyeballs of different individuals, eliminating systematic errors caused by fixed geometric models. Simultaneously, by fully utilizing the point cloud data from both the cornea and sclera, the obtained geometric optical axis direction has higher accuracy and robustness compared to single-region fitting.
[0157] As a complex optical organ, the eyeball has a definite correspondence between its geometric structure and physiological function: the cornea and sclera constitute an approximately double-sphere structure of the eyeball, and the direction of the line connecting their centers corresponds to the geometric optical axis of the eyeball; however, the actual gaze direction of the human eye (i.e., the visual axis) does not coincide with the geometric optical axis, and there is a physiological deviation angle between them, namely the Kappa angle. Directly using the geometric optical axis as the gaze direction will introduce systematic physiological errors. Therefore, after obtaining the geometric optical axis, it is necessary to combine it with individualized Kappa angle parameters for correction to obtain the true gaze direction. Therefore, in some embodiments, the gaze direction of the target eyeball is determined based on the corneal and scleral centers, including:
[0158] S1053-1. Determine the optical axis direction of the target eyeball based on the corneal center and scleral center.
[0159] For example, , Characterizes the direction of the optical axis of the target eyeball.
[0160] S1053-2. Based on the obtained Kappa angle parameters, the optical axis direction is corrected to obtain the line of sight direction of the target eyeball.
[0161] The Kappa angle parameter can be determined in the following ways: using the average value of the population as a preset value, such as 4° in the horizontal direction and 0.5° in the vertical direction; or, by having the subject fixate on a target at a known location (such as multiple target points at known locations on a monitor screen), measuring the deviation of the actual eye orientation from the optical axis, and calculating the individualized Kappa angle parameter; or, during continuous measurement, dynamically updating the Kappa angle parameter based on the deviation between the gaze estimation result and the known fixation point.
[0162] For example, a rotation matrix R_kappa is constructed based on the obtained Kappa angle parameters. The optical axis direction is then multiplied by the rotation matrix to obtain the line of sight direction of the target eye. The direction of the line of sight. This is the actual gaze direction of the target eye at the current moment, which is a three-dimensional unit vector pointing from the center of eye rotation to the gaze point.
[0163] This application's embodiments, by distinguishing between the geometric optical axis and the physiological visual axis, introduce an individualized Kappa angle correction mechanism, fundamentally eliminating the systematic physiological errors introduced by neglecting the Kappa angle. Based on the optical axis direction obtained through high-precision dual-sphere fitting, combined with accurate Kappa angle correction, the gaze direction estimation fully conforms to the actual anatomical structure and physiological function of the human eye, significantly improving the physiological accuracy and individual adaptability of eye tracking, and meeting the stringent requirements of sub-degree accuracy for clinical diagnosis and highly immersive applications.
[0164] During eye tracking, the tear film, as a thin liquid film covering the corneal surface, is not constant in thickness but dynamically changes over time. Factors such as the gradual rebuilding of the tear film after blinking, local thinning caused by tear evaporation, and tear film rupture in patients with dry eye syndrome can all cause instantaneous fluctuations in tear film thickness. If this dynamic change is not compensated for in time, it will be directly reflected in the 3D point cloud data, leading to a deviation between the reconstructed corneal surface and the actual physical shape, thereby affecting the accuracy and stability of the gaze direction estimation. Therefore, in some embodiments, based on the current tear film thickness distribution and the basic shape of the cornea, the reconstruction of 3D point cloud data representing the actual physical shape of the target eyeball surface includes the following steps:
[0165] S1041. Based on the change in the reflected light intensity of the target eyeball at the current moment relative to the reference reflected light intensity, and the tear film thickness distribution at the previous moment, dynamically compensate and update the tear film thickness distribution at the current moment.
[0166] It should be understood that changes in tear film thickness will cause corresponding changes in reflected light intensity. By continuously monitoring the difference between the current reflected light intensity and the baseline reflected light intensity, and integrating this difference over time, the change in tear film thickness can be accumulated. At the same time, combined with the tear film thickness distribution at the previous moment, dynamic compensation and updates of the tear film thickness at the current moment can be achieved through weighted fusion.
[0167] For example, obtain the reflected light intensity measured at the current moment. The light intensity value can be extracted from the structured light image data acquired in step S101; the difference between the reflected light intensity at the current moment and the reference reflected light intensity is calculated. This difference reflects the direction and magnitude of the tear film thickness shift relative to the baseline state at the current moment; integrating this difference over time yields the cumulative change in light intensity. Based on the dynamic tear film compensation model, the tear film thickness distribution at the current moment is updated: In the formula, The updated tear film thickness distribution at the current moment is the compensation. This represents the tear film thickness distribution at the previous moment; This is the integral gain coefficient, which controls the contribution weight of the current measurement value to the update, and its value range is, for example, 0.1-0.3; This is the historical attenuation coefficient, which controls the degree to which the thickness distribution at the previous moment is preserved. Its value range is, for example, 0.7-0.9. and satisfy ≈1, ensuring energy conservation. The larger the value, the faster the system response, but the greater the noise. The larger the value, the smoother the result, but the more pronounced the lag.
[0168] S1042. Based on the compensated and updated tear film thickness distribution and corneal basic shape, reconstruct the current three-dimensional point cloud data.
[0169] Basic corneal shape The ideal geometry of the corneal surface without tear film coverage was described, and the updated tear film thickness distribution was compensated for. This describes the instantaneous thickness change of the tear film on the corneal surface at the current moment. By superimposing the two data, the complete corneal surface morphology including the tear film can be obtained. Discrete sampling of this surface generates three-dimensional point cloud data.
[0170] For example, each sampling point contains the complete corneal surface height value of the tear film. Discrete sampling is performed on the complete corneal surface obtained after superposition to generate three-dimensional point cloud data, with each sampling point corresponding to a three-dimensional coordinate: ( ).
[0171] In this embodiment, by real-time monitoring of reflected light intensity differences and time integration, combined with the tear film thickness distribution of the previous moment for weighted fusion, dynamic compensation and updating of the tear film thickness distribution is achieved, effectively eliminating the interference of rapid tear film changes on three-dimensional reconstruction. The high-density three-dimensional point cloud data generated by superimposing the compensated and updated tear film thickness distribution with the basic shape of the cornea accurately reflects the true physical shape of the eyeball at the current moment, ensuring smooth transition of data between frames in continuous measurement, and significantly improving the accuracy and dynamic stability of three-dimensional reconstruction.
[0172] Based on the above embodiments, the eye-tracking method will be explained in detail with reference to a specific embodiment. Figure 2 Flowchart of the eye-tracking method provided in the embodiments of this application Figure 2 .like Figure 2 As shown, the eye-tracking method includes:
[0173] S201, System Calibration and Initialization.
[0174] This step involves display-camera system calibration, individual eye parameter initialization, and system geometry calibration. The aim of this step is to establish a unified coordinate reference and individualized geometric model for subsequent high-precision eye tracking through precise calibration of the measurement system and initial measurements of individual eye parameters.
[0175] The display-camera system calibration includes: using Zhang Zhengyou's calibration method, taking multiple images of the checkerboard calibration board at different positions and orientations, detecting the corner coordinates in the images, establishing the projection relationship between the world coordinate system and the image coordinate system, and solving the camera intrinsic parameters. Camera internal parameters This reflects the camera's own imaging geometry; it displays a checkerboard pattern with known pixel coordinates on the monitor, establishes a mapping between the display coordinates and camera coordinates through camera observations, and obtains the monitor's equivalent intrinsic parameters. This equivalent internal reference This describes the transformation relationship between the display pixel coordinates and the direction of its emitted light rays in space; based on camera intrinsic parameters. Equivalent internal parameters of the display Calculate the system's intrinsic parameters . System internal parameters It is stored in system memory for coordinate transformation during subsequent eye tracking.
[0176] Individual ocular parameter initialization includes: measuring the corneal radius of curvature, scleral radius of curvature, and corneal aspheric coefficient of the subject (target eye) using equipment such as corneal topography or anterior segment optical coherence tomography (OCT); establishing a basic corneal shape model reflecting individual characteristics based on the measured individual parameters to eliminate systematic errors caused by fixed geometric models in traditional methods; and calibrating the initial values of tear film thickness distribution of the target eye. All individual ocular parameters are stored in system memory for subsequent corneal and tear film compensation during eye tracking.
[0177] The system geometric calibration includes: placing a standard calibration object with known geometric dimensions in the measurement space, controlling the display to show the marked point pattern, and simultaneously acquiring multiple sets of images by the camera. By changing the position and orientation of the calibration object, multiple sets of corresponding point data covering the entire measurement field of view are obtained. Adjustment processing is performed on each set of acquired data points to establish observation equations. The initial estimates of the camera pose parameters are solved using the least squares principle, and the corresponding least squares error is calculated. With the goal of minimizing the reprojection error of all observation points, a nonlinear optimization algorithm is used to iteratively solve for the minimum point of the least squares error to obtain the optimal system geometric calibration result. In this embodiment, the geometric calibration result obtained includes intrinsic parameter calibration coefficients: first-order radial distortion coefficient k1, second-order radial distortion coefficient k2, and third-order radial distortion coefficient k3. These coefficients are used for distortion correction during subsequent image acquisition.
[0178] S202, Single phase deflection measurement.
[0179] This step aims to obtain the encapsulated phase distribution for subsequent 3D reconstruction by projecting a composite sinusoidal fringe pattern (composite structured light pattern) and acquiring reflected images, followed by phase demodulation.
[0180] Correspondingly, this step involves cross-sine pattern generation, stereo image acquisition, and phase calculation. Cross-sine pattern generation includes controlling the projection unit (display) to generate a composite structured light pattern carrying sinusoidal fringe components with at least two different spatial frequencies. For example, this composite structured light pattern superimposes sinusoidal fringes in the horizontal and vertical directions, enabling simultaneous encoding of phase information in both directions with a single projection.
[0181] Stereoscopic image acquisition includes: displaying the generated composite sinusoidal fringe pattern on a high-resolution display and projecting it onto the target eyeball surface; employing a dual-lens camera system (e.g., 6400 dpi). The system simultaneously acquires deformed images reflected from the surface of the target eyeball (480, frame rate 30Hz), preprocesses these deformed images to obtain structured light image data, including distortion correction, grayscale normalization, and region of interest extraction.
[0182] Phase resolution includes: using a phase deflection algorithm to demodulate the phase of the structured light image data and extracting the wrap-around phase distribution corresponding to at least two spatial directions. This wrap-around phase distribution is the reflected light phase information modulated by the corneal basic shape and tear film thickness distribution, and will be used as input data for subsequent corresponding point establishment and normal vector calculation.
[0183] S203, Target eyeball three-dimensional surface reconstruction and tear film compensation.
[0184] This step aims to reconstruct three-dimensional point cloud data containing tear film dynamic information by establishing a precise correspondence between camera pixels and display pixels, calculating surface normal vectors, and solving the tear film thickness distribution using an iterative optimization method based on the wrapping phase distribution obtained in step S202.
[0185] Correspondingly, this step involves surface normal vector calculation, depth map reconstruction, and tear film interference compensation. Among them, surface normal vector calculation includes: determining the instantaneous normal vector of each point on the surface of the target eyeball based on the encapsulated phase distribution, as detailed in step S102.
[0186] Depth map reconstruction includes: based on a pre-constructed coupled surface model containing the basic shape of the cornea and the tear film thickness distribution, the tear film thickness distribution at the current moment is iteratively determined according to the instantaneous normal vector and the initial value of the tear film thickness distribution of the target eyeball, as detailed in steps S1031~S1033; based on the tear film thickness distribution and the basic shape of the cornea at the current moment, three-dimensional point cloud data representing the actual physical shape of the target eyeball surface is reconstructed.
[0187] Tear film interference compensation includes: dynamically compensating and updating the tear film thickness distribution at the current moment based on the change in reflected light intensity of the target eyeball relative to the reference reflected light intensity and the tear film thickness distribution at the previous moment; and reconstructing the three-dimensional point cloud data at the current moment based on the compensated and updated tear film thickness distribution and the basic shape of the cornea. See steps S1041-S1042 for details.
[0188] S204, Eye movement parameter estimation and gaze direction output.
[0189] This step aims to calculate the precise gaze direction of the target eye at the current moment based on the 3D point cloud data reconstructed in step S203, through regional spherical fitting and Kappa angle correction.
[0190] This step involves corneal spherical fitting, scleral spherical fitting, optical axis calculation, Kappa angle correction, and gaze direction determination. Corneal spherical fitting involves using a robust estimation algorithm based on RANSAC to spherically fit the corneal region point cloud to obtain the corneal center of the target eye, as detailed in step S1051. Scleral spherical fitting is similar to corneal spherical fitting and will not be described further here. Optical axis calculation involves calculating the direction of the line connecting the corneal and scleral centers to determine the geometric optical axis direction of the eye, as detailed in step S1053. Kappa angle correction and gaze direction determination involve constructing a rotation matrix R_kappa based on the obtained Kappa angle parameters, and multiplying the optical axis direction by the rotation matrix to obtain the gaze direction of the target eye.
[0191] In summary, this application has at least the following advantages:
[0192] I. By acquiring structured light image data of the target eyeball surface reflection and extracting the encapsulated phase distribution modulated by the corneal basic shape and tear film thickness distribution, the instantaneous normal vector of each point on the eyeball surface can be determined, accurately capturing the influence of tear film dynamic changes on optical measurements. Based on this, using a pre-constructed coupled surface model incorporating the corneal basic shape and tear film thickness distribution, the tear film thickness distribution at the current moment is iteratively solved by combining the instantaneous normal vector and the initial value of the tear film thickness distribution. This effectively decouples the geometric contribution of the cornea from the optical contribution of the tear film, eliminating the systematic errors caused by ignoring the tear film or using a fixed geometric model in traditional methods. Simultaneously, by iteratively updating the tear film thickness distribution at the current moment, real-time optical measurement of tear film dynamic changes is achieved. The system employs a method that compensates for tear film thickness fluctuations, effectively eliminating the interference of tear film interference and thickness variations on gaze direction estimation. Based on the decoupled tear film thickness distribution and the basic corneal shape, it reconstructs the actual physical shape of the target eye's surface, achieving precise modeling and compensation for individual tear film differences. By reconstructing three-dimensional point cloud data of the entire ocular surface, it fully utilizes the rich geometric information of the eye's surface, significantly improving information density and surface reconstruction accuracy, providing more complete geometric constraints for gaze direction estimation. Finally, it determines the gaze direction based on three-dimensional point cloud data including tear film dynamic compensation. By fully considering individual corneal differences and dynamic changes in the tear film, it significantly improves the accuracy and robustness of eye tracking, meeting the stringent requirements of sub-degree gaze estimation in clinical diagnosis and highly immersive virtual reality applications.
[0193] Second, by fitting the spherical parameters of the cornea and sclera in different regions, the anatomical structure of the eyeball is fully considered. While obtaining the corneal and scleral centers with clear physiological significance, high-precision localization of the dual centers is achieved using the rich geometric constraints provided by high-density point clouds. Real-time fitting based on individual point cloud data can adapt to the geometric differences in the eyeballs of different individuals, eliminating systematic errors caused by fixed geometric models. At the same time, by fully utilizing the point cloud data from both the cornea and sclera, the obtained geometric optical axis direction has higher accuracy and robustness compared to single-region fitting.
[0194] Third, by distinguishing between the geometric optical axis and the physiological visual axis, an individualized Kappa angle correction mechanism is introduced, fundamentally eliminating the systematic physiological errors introduced by ignoring the Kappa angle. Based on the optical axis direction obtained by high-precision dual-sphere fitting, combined with accurate Kappa angle correction, the gaze direction estimation fully conforms to the real anatomical structure and physiological function of the human eye, significantly improving the physiological accuracy and individual adaptability of eye tracking, and meeting the stringent requirements of sub-degree accuracy for clinical diagnosis and highly immersive applications.
[0195] Fourth, by monitoring the difference in reflected light intensity in real time and performing time integration, and combining it with the tear film thickness distribution of the previous moment for weighted fusion, dynamic compensation and updating of the tear film thickness distribution is achieved, effectively eliminating the interference of rapid changes in the tear film on three-dimensional reconstruction; the high-density three-dimensional point cloud data generated by superimposing the compensated and updated tear film thickness distribution with the basic shape of the cornea accurately reflects the real physical shape of the eyeball at the current moment, ensuring smooth transition of data between frames in continuous measurement, and significantly improving the accuracy and dynamic stability of three-dimensional reconstruction.
[0196] Fifth, by employing the phase deflection algorithm, only a single frame of structured light image is needed to simultaneously extract the package phase distribution corresponding to at least two spatial directions. Compared with the traditional multi-step phase-shifting PMD, which requires the acquisition of 8-12 frames of images and introduces a system delay of more than 100 milliseconds, this application compresses the amount of data acquisition to a single frame, significantly reducing the time overhead of image acquisition and transmission, and compressing the system delay to below the millisecond level, meeting the real-time eye tracking requirements of 120Hz or even higher frame rates.
[0197] Next, this application provides an eye-tracking device, such as... Figure 3 As shown in the schematic diagram of the eye-tracking device provided in this embodiment, the eye-tracking device 30 includes: an acquisition module 31, a processing module 32, a determination module 33, and a reconstruction module 34. Wherein:
[0198] The acquisition module 31 is used to acquire structured light image data reflected from the surface of the target eyeball;
[0199] Processing module 32 is used to extract the wrapping phase distribution from the structured light image data, and based on the wrapping phase distribution, determine the instantaneous normal vector of each point on the surface of the target eyeball. The wrapping phase distribution represents the reflected light phase distribution after being modulated by the basic shape of the cornea and the tear film thickness distribution.
[0200] The determination module 33 is used to iteratively determine the tear film thickness distribution at the current moment based on a pre-constructed coupled surface model containing the basic shape of the cornea and the tear film thickness distribution, according to the instantaneous normal vector and the initial value of the tear film thickness distribution of the target eyeball.
[0201] Reconstruction module 34 is used to reconstruct three-dimensional point cloud data representing the actual physical shape of the target eyeball surface based on the tear film thickness distribution and corneal basic shape at the current moment.
[0202] The determination module 33 is also used to determine the direction of the target eye's gaze based on 3D point cloud data.
[0203] In one possible implementation, the coupled surface model satisfies the following form:
[0204]
[0205] in, This refers to the actual corneal surface height, including the tear film. For the basic shape of the cornea, Tear film thickness distribution; This is the residual term.
[0206] In one possible implementation, the determining module 33 is specifically used to: determine the model normal vector corresponding to the current estimated value of the tear film thickness distribution based on the coupled surface model; construct an objective function, which includes a difference term between the instantaneous normal vector and the model normal vector, as well as a physical rationality constraint term for the tear film thickness distribution, and adjust the relative importance of the difference term and the physical rationality constraint term through weight coefficients; and iteratively update the tear film thickness distribution with the goal of minimizing the objective function until the convergence condition is met, so as to obtain the tear film thickness distribution at the current moment.
[0207] In one possible implementation, the three-dimensional point cloud data includes corneal region point clouds and scleral region point clouds. The determination module 33 is further configured to: perform spherical fitting on the corneal region point cloud to obtain the corneal center of the target eye; perform spherical fitting on the scleral region point cloud to obtain the scleral center of the target eye; and determine the gaze direction of the target eye based on the corneal center and the scleral center.
[0208] In one possible implementation, the determining module 33 is further configured to: determine the optical axis direction of the target eyeball based on the corneal center and the scleral center; and correct the optical axis direction based on the obtained Kappa angle parameters to obtain the line of sight direction of the target eyeball.
[0209] In one possible implementation, the reconstruction module 34 is specifically used to: dynamically compensate and update the tear film thickness distribution at the current moment based on the change in the reflected light intensity of the target eyeball at the current moment relative to the reference reflected light intensity and the tear film thickness distribution at the previous moment; and reconstruct the three-dimensional point cloud data at the current moment based on the compensated and updated tear film thickness distribution and the basic shape of the cornea.
[0210] In one possible implementation, the processing module 32 is specifically used to: employ a phase deflection algorithm to perform phase demodulation on the structured light image data and extract the wrapping phase distribution corresponding to at least two spatial directions.
[0211] In one possible implementation, the structured light image data is obtained by projecting a composite structured light pattern onto the surface of the target eyeball, the composite structured light pattern containing at least two sinusoidal fringe components with different spatial frequencies.
[0212] Figure 4 This is a schematic diagram of the structure of the eye-tracking system provided in the embodiments of this application, as shown below. Figure 4As shown, the eye-tracking system includes: a projection unit 41, an image acquisition unit 42, and a control unit 43. Wherein:
[0213] Projection unit 41 is used to project incident light carrying structured light pattern 44 onto the surface of the target eyeball;
[0214] Image acquisition unit 42 is used to acquire structured light image data formed after reflection from the surface of the target eyeball;
[0215] The control unit 43 is electrically connected to the projection unit 41 and the image acquisition unit 42 respectively, and is used to execute various possible implementations in the above-described eye-tracking method embodiments.
[0216] For example, the structured light pattern 44 is, for instance, a composite sinusoidal stripe pattern, composed of superimposed sinusoidal stripes in the horizontal and vertical directions, used to simultaneously encode phase information in both directions. The projection unit 41 can be implemented using a high-resolution display or a digital light projector, projecting structured light onto the target eye surface according to the pattern signal generated by the control unit 43.
[0217] The target eyeball surface is covered by a tear film 45, beneath which lies the cornea 46. Incident light is reflected under the combined modulation of the tear film 45 and the cornea 46, forming a deformed stripe image that carries coupled information about the basic shape of the cornea and the thickness distribution of the tear film. The image acquisition unit 42 can be configured with a binocular camera, simultaneously triggering the left and right cameras to acquire eye reflection images, ensuring that the left and right views strictly correspond to the eyeball state at the same moment.
[0218] For example, the control unit 43 is used to: control the projection unit 41 to generate and project the structured light pattern 44; control the image acquisition unit 42 to synchronously acquire reflected images; preprocess the acquired image data, perform phase calculation, 3D reconstruction, and gaze direction estimation; and output the final eye-tracking result. The control unit 43 can be implemented using an embedded processor, FPGA, or industrial computer, and has real-time processing capabilities and a communication interface with external systems.
[0219] The aforementioned eye-tracking system enables fully automated eye-tracking from structured light projection and image acquisition to data processing, providing high-precision hardware support for clinical diagnosis, human-computer interaction, VR, and AR applications.
[0220] Figure 5 This is a schematic diagram of the structure of the control unit provided in the embodiments of this application, as shown below. Figure 5 As shown, the control unit 43 provided in this embodiment includes at least one processor 431 and a memory 432. Optionally, the control unit 43 further includes a communication component 433. The processor 431, memory 432, and communication component 433 are connected via a bus 434.
[0221] In a specific implementation, at least one processor 431 executes computer execution instructions stored in memory 432, causing at least one processor 431 to perform the above-described method.
[0222] The specific implementation process of processor 431 can be found in the above method embodiment, and its implementation principle and technical effect are similar. It will not be repeated here.
[0223] In the above embodiments, it should be understood that the processor can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), etc. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the method disclosed in this invention can be directly implemented by a hardware processor, or implemented by a combination of hardware and software modules within the processor.
[0224] The memory may include random access memory (RAM) and may also include non-volatile memory (NVM), such as at least one disk storage device.
[0225] The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. Buses can be categorized as address buses, data buses, control buses, etc. For ease of illustration, the buses shown in the accompanying drawings are not limited to a single bus or a single type of bus.
[0226] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the above-described method.
[0227] This application also provides a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, implement the above-described method.
[0228] The aforementioned readable storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk. The readable storage medium can be any available medium accessible to a general-purpose or special-purpose computer.
[0229] An exemplary readable storage medium is coupled to a processor, enabling the processor to read information from and write information to the readable storage medium. Of course, the readable storage medium can also be a component of the processor. The processor and the readable storage medium can reside in an Application Specific Integrated Circuit (ASIC). Alternatively, the processor and the readable storage medium can exist as discrete components in the device.
[0230] The division of units is merely a logical functional division; in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be indirect coupling or communication connection through some interfaces, devices, or units, and may be electrical, mechanical, or other forms.
[0231] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment, depending on actual needs.
[0232] In addition, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.
[0233] If a function is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0234] Next, this application embodiment also provides an eye-tracking device for integrating the aforementioned eye-tracking method and system into specific application products, thereby realizing the portable and scenario-based application of high-precision eye-tracking function.
[0235] One possible implementation is, such as Figure 6 As shown in the structural diagram of the eye-tracking device provided in this embodiment, the eye-tracking device 60 integrates the eye-tracking system described in the above embodiments. That is, in addition to the device body, the eye-tracking device also includes a projection unit 41, an image acquisition unit 42, and a control unit 43, forming a complete eye-tracking hardware platform. This integration method enables the eye-tracking device to independently complete the entire process of structured light projection, image acquisition, and data processing, making it suitable for application scenarios with high requirements for real-time performance and integration.
[0236] In another possible implementation, the eye-tracking device 60 integrates the control unit 43 described in the above embodiments, while the projection unit 41 and image acquisition unit 42 can be used as external modules or separate components. This approach is suitable for upgrading devices with existing imaging and projection hardware, enabling rapid deployment of eye-tracking functionality through the embedded control unit.
[0237] For example, the application forms of the eye-tracking device 60 are not limited to: VR or AR head-mounted display integrated devices for immersive interaction, including gaze point rendering and eye-tracking interaction; smart glasses terminals for eye-tracking interaction and attention analysis in everyday scenarios; in-vehicle human-machine interaction devices for safety warnings and intelligent control by detecting the driver's gaze status; and medical diagnostic auxiliary devices for eye movement monitoring in ophthalmic clinical examinations, neurological disease diagnosis, and rehabilitation training. These device forms all use the high-precision eye-tracking technology provided in this application as their core, and meet the needs of diverse application scenarios through different integration methods.
[0238] In the above embodiments, the descriptions of each embodiment have their own emphasis. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments. The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as the combination of these technical features does not contradict each other, it should be considered within the scope of this specification.
[0239] Other embodiments of this application will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of this application that follow the general principles of this application and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of this application are indicated by the following claims.
[0240] It should be understood that this application is not limited to the precise structure described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this application is limited only by the appended claims.
Claims
1. An eye-tracking method, characterized in that, include: Acquire structured light image data of the reflection from the surface of the target eyeball; The encapsulation phase distribution is extracted from the structured light image data, and the instantaneous normal vector of each point on the surface of the target eyeball is determined based on the encapsulation phase distribution. The encapsulation phase distribution represents the reflected light phase distribution after being modulated by the corneal basic shape and the tear film thickness distribution. The instantaneous normal vector carries information about the corneal basic shape and the tear film thickness distribution. The corneal basic shape is an ideal corneal geometry model without tear film coverage. Based on a pre-constructed coupled surface model containing the basic shape of the cornea and the tear film thickness distribution, the instantaneous normal vector is used as the constraint target, and the tear film thickness distribution at the current moment is used as the variable to be solved. The tear film thickness distribution is updated iteratively so that the model normal vector determined based on the coupled surface model approximates the instantaneous normal vector, thus obtaining the tear film thickness distribution at the current moment. Based on the tear film thickness distribution and the basic shape of the cornea at the current moment, reconstruct three-dimensional point cloud data representing the actual physical shape of the target eyeball surface; Based on the three-dimensional point cloud data, the gaze direction of the target eyeball is determined; The step of iteratively updating the tear film thickness distribution to make the model normal vector determined based on the coupled surface model approximate the instantaneous normal vector includes: Based on the coupled surface model, determine the model normal vector corresponding to the current estimated tear film thickness distribution; Construct an objective function, which includes a difference term between the instantaneous normal vector and the model normal vector, as well as a physical rationality constraint term for the tear film thickness distribution, and adjust the relative importance of the difference term and the physical rationality constraint term by weighting coefficients; The tear film thickness distribution is iteratively updated with the objective function minimized.
2. The eye-tracking method according to claim 1, characterized in that, The coupled surface model satisfies the following form: in, This refers to the actual corneal surface height, including the tear film. For the basic shape of the cornea, Tear film thickness distribution; This is the residual term.
3. The eye-tracking method according to claim 1 or 2, characterized in that, The three-dimensional point cloud data includes corneal region point clouds and scleral region point clouds. Determining the gaze direction of the target eyeball based on the three-dimensional point cloud data includes: Spherical fitting is performed on the point cloud of the corneal region to obtain the corneal center of the target eyeball; Spherical fitting is performed on the point cloud of the scleral region to obtain the scleral center of the target eyeball; The direction of vision of the target eyeball is determined based on the corneal center and the scleral center.
4. The eye-tracking method according to claim 3, characterized in that, Determining the line of sight direction of the target eyeball based on the corneal center and the scleral center includes: The optical axis direction of the target eyeball is determined based on the corneal center and the scleral center. Based on the obtained Kappa angle parameters, the optical axis direction is corrected to obtain the line of sight direction of the target eyeball.
5. The eye-tracking method according to claim 1 or 2, characterized in that, The three-dimensional point cloud data reconstructed based on the tear film thickness distribution and the basic corneal shape at the current moment, representing the actual physical shape of the target eyeball surface, includes: Based on the change in the reflected light intensity of the target eyeball at the current moment relative to the reference reflected light intensity, and the tear film thickness distribution at the previous moment, the tear film thickness distribution at the current moment is dynamically compensated and updated. Based on the compensated and updated tear film thickness distribution and the basic shape of the cornea, the three-dimensional point cloud data at the current moment is reconstructed.
6. The eye-tracking method according to claim 1 or 2, characterized in that, Extracting the encapsulated phase distribution from the structured light image data includes: The phase deflection algorithm is used to demodulate the structured light image data and extract the wrapping phase distribution corresponding to at least two spatial directions.
7. The eye-tracking method according to claim 1 or 2, characterized in that, The structured light image data is obtained by projecting a composite structured light pattern onto the surface of the target eyeball, the composite structured light pattern containing at least two sinusoidal fringe components with different spatial frequencies.
8. A control unit, characterized in that, include: Memory, processor; The memory stores computer-executed instructions; The processor executes computer execution instructions stored in the memory, causing the processor to perform the eye-tracking method as described in any one of claims 1 to 7.
9. An eye-tracking device, characterized in that, It integrates the control unit as described in claim 8.