Multifunctional visual function detection device and method based on VR and eye movement tracking
By combining high-precision eye tracking, multimodal sensing and adaptive VR rendering technology, the limitations of existing VR and eye tracking technologies in multifunctional integration, dynamic interaction accuracy and personalized adaptation are solved, and high-precision and personalized visual function detection is achieved, improving detection accuracy and application range.
Patent Information
- Application Number
- CN202510559155.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-30
- Publication Date
- 2025-08-15
AI Technical Summary
The existing VR and eye tracking technologies have significant limitations in multifunction integration, dynamic interaction accuracy and personalized adaptation, resulting in low detection accuracy, poor user experience, and equipment overheating problems affect the reliability of medical testing.
High-precision eye tracking, multimodal sensing and adaptive VR rendering technology are adopted, combined with binocular OLED display, Fresnel lens group, ToF depth sensor, nine-axis IMU and dual air duct cooling system, to realize the space-time synchronization of data and multi-source fusion, and the scene image is adjusted through adaptive VR rendering to dynamically adapt to user vision.
It significantly improves the accuracy and application range of visual function detection, solves the problem of equipment overheating, and provides personalized adaptation and high-precision multi-function visual function detection.
Smart Images

Figure CN120491809A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of visual function assessment, and more particularly to a multifunctional visual function detection device and method based on VR and eye tracking. Background Art
[0002] In recent years, with the rapid development of virtual reality (VR) and eye tracking technology, their applications in medical diagnosis and visual function assessment have gradually attracted attention. However, existing technologies still have significant limitations, especially in terms of multifunctional integration, dynamic interaction accuracy, and personalized adaptation, which restrict their clinical practicality and user experience. Specific shortcomings are as follows:
[0003] (1) Traditional vertigo detection equipment (such as electronystagmus) mainly focuses on vestibular function assessment and lacks the ability to comprehensively detect multi-dimensional visual functions such as stereoscopic vision, dynamic vision, and color vision.
[0004] (2) In dynamic vision assessment, the refresh rate (usually ≤90Hz) and delay (>10ms) of existing VR display modules are difficult to match high-speed eye movement responses, resulting in timing deviations between stimulation and feedback.
[0005] (3) The hardware design of existing equipment generally lacks adaptability to individual user differences. For example, the fixed-focal-length Fresnel lens cannot be dynamically adjusted according to the user's diopter, which means that myopic or hyperopic patients need to wear corrective glasses during the test, introducing additional interference.
[0006] (4) Long-term operation of high-power VR modules and multimodal sensors can easily lead to device overheating. Traditional single-duct cooling solutions are difficult to balance the temperature control requirements of different modules. For example, some commercial eye trackers have insufficient heat dissipation, and their temperatures exceed 50°C after one hour of continuous operation, causing data drift and even hardware failure, seriously affecting the reliability of medical testing.
[0007] Therefore, how to provide a multifunctional visual function detection device and method based on VR and eye tracking is a problem that needs to be solved urgently by those skilled in the art. Summary of the Invention
[0008] In view of this, the present invention provides a multifunctional visual function detection device and method based on VR and eye tracking. By integrating high-precision eye tracking, multimodal sensing and adaptive VR rendering technology, the above problems are solved, and the detection accuracy and application scope are significantly improved.
[0009] In order to achieve the above object, the present invention provides the following technical solutions:
[0010] A multifunctional visual function testing device based on VR and eye tracking, comprising:
[0011] Virtual reality (VR) display module, used to generate dynamic scene images and obtain eye tracking data and ambient light data;
[0012] Multimodal sensing module for acquiring depth information and posture data;
[0013] Data processing module, used for spatiotemporal synchronization and multi-source data fusion of eye tracking data, ambient light data, depth information, and posture data;
[0014] The adaptive VR rendering engine receives the data processed by the data processing module and dynamically adjusts the type of dynamic scene images of the virtual reality VR display module according to the user's vision data;
[0015] The evaluation and output module performs visual function evaluation according to the type of dynamic scene image and outputs the evaluation results.
[0016] Furthermore, the virtual reality VR display module includes a binocular OLED display screen, a Fresnel lens group binocular infrared camera and an ambient light sensor;
[0017] The Fresnel lens group, the binocular OLED display screen, and the binocular infrared camera are arranged in a common optical path through a beam splitter prism to achieve optical path multiplexing. The binocular infrared camera and the ambient light sensor respectively collect eye tracking data and ambient light data, and transmit them to the data processing module.
[0018] Furthermore, the multimodal sensing module includes a ToF depth sensor, a nine-axis IMU, and a dual-duct cooling system;
[0019] The ToF depth sensor and the nine-axis IMU respectively collect depth information and posture data and transmit them to the data processing module, and the dual-duct cooling system performs internal heat dissipation.
[0020] Furthermore, the data processing module uses a preset spatiotemporal synchronization model to achieve spatiotemporal synchronization. The expression of the spatiotemporal synchronization model is:
[0021]
[0022] Where θsync(t) is the head rotation angle after synchronization; t is the current timestamp; t i is the i-th gyroscope sampling time point; Δt is the gyroscope sampling interval; n is the interpolation order, n = 3; θgyro(t i ) is the gyroscope at time t i The measured angular velocity value.
[0023] Furthermore, the data processing module uses an improved cubature Kalman filter (CKF) to achieve multi-source data fusion, including:
[0024] The state vector X is:
[0025] X=[C L ,C R ,θ gyro ,a accel ,δ VR ] T
[0026] Equation of state:
[0027]
[0028] Observation equation:
[0029]
[0030] Where C L ,C R is the coordinate of the pupil center of gravity; θgyro is the angular velocity of the gyroscope; a accel is the three-axis acceleration value of the accelerometer; δ VR is the VR scene parallax; f is the nonlinear state transfer function; w k is the process noise, which obeys the Gaussian distribution with mean 0 and covariance matrix Q; h is the nonlinear observation function; v k is the observation noise, which obeys a Gaussian distribution with mean 0 and covariance matrix R.
[0031] Furthermore, the types of dynamic scene images of the virtual reality VR display module include:
[0032] Stereoscopic vision testing, dynamic vision analysis, color vision assessment and vestibular function testing.
[0033] A multifunctional visual function testing method based on VR and eye tracking, comprising:
[0034] S100: Generate dynamic scene images and obtain eye tracking data and ambient light data;
[0035] S200: Acquire depth information and posture data;
[0036] S300: performing spatiotemporal synchronization and multi-source data fusion processing of eye tracking data, ambient light data, depth information, and posture data;
[0037] S400: Dynamically adjusting the type of the dynamic scene image according to the processed data and the user's vision data;
[0038] S500: Perform visual function evaluation according to the type of dynamic scene image and output the evaluation result.
[0039] It can be seen from the above technical solution that compared with the existing technology, the present invention discloses a multifunctional visual function detection device and method based on VR and eye tracking, which deeply integrates VR scene interaction and multimodal sensory data to achieve high-precision eye tracking and nystagmus quantitative analysis, significantly improving the detection accuracy and application scope. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are merely embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying any creative work.
[0041] Figure 1 Schematic diagram of the device structure of the present invention;
[0042] Figure 2 Schematic diagram of the method of the present invention. DETAILED DESCRIPTION
[0043] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.
[0044] See also Figure 1 The embodiment of the present invention discloses a multifunctional visual function detection device based on VR and eye tracking, comprising:
[0045] Virtual reality (VR) display module, used to generate dynamic scene images and obtain eye tracking data and ambient light data;
[0046] Multimodal sensing module for acquiring depth information and posture data;
[0047] Data processing module, used for spatiotemporal synchronization and multi-source data fusion of eye tracking data, ambient light data, depth information, and posture data;
[0048] The adaptive VR rendering engine receives the data processed by the data processing module and dynamically adjusts the type of dynamic scene images of the virtual reality VR display module according to the user's vision data;
[0049] The evaluation and output module performs visual function evaluation according to the type of dynamic scene image and outputs the evaluation results.
[0050] In a specific embodiment, the virtual reality VR display module includes a binocular OLED display screen, a Fresnel lens group binocular infrared camera and an ambient light sensor;
[0051] The Fresnel lens group, the binocular OLED display screen, and the binocular infrared camera are arranged in a common optical path through a beam splitter prism to achieve optical path multiplexing. The binocular infrared camera and the ambient light sensor respectively collect eye tracking data and ambient light data, and transmit them to the data processing module.
[0052] Specifically, the VR display module includes a binocular 4K OLED display with a refresh rate of 120Hz;
[0053] The integrated design of Fresnel lens group and beam splitter prism has an optical path multiplexing efficiency of ≥90%;
[0054] The ambient light sensor dynamically adjusts the display brightness (range 1 to 1000 nit).
[0055] Specifically, the Fresnel lens system is located between the binocular OLED display and the human eye. Light from the OLED display is refracted by the Fresnel lens system, changing its direction and focusing the image on the display. This simultaneously amplifies the image size, allowing users to see a clearer picture from a wider viewing angle, creating an immersive experience similar to a giant-screen cinema.
[0056] Specifically, the Fresnel lens system and the binocular infrared camera share a common optical path via a beam-splitting prism. This design allows the infrared camera to utilize some of the light refracted or reflected by the Fresnel lens system, without affecting the display optical path, to achieve functions such as eye tracking. For example, the binocular infrared camera can more accurately monitor the user's eye position, gaze direction, and eye movement status by capturing infrared light reflected from the eyes and combining it with the Fresnel lens system's adjustment of the optical path.
[0057] In one specific embodiment, the multimodal sensing module includes a ToF depth sensor, a nine-axis IMU, and a dual-duct cooling system;
[0058] The ToF depth sensor and the nine-axis IMU respectively collect depth information and posture data and transmit them to the data processing module, and the dual-duct cooling system performs internal heat dissipation.
[0059] Specifically, it has a binocular infrared camera (resolution 640×480, frame rate 60Hz), equipped with a ring infrared fill light (wavelength 850nm); a nine-axis IMU (gyroscope sampling rate 200Hz, accelerometer range ±16g); and a ToF depth sensor (accuracy ±1mm, detection range 0.1~5m).
[0060] In a specific embodiment, a spatiotemporal synchronization model and an improved volumetric Kalman filter algorithm model are preset in the data processing module to support real-time fusion of eye movement, IMU and VR data.
[0061] Specifically, a preset space-time synchronization model is used to achieve space-time synchronization. The expression of the space-time synchronization model is:
[0062]
[0063] Where θsync(t) is the head rotation angle after synchronization; t is the current timestamp; t i is the i-th gyroscope sampling time point; Δt is the gyroscope sampling interval; n is the interpolation order, n = 3; θgyro(t i ) is the gyroscope at time t i The measured angular velocity value.
[0064] Specifically, the data processing module uses an improved cubature Kalman filter (CKF) to achieve multi-source data fusion, including:
[0065] The state vector X is:
[0066] X=[C L ,C R ,θ gyro ,a accel ,δ VR ] T
[0067] Equation of state:
[0068]
[0069] Observation equation:
[0070]
[0071] Where C L ,C R is the coordinate of the pupil center of gravity; θgyro is the angular velocity of the gyroscope; a accel is the three-axis acceleration value of the accelerometer; δ VR is the VR scene parallax; f is the nonlinear state transfer function; w k is the process noise, which obeys the Gaussian distribution with mean 0 and covariance matrix Q; h is the nonlinear observation function; v k is the observation noise, which obeys a Gaussian distribution with mean 0 and covariance matrix R.
[0072] In a specific embodiment, the types of dynamic scene images of the virtual reality VR display module include:
[0073] Stereoscopic vision testing, dynamic vision analysis, color vision assessment and vestibular function testing.
[0074] In a specific embodiment, the evaluation and output module embedded AI chip deploys a lightweight U-Net network with a single frame processing delay of <3ms;
[0075] More specifically, the method flow is as follows:
[0076] Pupil segmentation and motion compensation: Input binocular optical images into the U-Net network, output pupil mask and center of gravity coordinates; calculate the head rotation angle θ(t) through gyroscope integration, and convert pupil displacement into actual eye rotation value based on calibration parameters.
[0077] Nystagmus quantification and visual function testing: A three-dimensional spatiotemporal sequence model is constructed to calculate the nystagmus score; a VR scene generates random dot stereograms (RDS) and motion targets, and eye movement data is combined to evaluate stereoscopic vision and dynamic vision.
[0078] Specifically, the three-dimensional spatiotemporal sequence model integrates eye movement and inertial data to calculate the nystagmus intensity score:
[0079]
[0080] Where, f gyro (θ i ) is the eye rotation component predicted by the gyroscope; g accel (a i ) is the auxiliary correction term of the accelerometer (eliminating linear motion interference); θ i is the gyroscope bias; Δα i ,Δβ i is the actual eye rotation amount.
[0081] Specifically, N score ≥2.5 is considered pathological nystagmus.
[0082] Adaptive interaction and output:
[0083] Adjust lens focal length according to user's diopter;
[0084] Output diagnostic report including nystagmus grade, minimum discernible parallax and dynamic visual acuity gain.
[0085] More specifically, the dynamic vision assessment model can also be implemented by calculating the smooth pursuit gain, with a determination threshold of Gain ≤ 0.3, and the calculation formula is:
[0086]
[0087] Where Gain is the smooth pursuit gain (dimensionless), the smaller the value, the better the dynamic vision; θ˙ eye(t) is the angular velocity of the eyeball (unit: radians / second); θ˙ target (t) is the angular velocity of the virtual target (unit: radians / second); T is the total test duration (unit: seconds); is the maximum angular velocity of the target (unit: radians / second).
[0088] Specifically, the present invention provides a visual function testing device based on VR and eye tracking, which can be implemented in a form factor similar to an eye mask. As the primary auxiliary tool in nystagmus examinations, the eye mask's core function is to optimize eye movement recording in a darkroom environment. It is suitable for vestibular function assessment, central nervous system lesion localization, and balance disorder diagnosis. It can be used in conjunction with a vertigo diagnostic swivel chair, caloric test stimulator, visual target screen, dynamic and static balance table, head impulse test, and other scenarios where nystagmus can be used to assist in the diagnosis and treatment of vertigo.
[0089] On the other hand, see Figure 2 The embodiment of the present invention further discloses a multifunctional visual function detection method based on VR and eye tracking, comprising:
[0090] S100: Generate dynamic scene images and obtain eye tracking data and ambient light data;
[0091] S200: Acquire depth information and posture data;
[0092] S300: performing spatiotemporal synchronization and multi-source data fusion processing of eye tracking data, ambient light data, depth information, and posture data;
[0093] S400: Dynamically adjusting the type of the dynamic scene image according to the processed data and the user's vision data;
[0094] S500: Perform visual function evaluation according to the type of dynamic scene image and output the evaluation result.
[0095] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. Reference can be made to the common and similar parts between the various embodiments. For the devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple, and the relevant parts can be referred to the method description.
[0096] The above description of the disclosed embodiments is intended to enable one skilled in the art to implement or use the present invention. Various modifications to these embodiments will be readily apparent to one skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention is not limited to the embodiments shown herein but is intended to conform to the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A multifunctional visual function testing device based on VR and eye tracking, characterized in that: include: Virtual reality (VR) display module, used to generate dynamic scene images and obtain eye tracking data and ambient light data; Multimodal sensing module for acquiring depth information and posture data; Data processing module, used for spatiotemporal synchronization and multi-source data fusion of eye tracking data, ambient light data, depth information, and posture data; The adaptive VR rendering engine receives the data processed by the data processing module and dynamically adjusts the type of dynamic scene images of the virtual reality VR display module according to the user's vision data; The evaluation and output module performs visual function evaluation according to the type of dynamic scene image and outputs the evaluation results.
2. A multifunctional visual function detection device based on VR and eye tracking according to claim 1, characterized in that: The virtual reality VR display module includes a binocular OLED display screen, a Fresnel lens group, a binocular infrared camera and an ambient light sensor; The Fresnel lens group, the binocular OLED display screen, and the binocular infrared camera are arranged in a common optical path through a beam splitter prism to achieve optical path multiplexing. The binocular infrared camera and the ambient light sensor respectively collect eye tracking data and ambient light data, and transmit them to the data processing module.
3. The multifunctional visual function detection device based on VR and eye tracking according to claim 1, characterized in that: The multimodal sensing module includes a ToF depth sensor, a nine-axis IMU, and a dual-duct cooling system; The ToF depth sensor and the nine-axis IMU respectively collect depth information and posture data and transmit them to the data processing module, and the dual-duct cooling system performs internal heat dissipation.
4. The multifunctional visual function detection device based on VR and eye tracking according to claim 1, characterized in that: The data processing module uses a preset spatiotemporal synchronization model to achieve spatiotemporal synchronization. The expression of the spatiotemporal synchronization model is: Where θsync(t) is the head rotation angle after synchronization; t is the current timestamp; t i is the i-th gyroscope sampling time point; Δt is the gyroscope sampling interval; n is the interpolation order, n = 3; θgyro(t i ) is the gyroscope at time t i The measured angular velocity value.
5. The multifunctional visual function detection device based on VR and eye tracking according to claim 1, characterized in that: The data processing module uses an improved cubature Kalman filter (CKF) to achieve multi-source data fusion, including: The state vector X is: X=[C L ,C R ,i gyro ,a accel ,d VR ] T Equation of state: Observation equation: Where C L ,C R is the coordinate of the pupil center of gravity; θgyro is the angular velocity of the gyroscope; a accel is the three-axis acceleration value of the accelerometer; δ VR is the VR scene parallax; f is the nonlinear state transfer function; w k is the process noise, which obeys the Gaussian distribution with mean 0 and covariance matrix Q; h is the nonlinear observation function; v k is the observation noise, which obeys a Gaussian distribution with mean 0 and covariance matrix R.
6. The multifunctional visual function detection device based on VR and eye tracking according to claim 1, characterized in that: The types of dynamic scene images of the virtual reality VR display module include: Stereoscopic vision testing, dynamic vision analysis, color vision assessment and vestibular function testing.
7. A multifunctional visual function detection method based on VR and eye tracking, characterized in that: include: S100: Generate dynamic scene images and obtain eye tracking data and ambient light data; S200: Acquire depth information and posture data; S300: performing spatiotemporal synchronization and multi-source data fusion processing of eye tracking data, ambient light data, depth information, and posture data; S400: Dynamically adjusting the type of the dynamic scene image according to the processed data and the user's vision data; S500: Perform visual function evaluation according to the type of dynamic scene image and output the evaluation result.
Citation Information
Cited By
Self-adaptive visual stimulation presentation and eye movement signal quantitative analysis method and system based on mobile terminal
CN121845513A