A visual behavior simulation and analysis management system

CN122604299APending Publication Date: 2026-08-21AIER EYE HOSPITAL GRP CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610760061.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-29
Publication Date
2026-08-21

AI Technical Summary

Technical Problem

这些量表虽然能反映患者的主观感受,但在实际应用中存在显著短板:1、主观偏倚:依赖患者的回忆与自我感知,缺乏客观行为数据支撑;2、敏感度不足:难以捕捉病情的微小纵向变化,且常受困于“天花板效应”,导致对轻度受损或视力较好患者的评估失效;3、精度受限:原始量表的颗粒度较粗,难以满足精准医疗对评估精度的要求

Benefits of technology

[0017] The visual behavior simulation and analysis management system provided in this application collects the patient's full-body motion data through motion capture equipment to generate motion capture data. A VR headset allows the patient to interact with virtual objects in a virtual environment, collecting eye movement data, posture data, and task completion data. The doctor's management terminal performs time compensation and time axis alignment on the aforementioned multimodal data under a unified reference clock, displaying the patient's task execution process in a simulated life scenario from a third-person perspective, and obtaining evaluation results. This provides researchers with quantitative objective indicators of visual behavior as an objective basis for functional vision assessment, and provides data support for vision-related quality of life research, follow-up comparisons, or auxiliary analysis. Furthermore, through the doctor's management terminal, in a free-movement VR scene, using the rigid body of the head as a calibration bridge and the motion capture world coordinate system as a stable spatial reference, gaze space correction is performed to determine the corrected gaze point position, thereby improving the spatial accuracy of the gaze point overlay display.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122604299A_ABST
    Figure CN122604299A_ABST
Patent Text Reader

Abstract

The application provides a visual behavior simulation and analysis management system, comprising a motion capture device, a VR head-mounted device and a doctor management terminal. The motion capture device collects and generates motion capture data; the VR head-mounted device receives task instructions, state data and motion capture data for local visual rendering, displays a virtual environment, a virtual character and task information to a patient from a first perspective, and collects eye movement data, action posture data and task completion data; the doctor management terminal receives the motion capture data, the eye movement data and the action posture data, performs time compensation and unified time axis alignment on the multi-source heterogeneous data, and performs spatial correction on the eye movement data, displays the interaction process of the patient and the virtual environment from a third perspective, and outputs an evaluation result. The application can simultaneously collect and analyze visual data and behavior data in a simulated life scene, and in a free motion VR scene, improve the spatial accuracy of gaze point position calculation to meet the demand for accurate quantitative evaluation of functional vision.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of virtual reality (VR) interaction technology, and in particular to a visual behavior simulation and analysis management system. Background Technology

[0002] In assessing patients' functional vision and vision-related quality of life, clinicians often rely on patient-reported outcome measures such as the NEI VFQ-25 (National Eye Institute Visual Function Questionnaire - 25-item version) and the VALV VFQ-48 (Veterans Affairs Low Vision Visual Function Questionnaire - 48-item version). While these scales reflect patients' subjective feelings, they have significant shortcomings in practical application: 1. Subjective bias: They rely on patients' recollections and self-perceptions, lacking support from objective behavioral data; 2. Insufficient sensitivity: They struggle to capture subtle longitudinal changes in the condition and are often hampered by the "ceiling effect," leading to ineffective assessments for patients with mild vision impairment or good visual acuity; 3. Limited precision: The original scales have a coarse granularity, making it difficult to meet the precision requirements of precision medicine.

[0003] However, functional vision and vision-related quality of life are not directly equivalent to traditional static visual function values ​​such as visual acuity and visual field, nor are they equivalent to patients' subjectively reported quality of life scores. Instead, they are more accurately reflected in a patient's ability to acquire, allocate, and utilize visual information to complete tasks in scenarios that closely resemble daily life. This ability is typically accompanied by changes in visual search, fixation allocation, head-eye coordination, body movement compensation, and task completion performance. Therefore, to obtain objective indicators that reflect functional vision, it is necessary to simultaneously collect visual and behavioral data from patients during task execution in simulated life scenarios and conduct joint analysis of these multimodal data.

[0004] In existing technologies, some studies have attempted to use virtual reality (VR) technology to construct simulated life scenarios to assess patients' functional visual abilities. For example, some solutions combine VR headsets with eye-tracking to record patients' gaze behavior in virtual environments. However, such solutions typically rely solely on the VR headset's own SLAM (Simultaneous Localization and Mapping) system to obtain spatial pose information. SLAM systems are prone to cumulative drift when patients perform large-scale free movements, leading to a gradual decrease in the spatial accuracy of the gaze point position over time, making it difficult to meet the needs for precise quantitative assessment of functional vision.

[0005] Therefore, how to provide a technical system that can simultaneously collect and analyze visual and behavioral data in simulated life scenarios, and improve the spatial accuracy of gaze point position calculation in free-movement VR scenarios, so as to meet the need for accurate quantitative evaluation of functional vision, is an urgent problem to be solved in this field. Summary of the Invention

[0006] To address the aforementioned technical issues, this application provides a visual behavior simulation and analysis management system that can simultaneously collect and analyze visual and behavioral data in simulated life scenarios, and improve the spatial accuracy of gaze point position calculation in free-movement VR scenarios, thereby meeting the need for precise quantitative evaluation of functional vision.

[0007] The technical solution provided in this application is as follows: A visual behavior simulation and analysis management system includes: motion capture equipment, VR headset, and doctor management terminal, wherein: The motion capture device is used to collect the patient's whole-body motion data, generate corresponding motion capture data, and send the motion capture data to the doctor's management terminal and the VR headset. The doctor management terminal is used to issue test task instructions to the VR headset and to perform bidirectional synchronization of control instructions and status data with the VR headset. The VR headset is used to perform local real-time visualization rendering based on the test task instructions, the status data, and the motion capture data. It displays the virtual environment, the patient's corresponding virtual character, and task information to the patient in real time from a first-person perspective. It enables the patient to interact with virtual objects in the virtual environment through head and hand tracking and gesture recognition. At the same time, it displays a virtual character driven by the motion capture data and synchronized with the patient's real body movements. It records eye movement data, movement posture data, and task completion data related to the patient and sends them to the doctor's management terminal. The doctor management terminal is also used to perform time compensation and unified time axis alignment on the motion capture data, eye movement data and action posture data with a unified reference clock, and to perform local real-time visualization rendering based on the motion capture data and action posture data. It displays the body movement state of the virtual character corresponding to the patient driven by the motion capture data in the virtual environment, as well as the process of the patient interacting with virtual objects in the virtual environment, to the doctor in real time from a third perspective, and generates and displays evaluation results based on the task completion data. The doctor management terminal is also used to construct a gaze ray based on the eye position coordinates and eye gaze direction in the eye tracking data. Based on the head rigid body pose corresponding to the motion capture data, and the external parameter transformation relationship between the eye coordinate system, head display coordinate system, and head rigid body coordinate system, the gaze ray is sequentially transformed from the eye coordinate system to the head display coordinate system, head rigid body coordinate system, and motion capture world coordinate system. When the gaze ray needs to be displayed in a virtual scene, the gaze ray is transformed to the virtual scene world coordinate system according to the transformation relationship between the motion capture world coordinate system and the virtual scene world coordinate system to obtain the transformed gaze ray. The intersection point of the transformed gaze ray with the reference plane or virtual scene object in the virtual environment is calculated to obtain the corrected gaze point position, and the corrected gaze point position is marked and displayed.

[0008] Preferably, in the visual behavior simulation and analysis management system, when the doctor management terminal performs the calculation of the intersection point of the converted gaze ray and the reference plane or virtual scene object in the virtual environment to obtain the corrected gaze point position, it is specifically used for: When the reference plane is represented by a plane equation, the intersection point of the transformed gaze ray and the reference plane is obtained according to the plane equation, and the corrected gaze point position is obtained. When the virtual scene object is represented by a collider or a mesh model, ray detection is performed along the transformed gaze ray, and the hit point corresponding to the minimum positive path parameter is taken as the corrected gaze point position.

[0009] Preferably, in the visual behavior simulation and analysis management system, when the doctor management terminal executes the process of constructing a gaze ray based on the eye position coordinates and eye gaze direction in the eye-tracking data, and based on the head rigid body pose corresponding to the motion capture data, and the external parameter transformation relationship between the eye coordinate system, the head display coordinate system, and the head rigid body coordinate system, sequentially transforming the gaze ray from the eye coordinate system to the head display coordinate system, the head rigid body coordinate system, and the motion capture world coordinate system; and when it is necessary to display the gaze ray in a virtual scene, transforming the gaze ray to the virtual scene world coordinate system according to the transformation relationship between the motion capture world coordinate system and the virtual scene world coordinate system to obtain the transformed gaze ray, specifically it is used for: A gaze ray is constructed based on the eye position coordinates and gaze direction in the eye movement data. The specific expression of the gaze ray is as follows: ; in, Indicates the gaze ray, This represents the origin of the gaze ray in the eye coordinate system E. This represents the unit gaze direction in eye coordinate system E. Indicates the path parameters of the gaze ray; Based on the head rigid body pose corresponding to the motion capture data, and the external parameter transformation relationship between the eye coordinate system, head display coordinate system, and head rigid body coordinate system, the gaze ray is sequentially transformed from the eye coordinate system to the head display coordinate system, head rigid body coordinate system, and motion capture world coordinate system to obtain the origin of the gaze ray and the gaze ray direction vector in the motion capture world coordinate system. The specific expression is as follows: ; ; in, Indicates time The origin of the gaze ray in the world coordinate system M of the motion capture system. Indicates time The transformation matrix from the rigid body coordinate system R of the lower head to the motion capture world coordinate system M. This represents the calibration extrinsic parameter from the head-mounted display coordinate system H to the head rigid body coordinate system R. This represents the calibration extrinsic parameter from the eye coordinate system E to the head-mounted display coordinate system H. Indicates time The gaze ray direction vector in the world coordinate system M of the motion capture system. Let represent the rotation sub-matrix of the head rigid body coordinate system R relative to the motion capture world coordinate system M at time t. Let represent the rotation submatrix in the homogeneous transformation matrix from the head-mounted display coordinate system H to the head rigid body coordinate system R. This represents the rotation submatrix in the homogeneous transformation matrix from the eye coordinate system E to the head-mounted display coordinate system H; When the gaze ray needs to be displayed in a virtual scene, the gaze ray is transformed to the virtual scene world coordinate system according to the transformation relationship between the motion capture world coordinate system and the virtual scene world coordinate system, resulting in the transformed gaze ray. The specific expression of the transformed gaze ray is as follows: ; ; in, This indicates the converted gaze ray. This represents the origin of the gaze ray in the virtual scene world coordinate system U at time t. Let represent the gaze ray direction vector in the virtual scene world coordinate system U at time t. This represents the transformation relationship from the motion capture world coordinate system M to the virtual scene world coordinate system U. This represents the rotation submatrix in the homogeneous transformation matrix from the motion capture world coordinate system M to the virtual scene world coordinate system U.

[0010] Preferably, in the visual behavior simulation and analysis management system, when the doctor management terminal executes the step of calculating the intersection point of the transformed gaze ray and the reference plane according to the plane equation to obtain the corrected gaze point position, it is specifically used for: For the reference plane Suppose it satisfies the plane equation: ; in, It is a plane normal vector. For plane constants, This indicates that the plane equation n is satisfied on the reference plane Π. The coordinate vector of any point in three-dimensional space where x+c=0 Indicates the transpose operation; Find the intersection point between the transformed gaze ray and the reference plane to obtain the intersection point parameters. for: ; exist At that time, the corrected fixation point position is: ; in, This indicates the corrected fixation point position at time t.

[0011] Preferably, in the visual behavior simulation and analysis management system, the VR headset includes an eye-tracking module and a posture tracking module; The eye-tracking module is used to track the patient's eye movements to generate the eye-tracking data, which includes one or more of the following: eye gaze direction, eye position coordinates, eye movement trajectory, and timestamp. The posture tracking module is used to track the patient's head posture and hand posture to generate the motion posture data.

[0012] Preferably, in the visual behavior simulation and analysis management system, the doctor management terminal is further used for: Perform extrinsic parameter calibration to obtain the relative poses between each pair of the eye coordinate system, head display coordinate system, and head rigid body coordinate system; During the patient's free movement, the correction transformation from the SLAM world coordinate system to the motion capture world coordinate system is calculated based on the difference between the VR headset pose output by the SLAM-based headset and the rigid body pose of the head. The translation and rotation components of the correction transformation are updated smoothly, and the smoothed correction transformation is used to transform the head display pose, gaze point position or eye movement trajectory in the SLAM world coordinate system to obtain the corresponding head display pose, gaze point position or eye movement trajectory in the motion capture world coordinate system.

[0013] Preferably, in the visual behavior simulation and analysis management system, the doctor management terminal is further used for: Based on the head angular velocity sequence output by the VR headset and the head rigid body angular velocity sequence output by the motion capture device, cross-correlation analysis is performed under a unified reference time axis. The time offset corresponding to the maximum value of the cross-correlation function is used as the residual time delay correction amount, and the unified reference time of the eye movement data is corrected using the residual time delay correction amount.

[0014] Preferably, the visual behavior simulation and analysis management system further includes a motion capture data access terminal; The doctor management terminal is connected to the motion capture device through the motion capture data access terminal; The doctor management terminal is also used to record the local timestamp and frame number of the VR headset, the motion capture device or the motion capture data access terminal when generating or receiving data. The doctor management terminal is also used to record the source timestamp, source frame number, and access time of the motion capture data source when the motion capture data is accessed via a local SDK callback method.

[0015] Preferably, in the visual behavior simulation and analysis management system, the doctor management terminal is further used to periodically initiate synchronization requests to the VR headset and receive synchronization responses, so as to obtain the first sending time of the synchronization request and the first receiving time of the synchronization response. The VR headset is also used to record the second time of receiving the synchronization request and the second time of sending the synchronization response; The doctor management terminal is also used to establish a time mapping relationship between the local time of the VR data source and the unified reference clock based on the first sending time, the first receiving time, the second receiving time and the second sending time; The doctor management terminal is also used to establish a time mapping relationship for motion capture data sources using any of the following methods: When the motion capture data is accessed via a local SDK callback, a time mapping relationship between the local time of the motion capture data source and the unified reference clock is established based on the source timestamp, the source frame number, and the access time of the motion capture data source. When the motion capture data source supports synchronization requests and synchronization responses, a synchronization request is periodically initiated to the motion capture data source and a synchronization response is received to obtain the third sending time of the synchronization request and the third receiving time of the synchronization response, as well as the fourth receiving time of the synchronization request and the fourth sending time of the synchronization response recorded by the motion capture data source; based on the third sending time, the third receiving time, the fourth receiving time, and the fourth sending time, a time mapping relationship between the local time of the motion capture data source and the unified reference clock is established.

[0016] Preferably, in the visual behavior simulation and analysis management system, when the doctor management terminal performs time compensation and unified time axis alignment of the motion capture data, eye movement data, and action posture data with a unified reference clock, it is specifically used for: Based on the time mapping relationship between the local time of the VR data source and the motion capture data source and the unified reference clock, the motion capture data, the eye tracking data and the action posture data are converted to the unified reference time axis, and out-of-order rearrangement and packet loss detection are performed according to the frame sequence number, and a cache queue is established for each type of data. Based on the target rendering time, adjacent samples are selected from each cache queue and interpolated. Specifically, linear interpolation is performed on the position data in the adjacent samples, spherical linear interpolation is performed on the rotation data in the adjacent samples, and normalized interpolation is performed on the gaze direction vector in the adjacent samples, thereby obtaining synchronized sample data under a unified reference time axis.

[0017] The visual behavior simulation and analysis management system provided in this application collects the patient's full-body motion data through motion capture equipment to generate motion capture data. A VR headset allows the patient to interact with virtual objects in a virtual environment, collecting eye movement data, posture data, and task completion data. The doctor's management terminal performs time compensation and time axis alignment on the aforementioned multimodal data under a unified reference clock, displaying the patient's task execution process in a simulated life scenario from a third-person perspective, and obtaining evaluation results. This provides researchers with quantitative objective indicators of visual behavior as an objective basis for functional vision assessment, and provides data support for vision-related quality of life research, follow-up comparisons, or auxiliary analysis. Furthermore, through the doctor's management terminal, in a free-movement VR scene, using the rigid body of the head as a calibration bridge and the motion capture world coordinate system as a stable spatial reference, gaze space correction is performed to determine the corrected gaze point position, thereby improving the spatial accuracy of the gaze point overlay display.

[0018] In summary, the above technical solutions can simultaneously collect and analyze visual and behavioral data in simulated life scenarios, and improve the spatial accuracy of gaze point position calculation in free-movement VR scenarios, so as to meet the needs of accurate quantitative evaluation of functional vision. Attached Figure Description

[0019] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0020] Figure 1 A schematic diagram of the structure of the visual behavior simulation and analysis management system provided in the embodiments of this application; Figure 2 A schematic diagram illustrating the selection of a virtual environment provided in the embodiments of this application; Figure 3 A schematic diagram of a first-view perspective provided for an embodiment of this application; Figure 4 A schematic diagram of a third-view perspective provided for an embodiment of this application; Figure 5 A schematic diagram illustrating the evaluation results provided in the embodiments of this application; Figure 6(a) is a schematic diagram of the virtual environment of a supermarket shelf scene provided in an embodiment of this application; Figure 6(b) is a visualization of the gaze behavior features provided in the embodiments of this application; Figure 7 The diagram illustrates the gait parameter quantification analysis provided in the embodiments of this application, wherein (a) is a diagram of vertical height and heel strike identification, (b) is a diagram of step length distribution, (c) is a diagram of vertical height and toe off identification, and (d) is a diagram of core gait quantification indicators. Figure 8(a) is a schematic diagram of the change of fixation point spatial error over time provided in the embodiment of this application; Figure 8(b) is a schematic diagram of the eye movement trajectory correction effect provided in the embodiment of this application; Figure 8(c) is an error distribution histogram provided in an embodiment of this application; Figure 8(d) is a cumulative SLAM pose drift curve provided in the embodiment of this application; Figure 9(a) is a schematic diagram of the time alignment error of the motion capture (120Hz) UDP wired data stream provided in the embodiment of this application as a function of time; Figure 9(b) is a schematic diagram of the time alignment error of the eye-tracking (72Hz) UDP wireless data stream provided in the embodiment of this application as a function of time; Figure 9(c) is a schematic diagram of the time alignment error of the attitude (72Hz) KCP wireless data stream provided in the embodiment of this application as a function of time; Figure 9(d) is a schematic diagram of the error distribution of the motion capture (120Hz) UDP wired data stream provided in the embodiment of this application; Figure 9(e) is a schematic diagram of the error distribution of the eye-tracking (72Hz) UDP wireless data stream provided in the embodiment of this application; Figure 9(f) is a schematic diagram of the error distribution of the attitude (72Hz) KCP wireless data stream provided in the embodiment of this application; Figure 10(a) is a schematic diagram of the pre-compensation angular velocity sequence (partially enlarged) provided in an embodiment of this application; Figure 10(b) is a schematic diagram of the cross-correlation function provided in the embodiment of this application; Figure 10(c) is a schematic diagram of the compensated angular velocity sequence (partially enlarged) provided in the embodiment of this application; Figure 10(d) is a schematic diagram of the alignment mean square error comparison provided in the embodiments of this application; Figure 11(a) is a schematic diagram of motion capture position interpolation rendering (partial magnification) from 120Hz to 60Hz provided in an embodiment of this application; Figure 11(b) is a schematic diagram of the arrival interval of dynamic capture UDP packets provided in the embodiment of this application; Figure 11(c) is a schematic diagram of the comparison of inter-frame smoothness in rendering provided in the embodiments of this application; Figure 11(d) is a schematic diagram of eye-tracking fixation interpolation rendering (partial magnification) from 72Hz to 60Hz provided in an embodiment of this application; Figure 11(e) is a schematic diagram of the eye-tracking UDP packet arrival interval provided in an embodiment of this application; Figure 11(f) is a schematic diagram of the comparison of rendering accuracy RMSE provided in the embodiments of this application. Detailed Implementation

[0021] To enable those skilled in the art to better understand the technical solutions in this application, the technical solutions in the embodiments of this application will be clearly and completely described below. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0022] In the embodiments provided in this application, it should be understood that the disclosed system can be implemented in other ways. The system embodiments described below are merely illustrative. For example, the division of units and modules is only a logical functional division, and there may be other division methods in actual implementation. For example, multiple units or modules may be combined, integrated into another system, or some features may be ignored or not executed. In addition, the coupling, direct coupling, or communication connection between the various components shown or discussed may be through some interfaces, indirect coupling or communication connection of devices or modules, and may be electrical, mechanical, or other forms.

[0023] In addition, each functional unit in the various embodiments of this application can be integrated into a single processor, or each unit can be a separate device, or two or more units can be integrated into a single device; each functional unit in the various embodiments of this application can be implemented in hardware or in the form of hardware plus software functional units.

[0024] It should be understood that the use of terms such as "system," "device," "unit," and / or "module" in this application is merely one method of distinguishing different components, elements, parts, sections, or assemblies at different levels. However, if other terms can achieve the same purpose, they may be replaced by other expressions.

[0025] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature. In the description of this application, "a plurality of" or "several" means two or more, unless otherwise explicitly specified.

[0026] It should be noted that the structures, proportions, sizes, etc., shown in the accompanying drawings of this specification are only for the purpose of assisting those skilled in the art in understanding and reading the content disclosed in the specification, and are not intended to limit the conditions under which this application can be implemented. Therefore, they have no substantial technical significance. Any modifications to the structure, changes in the proportions, or adjustments to the size should still fall within the scope of the technical content disclosed in this application, provided that they do not affect the effects and purposes that this application can produce.

[0027] like Figure 1 As shown, this application provides a visual behavior simulation and analysis management system, including: a motion capture device 1, a VR headset 2, and a doctor management terminal 3, wherein: Motion capture device 1 is used to collect the patient's whole-body motion data, generate corresponding motion capture data, and send the motion capture data to the doctor's management terminal 3 and VR head-mounted display device 2; Specifically, motion capture device 1 can employ existing motion capture equipment, such as the Nokov system, connected to an infrared camera array. Motion capture device 1 can capture the patient's full-body motion data at a high sampling rate and low latency to generate motion capture data. The motion capture data sent to the doctor's management terminal 3 is used to drive the display of the patient's corresponding virtual character in the doctor's management terminal 3 and serves as the input source for third-person visualization rendering, temporal compensation, and spatial correction. The motion capture data sent to the VR headset device 2 is used to drive the display of the patient's corresponding virtual character in the VR headset device 2, thereby reducing end-to-end latency in patient-side motion display. Motion capture data can be sent to the doctor's management terminal 3 and the VR headset device 2 via UDP, TCP, SDK local callbacks, or a combination thereof. In one example implementation, motion capture device 1 can output motion capture data at a frequency of approximately 120Hz via the Nokov SDK; this application is not limited to this.

[0028] The doctor management terminal 3 is used to send test task instructions to the VR headset 2 and to perform bidirectional synchronization of control instructions and status data with the VR headset 2. Specifically, the doctor management terminal 3 can be a high-performance PC terminal, capable of sending test task instructions to the VR headset 2 via a wireless transmission protocol. These test task instructions may include virtual scene switching information (such as selecting a test scene), initial virtual scene layout information (such as scene layout and initial object positions), and task information to be executed (such as a task list). Control instructions may include test pause / resume instructions. Status data may be scene parameters used to drive virtual scene rendering and object state updates, including at least one or more of the following: the patient's corresponding virtual character's movement trajectory, posture state, task progress information, scene object positions, and scene object states. The doctor management terminal 3 serves as the doctor's operation control terminal, proactively sending test tasks to the VR headset 2, controlling the start and stop of the test process and scene switching; simultaneously, it interacts bidirectionally with the VR headset 2, sending control instructions and receiving the real-time status of the VR headset 2, achieving data synchronization between the two ends.

[0029] VR headset device 2 is used to perform local real-time visualization rendering based on test task instructions, status data and motion capture data. It displays the virtual environment, the virtual character corresponding to the patient and task information to the patient in real time from a first-person perspective. It enables the patient to interact with virtual objects in the virtual environment through head and hand tracking and gesture recognition. At the same time, it displays a virtual character driven by motion capture data and synchronized with the patient's real body movements. It records eye movement data, movement posture data and task completion data related to the patient and sends them to the doctor's management terminal 3. Specifically, the VR headset 2 can be an existing head-mounted VR device, such as the Meta Quest Pro headset. The motion trajectory and posture of the patient's corresponding virtual character can be provided directly from motion capture data, or generated in the doctor's management terminal 3 based on motion capture data and then synchronized to the VR headset 2. Upon receiving the test task instructions, status data, and motion capture data, the VR headset 2 performs local real-time 3D visualization rendering, presenting the virtual environment, the patient's corresponding virtual character, and task information to the patient from a first-person perspective. The patient's corresponding virtual character is the patient's physical avatar in the virtual scene, driven by motion capture data and updated synchronously with the patient's physical movements in the real environment. The patient's corresponding virtual character is used to present a virtual body synchronized with the patient's real physical movements from a first-person perspective, enhancing the sense of presence and immersion during task execution. The interaction between the patient and virtual objects in the virtual environment is handled by head and hand tracking, gesture recognition, collision detection, grasping logic, and a physics engine on the VR headset 2 side, enabling tasks such as picking up, grasping, placing, moving, settling accounts, or obstacle avoidance. Furthermore, the VR headset 2 can also collect eye movement data and motion posture data related to the patient during the interaction process based on the pre-set eye tracking and posture tracking functions, record the patient's task completion data, and upload this data to the doctor's management terminal 3 via wireless transmission protocols (such as TCP / UDP / KCP protocols). Among them, eye movement data may include the patient's eye position coordinates, eye gaze direction, eye movement trajectory, etc., and motion posture data may include head / hand posture data, etc., and this application is not limited to these.

[0030] In some embodiments, the virtual environment is either a supermarket shopping scenario or a street travel scenario. The virtual environment can be selected through the human-computer interaction interface of the doctor management terminal 3, such as... Figure 2 As shown, virtual scene switching information is generated. Taking a supermarket shopping scene as an example, a shopping list can be pre-set as the task information to be executed. Then, the virtual scene switching information, the initial virtual scene layout information, and the task information to be executed are sent as test task instructions to the VR headset 2. After receiving the test task instructions, status data, and motion capture data, the VR headset 2 performs local real-time 3D visualization rendering to show the patient the supermarket shopping scene, the patient's corresponding virtual character, and the task information (shopping list) from a first-person perspective. Figure 3 As shown, patients need to follow the task information instructions to find the corresponding items and complete the entire interactive process from grabbing to settling at the cashier.

[0031] In other embodiments, when the virtual environment is a supermarket shopping scenario, the task completion data includes checkout completion data and task completion time. The checkout completion data characterizes the patient's performance in performing interactive operations such as picking up, grasping, moving, placing, and checking out objects within the virtual environment through head and hand tracking and gesture recognition on both sides of the VR headset; these interactive operations can be processed by collision detection, grasping logic, and a physics engine within the virtual scene. The task completion time is the total time from when the patient receives the task instruction and begins executing the task to when all task requirements are met, serving as a key indicator for evaluating the quality of the patient's task completion.

[0032] The doctor management terminal 3 is also used to perform time compensation and unified time axis alignment for motion capture data, eye tracking data and action posture data with a unified reference clock, and to perform local real-time visualization rendering based on motion capture data and action posture data. It displays the body movement state of the virtual character corresponding to the patient driven by motion capture data in the virtual environment, as well as the process of the patient interacting with virtual objects in the virtual environment, to the doctor in real time from a third perspective, and generates and displays evaluation results based on task completion data. Specifically, after receiving motion capture data, eye tracking data, and posture data, the doctor management terminal 3 performs time compensation and time axis alignment on the motion capture data, eye tracking data, and posture data using a unified reference clock to obtain synchronized rendering data. When calculating the gaze point position, it selects or interpolates the corresponding head rigid body pose based on the unified reference time axis, and then performs local real-time 3D visualization rendering. This displays the patient's corresponding virtual character's body movement state in the virtual environment, driven by the motion capture data, and the process of the patient interacting with virtual objects in the virtual environment through the interaction mechanism of the VR headset 2. The doctor management terminal 3 and the VR headset 2 display the same virtual environment and virtual character, with completely identical spatial layout. When the patient wears the VR headset 2 for visual behavior assessment tasks, the doctor can flexibly switch between third-person and third-person perspectives on the doctor management terminal 3, observing the patient's corresponding virtual character's actions, interaction process, and task completion from a third-person perspective. Taking a supermarket shopping scenario as an example, the visualization view of the doctor management terminal 3 is as follows: Figure 4 As shown. The doctor management terminal 3 can generate evaluation results based on task completion data using a preset scoring algorithm. These results may include the number of tasks completed, the task completion rate, etc., but this application is not limited to these. The doctor management terminal 3 can also display the evaluation results. Taking a supermarket shopping scenario as an example, the evaluation result visualization view is as follows: Figure 5 As shown.

[0033] The doctor management terminal 3 is also used to construct a gaze ray based on the eye position coordinates and eye gaze direction in the eye tracking data. Based on the head rigid body pose corresponding to the motion capture data, and the external parameter transformation relationship between the eye coordinate system, head display coordinate system, and head rigid body coordinate system, the gaze ray is transformed from the eye coordinate system to the head display coordinate system, head rigid body coordinate system, and motion capture world coordinate system in sequence. When it is necessary to display the gaze ray in the virtual scene, the gaze ray is transformed to the virtual scene world coordinate system according to the transformation relationship between the motion capture world coordinate system and the virtual scene world coordinate system to obtain the transformed gaze ray. The intersection point of the transformed gaze ray with the reference plane or virtual scene object in the virtual environment is calculated to obtain the corrected gaze point position, and the corrected gaze point position is marked and displayed.

[0034] Specifically, a rigid head body is fixedly set on the VR headset 2, and the relative pose between the rigid head body and the headset remains unchanged. The motion capture device 1 outputs the position and posture of the rigid head body in the motion capture world coordinate system in real time. The doctor management terminal 3 aligns the pose of the rigid head body with the eye movement data according to a unified reference time axis. Before the experiment begins, the external parameter transformation relationship between the eye coordinate system, the headset coordinate system and the rigid head body coordinate system is obtained through calibration. After obtaining the eye position coordinates and eye gaze direction from the eye-tracking data output by the VR headset 2, a gaze ray is first constructed in the eye coordinate system E. Then, the doctor management terminal 3, based on the position and posture of the head rigid body in the motion capture world coordinate system at the same reference moment and the external parameter transformation relationship, sequentially transforms the gaze ray from the eye coordinate system to the headset coordinate system, the head rigid body coordinate system, and the motion capture world coordinate system. When the gaze ray needs to be displayed in the virtual scene, it is further transformed from the motion capture world coordinate system to the virtual scene world coordinate system, thereby obtaining the transformed gaze ray in the virtual scene world coordinate system. The intersection point of the transformed gaze ray with the reference plane or virtual scene object in the virtual environment is calculated to obtain the corrected gaze point position. The doctor management terminal 3 performs eye-tracking spatial correction based on motion capture world coordinates through the above implementation method to obtain the corrected gaze point position, which can improve the spatial accuracy of the gaze point position under the condition of free patient movement. The corrected gaze point position can be superimposed and displayed in the third-person perspective visualization interface, so that the doctor can observe the patient's gaze position, gaze object, and gaze transfer process in the virtual environment while observing the patient's body movements and interactive behavior. Doctors can use this data to simultaneously observe patients' movements, gaze positions, gaze shifts, and task completion performance, and generate evaluation results based on the task completion data. Figure 4For example, the doctor management terminal 3 can overlay fixation point position information on the visualization interface, thereby helping doctors to more accurately and intuitively identify the patient's visual attention allocation characteristics, visual search strategies, motion compensation characteristics, and task performance during task execution. The doctor management terminal 3 can also generate and display corrected eye movement trajectories based on the corrected fixation point positions at continuous time points.

[0035] In some embodiments, when performing the above-mentioned eye-tracking spatial correction based on motion capture world coordinates, the doctor management terminal 3 is specifically used for: Eye position coordinates and gaze direction are obtained from eye-tracking data, and a gaze ray is constructed in eye coordinate system E. Let the origin of the gaze ray in eye coordinate system E be... The unit gaze direction in eye coordinate system E is Then gaze at the ray , is represented as:

[0036] in, The path parameter representing the gaze ray (i.e., the dimensionless distance parameter along the positive direction of the ray), λ≥0 ensures that the value is in the positive direction of the ray, and is used to parameterize the position of any point on the gaze ray; preferably, , , This refers to the original gaze direction vector in the eye-tracking data. This indicates the transpose operation.

[0037] Define the eye coordinate system as E, the head-mounted display coordinate system as H, the head rigid body coordinate system as R, the motion capture world coordinate system as M, and the virtual scene world coordinate system as U. Let... Indicates from coordinate system To coordinate system The homogeneous transformation matrix, This represents the rotation submatrix in the homogeneous transformation matrix.

[0038] Before the experiment, the relative pose between the head-mounted display coordinate system H and the head rigid body coordinate system R was determined through external parameter calibration. It is assumed that data was collected during the calibration process. Group corresponding points ,in, The coordinates of the calibration point in the head-mounted display coordinate system H are... Let H be the corresponding coordinate in the head rigid body coordinate system R. Then the extrinsic parameters from the head display coordinate system H to the head rigid body coordinate system R are... It can be obtained through the following rigid body registration model:

[0039] And satisfy:

[0040] in, The identity matrix. The extrinsic parameters from the eye coordinate system E to the head-mounted display coordinate system H. It can be obtained from the gaze calibration process. This represents the rotation submatrix in the homogeneous transformation matrix from the head-mounted display coordinate system H to the head rigid body coordinate system R; This represents the translation sub-vector in the homogeneous transformation matrix from the head-mounted display coordinate system H to the head rigid body coordinate system R (i.e., the offset of the origin of the head-mounted display coordinate system H relative to the origin of the head rigid body coordinate system R). This represents the rotation matrix to be optimized during the external parameter calibration solution process (rotation component from the head-mounted display coordinate system H to the head rigid body coordinate system R). This represents the translation vector to be optimized during the external parameter calibration solution process (the translation component from the head-mounted display coordinate system H to the head rigid body coordinate system R). Indicates the current moment; This represents the transpose of the rotated submatrix; Let R be the determinant of the rotation matrix R. For a valid rotation matrix, its value is always +1, ensuring the orthogonality of the rotation matrix.

[0041] Based on the head rigid body pose corresponding to the motion capture data, and the external parameter transformation relationship between the eye coordinate system E, the head display coordinate system H, and the head rigid body coordinate system R, the gaze ray is transformed from the eye coordinate system E to the head display coordinate system H, the head rigid body coordinate system R, and the motion capture world coordinate system M in sequence, so as to obtain the origin of the gaze ray and the gaze ray direction vector in the motion capture world coordinate system M.

[0042] At any moment If the pose of the head rigid body output by motion capture device 1 in the motion capture world coordinate system M is... The origin and direction of the gaze ray can then be transformed sequentially from the eye coordinate system E to the head-mounted display coordinate system H, the head rigid body coordinate system R, and the motion capture world coordinate system M. The expression for the origin of the gaze ray in the motion capture world coordinate system M is:

[0043] in, Indicates time The origin of the gaze ray in the world coordinate system M of the motion capture system. Indicates time The transformation matrix from the rigid body coordinate system R of the lower head to the motion capture world coordinate system M. This represents the calibration extrinsic parameter from the head-mounted display coordinate system H to the head rigid body coordinate system R. This represents the calibration extrinsic parameter from the eye coordinate system E to the head-mounted display coordinate system H.

[0044] The expression for the gaze ray direction vector in the motion capture world coordinate system M is:

[0045] in, Indicates time The gaze ray direction vector in the world coordinate system M of the motion capture system. It represents the rotation sub-matrix of the head rigid body coordinate system R relative to the motion capture world coordinate system M at time t (i.e., the rotation component in the head rigid body posture output in real time by motion capture device 1). Let represent the rotation submatrix in the homogeneous transformation matrix from the head-mounted display coordinate system H to the head rigid body coordinate system R. This represents the rotation submatrix in the homogeneous transformation matrix from the eye coordinate system E to the head-mounted display coordinate system H; When it is necessary to display the gaze ray in the virtual scene, the transformation relationship between the motion capture world coordinate system M and the virtual scene world coordinate system U is used. Furthermore, we can obtain:

[0046] in, The origin of the gaze ray in the virtual scene world coordinate system U at time t is represented (obtained by transforming the origin of the ray in the motion capture world coordinate system M). It represents the gaze ray direction vector in the virtual scene world coordinate system U at time t (the unit vector after the gaze direction in the motion capture world coordinate system M is transformed to the virtual scene coordinate system U). This represents the rotation submatrix in the homogeneous transformation matrix from the motion capture world coordinate system M to the virtual scene world coordinate system U.

[0047] Transform the gaze ray to the virtual scene world coordinate system U to obtain the transformed gaze ray. The specific expression of the transformed gaze ray is as follows:

[0048] in, This represents the transformed gaze ray. Then, the transformed gaze ray is calculated. The corrected gaze point position is obtained by finding the intersection point with a reference plane or virtual scene object in the virtual environment.

[0049] When the reference plane is represented by a plane equation, the transformed gaze ray is obtained from the plane equation. The intersection with the reference plane yields the corrected fixation point position; For the reference plane Suppose it satisfies the plane equation:

[0050] in, It is a plane normal vector. For plane constants, This indicates that the plane equation n is satisfied on the reference plane Π. The coordinate vector of any three-dimensional point in space where x+c=0; Find the intersection point between the transformed gaze ray and the reference plane to obtain the intersection point parameters. for:

[0051] exist At that time, the corrected fixation point position is:

[0052] in, This indicates the corrected fixation point position at time t.

[0053] When virtual scene objects are represented using colliders or mesh models, along the transformed gaze ray... Ray detection is performed. In ray detection, the path parameter represents the distance from the origin of the gaze ray to the hit point. A positive value represents the forward propagation direction of the ray, and the minimum positive value is the intersection point of the virtual object that the ray first collides with. The hit point corresponding to the path parameter with the minimum positive value is taken as the corrected gaze point position. The specific calculation process is a well-known technique and will not be elaborated here.

[0054] The visual behavior simulation and analysis management system provided in the above embodiments collects the patient's full-body motion data through motion capture device 1 to generate motion capture data. A VR headset 2 enables the patient to interact with virtual objects in the virtual environment, collecting the patient's eye movement data, posture data, and task completion data. The doctor's management terminal 3 performs time compensation and time axis alignment under a unified reference clock on the aforementioned multimodal data, displaying the patient's task execution process in a simulated life scenario from a third-person perspective, obtaining evaluation results. This provides researchers with quantitative objective indicators of visual behavior as an objective basis for functional vision assessment, and provides data support for vision-related quality of life research, follow-up comparisons, or auxiliary analysis. Furthermore, through the doctor's management terminal 3, in a free-movement VR scene, using the rigid body of the head as a calibration bridge and the motion capture world coordinate system as a stable spatial reference, gaze space correction is performed to determine the corrected gaze point position, thereby improving the spatial accuracy of the gaze point overlay display.

[0055] In summary, the above embodiments can simultaneously collect and analyze visual and behavioral data in simulated life scenarios, and improve the spatial accuracy of gaze point position calculation in free-movement VR scenarios, so as to meet the needs of accurate quantitative evaluation of functional vision.

[0056] In other embodiments of this application, the VR headset device 2 includes an eye-tracking module and a posture tracking module; the eye-tracking module is used to track the patient's eye movements to generate eye movement data, which includes one or more of the following: eye gaze direction, eye position coordinates, eye movement trajectory and its timestamp; the posture tracking module is used to track the patient's head posture and hand posture to generate motion posture data.

[0057] Specifically, the eye-tracking module and posture tracking module can be implemented based on existing eye-tracking technology and head-mounted display posture tracking technology. The eye-tracking module is specifically used to collect the patient's eye gaze direction and eye position coordinates. Based on the eye gaze direction and eye position coordinates, it calculates the gaze point position under free head movement, continuously calculating multiple gaze point positions. Based on the data acquisition timestamps, it obtains multiple ordered gaze point positions, generates an eye movement trajectory based on these gaze point positions, and sends it to the doctor's management terminal 3. The posture tracking module, through head and hand tracking, generates head posture data and hand posture data, which are then sent as motion posture data to the doctor's management terminal 3.

[0058] In other embodiments of this application, the doctor management terminal 3 is further configured to: perform extrinsic parameter calibration to obtain the relative poses between each pair of the eye coordinate system, the head display coordinate system, and the head rigid body coordinate system; during the patient's free movement, calculate the correction transformation from the SLAM world coordinate system to the motion capture world coordinate system based on the difference between the head display pose and the head rigid body pose output by the VR head display device 2; perform smooth updates on the translation and rotation components of the correction transformation respectively, and use the smoothed correction transformation to convert the head display pose, gaze point position, or eye movement trajectory in the SLAM world coordinate system to obtain the corresponding head display pose, gaze point position, or eye movement trajectory in the motion capture world coordinate system.

[0059] Specifically, before the experiment begins, the relative poses between the eye coordinate system, the head display coordinate system, and the head rigid body coordinate system are determined through extrinsic parameter calibration. When the VR head display device 2 simultaneously outputs the head display pose based on SLAM (Simultaneous Localization and Mapping), the doctor management terminal 3 calculates the correction transformation from the SLAM world coordinate system to the motion capture world coordinate system by comparing the difference between the head display pose based on SLAM and the pose of the head rigid body in the motion capture world coordinate system. The correction transformation is continuously updated during the experiment to compensate for the translational and rotational drift of SLAM. The correction transformation is recalculated when switching scenes, transferring characters, reconnecting devices, or re-identifying the head rigid body. To reduce correction jitter caused by motion capture noise or instantaneous matching errors, the translation and rotation components of the correction transformation are smoothly updated. The smoothed correction transformation is then used to convert the head display pose, gaze point position, or eye trajectory in the SLAM world coordinate system, thereby correcting the eye-tracking data generated based on the SLAM world coordinate system to a unified motion capture world coordinate system. The specific calculation process is as follows: When VR headset 2 simultaneously outputs the headset pose based on the SLAM world coordinate system S At that time, the correction transformation based on the SLAM world coordinate system S to the motion capture world coordinate system M It can be represented as:

[0060] in, Indicates time The transformation matrix from the rigid body coordinate system R of the lower head to the motion capture world coordinate system M. This represents the calibration extrinsic parameter from the head-mounted display coordinate system H to the head rigid body coordinate system R.

[0061] Therefore, any point in the SLAM world coordinate system S can be... Transform to motion capture world coordinate system M:

[0062] in, Represents the coordinates of a point in the SLAM world coordinate system S. Transformed matrix After mapping, the three-dimensional homogeneous coordinates are represented in the motion capture world coordinate system M.

[0063] To reduce correction jitter caused by motion capture noise or instantaneous matching error, The translation and rotation components are updated smoothly respectively. Let the first... The translation components and rotation quaternions obtained from the instantaneous frame estimation are respectively and Then its smoothing result and for:

[0064]

[0065] in, The smoothing coefficient satisfies . Express the smoothing result of the translation component in the (k-1)th frame; Express the smoothing result of the rotation component in the (k-1)th frame; This represents a spherical linear interpolation function used to perform smooth interpolation at a constant speed between two rotated quaternions to ensure the continuity and geometric correctness of the rotational transition.

[0066] Using smoothed Transforming the head display pose, fixation point position, or eye movement trajectory in the SLAM world coordinate system S to obtain the corresponding head display pose, fixation point position, or eye movement trajectory in the motion capture world coordinate system M can further improve the stability and accuracy of eye movement spatial correction under the patient's free movement conditions, thereby improving the spatial accuracy of fixation point overlay display, eye movement trajectory analysis, and task behavior analysis.

[0067] In other embodiments of this application, the doctor management terminal 3 is also used to: perform cross-correlation analysis on the head angular velocity sequence output by the VR head display device 2 and the head rigid body angular velocity sequence output by the motion capture device 1 under a unified reference time axis, and use the time offset corresponding to the maximum value of the cross-correlation function as the residual time delay correction amount, and use the residual time delay correction amount to correct the unified reference time of the eye movement data.

[0068] Specifically, to compensate for the residual fixed time delay between the eye-tracking acquisition link and the motion capture acquisition link, the doctor management terminal 3 can also perform cross-correlation analysis based on the head angular velocity sequence output by the VR headset 2 and the head rigid body angular velocity sequence output by the motion capture device 1. Let the two sets of angular velocity magnitude sequences under a unified reference time axis be... and The residual delay correction amount It can be determined in the following way:

[0069] in, Represents the head angular velocity sequence. This represents the sequence of angular velocities of the rigid body of the head. The time delay parameter represents the time offset between the VR headset angular velocity sequence and the motion-captured head rigid body angular velocity sequence; corr(·) represents the cross-correlation function; argmax is used to find the index of the maximum value in the cross-correlation function. By analyzing the cross-correlation, the value of τ corresponding to the maximum value of the cross-correlation function corr(·) is found, which is the residual time delay estimate to be corrected.

[0070] The unified reference time for eye-tracking data was further revised as follows:

[0071] in, Indicates the unified reference time for eye-tracking data before correction; Indicates the unified reference time for the corrected eye-tracking data; The above methods can further improve the alignment accuracy of eye-tracking data and motion capture data in fast head-movement scenarios.

[0072] In other embodiments of this application, the above-mentioned visual behavior simulation and analysis management system further includes a motion capture data access terminal; a doctor management terminal 3, which is connected to the motion capture device 1 through the motion capture data access terminal; the doctor management terminal 3 is also used to record the local timestamp and frame sequence number of the VR head-mounted display device 2 and the motion capture device 1 or the motion capture data access terminal when generating or receiving data; the doctor management terminal 3 is also used to record the source timestamp, source frame sequence number and access time of the motion capture data source when the motion capture data is accessed through the local SDK callback method.

[0073] Specifically, using the local monotonic clock of the doctor management terminal 3 as a unified reference time axis, the VR headset 2 and motion capture device 1 record the corresponding local timestamps and frame numbers when generating eye-tracking data, motion posture data, and motion capture data, respectively; among them, the local timestamp is preferably a monotonically increasing time that does not jump back with system time modifications. The doctor management terminal 3 is also used to record the source timestamp, source frame number, and access time of the motion capture data source when motion capture data is accessed via local SDK callback. The motion capture data access terminal is the intermediate interface or module connecting the motion capture device 1 and the doctor management terminal 3, mainly responsible for receiving, transmitting, and preprocessing motion capture data, and is a key link in the data flow of the motion capture system.

[0074] In other embodiments of this application, the doctor management terminal 3 is further configured to periodically initiate synchronization requests to the VR headset device 2 and receive synchronization responses, so as to obtain the first sending time of the synchronization request and the first receiving time of the synchronization response; the VR headset device 2 is further configured to record the second receiving time of the synchronization request and the second sending time of the synchronization response; the doctor management terminal 3 is further configured to establish a time mapping relationship between the local time of the VR data source and the unified reference clock based on the first sending time, the first receiving time, the second receiving time and the second sending time. The doctor management terminal 3 is also used to establish a time mapping relationship for motion capture data sources using any of the following methods: When motion capture data is accessed via a local SDK callback, a time mapping relationship between the local time of the motion capture data source and the unified reference clock is established based on the source timestamp, source frame sequence number, and access time of the motion capture data source; When the motion capture data source supports synchronization requests and synchronization responses, a synchronization request is periodically initiated to the motion capture data source and a synchronization response is received to obtain the third sending time of the synchronization request and the third receiving time of the synchronization response, as well as the fourth receiving time of the synchronization request and the fourth sending time of the synchronization response recorded by the motion capture data source; Based on the third sending time, the third receiving time, the fourth receiving time, and the fourth sending time, a time mapping relationship between the local time of the motion capture data source and the unified reference clock is established.

[0075] Specifically, for the VR headset 2, the doctor management terminal 3 periodically sends synchronization request packets to the VR headset 2 and receives synchronization response packets. The doctor management terminal 3 records the first sending time of the synchronization request packet and the first receiving time of the synchronization response packet, while the VR headset 2 records the second receiving time of the synchronization request packet and the second sending time of the synchronization response packet. The doctor management terminal 3 calculates the clock offset and round-trip delay for the corresponding synchronization period based on the four times (first sending time, first receiving time, second receiving time, and second sending time), and removes abnormal samples with round-trip delays exceeding a preset threshold within multiple synchronization periods. Based on multiple synchronization periods, a time mapping relationship is established between the local time of the VR data source and the unified reference clock of the doctor management terminal 3. Preferably, the time mapping relationship is a linear mapping relationship or a piecewise linear mapping relationship, used to simultaneously compensate for clock offset and clock drift.

[0076] For motion capture data sources, when motion capture data is accessed via local SDK callback, a time mapping relationship can be established based on the source timestamp, source frame sequence number, and access time of the motion capture data source. When the motion capture data source supports synchronization requests and synchronization responses, the doctor management terminal 3 periodically sends synchronization request packets to the motion capture data source and receives synchronization response packets, recording the third sending time of the synchronization request packet and the third receiving time of the synchronization response packet. The motion capture data source records the fourth receiving time of the synchronization request packet and the fourth sending time of the synchronization response packet. The doctor management terminal 3 calculates the clock offset and round-trip delay of the corresponding synchronization period based on the four times (third sending time, third receiving time, fourth receiving time, and fourth sending time), and removes abnormal samples with round-trip delays exceeding a preset threshold within multiple synchronization periods. A time mapping relationship is established between the local time of the motion capture data source and the unified reference clock of the doctor management terminal 3 based on multiple synchronization periods.

[0077] Assume the local monotonic clock of the doctor management terminal 3 is the unified reference time axis. Let any data source The local monotonic clock of (such as VR headset device 2 or motion capture data source) is For the k-th synchronization cycle, the doctor management terminal 3 records the sending time of the synchronization request packet. and the timing of receiving the synchronization response packet Data source Record the time of receiving the synchronization request packet and the timing of sending the synchronization response packet .

[0078] Calculate the round-trip delay of the synchronization period based on the four times mentioned above. for:

[0079] And construct a set of approximately simultaneous time pairs under this synchronization period:

[0080] in, This indicates the median synchronization time under the 3 reference time axis in the doctor management terminal. Indicates in the data source The median time of synchronization under the local timeline.

[0081] For any data source Establish its local time To the unified reference time The time mapping relationship is as follows:

[0082] in, Indicates the clock drift compensation coefficient. This represents the clock offset compensation amount. When the test duration is long or the clock rate changes due to temperature drift, the time mapping relationship can also be expressed as a piecewise linear mapping.

[0083] Doctor Management Terminal 3 obtains time pairs based on multiple synchronization cycles Clock drift compensation coefficient and clock offset compensation amount An estimation is performed. Preferably, a weighted least squares method is used to solve the problem:

[0084]

[0085] in, To exclude valid synchronization samples whose round-trip latency exceeds a preset threshold, The weight of the k-th synchronization period, To prevent extremely small positive numbers with a denominator of zero, synchronous samples with shorter round-trip delays are given higher weights to reduce the impact of network jitter on the time mapping relationship. This represents an estimated value of the clock drift compensation coefficient, reflecting the drift ratio between the data source clock frequency and the reference clock. This represents an estimated value for clock offset compensation, reflecting the fixed initial offset between two clocks.

[0086] For the Each original sample, with its local timestamp as... Then the timestamp of the sample after conversion to a unified reference timeline for:

[0087] The weighted least squares time mapping described above can more accurately reflect the true mapping relationship between the local clock of the data source and the unified reference clock, thereby ensuring high-precision alignment of multi-source data such as VR headset 2 and motion capture device 1 on the unified reference time axis.

[0088] In other embodiments of this application, when the doctor management terminal 3 performs time compensation and unified time axis alignment of motion capture data, eye movement data, and action posture data with a unified reference clock, it is specifically used for: Based on the time mapping relationship between the local time of the VR data source and the motion capture data source and the unified reference clock, the motion capture data, eye tracking data, and action pose data are converted to the unified reference time axis. Out-of-order reordering and packet loss detection are performed according to the frame sequence number, and a cache queue is established for each type of data. According to the target rendering time, adjacent samples are selected from each cache queue, and interpolation is performed on the adjacent samples. Specifically, linear interpolation is performed on the position data in the adjacent samples, spherical linear interpolation is performed on the rotation data in the adjacent samples, and normalized interpolation is performed on the gaze direction vector in the adjacent samples, thereby obtaining synchronized sample data under the unified reference time axis.

[0089] Specifically, the doctor management terminal 3, based on the time mapping relationship between the local time of the VR data source and the unified reference clock, and the time mapping relationship between the local time of the motion capture data source and the unified reference clock, converts eye-tracking data, motion pose data, and motion capture data to the unified reference time axis according to their respective time mapping relationships. It then performs out-of-order rearrangement and packet loss detection based on frame sequence numbers, and establishes a cache queue for each type of data. Based on the target rendering time, it selects adjacent samples from each cache queue and interpolates these adjacent samples. When the time interval between adjacent samples is greater than a preset gap threshold, the interval is marked as invalid; when the time interval between adjacent samples is less than or equal to the preset gap threshold, interpolation compensation is performed on the interval.

[0090] At the time of target rendering Below, if the corresponding data sources have adjacent samples on a unified reference timeline as follows: and Then the interpolation coefficients for:

[0091] in, This represents the timestamp of the i-th sample on the unified reference time axis after time mapping; This represents the location data corresponding to the i-th sample. This represents the timestamp of the (i+1)th sample on the unified reference time axis; This represents the location data corresponding to the (i+1)th sample. For positional data in adjacent samples, linear interpolation is used:

[0092] in, Indicates the time of target rendering Below, the estimated position data obtained through linear interpolation; For rotated data in adjacent samples, spherical linear interpolation is used:

[0093] in, Indicates the time of target rendering Below is the result of rotational quaternion interpolation obtained through spherical linear interpolation (SLERP); Represents the rotation quaternion of the i-th sample on the unified reference time axis; Represents the rotation quaternion of the (i+1)th sample on the unified reference time axis; For gaze direction vectors in adjacent samples, normalized vector interpolation is used:

[0094] in, Indicates the time of target rendering Below is the result of unit vector interpolation of the gaze direction obtained through normalized vector interpolation; This represents the gaze direction vector corresponding to the i-th sample on the unified reference time axis; This represents the gaze direction vector corresponding to the (i+1)th sample on the unified reference time axis; The data obtained after the above interpolation process is used as the synchronous sample data under the unified reference time axis.

[0095] By using the above method, heterogeneous data collected by different devices can be uniformly mapped to the same reference time axis even in the presence of network latency, clock jitter, clock drift and packet loss, and synchronous sample data can be obtained for third-person rendering and subsequent analysis.

[0096] In other embodiments of this application, the aforementioned visual behavior simulation and analysis management system further includes an eye-tracking behavior research terminal, which is used to issue training task instructions to the VR headset 2; the VR headset 2 is also used to perform local visualization rendering according to the training task instructions, display the virtual environment to the patient from a first-person perspective, and collect eye-tracking data; the eye-tracking behavior research terminal is also used to generate eye-tracking indicators based on the eye-tracking data, extract eye-tracking trajectories and fixation point positions, and display fixation behavior characteristics.

[0097] The eye-tracking behavior research terminal can be a high-performance PC terminal, capable of sending training task instructions to the VR headset 2 via a wireless transmission protocol. These instructions may include virtual scene switching information and initial virtual scene layout information. Upon receiving the training task instructions, the VR headset 2 performs local 3D visualization rendering, presenting the virtual environment to the patient from a first-person perspective. Based on a pre-set eye-tracking function, it collects eye-tracking data relevant to the patient and uploads this data to the eye-tracking behavior research terminal. This eye-tracking data may include the patient's eye movement trajectory, fixation point position, etc., but this application is not limited to these. The eye-tracking behavior research terminal generates and visualizes eye-tracking indicators based on the eye-tracking data, extracts and visualizes eye movement trajectories and fixation point positions, and analyzes and visualizes fixation behavior characteristics based on these trajectories and fixation point positions. Taking a supermarket shopping scenario as an example, the eye-tracking behavior research terminal performs local real-time visualization rendering to display the virtual environment of a supermarket shelf scene, as shown in Figure 6(a). After receiving eye-tracking data uploaded by VR headset 2, the eye-tracking behavior research terminal can calculate eye-tracking indicators based on the data, including one or more of the following: dwell time, saccade time, average fixation time, average saccade time, average saccade velocity, maximum saccade velocity, average saccade amplitude, and regression count, and visualize them on a layer of the virtual environment. Furthermore, the eye-tracking behavior research terminal can extract eye-tracking trajectories and fixation point positions from the eye-tracking data, visualize them on a layer of the virtual environment, and analyze fixation behavior characteristics based on the eye-tracking trajectories and fixation point positions, such as visual fixation distribution results and eye-tracking heatmaps, as shown in Figure 6(b).

[0098] Through the eye-tracking behavior research platform, the patterns of patients' attention allocation in virtual scenes can be presented intuitively, providing researchers with quantitative basis for eye-tracking behavior analysis. It can also be used to extract contextualized objective indicators that reflect visual search, gaze allocation, and task execution strategies.

[0099] In other embodiments of this application, the motion capture data also includes a three-dimensional coordinate sequence of key markers in the pelvic region and key markers in the foot for gait quantification; the doctor management terminal 3 is further used to: calculate the pelvic center trajectory based on the key markers in the pelvic region, and select a continuous walking segment that meets preset spatial constraints as the steady-state straight-line analysis interval; perform abnormal data cleaning and smoothing filtering on the three-dimensional coordinate sequence of the key markers in the foot; identify heel-to-ground events and toe-to-ground events based on the vertical height sequences of the left and right heels and toes; calculate the stride length based on the heel-to-ground events, and remove abnormal stride lengths to obtain the cleaned effective stride length; and generate a gait parameter quantification report including average stride length, stride length standard deviation, stride length coefficient of variation, and average gait speed based on the cleaned effective stride length and the walking distance and duration of the steady-state straight-line analysis interval.

[0100] Specifically, the motion capture device 1 can output a three-dimensional coordinate sequence of key pelvic and foot markers collected during the patient's walking process for gait quantification. The three-dimensional coordinate sequence can be a C3D file or other equivalent temporal coordinate data format. The doctor management terminal 3 reads the above three-dimensional coordinate sequence and extracts key markers related to gait analysis, including key pelvic markers WaistLFront (left hip front), WaistRFront (right hip front), WaistLBack (left hip back), and WaistRBack (right hip back), as well as key foot markers LHeel (left heel), RHeel (right heel), LToe (left forefoot front), and RToe (right forefoot front).

[0101] The doctor management terminal 3 calculates the pelvic center trajectory based on key markers in the pelvic region. Preferably, the pelvic center trajectory can be obtained by averaging the three-dimensional coordinates of the four markers: WaistLFront, WaistRFront, WaistLBack, and WaistRBack. When key markers in the pelvic region are missing, it can also be approximated by the midpoints of the left and right heel markers. The doctor management terminal 3 further spatially segments the entire walking data based on the absolute coordinates of the pelvic center trajectory within preset X-axis and Z-axis spatial ranges, and selects continuous walking segments that meet preset spatial constraints. Preferably, the longest continuous segment that meets the preset spatial constraints is used as the steady-state straight-line analysis interval to reduce the impact of starting, deceleration, turning, or leaving the acquisition area on gait parameter statistics. The preset spatial constraints can be set based on actual needs.

[0102] After obtaining the steady-state straight-line analysis interval, the doctor management terminal 3 performs abnormal data cleaning and smoothing filtering on the three-dimensional coordinate sequences of the left heel, right heel, left toe, and right toe. Specifically, frames with a vertical height lower than a preset threshold can be identified as frames with missing marker points or frames with clipping abnormalities. These frames are then set to null values ​​and repaired using interpolation, preferably cubic spline interpolation. Subsequently, the repaired three-dimensional coordinate sequences are smoothed using low-pass filtering, preferably zero-phase low-pass filtering, to avoid phase shifts in the time positions of gait events such as heel contact and toe lift-off.

[0103] The doctor management terminal 3 further identifies gait events based on the time sequence of vertical height of foot markers. It can detect local minima of the vertical height sequence of the left and right heels to identify the heel strike event (HS); detect local minima of the vertical height sequence of the left and right toes to identify the toe off event (TO); and filter candidate events by combining the minimum peak distance and minimum peak significance threshold to filter out spurious events caused by local shaking. Figure 7 Figure (a) shows the changes in vertical height of the left and right heels over time and the identified heel contact events. Figure 7 (c) shows the change in vertical height of the left and right toes over time and the identified toe-off events.

[0104] After recognizing a heel contact event, the doctor management terminal 3 calculates the step length for each step based on the heel coordinate difference along the direction of travel. Preferably, when the Z-axis is used as the direction of travel in the motion capture world coordinate system, the i-th left-side step length is... It can be represented as:

[0105] j-th right-side step size It can be represented as:

[0106] in, This indicates the moment when the left heel touches the ground for the i-th time. This indicates the time when the right heel touches the ground for the jth time. This indicates the coordinate value of the left heel marker point in the Z-axis direction in the motion capture world coordinate system, used to detect the moment when the left heel touches the bottom; This indicates the coordinate value of the right heel marker point in the Z-axis direction of the motion capture world coordinate system, used to detect the moment when the right heel touches the ground.

[0107] After merging the left and right step lengths, the median step length is used as the steady-state benchmark. Abnormal step lengths that deviate from the preset proportional range of the median step length are removed to obtain the cleaned effective step length, thereby reducing the impact of boundary deceleration steps, fragmented steps, and misjudgment events on the statistical results. Figure 7Figure (b) shows the effective step size distribution after anomaly removal.

[0108] The doctor management terminal 3 generates a gait parameter quantification report based on the effective stride length after cleaning and the walking distance and duration of the steady-state straight-line analysis interval. Gait parameters include at least the average stride length, stride length standard deviation, stride length coefficient of variation, and average gait speed. The stride length coefficient of variation (CV) can be expressed as:

[0109] in, Indicates the effective step size standard deviation. This represents the average effective step length; average step speed. It can be represented as:

[0110] in, This represents the total distance traveled in the steady-state straight-line analysis interval. This indicates the total duration. Figure 7 (d) shows the core indicators for gait quantification.

[0111] Therefore, in addition to outputting eye-tracking related indicators, this system can also extract behavioral indicators reflecting gait stability, motor coordination, and movement compensation characteristics based on the whole-body motion data collected by the motion capture device, and perform joint analysis with eye-tracking indicators and task completion data.

[0112] In other embodiments of this application, the doctor management terminal 3 and the VR headset device 2 constitute a dual-end synchronization subsystem. Based on a network synchronization framework, they achieve virtual scene object state synchronization and bidirectional transmission of control and data flows. The motion capture device 1, as an independent motion capture terminal, together with the dual-end synchronization subsystem, forms a multi-end collaborative deployment structure. The following descriptions of the Mirror framework, Unity project, GameLogic module, Nokov SDK, and UDP / TCP / KCP protocol are merely exemplary project implementations, used to illustrate how this application can be implemented, and do not constitute a limitation on the scope of protection of this application.

[0113] Specifically, the aforementioned visual behavior simulation and analysis management system adopts a dual-terminal deployment mode of "management terminal (PC) + virtual reality client (VR)". Its core lies in the sharing of main program logic across both terminals and the division of labor through conditional compilation. The main program is generated from the same code project via conditional compilation, ensuring that the VR and PC versions run the same main logic, thus guaranteeing state consistency. The doctor management terminal 3 acts as the test control center, responsible for task generation, network synchronization management, and test control through the network communication layer; while the VR headset 2 focuses on immersive experience and data acquisition, undertaking the local processing and reporting of key behavioral data such as patient interaction, interface display, eye-tracking data, and posture data. At the communication architecture level, the doctor management terminal 3 and the VR headset 2 achieve direct mapping of Unity objects based on the Mirror framework, supporting bidirectional transmission of control and data flows, bidirectional verification at the technical layer, and achieving efficient collaboration between the doctor management terminal 3 and the VR headset 2 through a combination of layered control, hybrid protocols, and end-to-end mapping technologies. While ensuring low latency, the system's security and reliability are ensured through a main control verification mechanism. The working principle of the main control verification mechanism is as follows: VR headset device 2 is only responsible for reporting operations. The doctor management terminal 3 confirms the status change only after verifying the legality. VR headset device 2 must wait for authorization from the doctor management terminal 3 before updating the UI.

[0114] In a specific embodiment, the above-mentioned visual behavior simulation and analysis management system is divided into VR head-mounted display device 2, doctor management terminal 3, eye-tracking module (built-in hardware module of VR head-mounted display device 2) and motion capture device 1 according to the deployment platform; wherein, doctor management terminal 3 and VR head-mounted display device 2 constitute a dual-end synchronous subsystem, and motion capture device 1, as an independent motion capture terminal, together with the dual-end synchronous subsystem, constitutes a multi-end collaborative deployment structure, as shown in Table 1 below.

[0115] Table 1 Deployment Platform and Modules

[0116] The aforementioned visual behavior simulation and analysis management system adopts a multi-terminal collaborative deployment structure of "motion capture terminal + doctor management terminal + VR client". Motion capture device 1 outputs motion capture data simultaneously to doctor management terminal 3 and VR headset device 2 via Nokov SDK, UDP, TCP, SDK local callbacks, or combinations thereof. The motion capture data sent to doctor management terminal 3 directly drives the display of the patient's corresponding virtual character on doctor management terminal 3 and serves as the input source for third-person rendering, time compensation, spatial correction, and data recording. The motion capture data sent to VR headset device 2 directly drives the display of the patient's corresponding virtual character on VR headset device 2 to reduce the latency of patient-side action display. The interaction between the patient and virtual objects in the virtual environment is handled by head and hand tracking, gesture recognition, collision detection in the virtual scene, grasping logic, and the physics engine on VR headset device 2. Doctor management terminal 3, as the authoritative control terminal, is responsible for test task generation, global state arbitration, experimental process control, time compensation, spatial correction, and data persistence; VR headset device 2 is responsible for immersive interaction, first-person rendering, head and hand tracking, eye tracking, and behavioral data feedback. The doctor management terminal 3 and the VR headset device 2 achieve bidirectional synchronization of control commands and status data based on a network synchronization framework to maintain consistency in task progress, scene object position, object status, and virtual character performance between the two ends.

[0117] In some embodiments, the main program of the system can generate PC and VR versions based on the same project. Both ends share the main logic, but the doctor management end 3 undertakes the authority control and status arbitration functions, and the VR head-mounted display device 2 undertakes the immersive interaction and local rendering functions.

[0118] Furthermore, the doctor management terminal 3 and the VR headset device 2 can dynamically select UDP / TCP / KCP protocols based on data characteristics to achieve a balance between latency and reliability. For example, low-latency UDP protocol can be used to transmit eye-tracking data, reliable TCP protocol can be used to transmit task instructions, and KCP protocol can be used to transmit status synchronization data. This application is not limited to these.

[0119] In some embodiments, the VR headset 2 is further used to convert between the headset coordinate system and the virtual scene world coordinate system; the motion capture device 1 is further used to convert between the motion capture world coordinate system and the virtual scene world coordinate system; and the doctor management terminal 3 is further used to convert between the virtual scene world coordinate system and the screen coordinate system. In one example implementation, the virtual scene world coordinate system can be implemented using the Unity world coordinate system.

[0120] Specifically, VR headset 2 can convert coordinate data collected in the headset coordinate system into coordinate data in the virtual scene world coordinate system and assign it to the scene object; motion capture device 1 can convert coordinate data collected in the motion capture world coordinate system into coordinate data in the virtual scene world coordinate system and assign it to the scene object; the doctor management terminal 3 can convert the coordinate data in the virtual scene world coordinate system of the scene object into screen coordinates in the screen coordinate system for subsequent rendering and display. In one example implementation, the scene object can be a GameObject in Unity.

[0121] In an exemplary simulation verification, this application establishes a simulation model covering the entire link of "device clock - data sampling - network transmission - management terminal reception" for a multi-source data acquisition and synchronization system composed of heterogeneous devices, in order to quantitatively evaluate the technical effect of the above technical solution under near-real deployment conditions.

[0122] The simulation model reproduces the imperfect characteristics of data flow from heterogeneous devices from the following three aspects: At the device clock level, an independent clock model was established for each device, including initial offset (motion capture +12ms, VR headset -8.5ms), linear drift (motion capture 30ppm, VR headset 55ppm), temperature-induced periodic drift fluctuations (amplitude 3~8ppm, period about 120 seconds), and sample-by-sample random jitter (10~15 microseconds) to simulate clock inconsistencies caused by differences in crystal oscillator characteristics of different devices.

[0123] At the network transmission level, based on the hybrid protocol architecture actually adopted by this system, differentiated channel models were established for the three data streams: motion capture data is transmitted via a wired LAN using the UDP protocol (referred to as UDP wired in the figure), with a basic round-trip time of about 1ms, a jitter standard deviation of about 0.3ms, and a basic packet loss rate of about 0.5%; eye-tracking data is transmitted via WiFi 6 using the UDP protocol (referred to as UDP wireless in the figure), with a basic round-trip time of about 3ms, a jitter standard deviation of about 1.5ms, a basic packet loss rate of about 2%, and a probability of about 1% out-of-order delivery; attitude synchronization data is transmitted via WiFi 6 using the KCP protocol, with a basic round-trip time of about 5ms, a jitter standard deviation of about 2.0ms, and a basic packet loss rate of about 2%, and achieves effective zero packet loss in this simulation through a fast retransmission mechanism (retransmission timeout of 15ms). In addition, the packet loss model adopts the Gilbert-Elliott burst packet loss model, in which the probability of the next packet being lost increases significantly when the current packet is lost (15% for UDP wired, 35% for UDP wireless, and 30% for KCP) to restore the temporal correlation of burst packet loss in real networks.

[0124] At the data sampling level, motion capture device 1 samples the whole-body skeletal posture at 120Hz (measurement noise approximately 0.5mm RMS), while VR headset 2 samples eye-tracking gaze direction (measurement noise approximately 0.5 degrees RMS) and head / hand posture (measurement noise approximately 0.2 degrees RMS) at 72Hz. When these three data streams reach the doctor management terminal 3 via their respective network channels, they exhibit heterogeneous characteristics: different sampling frequencies (120Hz and 72Hz), different arrival delays (0.75ms, 2.76ms, and 4.54ms), different packet loss characteristics (0.59%, 2.86%, and 0%), and different out-of-order conditions. The doctor management terminal 3 performs third-person visualization at a rendering frame rate of 60Hz. A patient is simulated walking in a virtual supermarket scene for 300 seconds, including natural head and eye saccades. Based on the above simulation conditions, the four core technical effects of the proposed solution are quantitatively evaluated.

[0125] Figures 8(a) to 8(d) show the simulation results of eye movement spatial correction based on motion capture world coordinates. Figure 8(a) illustrates the trend of fixation point spatial error over time; the red curve represents the error (3-second sliding mean) relying solely on VR headset SLAM positioning, while the green curve represents the error after correction using motion capture world coordinates. Figure 8(b) shows the distribution of the eye movement trajectory (composed of 24 discrete fixation points) on the two-dimensional gaze plane, using an elliptical scan performed by the patient at time 130 as an example. The gray dots and black dashed ellipses represent the actual fixation point positions and scan paths, while the red crosses indicate fixation points located solely by SLAM (due to SLAM drift, the overall shift is approximately 1.7 degrees). In addition, the underlying measurement error of eye tracking is about 1 degree. The green rhombus represents the gaze point after motion capture world coordinate correction (eliminating SLAM drift, but retaining the inherent measurement error of eye tracking). The arrow indicates the saccade direction. This figure intuitively shows that spatial correction can eliminate the systematic offset caused by SLAM cumulative drift, but the measurement noise of the eye tracking module itself (about 1 degree RMS) still exists as an uneliminable underlying error. Figure 8(c) shows the error probability distribution of the two schemes in the entire 300-second test. Figure 8(d) shows the cumulative curve of VR headset SLAM pose drift.

[0126] Simulation results show that when relying solely on SLAM localization, the spatial error of the gaze point accumulates continuously over time, reaching an average error of 37.08 mm within 300 seconds, with a P95 error (95th percentile error) of 71.05 mm. The average error increases to 67.87 mm in the last 60 seconds, resulting in a final SLAM cumulative drift of 83.3 mm. By employing the eye-tracking spatial correction method based on motion-captured world coordinates proposed in this application, the average error is reduced to 0.86 mm, the P95 error to 1.36 mm, and the average error in the last 60 seconds to only 0.86 mm. Furthermore, the error does not accumulate over time, resulting in an approximately 43.3-fold improvement in spatial accuracy.

[0127] Figures 9(a) to 9(f) show the simulation results comparing the time synchronization accuracy of multi-source data from heterogeneous devices. Figures 9(a), 9(b), and 9(c) show the time alignment error of the three data streams—motion capture (120Hz) UDP wired, eye tracking (72Hz) UDP wireless, and attitude (72Hz) KCP wireless—as time, respectively. Red represents the simple offset method, and green represents the weighted least squares method of this scheme. Figures 9(d), 9(e), and 9(f) show the corresponding absolute error distributions.

[0128] Simulation results show that the simple offset method cannot compensate for clock drift, and the error increases linearly with time. Specifically, the average error for eye-tracking data from VR headsets with high drift rates reaches 7.535 ms, and for P95 it reaches 16.271 ms. In contrast, the weighted least squares method, through sliding window piecewise fitting and RTT (Round-Trip Time) weighting, effectively compensates for clock drift and temperature drift, reducing the eye-tracking data alignment error to 2.423 ms and the motion capture data to 2.54 ms. All three heterogeneous data streams achieved millisecond-level time alignment accuracy, and the error did not accumulate over time.

[0129] Figures 10(a) to 10(d) show the simulation results of residual delay compensation based on cross-correlation analysis. Figure 10(a) shows a magnified comparison of the motion capture and VR headset angular velocity sequences before compensation, revealing the signal shift caused by the residual delay. Figure 10(b) shows the cross-correlation function; the red dashed line represents the detected delay value, and the green dashed line represents the actual implanted residual delay value. Figure 10(c) shows the alignment effect of the two signals after compensation. Figure 10(d) quantifies the alignment accuracy before and after compensation using mean square error. Under the condition that a residual fixed delay of 7.2ms still exists after time synchronization, the cross-correlation method detects a delay of 5.56ms (detection error 1.64ms), with a peak cross-correlation value of 0.982. The mean square error of signal alignment after compensation decreases from 7.8304 to 4.0548, a reduction of 48.2%.

[0130] Figures 11(a) to 11(f) show the simulation results of smoothness in interpolation rendering of multi-frequency heterogeneous data (including the impact of packet loss). Among them, Figure 11(a) shows a local comparison between the most recent frame sampling (red step line) and linear interpolation (green curve) when the motion capture position data is downsampled from 120Hz to a rendering frame rate of 60Hz; Figure 11(b) shows the actual arrival interval of motion capture data packets after UDP transmission, and a sudden increase in interval due to packet loss can be observed (7 packets were lost in this 10-second segment); Figure 11(c) compares the inter-frame smoothness (inter-frame jitter standard deviation) of the two schemes on position and gaze data; Figure 11(d) shows the interpolation effect from 72Hz to 60Hz for eye movement gaze direction; Figure 11(e) shows the arrival interval of eye movement UDP data packets (17 packets were lost, and the packet loss rate is significantly higher than that of motion capture); Figure 11(f) compares the root mean square error of the two schemes.

[0131] Simulation results show that when heterogeneous data streams of motion capture at 120Hz and eye tracking at 72Hz are transmitted to a doctor management terminal at a rendering frame rate of 60Hz via different protocols, the effectiveness of the two schemes differs depending on the ratio of the sampling rate to the rendering frame rate. For motion capture position data (120Hz to 60Hz, with the sampling rate being an integer multiple of the rendering frame rate), the RMSE of both the most recent frame sampling and linear interpolation is approximately 1.000mm, and the standard deviation of inter-frame jitter is approximately 1.000mm, with no significant difference in numerical error between the two. Under this condition, the main contribution of interpolation is to eliminate the inter-frame discrete staircase effect, making the rendering trajectory smoother and more continuous, and reducing visual jumps caused by direct frame skipping. For eye tracking fixation data (72Hz to 60Hz, with the sampling rate and rendering frame rate being a non-integer ratio), the improvement effect of interpolation is more significant: the RMSE decreases from 0.3438 degrees to 0.2292 degrees (an improvement of approximately 33.3%), and the standard deviation of inter-frame jitter decreases from 0.8594 degrees to 0.8021 degrees (an improvement of approximately 6.7%). This is because the non-integer sampling ratio between 72Hz and 60Hz causes periodic time offset errors in the most recent frame sampling, while linear interpolation can effectively eliminate such errors and significantly improve the accuracy of gaze direction restoration. Furthermore, this scheme detects and marks data gaps (sampling intervals exceeding 50ms) caused by UDP packet loss, avoiding the generation of fake continuous data during interpolation in gap intervals. The above results demonstrate that the accuracy improvement of the interpolation scheme is particularly significant in scenarios where there is a non-integer ratio between the sampling rate and rendering frame rate of heterogeneous device data streams (e.g., eye tracking 72Hz → rendering 60Hz).

[0132] The simulation verification above shows that the technical solution of this application can achieve high-precision spatiotemporal unification and smooth rendering of multimodal behavioral data under actual conditions such as different sampling frequencies of heterogeneous devices (120Hz and 72Hz), different clock characteristics (drift rate of 30~55ppm), significant differences in transmission protocols (UDP wired low latency with a small amount of packet loss vs. UDP wireless high jitter and high packet loss rate vs. KCP fast retransmission with effective zero packet loss), and network conditions with sudden packet loss, out-of-order delivery, and latency jitter. This is achieved through techniques such as spatial correction based on motion capture world coordinates, weighted least squares time mapping, cross-correlation residual delay compensation, and multi-frequency buffer queue interpolation. This provides a reliable data fusion foundation for the joint analysis of multi-source data in functional visual assessment. The clock offset, drift rate, jitter amplitude, packet loss rate, sampling frequency, rendering frame rate, and other numerical parameters mentioned above are only exemplary simulation conditions and do not constitute a limitation on the scope of protection of this application.

[0133] The above description of the disclosed embodiments enables those skilled in the art to make or use this application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this application. Therefore, this application is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A visual behavior simulation and analysis management system, characterized in that, include: Motion capture equipment, VR headsets, and doctor management terminals, among which: The motion capture device is used to collect the patient's whole-body motion data, generate corresponding motion capture data, and send the motion capture data to the doctor's management terminal and the VR headset. The doctor management terminal is used to issue test task instructions to the VR headset and to perform bidirectional synchronization of control instructions and status data with the VR headset. The VR headset is used to perform local real-time visualization rendering based on the test task instructions, the status data, and the motion capture data. It displays the virtual environment, the patient's corresponding virtual character, and task information to the patient in real time from a first-person perspective. It enables the patient to interact with virtual objects in the virtual environment through head and hand tracking and gesture recognition. At the same time, it displays a virtual character driven by the motion capture data and synchronized with the patient's real body movements. It records eye movement data, movement posture data, and task completion data related to the patient and sends them to the doctor's management terminal. The doctor management terminal is also used to perform time compensation and unified time axis alignment on the motion capture data, eye movement data and action posture data with a unified reference clock, and to perform local real-time visualization rendering based on the motion capture data and action posture data. It displays the body movement state of the virtual character corresponding to the patient driven by the motion capture data in the virtual environment, as well as the process of the patient interacting with virtual objects in the virtual environment, to the doctor in real time from a third perspective, and generates and displays evaluation results based on the task completion data. The doctor management terminal is also used to construct a gaze ray based on the eye position coordinates and eye gaze direction in the eye tracking data. Based on the head rigid body pose corresponding to the motion capture data, and the external parameter transformation relationship between the eye coordinate system, head display coordinate system, and head rigid body coordinate system, the gaze ray is sequentially transformed from the eye coordinate system to the head display coordinate system, head rigid body coordinate system, and motion capture world coordinate system. When the gaze ray needs to be displayed in a virtual scene, the gaze ray is transformed to the virtual scene world coordinate system according to the transformation relationship between the motion capture world coordinate system and the virtual scene world coordinate system to obtain the transformed gaze ray. The intersection point of the transformed gaze ray with the reference plane or virtual scene object in the virtual environment is calculated to obtain the corrected gaze point position, and the corrected gaze point position is marked and displayed.

2. The system according to claim 1, characterized in that, The doctor management terminal, when performing the calculation of the intersection point between the converted gaze ray and the reference plane or virtual scene object in the virtual environment to obtain the corrected gaze point position, is specifically used for: When the reference plane is represented by a plane equation, the intersection point of the transformed gaze ray and the reference plane is obtained according to the plane equation, and the corrected gaze point position is obtained. When the virtual scene object is represented by a collider or a mesh model, ray detection is performed along the transformed gaze ray, and the hit point corresponding to the minimum positive path parameter is taken as the corrected gaze point position.

3. The system according to claim 2, characterized in that, The doctor management terminal, when executing the process of constructing a gaze ray based on the eye position coordinates and gaze direction in the eye-tracking data, and based on the head rigid body pose corresponding to the motion capture data, and the external parameter transformation relationship between the eye coordinate system, head-mounted display coordinate system, and head rigid body coordinate system, sequentially transforms the gaze ray from the eye coordinate system to the head-mounted display coordinate system, head rigid body coordinate system, and motion capture world coordinate system; when it is necessary to display the gaze ray in a virtual scene, it transforms the gaze ray to the virtual scene world coordinate system according to the transformation relationship between the motion capture world coordinate system and the virtual scene world coordinate system to obtain the transformed gaze ray, specifically for: A gaze ray is constructed based on the eye position coordinates and gaze direction in the eye movement data. The specific expression of the gaze ray is as follows: ; in, Indicates the gaze ray, This represents the origin of the gaze ray in the eye coordinate system E. This represents the unit gaze direction in eye coordinate system E. Indicates the path parameters of the gaze ray; Based on the head rigid body pose corresponding to the motion capture data, and the external parameter transformation relationship between the eye coordinate system, head display coordinate system, and head rigid body coordinate system, the gaze ray is sequentially transformed from the eye coordinate system to the head display coordinate system, head rigid body coordinate system, and motion capture world coordinate system to obtain the origin of the gaze ray and the gaze ray direction vector in the motion capture world coordinate system. The specific expression is as follows: ; ; in, Indicates time The origin of the gaze ray in the world coordinate system M of the motion capture system. Indicates time The transformation matrix from the rigid body coordinate system R of the lower head to the motion capture world coordinate system M. This represents the calibration extrinsic parameter from the head-mounted display coordinate system H to the head rigid body coordinate system R. This represents the calibration extrinsic parameter from the eye coordinate system E to the head-mounted display coordinate system H. Indicates time The gaze ray direction vector in the world coordinate system M of the motion capture system. Let represent the rotation sub-matrix of the head rigid body coordinate system R relative to the motion capture world coordinate system M at time t. Let represent the rotation submatrix in the homogeneous transformation matrix from the head-mounted display coordinate system H to the head rigid body coordinate system R. This represents the rotation submatrix in the homogeneous transformation matrix from the eye coordinate system E to the head-mounted display coordinate system H; When the gaze ray needs to be displayed in a virtual scene, the gaze ray is transformed to the virtual scene world coordinate system according to the transformation relationship between the motion capture world coordinate system and the virtual scene world coordinate system, resulting in the transformed gaze ray. The specific expression of the transformed gaze ray is as follows: ; ; in, This indicates the converted gaze ray. This represents the origin of the gaze ray in the virtual scene world coordinate system U at time t. Let represent the gaze ray direction vector in the virtual scene world coordinate system U at time t. This represents the transformation relationship from the motion capture world coordinate system M to the virtual scene world coordinate system U. This represents the rotation submatrix in the homogeneous transformation matrix from the motion capture world coordinate system M to the virtual scene world coordinate system U.

4. The system according to claim 3, characterized in that, When the doctor management terminal executes the step of calculating the intersection point of the transformed gaze ray and the reference plane based on the plane equation to obtain the corrected gaze point position, it is specifically used for: For the reference plane Suppose it satisfies the plane equation: ; in, It is a plane normal vector. For plane constants, This indicates that the plane equation n is satisfied on the reference plane Π. The coordinate vector of any point in three-dimensional space where x+c=0 Indicates the transpose operation; Find the intersection point between the transformed gaze ray and the reference plane to obtain the intersection point parameters. for: ; exist At that time, the corrected fixation point position is: ; in, This indicates the corrected fixation point position at time t.

5. The system according to claim 1, characterized in that, The VR headset includes an eye-tracking module and a posture tracking module; The eye-tracking module is used to track the patient's eye movements to generate the eye-tracking data, which includes one or more of the following: eye gaze direction, eye position coordinates, eye movement trajectory, and timestamp. The posture tracking module is used to track the patient's head posture and hand posture to generate the motion posture data.

6. The system according to claim 5, characterized in that, The doctor management terminal is also used for: Perform extrinsic parameter calibration to obtain the relative poses between each pair of the eye coordinate system, head display coordinate system, and head rigid body coordinate system; During the patient's free movement, the correction transformation from the SLAM world coordinate system to the motion capture world coordinate system is calculated based on the difference between the VR headset pose output by the SLAM-based headset and the rigid body pose of the head. The translation and rotation components of the correction transformation are updated smoothly, and the smoothed correction transformation is used to transform the head display pose, gaze point position or eye movement trajectory in the SLAM world coordinate system to obtain the corresponding head display pose, gaze point position or eye movement trajectory in the motion capture world coordinate system.

7. The system according to claim 1, characterized in that, The doctor management terminal is also used for: Based on the head angular velocity sequence output by the VR headset and the head rigid body angular velocity sequence output by the motion capture device, cross-correlation analysis is performed under a unified reference time axis. The time offset corresponding to the maximum value of the cross-correlation function is used as the residual time delay correction amount, and the unified reference time of the eye movement data is corrected using the residual time delay correction amount.

8. The system according to claim 1, characterized in that, It also includes the motion capture data access terminal; The doctor management terminal is connected to the motion capture device through the motion capture data access terminal; The doctor management terminal is also used to record the local timestamp and frame number of the VR headset, the motion capture device or the motion capture data access terminal when generating or receiving data. The doctor management terminal is also used to record the source timestamp, source frame number, and access time of the motion capture data source when the motion capture data is accessed via a local SDK callback method.

9. The system according to claim 8, characterized in that, The doctor management terminal is also used to periodically initiate synchronization requests to the VR headset and receive synchronization responses, so as to obtain the first sending time of the synchronization request and the first receiving time of the synchronization response. The VR headset is also used to record the second time of receiving the synchronization request and the second time of sending the synchronization response; The doctor management terminal is also used to establish a time mapping relationship between the local time of the VR data source and the unified reference clock based on the first sending time, the first receiving time, the second receiving time and the second sending time; The doctor management terminal is also used to establish a time mapping relationship for motion capture data sources using any of the following methods: When the motion capture data is accessed via a local SDK callback, a time mapping relationship between the local time of the motion capture data source and the unified reference clock is established based on the source timestamp, the source frame number, and the access time of the motion capture data source. When the motion capture data source supports synchronization requests and synchronization responses, a synchronization request is periodically initiated to the motion capture data source and a synchronization response is received to obtain the third sending time of the synchronization request and the third receiving time of the synchronization response, as well as the fourth receiving time of the synchronization request and the fourth sending time of the synchronization response recorded by the motion capture data source; based on the third sending time, the third receiving time, the fourth receiving time, and the fourth sending time, a time mapping relationship between the local time of the motion capture data source and the unified reference clock is established.

10. The system according to claim 9, characterized in that, When the doctor management terminal performs time compensation and unified time axis alignment on the motion capture data, eye movement data, and action posture data using a unified reference clock, it is specifically used for: Based on the time mapping relationship between the local time of the VR data source and the motion capture data source and the unified reference clock, the motion capture data, the eye tracking data and the action posture data are converted to the unified reference time axis, and out-of-order rearrangement and packet loss detection are performed according to the frame sequence number, and a cache queue is established for each type of data. Based on the target rendering time, adjacent samples are selected from each cache queue and interpolated. Specifically, linear interpolation is performed on the position data in the adjacent samples, spherical linear interpolation is performed on the rotation data in the adjacent samples, and normalized interpolation is performed on the gaze direction vector in the adjacent samples, thereby obtaining synchronized sample data under a unified reference time axis.