A VR-based high-altitude fall simulation system
By collecting real-time data on users' interpupillary distance and gaze, a high-altitude virtual scene is dynamically constructed and visual perturbations are applied. Combined with neural network calculation of anxiety index, adaptive adjustment of visual stimuli is achieved, solving the problem of high-altitude simulation that cannot be dynamically adjusted in existing technologies, and improving the effect of personalized immersive behavior assessment and adaptive training.
Patent Information
- Application Number
- CN202510693287.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-27
- Publication Date
- 2026-01-30
- Estimated Expiration
- 2045-05-27
AI Technical Summary
Existing high-altitude simulation technologies lack dynamic adjustment mechanisms based on users' real-time psychological response characteristics, making it impossible to achieve adaptive adjustment of visual stimuli. This results in an inability to effectively address users' emotional states such as fear and alertness, and makes it difficult to conduct personalized immersive behavioral assessments and adaptation training.
By collecting user interpupillary distance and initial gaze data, a high-altitude virtual scene is dynamically constructed using a 3D modeling engine. The gaze point position and gaze duration are captured in real time to generate gaze data. Based on the gaze data, visual perturbation is performed. Anxiety index and focus drift speed are calculated by combining a temporal convolutional neural network. Fuzzy logic control and adaptive gain adjustment methods are used to dynamically adjust the visual perturbation effect and generate an immersive experience report.
It achieves intelligent dynamic adjustment of the intensity and frequency of visual disturbances, constructs an immersive adaptation mechanism with closed-loop regulation capability of psychological state, accurately controls the stimulation level in high-altitude fall simulation, and improves the personalized matching capability of psychological and behavioral intervention and virtual high-altitude adaptation training.
Smart Images

Figure CN120585328B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of virtual reality simulation technology, and in particular to a high-altitude fall simulation system based on VR devices. Background Technology
[0002] With the continuous development of Virtual Reality (VR) technology, its applications in fields such as psychological intervention, behavioral training, immersive experiences, and crisis response are gradually deepening, especially in high-altitude environment simulation, where it demonstrates strong scene reconstruction capabilities. Existing high-altitude simulations primarily rely on 3D modeling engines and image rendering algorithms to construct virtual high-altitude scenes through preset scenes, viewpoint control, and stereoscopic display devices to induce a high level of fear in users, thereby facilitating related behavioral assessments or psychological adaptation training. Building upon this, some systems have introduced motion capture or simple physiological signal monitoring methods to passively record individual reactions. However, based on current publicly available technical literature and product solutions, the linkage mechanisms between immersion construction and interactive response in these systems remain relatively limited, especially in complex situations. They lack the ability to provide real-time feedback and intervention on changes in users' psychological states, making it difficult to support personalized adaptation processes.
[0003] The main shortcoming of existing high-altitude simulation technology lies in the lack of a linkage mechanism that dynamically adjusts based on the user's real-time psychological response characteristics. Specifically, most current virtual reality high-altitude simulation solutions only achieve static scene presentation or simple interactive triggers at the visual level, failing to combine the user's dynamic behavioral characteristics (such as eye movement trajectory and gaze pattern) and psychological reactions (such as anxiety level and attention drift) generated during the simulation for real-time identification and response. This disconnect prevents the simulation from being personalized when faced with the user's emotional states such as fear and alertness, and also makes it difficult to form an effective immersive behavioral assessment and adaptation training process. Therefore, how to construct a virtual reality high-altitude simulation solution that can perceive the user's psychological and behavioral characteristics in real time and dynamically adjust visual stimuli accordingly has become one of the key bottlenecks in current technological development. Summary of the Invention
[0004] In view of the aforementioned existing problems, the present invention is proposed.
[0005] Therefore, the present invention provides a VR-based high-altitude fall simulation system to solve the problem of the inability to achieve adaptive adjustment of visual stimuli based on the user's psychological and behavioral state.
[0006] To solve the above-mentioned technical problems, the present invention provides the following technical solution:
[0007] This invention provides a VR-based high-altitude fall simulation system, comprising: a data acquisition module for acquiring user interpupillary distance and initial gaze data; a gaze capture module for dynamically constructing a high-altitude virtual scene using a 3D modeling engine based on the initial gaze data, and capturing the user's gaze point position and gaze duration in real time to generate gaze data; a visual disturbance module for dynamically adjusting local visual disturbances in corresponding areas of the high-altitude virtual scene based on the gaze data to create a visual disturbance effect; a psychological index calculation module for calculating anxiety index and focus drift speed using a temporal convolutional neural network based on the visual disturbance effect and user eye movement behavior to obtain anxiety index and focus drift speed indicators; a disturbance adjustment module for adaptively adjusting the frequency and amplitude of the visual disturbance effect using a combination of fuzzy logic control and adaptive gain adjustment to obtain disturbance rhythm parameters; and a dynamic simulation module for triggering visual perspective transformation and viewpoint rotation using the disturbance rhythm parameters to simulate the visual effect of a high-altitude fall, and combining gaze data, anxiety index, and focus drift speed indicators to annotate behavioral characteristics in real time and generate an immersive experience report.
[0008] As a preferred embodiment of the VR-based high-altitude fall simulation system of the present invention, the specific steps for collecting the user's interpupillary distance and initial gaze data are as follows:
[0009] The system acquires images of the user's eyes, uses binocular geometric analysis to extract the center point positions of the left and right pupils, and calculates the distance between the left and right pupils to obtain the user's interpupillary distance.
[0010] The initial gaze data is generated by collecting the coordinates of the gaze origin, binocular disparity information, head posture, gaze stability parameters, and timestamps.
[0011] In a preferred embodiment of the VR-based high-altitude fall simulation system of the present invention, the specific steps for generating gaze data are as follows:
[0012] The user's interpupillary distance and initial gaze data are spatially transformed and superimposed, and then fused using a rotation matrix to obtain the gaze direction vector. This vector is then extended in virtual space to perform ray intersection to obtain the spatial coordinate information of the user's facing area.
[0013] Based on the spatial coordinates of the user-oriented area, the 3D modeling engine is driven to call high-altitude scene resources and generate a high-altitude virtual scene.
[0014] The spatial coordinates and duration of the user-facing area in the high-altitude virtual scene are recorded in real time and integrated into gaze data in chronological order.
[0015] As a preferred embodiment of the VR-based high-altitude fall simulation system described in this invention, the following steps are taken: Based on the spatial coordinate information of the user-facing area, a 3D modeling engine is driven to call upon high-altitude scene resources to generate a high-altitude virtual scene.
[0016] Using an octree spatial index and multi-level scene management method, environmental elements and detailed data of the corresponding area in the high-altitude scene resource library are retrieved and filtered to obtain a scene element dataset.
[0017] The scene element dataset is classified and sorted, and the graphics rendering interface is called to load and construct the three-dimensional environment of the high-altitude virtual scene, generating real-time rendering frame data.
[0018] The loading and rendering of high-altitude virtual scene details are dynamically adjusted based on real-time rendering frame data to generate high-altitude virtual scenes.
[0019] As a preferred embodiment of the VR-based high-altitude fall simulation system of the present invention, the following steps are taken: a gaze point heatmap analysis method is used to calculate the spatial density of gaze data to obtain a set of coordinates of gaze hotspot areas, and the visual elements of the corresponding areas in the high-altitude virtual scene are dynamically adjusted to form a visual disturbance effect.
[0020] As a preferred embodiment of the VR-based high-altitude fall simulation system of the present invention, the specific steps for obtaining the anxiety index and focus drift speed index are as follows:
[0021] Collect user eye movement behavior data such as gaze trajectory, fixation point position, pupil diameter changes and fixation duration, and assemble them into a multidimensional time-series data sequence in chronological order;
[0022] A multidimensional temporal data sequence is input into a temporal convolutional neural network to extract multidimensional temporal feature representations that reflect the response to visual perturbation effects.
[0023] The multidimensional temporal feature representation is correlated with the visual perturbation effect. By using time alignment and feature matching to perform synchronous matching on the time axis, the change correlation pattern is extracted, and the predicted values of anxiety index and focus drift speed are obtained.
[0024] The predicted values are normalized using the z-score standardization method and dynamically adjusted by combining multidimensional time-series feature representations to obtain the anxiety index and focus drift speed index.
[0025] As a preferred embodiment of the VR-based high-altitude fall simulation system of the present invention, the following steps are taken: A multi-dimensional temporal data sequence is input into a temporal convolutional neural network to extract multi-dimensional temporal feature representations reflecting the response to visual perturbation effects.
[0026] A multi-dimensional time-series data sequence is input into a multi-layer one-dimensional convolutional neural network. Local dynamic features are extracted through a sliding time window, and a preliminary feature tensor is output.
[0027] By employing multi-scale fusion and contextual modeling methods, combined with residual connections and attention mechanisms, the initial feature tensors are weighted, integrated, and temporally optimized to generate temporal feature sequences of behavioral change trends.
[0028] By using a fully connected layer, the temporal feature sequence is compressed in dimension and integrated to obtain a multidimensional temporal feature representation.
[0029] In a preferred embodiment of the VR-based high-altitude fall simulation system of the present invention, the specific steps for obtaining the disturbance rhythm parameters are as follows:
[0030] The anxiety index and focus drift speed index are fuzzified using a fuzzy logic control method and then input into a preset fuzzy rule for fuzzy aggregation and defuzzification to derive the fuzzy control output value.
[0031] Based on the fuzzy control output value, an adaptive gain adjustment method is used to adjust the frequency and amplitude of the visual disturbance effect to obtain the disturbance rhythm parameter.
[0032] As a preferred embodiment of the VR-based high-altitude fall simulation system of the present invention, the specific steps for simulating the visual effect of a high-altitude fall by triggering visual perspective transformation and viewpoint rotation using perturbation rhythm parameters are as follows:
[0033] Based on the perturbation rhythm parameters, the visual perspective transformation matrix is calculated using the matrix transformation method, and the angular velocity of the viewpoint rotation is obtained using the angular velocity calculation method.
[0034] The projection parameters of the high-altitude virtual scene are updated based on the visual perspective transformation matrix, and the position and orientation of the virtual camera are dynamically adjusted in combination with the angular velocity of the viewpoint rotation to generate high-altitude virtual scene transformation frames and viewpoint rotation trajectory.
[0035] By employing a time synchronization and visual sequence fusion method, the changing frames of the high-altitude virtual scene are dynamically fused with the viewpoint rotation trajectory to create a visual effect simulating a high-altitude fall.
[0036] In a preferred embodiment of the VR-based high-altitude fall simulation system of the present invention, the steps for generating the immersive experience report are as follows:
[0037] Based on fixation data, anxiety index, and focus drift speed index, multi-dimensional statistical analysis and time series analysis are performed to identify abnormal fixation patterns and anxiety changes, and output behavioral feature labels and behavioral change trend data.
[0038] By using the support vector machine algorithm, behavioral change trend data is classified and identified, psychological states and behavioral types are analyzed, and an immersive experience report is generated.
[0039] The beneficial effects of this invention are as follows: By integrating fuzzy logic and adaptive gain control, intelligent dynamic adjustment of the intensity and frequency of visual disturbances is achieved, constructing an immersive adaptation mechanism with closed-loop psychological state regulation capabilities. It senses changes in the user's anxiety state and gaze behavior in real time, dynamically optimizing the intensity of visual intervention while addressing the ambiguity of psychological responses, thereby precisely controlling the stimulation level in high-altitude fall simulations. This effectively avoids stress reactions caused by excessive stimulation or training failures due to insufficient stimulation, significantly improving the personalized matching ability and training effectiveness in psychological and behavioral interventions and virtual high-altitude adaptation training. Attached Figure Description
[0040] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the following description of the embodiments will be briefly introduced. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0041] Figure 1 This is a schematic diagram of the VR-based high-altitude fall simulation system of the present invention.
[0042] Figure 2 This is a flowchart of the acquisition of pupillary distance and line of sight data in this invention.
[0043] Figure 3 This is a flowchart illustrating the generation of the visual effect of falling from a height in this invention.
[0044] Figure 4 This is a flowchart of the dynamic construction of a high-altitude virtual scene in this invention. Detailed Implementation
[0045] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings.
[0046] Many specific details are set forth in the following description in order to provide a full understanding of the invention. However, the invention may also be practiced in other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the spirit of the invention. Therefore, the invention is not limited to the specific embodiments disclosed below.
[0047] Secondly, the term "one embodiment" or "embodiment" as used herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The phrase "in one embodiment" appearing in different places in this specification does not necessarily refer to the same embodiment, nor is it a single or selective embodiment that is mutually exclusive with other embodiments.
[0048] Reference Figures 1-4 This embodiment provides a high-altitude fall simulation system based on VR devices, including the following steps:
[0049] The data acquisition module collects the user's interpupillary distance and initial gaze data.
[0050] Images of the user's eyes are acquired, and the center points of the left and right pupils are extracted using binocular geometric analysis. The distance between the left and right pupils is then calculated to obtain the user's interpupillary distance.
[0051] Specifically, after acquiring images of both eyes, the left and right eye images are converted to grayscale to reduce interference from changes in lighting. In each image, edge detection is used to extract the eye contour, and then the boundaries of the left and right pupils are located using Hough circle transform or ellipse fitting. The coordinates of the center points of the left and right pupils are calculated separately. Using binocular geometric analysis, given the distance between the binocular cameras, and combining the positions of the center points of the left and right pupils in the image coordinate system, triangulation is used to calculate the three-dimensional coordinates of the two pupil center points in space. The user's interpupillary distance is obtained based on the Euclidean distance between these two spatial coordinate points.
[0052] The initial gaze data is generated by collecting the coordinates of the gaze origin, binocular disparity information, head posture, gaze stability parameters, and timestamps.
[0053] Specifically, a binocular camera is used to acquire a sequence of synchronized images of both eyes. Image processing methods are used to extract the coordinates of the center points of the left and right pupils. Combined with binocular geometric analysis, the focal position of the user's gaze in three-dimensional space is calculated as the gaze initiation point coordinates. Based on the known pixel coordinates of the pupil center points in the binocular images, and combined with the baseline length and focal length parameters of the binocular camera, the horizontal pixel difference between the gaze focal points of the left and right eyes on the image plane is calculated to obtain binocular disparity information. A three-dimensional facial feature point set is extracted using facial feature point detection methods. A pose estimation algorithm (such as the PNP method) is used to calculate the rotation and displacement vectors of the head in space as head pose data. A sequence of gaze point coordinates is continuously acquired within a short time window, and the spatial fluctuation range of the gaze point within the time window is calculated as a gaze stability parameter. A high-precision system clock time is synchronously recorded each time gaze data acquisition is completed, serving as the timestamp for the corresponding data. It should also be noted that all acquisition steps are completed within the same time frame to ensure that the gaze initiation point coordinates, binocular disparity information, head pose, gaze stability parameters, and timestamps correspond to a consistent gaze behavior state, thus forming initial gaze data.
[0054] The gaze capture module drives the 3D modeling engine to dynamically construct a high-altitude virtual scene based on the initial gaze data, and captures the user's gaze point position and gaze duration in real time to generate gaze data.
[0055] The user's interpupillary distance and initial gaze data are spatially transformed and superimposed, and then fused using a rotation matrix to obtain the gaze direction vector. This vector is then extended in virtual space to perform ray intersection to obtain the spatial coordinate information of the area the user is facing.
[0056] Specifically, the system collects the pupil center positions of the user's left and right eyes. Based on the geometric relationship between the pupils, it calculates the user's gaze center position, which serves as the initial input for the gaze data. Combining this with the posture angle information output from the user's head posture sensor, the initial gaze data is transformed into a unified coordinate system in 3D space. A rotation matrix is constructed to represent changes in head posture, and this rotation matrix is applied to the initial gaze data to achieve unified adjustment of the spatial coordinates. The rotated and fused initial gaze data is then expanded into a direction vector, and a gaze ray is generated in virtual space starting from the user's current head center position and following the direction of the starting point. Based on a pre-defined 3D scene model, the intersection of the ray and the scene surface is calculated. The coordinates of the intersection point are extracted and identified as the spatial region the user is facing. These intersection coordinates serve as the spatial coordinate information of the user's currently facing region, providing fundamental spatial reference data for subsequent visual feature analysis and dynamic content adjustment.
[0057] It should also be noted that the preset 3D scene model collects the geometric shape, spatial position, and boundary information of various static and dynamic objects in the virtual environment and converts them into a unified 3D coordinate description form. For example, it obtains the size parameters and relative positions of objects such as desktops and panels. It uses the boundary volume hierarchy structure or octree method to construct a spatial index for the 3D mesh data to improve the efficiency of ray intersection. It normalizes and aligns the 3D descriptions of all objects to integrate a complete 3D scene dataset. Finally, it loads the complete 3D scene dataset as a reference space for line-of-sight extension and spatial interaction calculations.
[0058] Based on the spatial coordinates of the user-oriented area, the 3D modeling engine is driven to call high-altitude scene resources and generate a high-altitude virtual scene.
[0059] Furthermore, by employing an octree spatial index and a multi-level scene management method, environmental elements and detailed data of the corresponding region in the high-altitude scene resource library are retrieved and filtered to obtain a scene element dataset.
[0060] Specifically, based on the spatial coordinates of the user-facing area, the target index region containing the spatial coordinates is located in the high-altitude scene resource library. The spatial coordinates are used as the query key, and the octree spatial index is called to perform position matching in three-dimensional space. In the octree spatial index, the voxel node to which the spatial coordinates belong is determined level by level. The process is traversed down from the root node to the leaf node to lock the smallest voxel unit containing the spatial coordinates. The scene element index information in the high-altitude scene resource library associated with the leaf node is read, and the process enters the multi-level scene management method. In the multi-level scene management method, scene elements are classified according to scene resource type (such as architecture, weather, terrain, etc.) and level identifier, and irrelevant scene elements are filtered according to the level of the current user-facing area. Within the level, the resource path is called, and the corresponding environmental elements and detailed data that coincide with or are adjacent to the spatial coordinates in the high-altitude scene resource library are read. The detailed data includes texture information, attached structural models, dynamic simulation parameters, and environmental interaction values used to enhance scene accuracy and realism. A local scene element set is formed; based on indicators such as spatial distance and viewpoint relevance, resource elements in the local scene element set that are highly relevant to the user's gaze direction are further filtered, and non-visible elements not involved in the current frame rendering are removed; finally, the environmental elements and detailed data retained after the above processing are integrated to form a scene element dataset.
[0061] It should also be noted that the specific steps for constructing the high-altitude scene resource library are as follows: representative overhead view scene images are extracted from existing high-resolution remote sensing images and oblique photogrammetry results, and then, through 3D reconstruction and semantic annotation, they are organized into a high-altitude scene resource library for virtual interaction.
[0062] In the multi-level scene management method, resource metadata is extracted based on the spatial attributes and semantic types of scene resources, and categorized into types such as architecture, terrain, and meteorology. Then, according to the resource category and spatial scale, a hierarchical identifier is assigned in a pre-defined hierarchical coding system; for example, architecture resources are assigned L2, and meteorological resources are assigned L3. Through these hierarchical identifiers, all scene resources are managed hierarchically. After the user-oriented region is determined, resources irrelevant to the current hierarchical level are filtered out, thereby achieving hierarchical classification and efficient retrieval of scenes.
[0063] The scene element dataset is classified and sorted, and the graphics rendering interface is called to load and construct the 3D environment of the high-altitude virtual scene, generating real-time rendering frame data.
[0064] Specifically, the scene element dataset is categorized based on element type information, with architectural elements, meteorological elements, terrain elements, texture elements, and dynamic effects elements each assigned to a corresponding resource queue. These are then sorted in ascending order based on their relative distance to the user-facing area's spatial coordinates. For the sorted resource queues, the graphics rendering interface is invoked to load the corresponding 3D model files, texture maps, and animation control parameters sequentially. During resource loading, the graphics rendering interface's management functions are first used to load triangular mesh models for architectural and terrain elements, followed by... The texture mapping function matches the corresponding texture elements with the loaded triangular mesh model and binds the texture maps; the lighting calculation function of the graphics rendering interface simulates the lighting and shadow of the 3D mesh model according to the scene time and the direction of the virtual light source; then, the virtual camera parameters are set according to the spatial coordinate information of the user-facing area, including the viewing position, orientation, near clipping plane and far clipping plane, and perspective projection transformation is performed; the processed elements are rendered into 3D image frames, and dynamic effects elements are dynamically loaded and inserted into the current frame image in a frame-by-frame loop; finally, real-time rendering frame data including buildings, terrain, weather and effects content is output.
[0065] It should also be noted that in the process of sorting in ascending order according to the relative distance between the spatial location and the spatial coordinates of the user-facing area, the relative distance refers to the three-dimensional spatial distance between the position coordinates of each environmental element in the high-altitude scene resource library and the spatial coordinates corresponding to the user-facing area. The three-dimensional spatial distance can be calculated based on the difference in three-dimensional coordinates and is used to reflect the proximity of scene elements relative to the user's viewing area. By sorting the three-dimensional spatial distance in ascending order, the priority of resource elements in a local scene is arranged from near to far, providing a reference for the subsequent scene loading order, rendering processing order, or interaction response order.
[0066] The loading and rendering of high-altitude virtual scene details are dynamically adjusted based on real-time rendering frame data to generate high-altitude virtual scenes.
[0067] Specifically, the system collects image resolution, frame rate, and rendering time from real-time rendered frame data as performance metrics, provided by the graphics rendering interface and rendering timestamps. The collected performance metrics are compared with preset performance thresholds to determine if adjustments are needed. Based on the user's facing area spatial coordinates and the virtual camera's position, the spatial distances of environmental elements are calculated, and an octree spatial index is used to quickly locate the spatial unit where the element is located, evaluating spatial priority. Combining performance metrics and spatial priority, multi-level texture maps and different triangular mesh model files are dynamically switched to adjust texture resolution and mesh complexity. The graphics rendering interface is called to reload texture and mesh data, and lighting effects are recalculated when scene time or lighting changes. The virtual camera's observation parameters are updated in real-time based on the user's head pose and viewpoint parameters. Finally, the adjusted scene elements are rendered to the frame buffer through the graphics rendering interface, outputting the current high-altitude virtual scene.
[0068] It should also be noted that the specific steps for setting the performance threshold are as follows: collect frame rate and rendering time data of the target device under typical rendering tasks; evaluate processor performance parameters (such as clock speed and number of cores) and graphics processing capability parameters (such as GPU computing power and video memory bandwidth); combine the subjective acceptable range of screen smoothness and response latency in the preset performance threshold to determine the user's acceptable performance lower limit; based on the statistical analysis results, set the frame rate threshold to be no less than 30 frames / second as in the example, and the rendering time threshold to be no more than 33 milliseconds as in the example.
[0069] The specific steps for comparing the collected performance metrics with preset performance thresholds are as follows: Compare the frame rate in the collected real-time rendering frame data with the preset frame rate threshold. If the frame rate is lower than the preset frame rate threshold, the performance is deemed not to meet the requirements. At the same time, compare the collected single-frame rendering time with the preset rendering time threshold. If the single-frame rendering time exceeds the preset rendering time threshold, the performance is deemed not to meet the requirements. When the frame rate is not lower than the preset frame rate threshold and the single-frame rendering time does not exceed the preset rendering time threshold, the performance is deemed to meet the requirements.
[0070] The spatial coordinates and duration of the user-facing area in the high-altitude virtual scene are recorded in real time and integrated into gaze data in chronological order.
[0071] Specifically, the system acquires the spatial coordinates of the user's facing area in real time and records the timestamp of each spatial coordinate change; it calculates the duration between adjacent timestamps to determine the length of time the user gazes at the same spatial coordinate; it merges the time periods of continuous gazing at the same spatial coordinate to generate a gaze duration record; and it sorts all gaze data by time sequence to form a complete gaze data sequence. For example, if a user gazes at a spatial coordinate continuously for more than 0.5 seconds, it is considered a valid gaze event. Multiple sampled gaze data within the time period are merged, and the final output is gaze data containing the spatial coordinate and the corresponding gaze duration.
[0072] It should also be noted that the specific steps for obtaining the spatial coordinate information of the user's facing area in real time are as follows: continuously collect the user's head posture data, including position and orientation information, through a head tracking device; calculate the corresponding spatial coordinates based on the head orientation by combining the virtual camera parameters; record the calculated spatial coordinates and corresponding timestamps every fixed time interval (e.g., 10 milliseconds in the example); and transmit the collected data stream to the processing unit in real time to ensure that the spatial coordinate information of the user's currently facing area is obtained continuously and without loss.
[0073] The visual perturbation module dynamically adjusts the local visual perturbation of the corresponding area in the virtual scene based on gaze data to create a visual perturbation effect.
[0074] By employing a gaze point heatmap analysis method, spatial density calculations are performed on gaze data to obtain a set of coordinates for gaze hotspot regions. Visual elements in corresponding regions of the virtual scene are then dynamically adjusted to create a visual perturbation effect.
[0075] Specifically, the spatial coordinates of each gaze point in the gaze data are mapped to a pre-divided spatial grid cell, and the number of gaze points in each grid cell is counted to calculate the spatial density. A preset gaze density threshold is used to filter out grid cells with high gaze density, and the center coordinates are extracted as the gaze hotspot region coordinate set. Based on the gaze hotspot region coordinate set, the visual elements in the corresponding virtual scene are determined. According to the set visual perturbation rules, the visual elements are dynamically adjusted, including changing attributes such as color, brightness, or texture details, thereby creating a visual perturbation effect.
[0076] It should also be noted that the specific steps for setting the fixation density threshold are as follows: count the number of all fixation points in the spatial grid cell and calculate the fixation density of each grid cell; collect the fixation density data of all grid cells and sort them according to the numerical value; set the fixation density threshold based on the statistical analysis results, for example, select the top 10% of the fixation density distribution as the fixation density threshold; filter out the grid cells with fixation density higher than the fixation density threshold and extract the center coordinates as the fixation hotspot area.
[0077] The specific steps for setting visual perturbation rules are as follows: For visual elements corresponding to the gaze hotspot area, obtain the current color, brightness, and texture detail parameters; define the perturbation amplitude, such as the percentage increase or decrease in color brightness and the adjustment level of texture detail; dynamically change the color value, brightness value, or replace it with texture maps of different detail levels according to the perturbation amplitude; apply the adjusted visual attributes to the corresponding visual elements to complete the dynamic update of the visual perturbation effect.
[0078] The psychological index calculation module uses a temporal convolutional neural network to calculate the anxiety index and focus drift speed based on the visual perturbation effect and the user's eye movement behavior, thus obtaining the anxiety index and focus drift speed indicators.
[0079] User eye movement behavior data, including gaze trajectory, fixation point position, pupil diameter changes, and fixation duration, are collected and arranged into a multidimensional time-series data sequence.
[0080] Specifically, a high-frequency sampling device is used to collect user gaze trajectory data at a frequency of 60 times per second (example value), recording the three-dimensional spatial coordinates of each sampling time point. A gaze detection algorithm is used to identify valid gaze events, collecting the gaze point position and gaze duration, retaining only valid gaze events with a gaze duration greater than or equal to 200 milliseconds (example value). An optical pupil tracking device is used to record the pupil diameter at each sampling time point, with a measurement accuracy of 0.01 mm (example value), and the change in pupil diameter between adjacent time points is calculated, extracting the pupil diameter change value. The time-stamped data of gaze trajectory, gaze point position, pupil diameter change, and gaze duration are stored in their respective time series lists. All timestamps are extracted, merged, and sorted in ascending order. For each ascending-sorted timestamp, the corresponding multidimensional time series data is searched from each time series. If missing, nearest neighbor interpolation or previous values are used for filling. The multidimensional time series data items are combined in time-stamp order to form a synchronized multidimensional time series data sequence, with the timestamp error controlled within 1 millisecond (example value).
[0081] Multidimensional temporal data sequences are input into a temporal convolutional neural network to extract multidimensional temporal feature representations that reflect the response to visual perturbation effects.
[0082] Furthermore, multidimensional time-series data sequences are input into a multilayer one-dimensional convolutional neural network, and local dynamic features are extracted through a sliding time window to output a preliminary feature tensor.
[0083] Specifically, the multidimensional temporal data sequence is divided into sliding time windows according to a set length and stride. In the example, the time window length is set to 300 milliseconds and the stride is 100 milliseconds. Multiple overlapping windows can be divided within a unit of time. For the multidimensional temporal data within each divided sliding time window, the first convolutional layer of a multi-layer one-dimensional convolutional neural network is input. In the example, the kernel size is 30, the stride is 1, the number of output channels is 32, and the padding method is "same". One-dimensional convolution is performed to extract local temporal features. The convolution results of each window are then passed to the next convolutional layer to continue extracting higher-dimensional temporal features. In the example, the second convolutional layer has a kernel size of 15 and 64 channels, and the third convolutional layer has a kernel size of 7 and 128 channels. The convolution outputs of all time periods are concatenated in chronological order to obtain a preliminary feature result with a three-dimensional structure. The dimensions correspond to the number of time periods, the data length within each time period, and the number of temporal features in each data period.
[0084] By employing multi-scale fusion and contextual modeling methods, combined with residual connections and attention mechanisms, the preliminary feature tensors are weighted, integrated, and temporally optimized to generate temporal feature sequences that reflect behavioral change trends.
[0085] Specifically, the initial feature tensor is input into three one-dimensional convolutional paths with different kernel sizes of 3, 5, and 7, respectively, and the number of output channels for each convolutional path is set to 64. The output features of the three convolutional paths are concatenated along the channel dimension to form a fused feature tensor. The fused feature tensor is then input into a bidirectional long short-term memory network for contextual temporal modeling, with the hidden state dimension set to 128, outputting a contextual feature sequence. A channel attention mechanism is applied to the contextual feature sequence, and channel statistics are extracted through global average pooling. Two fully connected layers are then connected to calculate the channel weights, and each channel of the contextual feature sequence is multiplied by its corresponding weight coefficient for weighting. The weighted contextual feature sequence is then added element-wise to the initial feature tensor along the same dimension to complete the residual connection operation, resulting in a fused temporal feature tensor. The fused temporal feature tensor is then input into a one-dimensional convolutional layer for channel compression, with a kernel size of 3 and an output channel count of 64, ultimately generating a temporal feature sequence representing behavioral change trends.
[0086] It should also be noted that the training and modeling process of inputting the fused feature tensor into the bidirectional long short-term memory network for contextual temporal modeling is as follows: The fused feature tensor is arranged in chronological order to form an input sequence, which is then input into the forward and backward hidden layers of the bidirectional long short-term memory network. The forward hidden layer processes the sequence information in forward chronological order, and the backward hidden layer processes the sequence information in reverse chronological order. During the training phase, training samples with labeled behavioral or psychological state categories are used. By minimizing the cross-entropy loss function between the predicted category and the true label, combined with the backpropagation algorithm and the temporal backpropagation algorithm, the weight parameters and bias parameters in the bidirectional long short-term memory network are iteratively updated. After training is completed, the trained bidirectional long short-term memory network is used to infer the newly input fused feature tensor, calculate the forward and backward hidden states respectively, and then concatenate or weightedly fuse the bidirectional hidden states to extract a state representation containing complete temporal contextual dependencies for subsequent behavior recognition or psychological state analysis.
[0087] By using a fully connected layer, the temporal feature sequence is compressed in dimension and integrated to obtain a multidimensional temporal feature representation.
[0088] Specifically, each context feature vector in the context feature sequence is taken as input to a time step and sequentially fed into a fully connected layer for linear transformation. An example of the linear transformation weight matrix dimension is 256×64, and an example of the bias vector dimension is 64. Each context feature vector is multiplied by its corresponding weight matrix and the bias is added to obtain the compressed time step feature vector. The compressed feature vectors of all time steps are arranged in chronological order to form a compressed temporal feature sequence. The compressed temporal feature sequence is then pooled along the time dimension using either average pooling or max pooling, with a pooling window of the entire sequence length, to extract the global feature vector. Finally, the global feature vector is concatenated with the compressed temporal feature sequence to form a multidimensional temporal feature representation.
[0089] By performing correlation analysis between multidimensional temporal feature representation and visual perturbation effect, and using time alignment and feature matching to perform synchronous matching on the time axis, change correlation patterns are extracted to obtain predicted values of anxiety index and focus drift speed.
[0090] Specifically, the multidimensional temporal feature representation and the time-stamped visual perturbation effect data are synchronized and aligned according to the timestamps. For parts with mismatched timestamps, linear interpolation or previous value padding is used to fill in the data. At each time point, the rate of change of the fixation point spatial coordinates in the multidimensional temporal feature representation is calculated as the focus drift velocity. At the same time, the weighted average of the pupil diameter change and fixation duration change is calculated as the original anxiety index. A sliding time window method is used to smooth the focus drift velocity and the original anxiety index. The window size is 5 time points in the example, and the step size is 1 time point in the example. According to the aligned time points, the correlation coefficient between the multidimensional temporal feature representation and the change in the intensity of the visual perturbation effect is calculated within each time window, and time periods with significant correlation are selected. The average values of the focus drift velocity and the original anxiety index within the time periods with significant correlation are calculated as the predicted values of focus drift velocity and anxiety index for the corresponding time periods.
[0091] The predicted values are normalized using the z-score standardization method and dynamically adjusted by combining multidimensional time-series feature representations to obtain the anxiety index and focus drift speed index.
[0092] Specifically, the mean and standard deviation of the predicted anxiety index and focus drift velocity sequences are calculated respectively. The z-score standardization method is used to obtain a normalized standardized predicted value sequence. This normalized predicted value sequence is then combined with the multidimensional time-series feature representations at corresponding time points. A weighted average or linear mapping method is used to dynamically adjust the standardized predicted value sequence. The adjustment result is used to correct the original anxiety index and focus drift velocity sequences, resulting in dynamically corrected anxiety index and focus drift velocity sequences. Finally, the dynamically corrected sequence is output in chronological order as the final anxiety index and focus drift velocity indicators.
[0093] The disturbance adjustment module uses a combination of fuzzy logic control and adaptive gain adjustment to adaptively adjust the frequency and amplitude of visual disturbances to obtain disturbance rhythm parameters.
[0094] The anxiety index and focus drift speed index are fuzzified using a fuzzy logic control method and then input into a preset fuzzy rule for fuzzy aggregation and defuzzification to derive the fuzzy control output value.
[0095] Specifically, a fuzzy logic control method is used to fuzzify the anxiety index and focus drift speed index. The fuzzification process includes mapping the anxiety index and focus drift speed to a set of fuzzy linguistic variables based on preset membership functions. For example, the anxiety index is divided into three fuzzy linguistic values: "low," "medium," and "high," and the focus drift speed is divided into three fuzzy linguistic values: "slow," "moderate," and "fast." The fuzzification results are then substituted into a preset set of fuzzy rules as input variables, and the output trends corresponding to different combinations of input variables are described in the form of "if-then." The activated fuzzy rules are then subjected to fuzzy aggregation to integrate the output trends, and finally defuzzified using a weighted average method to obtain the fuzzy control output value used to characterize the state response.
[0096] It should also be explained that the steps for setting the preset membership functions and the preset fuzzy rule set are as follows: Construct membership functions for the anxiety index and the focus drift speed index. For the anxiety index, set the input range to the example [0, 1], and define three membership functions, corresponding to the fuzzy linguistic variables "low", "medium", and "high" respectively. The membership functions can be described using triangular or trapezoidal functions. For example, "low" can correspond to a triangular function with a center value of 0.2 and a left-right expansion range of ±0.1; "medium" corresponds to a triangular function with a center value of 0.5 and a left-right expansion range of ±0.15; and "high" corresponds to a triangular function with a center value of 0.8 and a left-right expansion range of ±0.1. For the focus drift speed index, the input range is set to example [0, 10] (unit is degrees / second), and three membership functions are defined to correspond to the fuzzy linguistic variables "slow", "moderate" and "fast". For example, "slow" corresponds to a center value of 2 and an expansion range of ±1; "moderate" corresponds to a center value of 5 and an expansion range of ±1.5; "fast" corresponds to a center value of 8 and an expansion range of ±1.
[0097] A pre-defined set of fuzzy rules is constructed. Each fuzzy rule represents the mapping relationship between the combination of input variables and the output value using an "if-then" statement. For example: if the anxiety index is "low" and the focus drift speed is "slow", the output is "stable"; if the anxiety index is "medium" and the focus drift speed is "moderate", the output is "alert"; if the anxiety index is "high" and the focus drift speed is "fast", the output is "abnormal"; if the anxiety index is "medium" and the focus drift speed is "fast", the output is "alert", and so on. By enumerating all possible combinations of fuzzy linguistic variables, the fuzzy rule set is established and used for subsequent fuzzy inference and control output calculation.
[0098] Based on the fuzzy control output value, an adaptive gain adjustment method is used to adjust the frequency and amplitude of the visual disturbance effect to obtain the disturbance rhythm parameter.
[0099] Specifically, based on the fuzzy control output value and the preset frequency and amplitude threshold, the adaptive gain coefficient is calculated. The adaptive gain coefficient is adjusted according to the deviation between the fuzzy control output value and the preset frequency and amplitude threshold. A discrete update method with a time step of 1 second is used to iteratively update the frequency and amplitude parameters until both the frequency and amplitude parameters are stable within the preset frequency and amplitude threshold range. The frequency and amplitude of the adjusted visual disturbance effect are then output as the disturbance rhythm parameters.
[0100] It should also be noted that the steps for setting the frequency and amplitude thresholds are as follows: collect a large amount of visual disturbance experimental data, including users' subjective comfort evaluations at different frequencies and amplitudes; perform statistical analysis on the experimental data to calculate the distribution of user comfort scores corresponding to the frequency and amplitude combinations; determine the frequency and amplitude ranges where the comfort scores reach the preset standard (e.g., 80% of users have no discomfort); and finally set the minimum and maximum values within the range as the upper and lower limits of the frequency threshold and amplitude threshold, respectively. For example, the frequency threshold is 0.5 Hz to 5 Hz, and the amplitude threshold is 0 to 1.
[0101] The dynamic simulation module uses perturbation rhythm parameters to trigger visual perspective transformation and viewpoint rotation to simulate the visual effect of falling from a height. Combined with gaze data, anxiety index and focus drift speed indicators, it annotates behavioral characteristics in real time and generates an immersive experience report.
[0102] Based on the perturbation rhythm parameters, the visual perspective transformation matrix is calculated using the matrix transformation method, and the angular velocity of the viewpoint rotation is obtained using the angular velocity calculation method.
[0103] Specifically, based on the perturbation rhythm parameters, a two-dimensional or three-dimensional visual perspective transformation matrix is constructed. The specific steps include: calculating the corresponding translation, rotation, and scaling transformation matrix elements based on the frequency and amplitude information in the perturbation rhythm parameters; for example, using the amplitude parameter to determine the translation offset in the transformation matrix, and using the frequency parameter to determine the change amplitude of the rotation angle; and combining the various basic transformation matrices sequentially through matrix multiplication to obtain the complete visual perspective transformation matrix.
[0104] The angular velocity calculation method involves the following steps: collecting viewpoint data at consecutive moments, calculating the angular difference between adjacent viewpoints (in radians or degrees); calculating the angular velocity based on the time interval to obtain the viewpoint rotation speed (e.g., if the viewpoint difference is 5 degrees and the time difference is 0.1 seconds, the angular velocity is 50 degrees per second); calculating the corresponding angular velocities for viewpoint rotation in multiple directions; and outputting the visual perspective transformation matrix and angular velocity to describe the spatial transformation and rotation dynamics under visual perturbation.
[0105] The projection parameters of the high-altitude virtual scene are updated based on the visual perspective transformation matrix, and the position and orientation of the virtual camera are dynamically adjusted in combination with the angular velocity of the viewpoint rotation to generate high-altitude virtual scene transformation frames and viewpoint rotation trajectory.
[0106] Specifically, the near-plane and far-plane distances of the view frustum are calculated by extracting specific elements from the visual perspective transformation matrix. The horizontal viewing angle and aspect ratio are then inferred from the relevant values in the matrix to obtain the corresponding projection matrix. The current viewing angle rotation angular velocity of the virtual camera is collected and represented as an angular velocity vector around a fixed coordinate axis. Based on the viewing angle rotation angular velocity, the rotation increment of the virtual camera at the current time step is calculated using Euler's integral or quaternion interpolation. This rotation increment is then applied to the orientation matrix of the virtual camera at the previous time step to obtain the updated orientation matrix. The position change of the virtual camera in the three-dimensional coordinate system is calculated synchronously. For example, the displacement vector is obtained by multiplying the rotation axis direction by the movement step size and then superimposed with the current coordinates to obtain the new position. A view matrix is constructed based on the updated virtual camera position and orientation. The view matrix is multiplied by the projection matrix to obtain the high-altitude virtual scene transformation frame. The virtual camera position and orientation in each frame are recorded in chronological order to generate the viewing angle rotation trajectory.
[0107] By employing a time synchronization and visual sequence fusion method, the changing frames of the high-altitude virtual scene are dynamically fused with the viewpoint rotation trajectory to create a visual effect simulating a high-altitude fall.
[0108] Specifically, each frame of the high-altitude virtual scene transformation is sorted by its generation timestamp, and the viewpoint rotation trajectory points formed by the virtual camera position and orientation at the corresponding time point are extracted to ensure a one-to-one correspondence between the high-altitude virtual scene transformation frames and the viewpoint rotation trajectories on the time axis. At each time point, the corresponding viewpoint rotation trajectory points are used as reference viewpoint parameters for the current frame, and the image content of the high-altitude virtual scene transformation frame for the current frame is read. Using a pose-based image synthesis method, the virtual camera pose (including 3D coordinate position and 3D Euler angle direction) in the current viewpoint rotation trajectory points is applied to the current frame, and viewpoint distortion correction and rotation transformation are performed on the image within the frame to generate a simulated image obtained from the pose observation. All frames that have undergone pose fusion processing are combined into an image sequence in chronological order to form a dynamically fused high-altitude virtual scene transformation frame sequence. Based on the amplitude of virtual camera pose changes between consecutive frames, a corresponding motion blur kernel is generated, and dynamic blur processing is further applied to each frame in the image sequence to output the synthesized simulated high-altitude fall visual effect.
[0109] Based on gaze data, anxiety index, and focus drift speed index, multi-dimensional statistical analysis and time series analysis are performed to identify abnormal gaze patterns and anxiety changes, and output behavioral feature labels and behavioral change trend data.
[0110] Specifically, fixation data is arranged according to timestamps, and basic features such as fixation position coordinates, fixation duration, and fixation point change frequency are extracted for each moment. Anxiety index and focus drift velocity index are aligned according to their corresponding times to form a three-dimensional time series dataset consisting of fixation data, anxiety index, and focus drift velocity index. In this three-dimensional time series dataset, sliding window statistical processing is performed to calculate statistical features such as average fixation duration, fixation point jump frequency, mean and standard deviation of anxiety index, and mean and maximum focus drift velocity within each window. Frequency amplitude thresholds are used, such as an anxiety index standard deviation exceeding 0.25 or a fixation jump frequency exceeding 5 times per second, to mark the corresponding window as an abnormal state. The marked results are clustered and merged in chronological order to form continuous behavioral change segments. For each behavioral change segment, behavioral feature labels are assigned based on the distribution of statistical indicators, such as "tension," "avoidance," and "focus shift." Simultaneously, curve fitting is performed on the anxiety index and focus drift velocity over continuous time periods to extract the trend of the first derivative, which is used to characterize the behavioral change trend data. Finally, the behavioral feature labels and behavioral change trend data for the corresponding time period are output.
[0111] By using the support vector machine algorithm, behavioral change trend data is classified and identified, psychological states and behavioral types are analyzed, and an immersive experience report is generated.
[0112] Specifically, the system collects output behavioral change trend data, including behavioral feature labels, corresponding timestamps, and feature dimension values (such as fixation concentration, anxiety index change rate, and average focus drift speed). It constructs input feature vectors for the support vector machine algorithm according to a preset classification objective. These feature vectors include numerical combinations of the three dimensions mentioned above for each behavioral time slice, and are labeled with known psychological state category labels (e.g., tension, relaxation, focus, anxiety) and behavioral type category labels (e.g., frequent gaze skipping, stable gaze, persistent gaze wandering) from the training dataset. The labeled training samples are then input into the support vector machine algorithm for processing. Training involves constructing a hyperplane using linear kernel functions or radial basis functions, and determining the optimal parameter values C and γ through cross-validation (e.g., C=1.2, γ=0.1). The behavioral change trend data to be identified is input into a pre-trained support vector machine algorithm to obtain classification results for corresponding psychological states and behavioral types. These classification results are then merged according to the timeline to generate corresponding psychological state evolution sequences and behavioral type evolution sequences. The evolution sequences are mapped to immersive experience report content, which includes time period divisions, dominant psychological state annotations, typical behavioral type annotations, and fluctuation event identifiers.
[0113] It should also be noted that the preset classification target is determined based on the needs of behavior recognition and the purpose of psychological state analysis. Specifically, it is achieved by collecting a large amount of training sample data labeled with behavior types and psychological states, and using statistical analysis methods to classify behavior categories and psychological state categories. For example, behavior categories are set as "calm", "tense" and "anxious", and psychological state categories are set as "normal", "mildly abnormal" and "severely abnormal". Based on the classification target, the classification label categories and corresponding relationships of the support vector machine algorithm are clarified to form a standardized target set for training and classification.
[0114] In summary, this invention achieves intelligent dynamic adjustment of the intensity and frequency of visual disturbances by integrating fuzzy logic and adaptive gain control, constructing an immersive adaptation mechanism with closed-loop psychological state regulation capabilities. It senses changes in user anxiety and gaze behavior in real time, dynamically optimizing the intensity of visual intervention while addressing the ambiguity of psychological responses, thereby precisely controlling the stimulation level in high-altitude fall simulations. This effectively avoids stress reactions caused by excessive stimulation or training failures due to insufficient stimulation, significantly improving the personalized matching ability and training effectiveness in psychological and behavioral interventions and virtual high-altitude adaptation training.
[0115] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.
Claims
1. A VR device based high fall simulation system, characterized by: Comprising, a data acquisition module for acquiring user interpupillary distance and initial line-of-sight data; a gaze capture module for driving a three-dimensional modeling engine to dynamically construct a high-altitude virtual scene according to the initial line-of-sight data, and for capturing user gaze point positions and gaze durations in real time to generate gaze data; a visual disturbance module for dynamically adjusting local visual disturbances in corresponding regions of the high-altitude virtual scene based on the gaze data to form a visual disturbance effect; a psychological index calculation module for calculating an anxiety index and a focus drift speed using a time series convolutional neural network according to the visual disturbance effect and user eye movement behavior, to obtain anxiety index and focus drift speed indicators, the specific steps being as follows, acquiring user eye movement behavior data including line-of-sight trajectories, gaze point positions, pupil diameter changes, and gaze durations, and forming a multi-dimensional time series data sequence in chronological order; inputting the multi-dimensional time series data sequence into the time series convolutional neural network to extract multi-dimensional time series feature representations reflecting responses to the visual disturbance effect; synchronizing and aligning the multi-dimensional time series feature representations with the visual disturbance effect data with timestamps according to the timestamps, and using linear interpolation to fill in data for parts that do not match the timestamps; at each time point, calculating the gaze point spatial coordinate change rate in the multi-dimensional time series feature representations as the focus drift speed, and calculating the weighted average of the pupil diameter change amplitude and the gaze duration change amplitude as the anxiety index original indicator; using a sliding time window method to smooth the focus drift speed and the anxiety index original indicator; according to the aligned time points, calculating the correlation coefficients of the multi-dimensional time series feature representations and the visual disturbance effect intensity changes within each time window, and selecting time periods with significant correlation; calculating the average values of the focus drift speed and the anxiety index original indicator within the time periods with significant correlation as the focus drift speed prediction value and the anxiety index prediction value output for the corresponding time periods; using the z-score standardization method to normalize the prediction values, and combining the multi-dimensional time series feature representations to dynamically adjust, to obtain the anxiety index and focus drift speed indicators; a disturbance adjustment module for adaptively adjusting the frequency and amplitude of the visual disturbance effect using a method combining fuzzy logic control and adaptive gain adjustment, to obtain disturbance rhythm parameters; a dynamic simulation module for triggering visual perspective transformation and view rotation to simulate high-altitude falling visual effects using the disturbance rhythm parameters, and for real-time labeling of behavior characteristics in combination with the gaze data, anxiety index, and focus drift speed indicators, to form an immersive experience report.
2. The VR device-based high fall simulation system of claim 1, wherein: The specific steps for acquiring user interpupillary distance and initial line-of-sight data are as follows, acquiring user binocular images, extracting left and right eye pupil center positions using binocular geometry analysis, and calculating the distance between the left and right eye pupils to obtain the user interpupillary distance; acquiring gaze starting point coordinates, binocular parallax information, head posture, gaze stability parameters, and timestamps to form initial line-of-sight data.
3. The VR device-based high fall simulation system of claim 1, wherein: The specific steps for generating gaze data are as follows, performing spatial transformation and superposition on the user interpupillary distance and initial line-of-sight data, using a rotation matrix for fusion to obtain a line-of-sight direction vector, and performing ray intersection in a virtual space to obtain spatial coordinate information of the user-facing region; According to the spatial coordinate information of the user-facing area, a three-dimensional modeling engine is driven to call high-altitude scene resources to generate a high-altitude virtual scene. The spatial coordinate information of the user-facing area and the duration in the high-altitude virtual scene are recorded in real time, and are integrated into gaze data in chronological order.
4. The VR device-based high fall simulation system of claim 3, wherein: According to the spatial coordinate information of the user-facing area, a three-dimensional modeling engine is driven to call high-altitude scene resources to generate a high-altitude virtual scene, and the specific steps are as follows, An octree spatial index and a multi-level scene management method are used to retrieve and filter the environmental elements and detail data of the corresponding area in the high-altitude scene resource library to obtain a scene element dataset. The scene element dataset is classified and sorted, a graphics rendering interface is called to load and build a three-dimensional environment of the high-altitude virtual scene, and real-time rendering frame data is generated. Based on the real-time rendering frame data, the loading and rendering of the details of the high-altitude virtual scene are dynamically adjusted to generate the high-altitude virtual scene.
5. The VR device-based high fall simulation system of claim 1, wherein: A gaze point heat map analysis method is used to calculate the spatial density of the gaze data to obtain a gaze hotspot area coordinate set, and the visual elements of the corresponding area in the high-altitude virtual scene are dynamically adjusted to form a visual disturbance effect.
6. The VR device-based high fall simulation system of claim 1, wherein: A multi-dimensional time series data sequence is input into a time series convolutional neural network to extract a multi-dimensional time series feature representation reflecting the response of the visual disturbance effect, and the specific steps are as follows, A multi-dimensional time series data sequence is input into a multi-layer one-dimensional convolutional neural network to extract local dynamic features through a sliding time window, and output a preliminary feature tensor. A multi-scale fusion and context modeling method is used to combine residual connections and attention mechanisms to perform weighted integration and time series optimization processing on the preliminary feature tensor to generate a time series feature sequence of behavior change trends. Through a fully connected layer, the time series feature sequence is dimensionally compressed and feature integrated to obtain a multi-dimensional time series feature representation.
7. The VR device-based high fall simulation system of claim 1, wherein: The disturbance rhythm parameters are obtained by the following specific steps, A fuzzy logic control method is used to fuzz the anxiety index and the focal drift speed index, and input them into a preset fuzzy rule for fuzzy aggregation and defuzzification to derive a fuzzy control output value. According to the fuzzy control output value, an adaptive gain adjustment method is used to adjust the frequency and amplitude of the visual disturbance effect to obtain the disturbance rhythm parameters.
8. The VR device-based high fall simulation system of claim 1, wherein: The visual perspective transformation and the visual angle rotation are triggered by the disturbance rhythm parameters to simulate the high-altitude falling visual effect, and the specific steps are as follows, According to the disturbance rhythm parameters, a matrix transformation method is used to calculate the visual perspective transformation matrix, and an angular velocity calculation method is used to obtain the angular velocity of the visual angle rotation. Based on the visual perspective transformation matrix, the projection parameters of the high-altitude virtual scene are updated, and the position and orientation of the virtual camera are dynamically adjusted in combination with the angular velocity of the visual angle rotation to generate high-altitude virtual scene transformation picture frames and visual angle rotation trajectories. A time synchronization and visual sequence fusion method is used to dynamically fuse the high-altitude virtual scene transformation picture frames and the visual angle rotation trajectories to form a simulated high-altitude falling visual effect.
9. The VR device-based high fall simulation system of claim 1, wherein: The immersive experience report is formed by the following specific steps, Based on the gaze data, the anxiety index, and the focal drift speed index, multi-dimensional statistical analysis and time series analysis are performed to identify abnormal gaze patterns and anxiety changes, and output behavior characteristic labels and behavior change trend data. The behavior change trend data is classified and recognized by using a support vector machine algorithm, psychological states and behavior types are analyzed, and an immersive experience report is formed.
Citation Information
Patent Citations
Psychological intervention system based on eye movement data
CN108030498A
High-altitude falling simulation system based on VR equipment
CN118691769A