Generative AI training scene synthesis method for vestibular rehabilitation

By using a generative AI model to transform the patient's vestibular function assessment results into VR scene control vectors, a 3D VR scene that meets individual needs is generated. This solves the problems of fixed scenes and uncontrollable stimulation in existing VR rehabilitation systems, and enables efficient and personalized VR rehabilitation training.

CN122023735APending Publication Date: 2026-05-12XIAN HONGYI TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
XIAN HONGYI TECHNOLOGY CO LTD
Filing Date
2025-12-25
Publication Date
2026-05-12

AI Technical Summary

Technical Problem

Existing VR rehabilitation methods suffer from problems such as fixed scenarios, uncontrollable stimulation, and high development costs, making it impossible to provide customized training for individual vestibular defects and lacking a physiological closed loop.

Method used

A generative AI model is used to generate personalized VR training scenarios based on the patient's vestibular function assessment results. The generative AI model quantifies the vestibular function defect features into scene control vectors, and combines conditional diffusion model and generative adversarial network to generate 3D VR scenes that conform to physical laws.

Benefits of technology

It achieves unlimited scene generation capabilities, precise physiological adaptation, significantly reduces development costs, constructs a diagnosis-rehabilitation closed loop, and improves rehabilitation efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122023735A_ABST
    Figure CN122023735A_ABST
Patent Text Reader

Abstract

The invention discloses a generative AI training scene synthesis method for vestibular rehabilitation. The invention belongs to the technical field of medical artificial intelligence, and the method comprises the steps: obtaining vestibular function defect features of a subject, the features comprising a cold and hot test asymmetry ratio, a vHIT gain value or spontaneous nystagmus intensity; the features are coded into scene control vectors, the scene control vectors comprise the visual motion direction, the angular velocity amplitude, the acceleration change rate and the space complexity, and the angular velocity amplitude is dynamically calculated according to the patient compensation threshold value and the function retention rate and covers 60%-90% of the patient compensation threshold value and the function retention rate; inputting the vector into a pre-trained conditional diffusion model or a conditional generative adversarial network (cGAN) or other generative AI models to generate a customized VR scene; and rendering the scene to a head-mounted display device for rehabilitation training. According to the invention, one person has one scene and accurate stimulation is realized, the problems of fixed scene, uncontrollable stimulation, high development cost and the like of the existing VR rehabilitation scene are solved, and the training safety and effectiveness are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the intersection of medical rehabilitation technology and artificial intelligence, specifically to a method for generating virtual reality (VR) scenes for rehabilitation training of patients with vestibular dysfunction, and more particularly to a method for dynamically synthesizing safe and effective VR training scenes based on a generative artificial intelligence model and according to the individualized vestibular defect characteristics of patients. Background Technology

[0002] The vestibular system is a core organ for maintaining balance and spatial orientation. When vestibular function is impaired (e.g., benign paroxysmal positional vertigo (BPPV), vestibular neuritis, Meniere's disease), patients often experience symptoms such as dizziness, instability, and nausea, severely impacting their quality of life. Vestibular rehabilitation therapy (VRT), through repeated exposure to specific sensory conflict stimuli, promotes compensatory mechanisms in the central nervous system and is an internationally recognized non-pharmacological intervention.

[0003] In recent years, virtual reality (VR) technology has been widely used in vestibular rehabilitation due to its immersiveness, controllability, and safety. For example, the platform provided by Neuro Rehab VR in the United States includes more than ten preset scenarios such as "tightrope walking" and "supermarket shopping"; some hospitals in China have also introduced balance training games developed based on the Unity engine. However, such systems have the following significant drawbacks:

[0004] The scenario content is fixed and limited: all patients use the same scenario library, making it impossible to customize for individual vestibular defect patterns (such as left horizontal semicircular canal dysfunction, bilateral vestibular disease);

[0005] Uncontrollable stimulation parameters: key parameters such as visual motion speed, direction, and acceleration are mostly preset values ​​or coarse-grained adjustments (such as "low / medium / high" levels), making it difficult to accurately match the patient's compensatory threshold. Too strong a stimulus can easily cause nausea and vomiting, while too weak a stimulus cannot effectively activate neural plasticity.

[0006] High development costs: Each new scene requires artists to create models, programmers to code, and clinicians to verify them, which is time-consuming, costly, and difficult to continuously update.

[0007] Lack of physiological closed loop: The existing system does not use the results of vestibular function tests (such as caloric test, video head impulse test vHIT) as input for scene generation, resulting in a disconnect between "diagnosis" and "rehabilitation".

[0008] While some studies have attempted to adjust scene difficulty using rule engines (such as accelerating background movement based on center of gravity shifts), these efforts remain limited by the fixed structure of the underlying scene, failing to generate entirely new spatial layouts or motion logic. In recent years, generative AI has made breakthroughs in image and 3D scene synthesis, with diffusion models and generative adversarial networks (GANs) capable of generating highly realistic and diverse visual content. However, no published literature or patents have yet applied generative AI to the specific medical scenario of vestibular rehabilitation, and in particular, no mapping mechanism between "vestibular defect features → VR scene parameters" has been established.

[0009] Therefore, there is an urgent need for a technical solution that can automatically generate physiologically adapted, safe, controllable, and infinitely diverse VR training environments based on the patient's vestibular function assessment results, in order to improve rehabilitation efficiency and reduce content development costs. Summary of the Invention

[0010] The purpose of this invention is to overcome the problems of fixed VR rehabilitation scenes, uncontrollable stimulation, and high development costs in the existing technology, and to provide a generative AI training scene synthesis method for vestibular rehabilitation, so as to realize intelligent rehabilitation training with "one scene per person, precise stimulation, and safety and effectiveness".

[0011] Technical solution:

[0012] To achieve the above objectives, the present invention provides the following technical solution:

[0013] A generative AI training scene synthesis method for vestibular rehabilitation includes the following steps:

[0014] First, the vestibular function defect characteristics of the subject are obtained, including at least one semicircular canal functional status index, which is selected from at least one of the following: slow phase velocity asymmetry ratio of the cold and heat test, video head pulse gain value, or spontaneous nystagmus intensity.

[0015] Subsequently, the defect features are encoded into a scene control vector, which includes four dimensions: visual motion direction, angular velocity amplitude, acceleration rate of change, and spatial complexity. The angular velocity amplitude is mapped to the range of 60% to 90% of the patient's compensation threshold based on the functional state index. The compensation threshold is determined by the maximum tolerable angular velocity in historical training data or by dynamic calibration through initial adaptive testing.

[0016] Next, the scene control vector is input into a pre-trained generative AI model, which is a conditional diffusion model or a conditional generative adversarial network (cGAN). During the training phase, it is optimized using a labeled dataset, which includes vestibular defect labels, corresponding safe VR scene samples, and clinicians' ratings of scene stimulus intensity. The training objective of the model is to generate a 3D VR scene that conforms to physical laws and has an immersive feel.

[0017] Finally, the generative AI model outputs customized VR scene data, which includes scene geometry, texture mapping, dynamic object trajectories, and lighting parameters, and renders it in real time to a head-mounted display device for subjects to perform vestibular rehabilitation training.

[0018] Furthermore, the compensation threshold is determined by the maximum tolerable angular velocity in historical training data, or dynamically calibrated through initial adaptive testing.

[0019] Furthermore, the generative AI model uses a labeled dataset during the training phase, which includes: vestibular defect labels, corresponding safe VR scene samples, and clinicians' ratings of the intensity of scene stimuli.

[0020] Beneficial effects:

[0021] Compared with the prior art, the present invention has the following significant advantages:

[0022] Unlimited scene generation capability: Free from the limitations of the preset scene library, it can generate any spatial layout (such as corridor, square, rotating room) and motion mode (linear translation, angular acceleration, compound trajectory) as needed to meet diverse training needs;

[0023] Physiologically precise adaptation: The vestibular dysfunction is quantified into a control vector, dynamically calculated based on functional status indicators and the patient's compensation threshold, and mapped to the 60% to 90% range of the compensation threshold, taking into account both efficacy and comfort.

[0024] Significantly reduces development costs: No manual modeling or programming is required; AI automatically generates high-quality VR content, supporting rapid iteration and personalized expansion.

[0025] Building a closed loop of diagnosis and rehabilitation: Directly using clinical examination data to drive scenario generation, realizing "assessment as prescription", and improving the accuracy of rehabilitation;

[0026] Support for continuous model evolution: As new case data accumulates, the generated model can be continuously optimized through online learning, forming a positive cycle of "data-model-treatment effect". Attached Figure Description

[0027] Figure 1 This is a schematic diagram of the overall process of an embodiment of the present invention;

[0028] Figure 2 This is a diagram illustrating the training and inference architecture of the generative AI model of this invention.

[0029] Figure 3 is a comparative diagram of a traditional preset VR scene and the scene generated by this invention, wherein... Figure 3-1 This is an example of a low-complexity scenario for the present invention; Figure 3-2 This is an example of a high-complexity scenario for the present invention;

[0030] Figure 4 This is an example diagram illustrating the mapping relationship between scene control vectors and vestibular defect features. Detailed Implementation

[0031] The present invention will now be described in further detail with reference to the accompanying drawings. It should be noted that the following embodiments are for illustrative purposes only and are not intended to limit the scope of protection of the invention. Those skilled in the art can make equivalent substitutions or improvements to the technical details without departing from the spirit and scope of the invention.

[0032] - Overall system workflow:

[0033] like Figure 1 As shown, the overall process of the method of the present invention includes four consecutive stages: vestibular function assessment → defect feature quantification → scene control vector generation → generative AI scene synthesis and rendering. Each stage is coordinated and executed by a central controller, which can be an embedded Linux device or a cloud server.

[0034] The first step is the vestibular function assessment:

[0035] Subjects first completed a vestibular function test in a standard clinical setting, including:

[0036] Hot and cold air test: Use an infrared video eye tracker (such as Tobii Pro Fusion) to record the slow phase velocity (SPV) of both eyes and calculate the left and right ear asymmetry ratio A:

[0037] A = |RL| / ((R+L) / 2)×100%;

[0038] Where R is the slow phase velocity (SPV) of the right ear caloric test, in ° / s (degrees per second); L is the slow phase velocity (SPV) of the left ear caloric test, in ° / s (degrees per second).

[0039] Video Head Pulse Test (vHIT): The horizontal semicircular canal gain value G is obtained by synchronizing head angular velocity and eye movement gain using a high sampling rate IMU (such as Bosch BN0055).

[0040] Static eye-tracking: Spontaneous nystagmus was recorded for 5 minutes in a dark environment, and the mean slow-phase velocity Vsp was extracted.

[0041] Then, the defect feature quantification stage begins:

[0042] The above clinical indicators are converted into a structured defect feature vector d:

[0043] If A > 25% and G < 0.7, it is marked as "unilateral functional impairment";

[0044] Record the affected side (left / right);

[0045] Calculate the function retention rate R = G 受损侧 / G 健侧 =0.58 / 1.02≈0.57.

[0046] The vector d = [left, 0.57, 3.2] serves as the input for subsequent encoding.

[0047] Next, we proceed to the scene control vector generation stage:

[0048] like Figure 4 As shown, a mapping rule base (pre-stored in the system database) is established:

[0049] Visual movement direction: set to the opposite direction to the affected side (i.e., moving from right to left) to simulate the visual-motor stimulation when the head turns to the affected side;

[0050] Target angular velocity amplitude ω target The patient's maximum tolerable angular velocity ωmax was determined through initial adaptation testing.

[0051] ω target =ω max ×(0.6+0.3×R)

[0052] This formula ensures that the weaker the function, the weaker the stimulus, but not below the lower limit of the compensation threshold.

[0053] The rate of change of acceleration is fixed at 15-25° / s to avoid abrupt changes.

[0054] Space complexity: Based on Vsp settings: if Vsp > 2° / s, set to "low" (no dynamic objects); otherwise set to "medium".

[0055] The final scene control vector is generated as c = [right-to-left, 58, 20, low].

[0056] Next, we move on to the generative AI scene synthesis and rendering stage:

[0057] like Figure 2 As shown, the system calls the pre-trained Conditional Diffusion Model:

[0058] Model structure:

[0059] Encoder: Maps the control vector c to a 128-dimensional embedding vector through a fully connected layer;

[0060] Backbone network: Based on the U-Net architecture, it performs 50 denoising iterations in the latent space;

[0061] Decoder: Converts the latent representation into a 3D scene asset package in GLTF format, including meshes (.glb), materials (.png), and animation tracks (.json).

[0062] Training process:

[0063] The model was trained using a 500-example labeled dataset, with each example containing:

[0064] Input: Vestibular defect label d;

[0065] Output: A clinically validated and safe VR scene;

[0066] Monitoring signal: Physician's rating of stimulus intensity on a scale of 1-5 (≥4 points is considered acceptable).

[0067] The loss function is:

[0068] L=λ 1·Lrecon +λ 2·Lperceptual

[0069] Where λ1 = 1.0, λ2 = 0.5, Lrecon is the reconstruction loss, and Lperceptual is the perception loss.

[0070] Reasoning and Output:

[0071] After receiving the control vector c, the model generates a GLTF scene asset package within 2 seconds. The asset package includes scene geometry, texture maps, dynamic object trajectories, and lighting parameters.

[0072] like Figure 3-1 As shown on the right, the scene can be represented as a corridor that moves slowly from right to left, with vertical stripes on the walls to enhance the light flow effect, a flat ground, and no dynamic pedestrians (due to the complexity being set to "low").

[0073] like Figure 3-2As shown on the right, in the scenario of this invention, when the spatial complexity is set to "high," dynamic pedestrians and randomly distributed obstacles are added in addition to the basic corridor structure and optical flow effect. The addition of these elements not only increases the visual complexity of the scene but also provides users with a richer interactive experience. Obstacles such as roadblocks or furniture are added in the slightly forward area of ​​the middle of the corridor. These obstacles do not block the main passage but enhance the patient's perception of spatial relationships by introducing visual distractions. Simultaneously, they work in conjunction with dynamic pedestrians to create a multi-dimensional stimulating environment.

[0074] Please note, Figure 3-1 What is shown is a scene with low complexity settings, and Figure 3-2 There are clear differences between the two, which aim to emphasize the system's ability to flexibly adjust the scene content according to user needs.

[0075] The asset package can be loaded into a head-mounted display device via a general VR engine, which is then worn on the subject's head for vestibular rehabilitation training.

[0076] -Risk assessment and response strategies:

[0077] Despite the high degree of automation offered by this solution, in practical applications:

[0078] To prevent the generated scene from aggravating dizziness, the system is equipped with a real-time physiological monitoring module: if the eye tracker detects that the proportion of time with eyes closed exceeds 40%, or the heart rate suddenly increases by more than 20%, the current VR scene will be paused immediately and switched to a static screen.

[0079] To avoid generating unreasonable geometric structures (such as floating objects) in the model, a physical rationality verification module is added after the decoder to eliminate scenes that violate gravity or collision rules;

[0080] To address the bias in individual compensation threshold estimation, a "manual fine-tuning" interface is provided, allowing therapists to adjust the angular velocity within a range of ±10° / s and feed the adjustment results back to the model for online learning;

[0081] To prevent data privacy leaks, all physiological data is processed locally on edge devices, with only the anonymized control vector c uploaded to the cloud.

[0082] Example 1: Scene generation for patients with unilateral vestibular hypofunction

[0083] This embodiment uses a patient with significantly impaired function of the left horizontal semicircular canal as an example to illustrate the complete execution process of the method of the present invention.

[0084] First, vestibular function was assessed. Standardized clinical testing procedures were used to obtain vestibular function parameters: 40°C warm air was alternately injected into both ears of the subject using an infrared video eye tracker (for 30 seconds), and the slow phase velocity (SPV) of both eyes was recorded. The SPV of the right ear was measured to be 18° / s, and the SPV of the left ear was 9° / s. The asymmetry ratio was calculated.

[0085] A=|RL| / ((R+L) / 2) 100%=66.7%.

[0086] Next, head angular velocity and eye movement response were simultaneously measured using a video head-pulse test (vHIT) to obtain the vestibular-ocular reflex gains on both sides. The measured gain for the left side was 0.58, and for the right side it was 1.02. The function retention rate was then calculated.

[0087] R = G 受损侧 / G 健侧 ≈0.57.

[0088] In addition, spontaneous nystagmus was recorded for 5 minutes in a dark environment, and the mean slow phase velocity Vsp = 3.2° / s was extracted.

[0089] Based on the above results, the clinical diagnosis is "significant dysfunction of the left horizontal semicircular canal".

[0090] Based on the above assessment data, a structured defect feature vector d is constructed, which includes the damaged side (left), the functional retention rate (0.57), and the spontaneous nystagmus intensity (3.2° / s), i.e., d = [left, 0.57, 3.2].

[0091] Then, a compensation threshold calibration was performed: a rotating visual scene with angular velocity linearly increasing from 0° / s to 100° / s was played in the VR headset. The test was terminated when the subject reported "mild discomfort," and the maximum tolerable angular velocity was recorded.

[0092] ω max =75° / s, used as the individualized compensation threshold.

[0093] Based on this, the defect features are converted into scene control vector c according to a preset mapping rule library. The visual motion direction is set to "from right to left" to stimulate the left vestibular pathway;

[0094] The target angular velocity is calculated as: ω target =ω max ×(0.6+0.3×R)=75×(0.6+0.3×0.57)≈58° / s.

[0095] The rate of change of acceleration is fixed at 20° / s 2 (within the safe range of 15-25° / s) 2 Inside).

[0096] Space complexity: Since Vsp = 3.2° / s > 2° / s, it is set to "low" (no dynamic objects), that is, it does not contain dynamic objects.

[0097] The final generated control vector c is: c = [right-to-left, 58, 20, low].

[0098] Next, a pre-trained conditional diffusion model is used to generate a VR scene. This model consists of an encoder, a U-Net backbone network, and a decoder. It maps the control vector c to a 128-dimensional embedding, performs 50 denoising iterations in the latent space, and outputs a GLTF-formatted 3D asset package containing meshes, materials, and animation trajectories. The model is trained using 500 labeled samples validated by clinicians, with the loss function being:

[0099] L=1.0·Lrecon+0.5·Lperceptual,

[0100] Where Lrecon is the geometric reconstruction loss and Lperceptual is the VGG-16-based perceptual loss.

[0101] Inference output: During the inference phase, a corridor scene is generated slowly from right to left within 2 seconds. The walls have vertical stripes to enhance the optical flow effect, the ground is flat, and there are no dynamic pedestrians or obstacles. The scene is then loaded onto the head-mounted display device through a general VR engine, with the frame rate stabilized at 72 Hz.

[0102] The system further integrates multiple risk control mechanisms. If real-time physiological monitoring detects that the time spent with eyes closed exceeds 40% or the heart rate suddenly increases by more than 20%, the scene is immediately paused and switched to a static neutral screen, while prompting the therapist to intervene; scene elements that violate gravity or collision rules are removed through the physics engine module; and a compensation threshold fine-tuning interface is provided, allowing the therapist to manually adjust the angular velocity within a range of ±10° / s, and the adjustment results are used for online model learning.

[0103] To verify the effectiveness of this method, a controlled experiment was conducted: 30 patients with unilateral vestibular hypofunction (aged 55-75 years, mean age 62 years) were recruited. 15 patients received the personalized scenario generated by this invention, while the other 15 received a traditional preset scenario. The results showed that the average training time per session in the personalized group was 8.3±1.1 minutes, significantly higher than the 5.4±1.7 minutes in the control group (p<0.01); the self-rating rate of "moderate stimulus intensity" reached 94%, while only 65% ​​in the control group (p<0.05); after three weeks of training, the average improvement in DHI score was -18.3±4.0 points, better than the -12.6±5.4 points in the control group (p<0.01). All data were analyzed using independent samples t-test (SPSS 26), and p<0.05 was considered statistically significant.

[0104] Alternative implementation methods are described below:

[0105] Hardware replacement: The Tobii Pro Fusion, Bosch BN0055, Meta Quest 3 and other devices used in the above embodiments can be replaced with other brand devices that meet the same performance indicators (such as Sensomotoric Instruments eye tracker, InvenSense IMU module, HTC Vive headset) without affecting the implementation of the method of the present invention.

[0106] Model extension: The conditional diffusion model can be replaced by Stable Diffusion v1.4, ControlNet-3D, or other generative AI models with controllable 3D scene generation capabilities, as long as they can receive scene control vectors and output structured VR scene assets.

[0107] Application Scenarios Extension: The core mechanism of this invention—generating personalized dynamic visual stimulation scenarios based on physiological defect characteristics—can also be applied to other neurorehabilitation or functional training scenarios, such as balance disorder rehabilitation, cognitive-motor coordination training, and exposure therapy for anxiety disorders. In these scenarios, only the input definition of vestibular defect characteristics and the corresponding scene control dimensions (such as spatial complexity, movement speed, and social element density) need to be adjusted to reuse the framework of this method.

[0108] System Upgradeability: The system architecture of this invention adopts a modular design, supporting independent component upgrades as technology advances. For example, the eye-tracking module, inertial measurement unit, VR rendering engine, or generative AI model can be replaced or iterated individually without changing the overall process logic, thereby ensuring that the system maintains its technological advancement and clinical applicability in the long term.

Claims

1. A generative AI training scene synthesis method for vestibular rehabilitation, characterized in that, Includes the following steps: The vestibular function deficit characteristics of the subject are obtained, the deficit characteristics include at least one semicircular canal functional status index, the functional status index being selected from at least one of the following: slow phase velocity asymmetry ratio of caloric test, video head pulse gain value, or spontaneous nystagmus intensity. The defect features are encoded as scene control vectors, which include four dimensions: visual motion direction, angular velocity amplitude, acceleration rate of change, and spatial complexity. The angular velocity amplitude is dynamically calculated based on the functional state indicators and the patient's compensation threshold, and mapped to the 60% to 90% range of the compensation threshold. The scene control vector is input into a pre-trained generative AI model, which is a conditional diffusion model or a conditional generative adversarial network (cGAN), and is configured to receive the scene control vector and generate a 3D VR scene that conforms to physical laws and has an immersive feel. The generative AI model outputs customized VR scene data, which includes scene geometry, texture mapping, dynamic object trajectories and lighting parameters, and renders it in real time to a head-mounted display device for the subject to perform rehabilitation training. The method also includes a safety monitoring mechanism based on physiological signals, which is used to adjust or pause the training scenario when abnormal physiological responses are detected.

2. The method as described in claim 1, characterized in that, The compensation threshold is determined by the maximum tolerable angular velocity in historical training data, or by dynamic calibration through initial adaptive testing.

3. The method as described in claim 1, characterized in that, The generative AI model uses a labeled dataset during the training phase, which includes: vestibular defect labels, corresponding safe VR scene samples, and clinicians' ratings of the intensity of scene stimuli.

4. The method as described in claim 1, characterized in that, The direction of visual movement is determined based on the anatomical location of the hypofunctional semicircular canal and is used to provide visual-motor stimulation in the opposite direction to the defective side in order to induce vestibular compensation.

5. The method as described in claim 1, characterized in that, The formula for calculating the amplitude of the angular velocity is: oh target =ω max ×(0.6+0.3×R); Where, ω max Let R be the maximum tolerable angular velocity determined through initial fitness testing, and R be the functional retention rate, defined as: R=G 受损侧 / G 健侧 。 6. The method as described in claim 1, characterized in that, The space complexity in the scene control vector is determined based on the spontaneous nystagmus intensity Vsp: If Vsp > 2° / s, then set it to "Low" (no dynamic objects); Otherwise, set it to "Medium" (containing a small number of moving objects).

7. The method as described in claim 1, characterized in that, The generative AI model is trained using a loss function: L=λ1·Lrecon+λ2·Lperceptual; Where λ1 = 1.0, λ2 = 0.5, Lrecon is the reconstruction loss, and Lperceptual is the perception loss.

8. The method as described in claim 1, characterized in that, The safety monitoring mechanism includes: when it is detected that the subject's eyes are closed for more than 40% of the time or the heart rate rises sharply by more than 20%, the current scene is paused and switched to a static screen.