A sound-driven VR serious game interaction control method for semi-blind people
By constructing a 3D soundscape space and dynamic potential energy field, and combining multimodal interaction feature acquisition and boundary logic judgment, the problem of coordinate drift and perception delay for semi-blind users in virtual reality was solved, achieving high-precision spatial guidance and deep emotional interaction, and improving the user's autonomous perception and rehabilitation training effect.
Patent Information
- Application Number
- CN202610530974.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-04-21
- Publication Date
- 2026-08-25
AI Technical Summary
Existing voice-guided interaction solutions suffer from coordinate drift and logical loopholes when dealing with complex spatial boundary logic, making it difficult to achieve high-precision synchronization in three-dimensional space. This results in perceptual delays and insufficient immersive experience for semi-blind users in serious virtual reality games.
By constructing a 3D soundscape space, combining dynamic potential energy fields and head-related transfer function algorithms, the system collects multimodal interaction features of users in real time, dynamically adjusts sound parameters, and combines haptic feedback terminals and multi-boundary logic judgment to achieve high-precision spatial guidance and safe interaction under non-visual conditions.
It improved the positioning accuracy of semi-blind users in virtual space, ensured physical and psychological safety, enhanced autonomous perception, and achieved deep emotional interaction and self-empowerment experience through multimodal interaction, thereby improving the immersive experience and rehabilitation training effect.
Smart Images

Figure CN122633019A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of virtual reality technology, specifically relating to a sound-driven VR serious game interaction control method for people with partial blindness. Background Technology
[0002] With the continuous evolution of virtual reality (VR) technology, its application in assisted rehabilitation and serious gaming has become a key pathway to improve the quality of life and motor perception abilities of special populations. By constructing highly immersive virtual environments, VR systems can simulate complex physical interaction scenarios, providing safe and controlled rehabilitation training spaces for vulnerable groups such as the visually impaired. Against this backdrop, how to utilize multimodal interaction technologies to enhance users' spatial cognition and emotional expression has become a core issue in current interdisciplinary research on human-computer interaction and art therapy.
[0003] Among them, the voice-driven interactive control method for partially blind people aims to compensate for the lack of visual feedback through dynamic sound field construction and non-visual guidance mechanisms. This type of technology uses real-time changes in three-dimensional soundscapes and audio parameters to map physical motion trajectories, guiding users to follow rhythms and move their bodies in a virtual space. By combining voice-guided logic with motion perception, the system can achieve a synergy between user emotional release and self-empowerment while ensuring interactive safety, thereby improving their level of autonomous perception in complex environments.
[0004] However, existing voice-guided interaction solutions have significant limitations when handling complex spatial boundary logic. Traditional methods lack a rigorous logical judgment framework when verifying the coordinate relationship between the effective motion area and the obstacle avoidance area, making coordinate drift or logical loopholes prone to occur during dynamic interaction. Furthermore, for handling random collisions in three-dimensional space, existing technologies struggle to maintain high-precision synchronization between axial constraints and rotation directions, frequently leading to boundary errors under multi-conditional branches. In addition, due to insufficient capture of non-linear changes in the interaction state, there is a perceptual delay between sound feedback and motion commands, preventing accurate correlation between data and scene semantics, severely impacting the immersive experience and motion intervention effectiveness for semi-blind users.
[0005] Therefore, a voice-driven VR serious game interaction control solution is desired for the partially blind population. Summary of the Invention
[0006] The purpose of this invention is to provide a sound-driven virtual reality serious game interaction control method for semi-blind people, which can effectively solve the problems in the background art.
[0007] To achieve the above objectives, the technical solution adopted by the present invention is as follows: A voice-driven virtual reality serious game interaction control method for partially blind people includes the following specific steps: Step 1: Constructing a virtual reality 3D soundscape space and its motion guidance field: A 3D Cartesian coordinate system is established in the virtual environment. Based on the preset serious game scene logic, multiple virtual sound sources with specific acoustic properties are defined. The propagation, reflection and attenuation process of sound waves in space is simulated through the head-related transfer function algorithm to form a dynamic sound scene with spatial orientation and distance. At the same time, based on the gradient field theory, a potential energy field is generated in the virtual space to guide the user's movement, and the movement path is transformed into a nonlinear change law of sound frequency and intensity. Step 2: Real-time acquisition and calculation of user multimodal interaction features: Using a wearable inertial measurement unit and a handheld interactive terminal, the user's head rotation angular velocity, limb motion acceleration and hand touch signal are acquired at a preset sampling frequency. The extended Kalman filter algorithm is applied to denoise and fuse the raw sensor data, calculate the user's real-time pose coordinates and motion vector in the virtual space, and convert the collected physiological feedback signals into emotional state feature vectors. Step 3: Implement dynamic sound guidance and interactive feedback control. The user's pose coordinates obtained in step 2 are compared in real time with the motion guidance field preset in step 1. The orientation and distance deviations between the user's current position and the target path point are calculated. The volume balance, pitch and rhythm of the virtual sound source are dynamically adjusted. Correction instructions are output to the user through the auditory channel. Combined with the vibration frequency changes of the tactile feedback terminal, spatial orientation guidance under non-visual conditions is achieved. Step 4: Perform complex boundary condition verification and region logic determination. The system divides the virtual space into an effective movement area, a transition warning area, and an obstacle avoidance prohibited area. It performs multiple logical checks on the user's real-time coordinates and determines whether the user is within a safe movement range by calculating the Euclidean distance between the coordinate point and the boundary surface of the area. When the coordinates are detected to be close to the obstacle avoidance prohibited area, a high-priority warning sound effect is triggered and invalid movement command input is blocked to ensure the safety and logical integrity of the interaction process. Step 5 handles 3D space collision synchronization and random flipping constraints: For dynamic obstacles or interactive objects in the virtual environment, a collision detection model based on axial parallel bounding boxes is established. When random flipping or axial rotation occurs, a quaternion interpolation algorithm is used to maintain real-time synchronization between the physical collider and the sound rendering node. A multi-branch logic judgment matrix is used to solve boundary errors in the collision process and ensure the consistency of audio-visual synchronization. Step 6: Mapping Emotional Expression and Motion-Enabling Feedback A mapping model between emotional parameters and sound art features is established. Based on the user's speed, smoothness, and physiological feature vectors during exercise, background music or sound effects with motivational properties are synthesized in real time. Through an immersive audio experience, users are guided to follow the rhythm, achieving synergy between psychological intervention and sports rehabilitation.
[0008] Preferably, in the first step of constructing the virtual reality 3D soundscape space, a preset audio sampling rate and a predetermined quantization bit depth are used. Multiple virtual sound sources are employed, and the refresh rate of the sound field reconstruction is not lower than a preset refresh threshold to ensure the smoothness and positioning accuracy of the sound signal. The head-related transfer function algorithm employs a personalized database matching mechanism, selecting the transfer function model with the highest matching degree from multiple preset samples based on the user's head circumference and auricular feature parameters. Its spatial positioning error is within a preset allowable error range in both the horizontal and vertical directions. By setting the sound absorption coefficient, reflection coefficient, and transmission coefficient in the virtual environment, acoustic feedback in a real physical environment is simulated, enhancing the spatial immersion of semi-blind users.
[0009] Preferably, in step 2, when real-time acquisition of user multimodal interaction features, the accelerometer resolution and gyroscope zero-bias stability of the inertial measurement unit both reach a preset accuracy standard. Through the fusion of 3-axis acceleration, 3-axis angular velocity, and 3-axis geomagnetic intensity data, a predetermined tracking accuracy for user movements is achieved. The acquisition of physiological feedback signals includes heart rate variability data obtained through a photoplethysmography (PPG) sensor and skin conductance activity data obtained through a skin conductance sensor. The feature extraction period is set to a predetermined time interval, and a sliding window algorithm is used to achieve real-time classification of the user's emotional states such as anxiety, excitement, or fatigue, with the classification accuracy reaching a preset target value.
[0010] Preferably, in step 3, the dynamic sound guidance control uses a proportional-integral-derivative (PID) control algorithm to adjust the sound parameters. When the azimuth deviation exceeds a preset angle threshold, the system increases the sound level difference between the two ears, ensuring the difference remains within a preset range. When the distance to the target point is less than a preset distance threshold, the system enhances the urgency of the sound by increasing high-frequency harmonic components, guiding the user to accurately reach the designated location. The dynamic sound art guidance mechanism also includes real-time modulation of the timbre, using subtractive synthesis technology to change the filter cutoff frequency according to the user's deviation from the target, allowing the sound to smoothly switch between bright and deep tones, providing intuitive feedback.
[0011] Preferably, in step 3, the haptic feedback terminal includes linear resonant motors distributed at the user's fingertips and wrists. The vibration frequency dynamically changes within a preset frequency range, the vibration intensity is inversely proportional to the distance to obstacles in the virtual space, and the response delay is controlled within a preset delay threshold. The haptic feedback and auditory feedback are highly synchronized on the time axis. By synchronously increasing the vibration intensity and sound loudness as the user approaches the target point, a multimodal collaborative interaction channel is constructed to compensate for the lack of visual information in partially blind individuals.
[0012] Preferably, in step 4, the complex boundary condition verification uses ray projection for region determination. Multiple virtual detection rays are emitted in the user's movement direction each frame, and the coordinates of the intersection points between the rays and the boundary geometry are calculated. When the distance between the intersection points is less than a preset safety distance, the system automatically corrects the user's virtual pose to prevent them from penetrating the virtual wall or entering unauthorized areas. The region logic determination also includes curvature analysis of the movement trajectory. If the user's movement path exhibits abnormal nonlinear fluctuations, the system will determine it as a perceptual disorientation state and automatically switch to a low-speed guidance mode, stabilizing the user's emotions by enhancing low-frequency ambient sound effects.
[0013] Preferably, the coordinate verification logic for the effective motion area in step 4 includes checking the continuity of the coordinate sequence. By setting a maximum motion step size threshold, abnormal coordinate data caused by sensor jumps are filtered out to ensure the stability of path guidance. The system maintains a dynamic coordinate cache queue in the background. By linearly weighting a preset number of recent frame coordinate data, the system predicts the user's position at the next moment. If the predicted position falls into the obstacle avoidance prohibited area, the deceleration intervention logic is initiated a predetermined number of frames in advance.
[0014] Preferably, in step 5, the 3D space collision processing employs hierarchical bounding box technology, decomposing complex virtual objects into multiple simple combinations of spheres or cuboids, thus controlling the time complexity of collision detection within a predetermined level. When handling violent flips, a quaternion spherical linear interpolation algorithm is used to achieve a smooth transition of the colliding body coordinates, eliminating gimbal lock-up. The system establishes a collision response lookup table and presets various physical collision parameters for virtual objects of different materials, ensuring that the rebound trajectory and momentum exchange after a collision conform to physical laws.
[0015] Preferably, the collision sound effect synthesis in step 5 is based on a physical modeling algorithm. The system calculates the spectral envelope of the sound wave in real time according to the material properties, contact area, and impact velocity of the colliding objects, ensuring realistic collision feedback and that the audio-visual synchronization error is less than a preset synchronization threshold. For randomly generated dynamic obstacles, the system uses a particle system to simulate the visual shattering effect at the moment of collision and converts it into spatially distributed particle-synthesized audio signals. Through multi-channel surround sound rendering, the system provides the user with precise feedback on the location of the collision.
[0016] Preferably, the emotion expression mapping model in step 6 employs a deep belief network, taking physiological feature vectors as input and outputting the key, tempo, and orchestration parameters of the background music. When a user's low mood is detected, the system automatically increases the music tempo and the proportion of bright timbres according to a predetermined ratio to achieve psychological empowerment. The exercise empowerment feedback also includes real-time voice broadcast of the user's exercise performance, using natural language processing technology to transform abstract exercise data into encouraging, conversational feedback, enhancing the user's sense of accomplishment.
[0017] Preferably, the method utilizes a high-performance graphics processor for real-time parallel computing, supporting collaborative interaction among multiple users within the same local area network environment. Network transmission latency is below a preset latency threshold, and the overall system frame rate remains stable above a preset frame rate threshold. The system architecture employs a distributed processing mode, allocating audio rendering tasks and logical operation tasks to different processor cores. A shared memory mechanism enables ultra-fast data exchange, ensuring system stability under complex interactive loads.
[0018] Preferably, the serious game scenarios include a rhythm-guided mode, a spatial exploration mode, and an art creation mode. In rhythm-guided mode, the beat error of the sound guidance is controlled within a preset error range; in spatial exploration mode, the dynamic range of the environmental soundscape meets the preset dynamic range requirements; in art creation mode, the user's physical movements are mapped to real-time generated electronic music melodies, achieving a deep integration of art therapy and rehabilitation training. The sound guidance strategies in different modes can be graded and adjusted according to the user's degree of visual impairment, providing difficulty settings including multiple difficulty levels.
[0019] Preferably, the method further includes establishing a motion perception database containing a large sample size of semi-blind users, and using a long short-term memory neural network to conduct long-term tracking and analysis of users' motion trajectories, reaction times, and emotional fluctuations. Through in-depth mining of historical data in the database, the system can automatically optimize personalized guidance parameters for each user, predict lead time to reach a preset duration, thereby achieving proactive avoidance of motion risks.
[0020] Preferably, the method is applied to a dedicated virtual reality rehabilitation device, which is equipped with a high-fidelity headset, a multi-sensor fusion limb tracker, and an interactive controller with force feedback function. The overall power consumption of the device is less than a preset power consumption threshold, the continuous working time is greater than or equal to the predetermined working time, and it has an automatic calibration function, which can complete the setting of the user's initial posture within a predetermined initialization time.
[0021] Compared with the prior art, the present invention has the following beneficial effects: 1. Extremely high spatial perception and guidance accuracy This invention achieves high-precision spatial orientation guidance under non-visual conditions by constructing a 3D soundscape that combines a dynamic potential energy field with a head-related transfer function algorithm. Compared to traditional guidance methods that rely solely on simple volume changes, this invention provides a highly directional spatial reference, significantly improving the positioning accuracy of semi-blind users in virtual space and greatly enhancing their autonomous perception capabilities in complex environments.
[0022] 2. Robust border defense and interactive security The multi-boundary logic judgment framework and real-time ray detection mechanism established in this invention completely solve the coordinate drift and penetration error problems that are prone to occur in existing technologies. By strictly dividing the effective motion area and the obstacle avoidance prohibited area, and by verifying the continuity of nonlinear motion states, the system can predict and intercept most potential collision risks in advance, ensuring the physical and psychological safety of special groups during rehabilitation training.
[0023] 3. Deep emotional interaction and self-empowerment experience This invention deeply integrates multimodal interaction data with physiological feedback signals, achieving real-time capture and artistic feedback of the user's emotional state through a deep learning model. This serious game design based on sound art not only helps users complete necessary physical exercise training, but also achieves psychological release and self-empowerment through rhythm following and emotional resonance, expanding the technological boundaries of art therapy in the field of virtual reality.
[0024] 4. Excellent system performance and real-time response capability This invention employs a high-performance parallel computing architecture and a quaternion interpolation algorithm to control system interaction latency and collision synchronization error within a preset, extremely low threshold. This high real-time feedback mechanism effectively eliminates the perceived latency in traditional sound-guided solutions, ensuring precise correlation between sound feedback and motion commands, and significantly improving the immersive experience for semi-blind users and the intervention effect in serious games. Attached Figure Description
[0025] Figure 1 This is a schematic diagram of the overall technical solution architecture of the sound-driven VR serious game interaction control method for semi-blind people proposed in this invention; Figure 2 This is a schematic diagram of the core principle framework of sound guidance and interactive feedback based on dynamic soundscape space and motion guidance field in this invention; Figure 3 This is a flowchart illustrating the main stages of real-time acquisition, denoising and fusion, and emotional state calculation of user multimodal interaction features in this invention. Figure 4This is a schematic diagram of the multi-level interaction relationship and data flow of virtual space complex boundary verification, regional logic determination and 3D space collision synchronization in this invention; Figure 5 This is a schematic diagram illustrating the technical effect principle of motion-empowered feedback based on the mapping model of emotional parameters and sound art features in this invention; Detailed Implementation Example 1
[0026] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to specific embodiments.
[0027] In the aforementioned sound-driven virtual reality serious game interaction control method for partially blind people, the first step of constructing a virtual reality 3D soundscape space and its motion guidance field is achieved by initializing a high-precision 3D Cartesian coordinate system in the virtual environment computing unit. The 3D Cartesian coordinate system uses the ground projection point of the user's initial standing position as the origin, with the horizontal rightward direction as the positive X-axis, the horizontal forward direction as the positive Y-axis, and the vertical upward direction as the positive Z-axis. After the 3D Cartesian coordinate system is established, the system configures multiple virtual sound sources with specific acoustic properties in the space according to the preset serious game scene logic. These virtual sound sources not only contain position coordinate information but also a preset audio sampling rate and a predetermined quantization bit depth. The audio sampling rate is set to 48kHz, and the quantization bit depth is set to 24bit to ensure high fidelity of the sound signal. Specifically, in the first step of constructing a virtual reality 3D soundscape space, the number of virtual sound sources is set to multiple according to the complexity of the scene, and the refresh rate of the sound field reconstruction is not less than 60Hz to ensure the smoothness and positioning accuracy of the sound signal during spatial movement.
[0028] To achieve realistic auditory spatialization, this embodiment employs the Head Related Transfer Function (HRTF) algorithm, which simulates the propagation, reflection, and attenuation of sound waves in space. The HRTF algorithm further integrates a personalized database matching mechanism. The system pre-stores a personalized database containing multiple sets of samples and selects the transfer function model with the highest matching degree from the database based on the user's head circumference and auricular feature parameters (such as helix radius and antihelix height). In actual operation, the spatial positioning error generated by this algorithm is less than 15 degrees in the horizontal direction and less than 20 degrees in the elevation direction, ensuring extremely high fidelity in sound location.
[0029] While constructing the soundscape, a potential energy field is generated within the virtual space based on gradient field theory to guide the user's movement. This potential energy field is defined by a scalar function. To indicate, among which This represents the user's coordinates in the virtual space. The potential energy field transforms the motion path into a nonlinear variation of sound frequency and intensity. Specifically, the system determines the direction and magnitude of the motion-guiding force by calculating the negative gradient of the potential energy function, the mathematical expression of which is as follows:
[0030] in, Indicates at coordinate point The guiding force vector at the location is mapped to real-time processing parameters in the audio engine. When the user deviates from the preset path, the increase in gradient value will directly lead to an increase in sound frequency and a non-linear increase in sound intensity. In addition, the system presets the sound absorption coefficient, reflection coefficient, and transmission coefficient in the virtual environment, and simulates the acoustic feedback in the real physical environment through a real-time convolutional reverberation algorithm, such as simulating the high-frequency absorption effect of a wall or the long-distance attenuation effect of an open space, thereby enhancing the spatial immersion of semi-blind users.
[0031] In the aforementioned sound-driven virtual reality serious game interaction control method for semi-blind people, the second step involves real-time acquisition and calculation of the user's multimodal interaction features.
[0032] The system is equipped with a wearable inertial measurement unit (IMU) and a handheld interactive terminal. The IMU includes a 3-axis accelerometer, a 3-axis gyroscope, and a 3-axis magnetometer. The accelerometer resolution is set to 0.001g, and the gyroscope's zero-bias stability is set to 10 degrees / hour. The system acquires the user's head rotation angular velocity, limb motion acceleration, and hand touch signals at a preset sampling frequency of 200Hz. The calculation process employs the Extended Kalman Filter (EKF) algorithm, which denoises and fuses the raw sensor data. In the prediction phase, the pose at the current moment is predicted using the previous moment's pose state and the current kinematic equations. In the update phase, the predicted values are corrected using the observations from the accelerometer and magnetometer, thereby calculating the user's real-time pose coordinates and motion vector in virtual space. The real-time pose coordinates include three-dimensional position. The attitude angle is represented by a quaternion.
[0033] Simultaneously, the system collects the user's physiological feedback signals and converts them into emotional state feature vectors. The physiological feedback signals include heart rate variability (HRV) data acquired through a photoplethysmography (PPG) sensor and skin conductance activity data acquired through a skin conductance sensor (GSR). The feature extraction cycle is set to 2 seconds, and a sliding window algorithm is used to slice the raw physiological signals in real time. The system utilizes a pre-trained support vector machine (SVM) or random forest classifier to classify the user's current state as anxious, excited, fatigued, or calm based on the extracted time-domain and frequency-domain features (such as mean heart rate, number of peak skin conductance responses, etc.). In this embodiment, the accuracy threshold for emotional state classification is set to 91%.
[0034] In the aforementioned sound-driven virtual reality serious game interaction control method for semi-blind people, step 3 executes dynamic sound guidance and interactive feedback control.
[0035] The system compares the user's pose coordinates obtained in step 2 with the preset motion guidance field in step 1 in real time, calculating the azimuth and distance deviations between the user's current position and the target path point. Dynamic sound guidance control uses a proportional-integral-derivative (PID) control algorithm to adjust the sound parameters.
[0036] When the azimuth deviation exceeds a preset angle threshold of 15 degrees, the system adjusts the binaural audio gain to increase the sound level difference between the two ears. This sound level difference is adjusted within a range of 3dB to 12dB to generate clear lateral azimuth guidance. When the distance between the user's current location and the target path point is less than a preset distance threshold of 0.5 meters, the system increases the high-frequency harmonic components in the synthesized timbre and shortens the trigger interval of the pulse sound to enhance the urgency of the sound. The output of the PID controller directly acts on the control bus of the audio middleware, and its control logic is as follows:
[0037] in, This represents azimuth or distance deviation. This represents the adjustment amount of sound parameters (such as volume, pitch, and rhythm). This is achieved by appropriately setting the scaling factor. Integral coefficient and differential coefficients This ensures smooth sound feedback and fast response.
[0038] Furthermore, the dynamic sound art guidance mechanism also includes real-time modulation of timbre. The system utilizes subtractive synthesis technology to adjust the cutoff frequency of the low-pass or high-pass filter in real time based on the user's deviation from the target. When the user approaches the target, the cutoff frequency shifts to higher frequencies, making the sound brighter; when the user deviates from the target, the cutoff frequency shifts to lower frequencies, making the sound deeper.
[0039] The interactive feedback also incorporates a haptic feedback terminal.
[0040] The haptic feedback terminal includes linear resonant motors (LRAs) distributed across the user's fingertips and wrist. The vibration frequency of the motors dynamically varies within a preset frequency range of 150Hz to 250Hz. The vibration intensity is inversely proportional to the distance to obstacles in the virtual space; as the distance decreases, the vibration intensity increases linearly. The system controls the response delay of the haptic feedback to within 15ms and ensures that the haptic feedback and auditory feedback remain highly synchronized on the time axis. By synchronously increasing the vibration intensity and sound loudness as the user approaches the target point, a multimodal collaborative interaction channel is constructed.
[0041] In the aforementioned sound-driven virtual reality serious game interaction control method for semi-blind people, step 4 implements complex boundary condition verification and regional logic determination.
[0042] The system pre-divides the virtual space into effective movement areas, transition warning areas, and obstacle avoidance prohibited areas. To ensure real-time and accurate judgment, a ray projection method is employed. In each frame of computation, the system emits 32 virtual detection rays from the center point of the user's virtual pose, extending towards the user's current movement vector direction and preset angle directions on both sides. The system calculates the intersection coordinates of each ray with boundary geometry (such as virtual walls and obstacle edges) in real time. When the intersection distance is less than a preset safe distance of 0.3 meters, the system immediately activates boundary defense logic. The area logic judgment also includes curvature analysis of the movement trajectory. The system stores the user's movement coordinates for the most recent 30 frames in real time and calculates the instantaneous curvature of the trajectory. If the curvature exceeds a preset threshold, it indicates that the user's movement path exhibits abnormal nonlinear fluctuations, and the system determines that the user is in a perceptual disorientation state and automatically switches to a low-speed guidance mode. In low-speed guidance mode, the system stabilizes the user's emotions by enhancing low-frequency environmental sound effects (such as simulated ocean waves or wind sounds) while reducing the BPM (beats per minute) of the background music.
[0043] The valid motion area coordinate verification logic also includes checking the continuity of the coordinate sequence. The system sets a maximum motion step size threshold. If the displacement between two adjacent frames exceeds this threshold, it is determined to be an abnormal coordinate caused by sensor jumps and is filtered out. The system maintains a dynamic coordinate cache queue in the background. By linearly weighting the coordinate data of the most recent 10 frames, it predicts the user's position at the next moment. If the predicted position falls into the obstacle avoidance prohibited area, the system will initiate deceleration intervention logic 5 frames in advance, trigger a high-priority warning sound effect, and block invalid motion command input.
[0044] In the aforementioned sound-driven virtual reality serious game interaction control method for semi-blind people, step 5 handles 3D space collision synchronization and random flipping constraints.
[0045] For dynamic obstacles or interactive objects in a virtual environment, the system establishes a hierarchical collision detection model based on axial parallel bounding boxes (AABB). This model decomposes complex virtual objects into combinations of multiple simple spheres or cuboids, thereby keeping the time complexity of collision detection at a low level.
[0046] When handling random flips or axial rotations of virtual objects, the system employs the quaternion spherical linear interpolation (Slerp) algorithm. When an object undergoes a violent flip, this algorithm enables a smooth transition between the physical collider and the sound rendering node, eliminating gimbal lock-up and ensuring real-time synchronization between physical properties and acoustic performance. The system resolves boundary errors during collisions using a multi-branch logic decision matrix. This matrix calculates the rebound trajectory and momentum exchange after the collision in real time based on the collision normal vector, incident velocity vector, and object surface friction coefficient.
[0047] The collision sound effects are synthesized based on a physical modeling algorithm. The system calculates the spectral envelope of the sound waves in real time using modal synthesis technology, based on the material properties of the colliding objects (such as metal, wood, and stone), contact area, and impact velocity. For example, when a high-speed collision is detected, the system automatically increases the amplitude of the excitation signal and extends the high-frequency decay time to ensure realistic collision feedback. The system ensures that the audio-visual synchronization error is less than 30ms. For randomly generated dynamic obstacles, the system uses a particle system to simulate the visual shattering effect at the moment of collision and converts it into spatially distributed particle-synthesized audio signals. Through multi-channel surround sound rendering, the system provides the user with precise feedback on the location of the collision.
[0048] In the aforementioned sound-driven virtual reality serious game interaction control method for semi-blind people, step 6 maps emotional expression and motion empowerment feedback.
[0049] The system establishes an emotion mapping model based on a deep belief network (DBN). The model takes the physiological feature vectors (including heart rate, skin conductance, respiratory rate, etc.) calculated in step 2 as input.
[0050] The deep belief network is composed of multiple stacked Restricted Boltzmann Machines (RBMs). Through layer-by-layer unsupervised pre-training and supervised fine-tuning, it achieves a nonlinear mapping from low-level physiological data to high-level emotional states. The model outputs the key (e.g., major, minor), tempo (BPM), and orchestration parameters (e.g., string proportion, percussion intensity) of the background music. When the system detects that the user is in a low mood or fatigued state, it automatically increases the tempo of the background music by a predetermined percentage of 10% to 20% and increases the proportion of bright timbres (e.g., flute, synthesizer lead) to achieve psychological empowerment. Motion empowerment feedback also includes real-time voice broadcast of the user's exercise performance. The system uses Natural Language Processing (NLP) technology to transform abstract motion data (e.g., average speed, path overlap, completion percentage) into encouraging, conversational feedback. For example, when a user completes a complex rhythmic following action, the system synthesizes voice commands such as "The movement is very smooth, keep the rhythm" in real time, significantly enhancing the user's sense of accomplishment.
[0051] The system's overall architecture adopts a distributed processing mode, utilizing high-performance graphics processing units (GPUs) for real-time parallel computing.
[0052] Audio rendering tasks (such as HRTF convolution and physical modeling synthesis) and logic operation tasks (such as collision detection and trajectory prediction) are assigned to different processor cores. The system achieves ultra-fast data exchange through a shared memory mechanism, ensuring stability under complex interactive loads. When multiple users interact collaboratively in the same local area network environment, network transmission latency is controlled below 20ms, and the overall system frame rate remains stable above 90Hz.
[0053] The serious game scenarios covered in this embodiment include rhythm-guided mode, spatial exploration mode, and artistic creation mode.
[0054] In rhythm-guided mode, the system focuses on monitoring the beat error of the sound guidance, ensuring it is controlled within 50ms. In spatial exploration mode, the system guides users to explore complex virtual mazes autonomously by enhancing the dynamic range of the ambient soundscape (reaching over 80dB). In art creation mode, the user's body movement trajectory is mapped in real time to the melodic lines of electronic music, achieving a deep integration of art therapy and rehabilitation training.
[0055] The system also establishes a motion perception database containing a large sample of semi-blind users, utilizing a Long Short-Term Memory (LSTM) neural network to conduct long-term tracking and analysis of users' motion trajectories, reaction times, and emotional fluctuations. Through in-depth mining of historical data in the database, the system can automatically optimize personalized guidance parameters for each user, such as automatically adjusting the lead time of voice guidance based on the user's reaction delay, with a prediction lead time of 0.5 to 1.5 seconds, thereby achieving proactive avoidance of motion risks. This embodiment is applied to a dedicated virtual reality rehabilitation device, which is equipped with high-fidelity headphones, a multi-sensor fusion limb tracker, and an interactive handle with force feedback function. The overall power consumption of the device is less than 15W, the continuous working time is greater than or equal to 4 hours, and it has an automatic calibration function, which can complete the initial posture setting of the user within a predetermined initialization time of 30 seconds. Example 2
[0056] Based on Example 1, this example provides an enhanced voice-driven interactive control scheme for people with severe visual impairment.
[0057] In this scheme, the virtual reality 3D soundscape constructed in step 1 incorporates a "sonar detection" mechanism. Specifically, the system periodically emits virtual sound wave pulses in all directions from the user within the virtual space. When these pulses encounter virtual obstacles or boundaries, they generate echoes based on the geometric characteristics and physical materials of the obstacles. The timbre, delay, and attenuation of these echoes strictly follow an acoustic physics model. In this way, semi-blind users can actively trigger sound wave pulses and construct a mental spatial map using echolocation principles, further enhancing their spatial perception capabilities.
[0058] In the dynamic sound guidance control of step 3, this embodiment employs a multi-level audio coding strategy.
[0059] The system encodes navigation instructions into audio carriers of different frequencies. For example, low-frequency carriers (200Hz-500Hz) represent broad-range directional guidance, while high-frequency carriers (2kHz-5kHz) represent precise target point positioning. By discerning changes in sound intensity across different frequency bands, users can simultaneously obtain feedback on both the macroscopic path and the microscopic location.
[0060] For the boundary verification in step 4, this embodiment introduces an adaptive safety margin mechanism.
[0061] The system dynamically adjusts the boundary of the effective movement area based on the user's current movement speed. When the user's movement speed is high, the system automatically expands the width of the transition warning area and triggers a warning sound effect in advance. Specifically, the width of the warning area... With speed The relationship satisfies the following linear model: in, The initial width, This is a preset proportional coefficient. This dynamic adjustment mechanism ensures that even at high speeds, the user still has enough reaction time to avoid obstacles.
[0062] In the collision processing of step 5, this embodiment adds detailed simulation of the collision and shattering sound.
[0063] When a user collides with a virtual object, the system not only synthesizes the overall collision sound effect but also generates thousands of tiny audio particles based on the collision intensity. Each particle represents a fragment with an independent 3D trajectory and decay pattern. Through this highly precise audio rendering, users can perceive the extent and direction of the obstacle's fragmentation, thus gaining a more realistic physical interaction experience.
[0064] In the emotion expression mapping of step 6, this embodiment introduces a collaborative creation mechanism.
[0065] In multi-user interactive mode, the system fuses the motion feature vectors of different users to jointly drive the generation of background music. For example, one user's speed determines the tempo, while another user's motion fluidity determines the harmony of the melody. This social interaction based on sound art not only enhances the fun of rehabilitation training but also promotes emotional exchange and social integration among disadvantaged groups. Example 3
[0066] This embodiment details the system architecture implementation of a sound-driven VR serious game interaction control method for partially blind people.
[0067] The system hardware layer consists of a high-performance mobile computing terminal, a high-fidelity audio output unit, a multimodal sensor array, and a force feedback actuator. The multimodal sensor array employs a distributed synchronous bus protocol to ensure that the timestamp error of all sensor data is within 1ms.
[0068] At the software architecture level, the system adopts a microkernel-based real-time operating system, dividing tasks into a high-priority real-time interaction layer and a medium-priority logic operation layer. The real-time interaction layer is responsible for the rapid reading of sensor data, EKF calculation, and real-time mixed output of audio streams. The logic operation layer is responsible for potential field updates, DBN emotion model inference, and NLP speech synthesis.
[0069] Specifically, the audio rendering engine employs an object-based audio processing approach. Each virtual sound source is treated as an independent data object, containing its complete spatial coordinates, directivity function, and spectral characteristics. The engine performs independent HRTF convolution operations on each object through parallel processing units. To conserve computational resources, the system uses a near-field switching algorithm: for sound sources farther from the user, a simplified reverberation model is used; for near-field sound sources, personalized HRTF processing at full sampling rate is initiated.
[0070] In terms of data storage and transmission, the system employs a highly efficient binary serialization protocol to upload the user's movement trajectory, physiological data, and interaction logs to the cloud-based rehabilitation database in real time. The cloud server utilizes big data analytics to quantitatively assess the user's rehabilitation progress and generate a visualized health report. For boundary determination in step 4, this embodiment also introduces a voxel-based spatial partitioning technique. The system divides the virtual motion space into millions of tiny voxel units, each storing its associated region attributes (valid, warning, prohibited). When the user's real-time coordinates change, the system can obtain the region attributes of the current location with a single lookup operation, significantly reducing the computational overhead of collision detection for complex geometries.
[0071] In step 6, the exercise empowerment feedback, the system also integrates a rule-based expert system. Based on the advice of exercise rehabilitation experts, this system pre-sets a series of intervention strategies. For example, when a user remains in a stagnant state for an extended period, the expert system instructs the audio engine to play a guiding melody clip to encourage the user to continue exploring. This intelligent human-computer interaction design gives serious games a stronger therapeutic guidance significance.
[0072] Through the collaborative efforts of the aforementioned embodiments, this invention not only achieves high-precision sound guidance and safety protection for partially blind individuals, but also constructs a comprehensive and immersive space for motion perception and rehabilitation training through deep emotional interaction and synchronized physical collisions. While ensuring real-time performance and stability, the system greatly enriches the technical means of non-visual interaction, providing comprehensive technical support for the application of serious games in the field of caring for vulnerable groups.
[0073] The foregoing has shown and described the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The embodiments and descriptions in the specification are merely illustrative of the principles of the invention. Various changes and modifications can be made to the invention without departing from its spirit and scope, and all such changes and modifications fall within the scope of the present invention as claimed. The scope of protection of this invention is defined by the appended claims and their equivalents.
Claims
1. A voice-driven virtual reality serious game interaction control method for partially blind people, characterized in that, Includes the following steps: Constructing a 3D soundscape space and its motion guidance field in virtual reality: A 3D Cartesian coordinate system is established in the virtual environment. Based on the preset serious game scene logic, multiple virtual sound sources with specific acoustic properties are defined. The propagation, reflection and attenuation process of sound waves in space is simulated through the head-related transfer function algorithm to form a dynamic soundscape with spatial orientation and distance. At the same time, based on gradient field theory, a potential energy field is generated in the virtual space to guide the user's movement, and the movement path is transformed into a nonlinear change law of sound frequency and intensity. Real-time acquisition and calculation of user multimodal interaction features: Through wearable inertial measurement unit and handheld interactive terminal, the user's head rotation angular velocity, limb motion acceleration and hand touch signal are acquired at a preset sampling frequency. The extended Kalman filter algorithm is applied to denoise and fuse the raw sensor data to calculate the user's real-time pose coordinates and motion vector in virtual space, and the collected physiological feedback signals are converted into emotional state feature vectors. Execute dynamic sound guidance and interactive feedback control: Based on the acquired user pose coordinates, compare them with the preset motion guidance field in real time, calculate the orientation and distance deviation between the user's current position and the target path point, dynamically adjust the volume balance, pitch and rhythm of the virtual sound source, output correction instructions to the user through the auditory channel, and combine the vibration frequency changes of the haptic feedback terminal to realize spatial orientation guidance under non-visual conditions. Implement complex boundary condition verification and regional logic determination: Divide the virtual space into effective movement area, transition warning area and obstacle avoidance prohibited area, perform multiple logical verifications on the user's real-time coordinates, and determine whether the user is in a safe movement range by calculating the Euclidean distance between the coordinate point and the boundary surface of the area. When the coordinates are detected to be close to the obstacle avoidance prohibited area, trigger a high-priority warning sound effect and block invalid movement command input to ensure the safety and logical integrity of the interaction process. Handling 3D spatial collision synchronization and random flip constraints: For dynamic obstacles or interactive objects in the virtual environment, a collision detection model based on axial parallel bounding boxes is established. When random flips or axial rotations occur, quaternion interpolation algorithm is used to maintain real-time synchronization between physical colliders and sound rendering nodes. A multi-branch logic judgment matrix is used to resolve boundary errors in the collision process, ensuring the consistency of audio-visual synchronization. Mapping Emotional Expression and Motion-Enabling Feedback: Establishing a mapping model between emotional parameters and sound art features, and synthesizing motivating background music or sound effects in real time based on the user's speed, fluency, and physiological characteristic vectors during exercise. Through an immersive audio experience, users are guided to follow the rhythm, achieving synergy between psychological intervention and exercise rehabilitation.
2. The sound-driven virtual reality serious game interaction control method for partially blind people according to claim 1, characterized in that, The construction of the virtual reality 3D soundscape space and its motion guidance field includes: establishing a 3D Cartesian coordinate system with the ground projection point of the user's initial standing position as the origin; the virtual sound source has a preset audio sampling rate and a predetermined quantization bit depth, and the refresh frequency of the sound field reconstruction is not lower than a preset refresh frequency threshold; the head-related transfer function algorithm integrates a personalized database matching mechanism, selects the transfer function model with the highest matching degree from multiple preset samples based on the user's head circumference and auricular feature parameters, and controls the spatial positioning error within a preset allowable error range; configuring preset sound absorption coefficient, reflection coefficient, and transmission coefficient in the virtual environment, simulating acoustic feedback in a real physical environment through a real-time convolutional reverberation algorithm, and calculating the negative gradient of the potential energy function to determine the direction and magnitude of the motion guidance force by defining a scalar function to represent the potential energy field, and mapping the gradient value to real-time processing parameters in the audio engine.
3. The sound-driven virtual reality serious game interaction control method for partially blind people according to claim 1, characterized in that, The real-time acquisition and calculation of user multimodal interaction features includes: using accelerometers, gyroscopes, and magnetometers with preset resolution and zero bias stability to acquire angular velocity, acceleration, and magnetic field strength data of the user in three axes; the extended Kalman filter algorithm uses the previous pose state and kinematic equations to predict the current pose in the prediction phase, and uses the observations of the accelerometer and magnetometer to correct the predicted values in the update phase, calculating the real-time pose coordinates including three-dimensional position and quaternion-represented pose angles; the extraction of emotional state feature vectors includes acquiring heart rate variability data through a photoplethysmography (PPG) sensor and skin conductance data through a skin conductance sensor, and performing signal slicing using a sliding window algorithm within a predetermined feature extraction period, and using a preset classifier to achieve real-time classification of anxiety, excitement, or fatigue emotional states based on the extracted time-domain and frequency-domain features, with the classification accuracy within a preset target range.
4. The sound-driven virtual reality serious game interaction control method for partially blind people according to claim 1, characterized in that, The dynamic sound guidance and interactive feedback control includes: applying a proportional-integral-derivative control algorithm to adjust sound parameters, using azimuth deviation or distance deviation as input to the controller, and outputting the adjustment amount of sound parameters; when the azimuth deviation is greater than a preset angle threshold, adjusting the binaural audio gain to keep the sound level difference between the two ears within a preset range, generating lateral azimuth guidance; when the distance between the user's current position and the target path point is less than a preset distance threshold, increasing the high-frequency harmonic components in the synthesized timbre and shortening the trigger interval of the pulse sound to enhance the sense of urgency of the sound; the dynamic sound art guidance mechanism also includes real-time modulation of the timbre, using subtractive synthesis technology to change the cutoff frequency of the filter in real time according to the degree of user deviation from the target, so that the sound smoothly switches between bright and deep.
5. The sound-driven virtual reality serious game interaction control method for partially blind people according to claim 4, characterized in that, The interactive feedback control also involves the coordination of haptic feedback terminals, including: utilizing linear resonant motors distributed on the user's fingertips and wrists to dynamically change the vibration frequency of the motors within a preset frequency range, and making the vibration intensity inversely proportional to the distance to obstacles in the virtual space; the haptic feedback and auditory feedback are highly synchronized on the time axis, and by synchronously increasing the vibration intensity and sound loudness when the user approaches the target point, a multimodal collaborative interactive channel is constructed, and the response delay of the haptic feedback is controlled within a preset delay threshold; the haptic feedback terminal triggers vibration excitation in the corresponding direction in real time according to the pose coordinates calculated by the sensor, and cooperates with the sound signal to achieve qualitative feedback on the position of spatial obstacles.
6. The sound-driven virtual reality serious game interaction control method for partially blind people according to claim 1, characterized in that, The implementation of complex boundary condition verification and regional logic determination includes: using ray projection method for regional determination, in each frame operation, starting from the center point of the user's virtual pose, emitting multiple virtual detection rays in the direction of the motion vector and the preset angle directions on both sides, and calculating the coordinates of the intersection point of each ray with the boundary geometry; when the distance between the intersection points is less than the preset safety distance, the boundary defense logic is activated and the user's virtual pose is automatically corrected; the regional logic determination also includes curvature analysis of the motion trajectory, storing the motion coordinates of a preset number of frames in real time and calculating the instantaneous curvature of the trajectory, if the curvature exceeds a preset threshold, it is determined that the user is in a state of perceptual disorientation and automatically switches to a low-speed guidance mode, which stabilizes the user's emotions by enhancing low-frequency environmental sound effects, while reducing the number of beats per minute of the background music.
7. The voice-driven virtual reality serious game interaction control method for partially blind people according to claim 6, characterized in that, The coordinate verification logic of the effective motion area includes checking the continuity of the coordinate sequence and filtering abnormal coordinate data generated by sensor jumps by setting a maximum motion step size threshold. A dynamic coordinate cache queue is maintained in the system background. By linearly weighting the coordinate data of the most recent frame by a preset number of frames, the position of the user at the next moment is predicted. If the predicted position falls into the obstacle avoidance prohibited area, the deceleration intervention logic is initiated in advance for a predetermined number of frames, and a high-priority warning sound effect is triggered, while blocking invalid motion command input. The boundary determination also introduces a voxel-based spatial partitioning technology, which divides the virtual motion space into multiple voxel units and stores the region attributes. The region category to which the real-time coordinates belong is obtained through a lookup operation.
8. The sound-driven virtual reality serious game interaction control method for partially blind people according to claim 1, characterized in that, The processing of 3D spatial collision synchronization and random flipping constraints includes: decomposing complex virtual objects into multiple spheres or cuboids to establish a hierarchical collision detection model; using a quaternion spherical linear interpolation algorithm to achieve a smooth transition between physical colliders and sound rendering nodes, eliminating gimbal lock-up, and calculating the rebound trajectory and momentum exchange after the collision in real time through a multi-branch logic judgment matrix based on the collision normal vector, incident velocity vector, and object surface friction coefficient; the synthesis of the collision sound effects is based on a physical modeling algorithm, calculating the spectral envelope of the sound waves in real time according to the material properties, contact area, and impact velocity of the colliding objects to ensure the realism of the collision feedback, and keeping the audio-visual synchronization error within a preset synchronization threshold range; for dynamic obstacles, a particle system is used to simulate the visual fragmentation effect at the moment of collision, and this is converted into spatially distributed particle-synthesized audio signals for surround sound rendering.
9. A sound-driven virtual reality serious game interaction control method for partially blind people according to claim 1, characterized in that, The mapping of emotional expression and motion-empowering feedback includes: establishing an emotion mapping model based on a deep belief network, which is composed of multiple stacked restricted Boltzmann machines. Through layer-by-layer unsupervised pre-training and supervised fine-tuning, the model takes physiological feature vectors as input and outputs the key, tempo, and orchestration parameters of the background music; when the user's mood is detected to be low or fatigued, the tempo of the background music is increased and the proportion of bright timbre is increased according to a predetermined ratio to achieve psychological empowerment; the motion-empowering feedback also includes using natural language processing technology to convert abstract motion data into encouraging, conversational voice commands for real-time broadcast; the system integrates an expert system to play guiding melody fragments when the user is at rest according to a preset intervention strategy.
10. A sound-driven virtual reality serious game interaction control method for partially blind people according to claim 1, characterized in that, The method also involves distributed processing and long-term data analysis of the system architecture, including: using high-performance graphics processors for real-time parallel computing, allocating audio rendering tasks and logic operation tasks to different processor cores, achieving data exchange through a shared memory mechanism, ensuring network transmission latency is below a preset latency threshold, and maintaining the overall system frame rate above a preset frame rate threshold; establishing a motion perception database containing a large sample of users, using long short-term memory neural networks to perform long-term tracking and analysis of users' motion trajectories, reaction times, and emotional fluctuations, deeply mining historical data, and optimizing personalized guidance parameters for each user, with prediction lead time reaching a preset duration; the serious game scenarios include rhythm guidance mode, spatial exploration mode, and artistic creation mode, with sound guidance strategies in different modes being graded and adjusted according to the user's degree of visual impairment.