Immersive virtual reality experience-oriented multi-sense hybrid interaction system and method

CN121879568APending Publication Date: 2026-04-17PIMAX TECH (SHANGHAI) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
PIMAX TECH (SHANGHAI) CO LTD
Filing Date
2025-12-31
Publication Date
2026-04-17

AI Technical Summary

Technical Problem

Existing immersive virtual reality multi-sensory interaction methods struggle to achieve high-precision synchronization, adaptive feedback generation, and continuous evolution of virtual object states, resulting in fragmented multi-sensory experiences that lack coherence and realism, thus limiting the depth of high-end application scenarios.

Method used

By loading intelligent meta-objects, multimodal interaction data is captured in real time, interaction events are predicted, multisensory feedback strategies are generated, and timing orchestration and hardware latency compensation are performed to form synchronous control commands that drive multisensory hardware feedback and monitor actual interaction data to update object attributes.

Benefits of technology

It achieves cross-modal perception fusion, improves the timing synchronization accuracy and perception consistency of multi-channel feedback, and enhances the naturalness of interaction, the coherence of immersion, and the realism of the virtual world.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121879568A_ABST
    Figure CN121879568A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-sense mixed interaction system and method for immersive virtual reality experience, and relates to the technical field of virtual reality man-machine interaction, and the method comprises the steps: loading and packaging an intelligent meta-object virtual environment with physical attributes and feedback rules; capturing user multi-modal data in real time and predicting an interaction event; according to the prediction event and the associated meta-object attributes, dynamically generating a multi-sensory feedback strategy including a touch and smell regulation strategy; performing prospective time sequence arrangement according to delay characteristics of each feedback channel to form a synchronous control instruction; executing instruction driving hardware to output feedback and monitor actual interaction data; according to the method, a closed-loop interaction system integrating perception, decision making, execution and evolution is constructed, high-precision synchronization of multi-sensory feedback and dynamic evolution of a virtual environment are achieved, and the reality sense and consistency of immersion experience are remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of virtual reality human-computer interaction technology, and in particular to a multi-sensory hybrid interaction system and method for immersive virtual reality experiences. Background Technology

[0002] In recent years, significant advancements have been made in technologies such as high-precision motion capture, distributed haptic feedback arrays, and microfluidic odor synthesis, providing the hardware foundation for building more comprehensive sensory simulations. Virtual reality (VR) technology is evolving from creating immersive experiences primarily based on vision and hearing to a deeper integration of multi-sensory collaborative interaction that incorporates touch, force, temperature, and even smell. The industry consensus is that the next generation of immersive experiences relies on breaking down barriers between sensory channels to achieve unified management and real-time synchronization of cross-modal information.

[0003] Existing technical solutions mostly adopt a rigid "event-response" architecture, with feedback logic typically predefined and isolated. This makes it difficult for the system to cope with dynamically changing interaction contexts and unable to generate adaptive multi-sensory fusion signals based on real-time user intent and object states. Specifically, existing methods have significant limitations in cross-modal temporal synchronization accuracy, real-time feedback synthesis based on physical attributes, and persistent state evolution of virtual objects due to interaction. This results in fragmented multi-sensory experiences lacking coherence and realism, limiting their application depth in high-fidelity simulations, professional training, and other high-end scenarios. Summary of the Invention

[0004] In view of the aforementioned existing problems, the present invention is proposed.

[0005] Therefore, this invention provides a multi-sensory hybrid interaction method for immersive virtual reality experiences, which solves the problem that existing immersive virtual reality multi-sensory interaction methods are difficult to achieve high-precision synchronization, adaptive feedback generation, and continuous evolution of virtual object states when dealing with dynamic interactions.

[0006] To solve the above-mentioned technical problems, the present invention provides the following technical solution: In a first aspect, the present invention provides a multi-sensory hybrid interaction method for immersive virtual reality experiences, characterized by comprising the following steps: Load a virtual environment containing smart meta-objects, which encapsulate physical properties and feedback rules; Capture users' multimodal interaction data in real time and predict upcoming interaction events; Based on the predicted interaction events and their associated intelligent meta-object attributes, generate corresponding multi-sensory feedback strategies. Based on the multi-sensory feedback strategy and the delay characteristics of each feedback channel, timing is arranged to form synchronous control commands; Execute synchronous control commands, drive multi-sensory hardware output feedback, and monitor actual interaction data; The attributes of the smart meta-object are adaptively updated based on the actual interaction data monitored.

[0007] As a preferred embodiment of the multi-sensory hybrid interaction method for immersive virtual reality experiences described in this invention, the loading of a virtual environment containing intelligent meta-objects, wherein the intelligent meta-objects encapsulate physical attributes and feedback rules, specifically includes the following steps: Read the topology of the virtual environment and the identifiers of all objects from the scene database; Based on the identifier, a corresponding data entity is instantiated from the intelligent meta-object library. Each data entity contains at least a geometric model, surface material spectrum, thermodynamic coefficient, and odor feature encoding. Register all instantiated smart meta-objects to the global state management list according to their spatial coordinates; Based on the physical attributes of each object in the global state management list, a base sensory field covering the entire virtual space is initialized. The base sensory field includes a preliminary temperature gradient distribution and environmental odor background.

[0008] As a preferred embodiment of the multi-sensory hybrid interaction method for immersive virtual reality experiences described in this invention, the specific steps of real-time capture of the user's multimodal interaction data and prediction of upcoming interaction events are as follows: The device uses built-in sensors to acquire the three-dimensional coordinates of the user's gaze point and pupil change data. The continuous position, motion velocity vector, and finger bending angle of each joint of both hands are obtained through a hand tracker. When no physical collision is detected, the baseline pressure micro-changes of the fingertip pressure sensors of the haptic glove are continuously read; The gaze point coordinates, motion velocity vector, and baseline pressure micro-changes are input into a trained temporal prediction model, which outputs one or more probabilistic event descriptions of specific types of contact between a user's limb and a specific intelligent meta-object within a predetermined future time window.

[0009] As a preferred embodiment of the multi-sensory hybrid interaction method for immersive virtual reality experiences described in this invention, the specific steps of generating corresponding multi-sensory feedback strategies based on predicted interaction events and their associated intelligent meta-object attributes are as follows: The probabilistic event description is parsed to extract the target intelligent element object identifier, predicted contact location, and contact mechanical parameters; Query the global state management list to obtain the surface material spectrum and odor feature encoding currently stored in the target intelligent meta-object; The surface material spectrum and the contact mechanical parameters are input into a tactile waveform generation network, which outputs a driving signal waveform that simulates the tactile sensation of the material under the contact parameters. Based on the odor feature code and the contact type, a cross-sensory modulation mapping table is queried to retrieve the corresponding modulating odor formula number and intensity modulation coefficient, which together constitute the multi-sensory feedback strategy.

[0010] As a preferred embodiment of the multi-sensory hybrid interaction method for immersive virtual reality experiences described in this invention, the step of timing arrangement based on multi-sensory feedback strategies and the latency characteristics of each feedback channel to form synchronous control commands includes the following specific steps: Receive the tactile drive signal waveform, the adjustable odor formula number, and the intensity modulation coefficient; Query the hardware latency database to obtain the current fixed response latency values ​​of the haptic actuator and odor release device; Based on the display time of the visual rendering frame, the start playback command of the tactile drive signal waveform is sent to the tactile driver by a first offset; the release command of the regulatory odor is sent to the odor controller by a second offset, wherein the second offset is greater than the first offset. All time-calibrated hardware control commands are packaged to generate the synchronization control instruction sequence.

[0011] As a preferred embodiment of the multi-sensory hybrid interaction method for immersive virtual reality experiences described in this invention, the specific steps of executing synchronous control commands, driving multi-sensory hardware to output feedback, and monitoring actual interaction data are as follows: The synchronization control command sequence is distributed to the corresponding hardware communication interface; The tactile actuator plays the drive signal waveform according to the instruction timestamp, and the odor controller activates the designated odor capsule and controls the release concentration according to the instruction. During the feedback output process, the pressure sensor and temperature sensor continuously collect data on the actual pressure distribution and temperature changes on the user's skin surface; The actual pressure distribution and temperature change data are compared with the predicted contact mechanical parameters in real time to generate a sensory consistency verification signal.

[0012] As a preferred embodiment of the multi-sensory hybrid interaction method for immersive virtual reality experiences described in this invention, the specific steps of adaptively updating the attributes of the intelligent meta-object based on monitored actual interaction data are as follows: If the sensory consistency verification signal is received, and the verification signal indicates that the actual interaction force continuously exceeds the predicted value by a threshold, then the interaction is determined to be an enhanced contact. Based on the mechanical data of the enhanced contact, calculate a property wear increment; Based on the attribute wear increment, the surface material spectrum of the target intelligent meta-object stored in the global state management list is attenuated and corrected. The corrected surface material spectrum is saved as a new attribute state of the smart meta-object for use in generating tactile waveforms for subsequent interactions.

[0013] Secondly, the present invention provides a multi-sensory hybrid interaction system for immersive virtual reality experiences, comprising, The meta-object loading module loads a virtual environment containing smart meta-objects, which encapsulate physical attributes and feedback rules. The interaction prediction module captures users' multimodal interaction data in real time and predicts upcoming interaction events; The strategy generation module generates corresponding multi-sensory feedback strategies based on the predicted interactive events and their associated intelligent meta-object attributes. The synchronization orchestration module performs timing orchestration based on the multi-sensory feedback strategy and the delay characteristics of each feedback channel to form synchronous control commands; The execution monitoring module executes synchronous control commands, drives multi-sensory hardware output feedback, and monitors actual interaction data; The attribute update module adaptively updates the attributes of smart meta objects based on the monitored actual interaction data.

[0014] Thirdly, the present invention provides a computer device including a memory and a processor, wherein the memory stores a computer program, wherein: when the computer program is executed by the processor, it implements any step of the multi-sensory hybrid interaction method for immersive virtual reality experience as described in the first aspect of the present invention.

[0015] Fourthly, the present invention provides a computer-readable storage medium having a computer program stored thereon, wherein: when the computer program is executed by a processor, it implements any step of the multi-sensory hybrid interaction method for immersive virtual reality experience as described in the first aspect of the present invention.

[0016] The beneficial effects of this invention are as follows: By introducing intelligent meta-objects with encapsulated attributes, generating predictive dynamic feedback strategies, and cross-modal temporal orchestration, a closed-loop interactive system integrating perception, decision-making, execution, and evolution is constructed. This fundamentally changes the rigidity of the traditional "event-response" architecture. The system can dynamically synthesize and weave multi-sensory feedback signals based on the user's real-time interaction intentions and the embedded physical states of virtual objects, realizing the transformation from isolated sensory superposition to cross-modal perception fusion. Through forward-looking prediction and hardware latency compensation, the temporal synchronization accuracy and perceptual consistency of multi-channel feedback such as vision, touch, and smell are effectively improved. At the same time, the system can continuously evolve and update the attributes of meta-objects based on actual interaction data, enabling the virtual environment to have dynamic response and memory capabilities, thereby significantly enhancing the naturalness of interaction, the coherence of immersion, and the realism and adaptability of the virtual world as a whole. Attached Figure Description

[0017] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the following description of the embodiments will be briefly introduced. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0018] Figure 1 This is a flowchart of the multi-sensory hybrid interaction method for immersive virtual reality experiences in Example 1. Detailed Implementation

[0019] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings.

[0020] Many specific details are set forth in the following description in order to provide a full understanding of the invention. However, the invention may also be practiced in other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the spirit of the invention. Therefore, the invention is not limited to the specific embodiments disclosed below.

[0021] Secondly, the term "one embodiment" or "embodiment" as used herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The phrase "in one embodiment" appearing in different places in this specification does not necessarily refer to the same embodiment, nor is it a single or selective embodiment that is mutually exclusive with other embodiments.

[0022] Example 1, referring to Figure 1 This is the first embodiment of the present invention, which provides a multi-sensory hybrid interaction method for immersive virtual reality experiences, characterized by including the following steps: Load a virtual environment containing smart meta-objects, which encapsulate physical properties and feedback rules; Capture users' multimodal interaction data in real time and predict upcoming interaction events; Based on the predicted interaction events and their associated intelligent meta-object attributes, generate corresponding multi-sensory feedback strategies. Based on the multi-sensory feedback strategy and the delay characteristics of each feedback channel, timing is arranged to form synchronous control commands; Execute synchronous control commands, drive multi-sensory hardware output feedback, and monitor actual interaction data; The attributes of the smart meta-object are adaptively updated based on the actual interaction data monitored.

[0023] It should be noted that it dynamically generates feedback signals that match physical properties and achieves cross-sensory synchronous output through forward-looking timing orchestration. The system further introduces a sensor monitoring closed loop, continuously optimizing and evolving the attributes of virtual objects based on actual interaction effects, thus forming a complete interactive closed loop from perception, decision-making, execution to adaptation, significantly improving the realism, consistency, and dynamic responsiveness of the immersive experience.

[0024] Specifically, the loading of the virtual environment containing intelligent meta-objects, which encapsulate physical attributes and feedback rules, involves the following steps: Read the topology of the virtual environment and the identifiers of all objects from the scene database; Based on the identifier, a corresponding data entity is instantiated from the intelligent meta-object library. Each data entity contains at least a geometric model, surface material spectrum, thermodynamic coefficient, and odor feature encoding. Register all instantiated smart meta-objects to the global state management list according to their spatial coordinates; Based on the physical attributes of each object in the global state management list, a base sensory field covering the entire virtual space is initialized. The base sensory field includes a preliminary temperature gradient distribution and environmental odor background.

[0025] It should be noted that the system accesses a pre-built scene database by calling an application programming interface (API), which stores data in a specific format. The system sends a query command to the database, requesting a 3D spatial topology diagram of the current virtual environment and simultaneously retrieving a list of unique numerical identifiers for each interactive object in the environment. These identifiers serve as search keys, used to accurately match the corresponding detailed data entities in subsequent steps. This provides the system with a structured blueprint and object index for the scene, establishing a bridge from abstract environmental description to concrete data entity retrieval, ensuring the accuracy and efficiency of subsequent data loading, and serving as a reliable data input starting point for the entire multi-sensory interaction process.

[0026] Based on the acquired list of identifiers, the system accesses another independent intelligent meta-object library. For each identifier in the list, the system locates its corresponding data storage unit in the meta-object library and reads the stored static attribute data into the runtime memory, creating a data entity instance that can be directly manipulated by the program. This instantiation process ensures that each virtual object carries its own unique geometric shape data, surface material spectrum describing tactile characteristics, thermodynamic coefficients determining temperature changes, and odor feature encoding for odor feedback. By binding and encapsulating discrete sensory attributes with the object model, the system creates an intelligent data unit with complete physical and sensory expression capabilities, providing a direct and unified raw data source for generating multimodal feedback based on real physical laws.

[0027] After instantiating all required intelligent meta-objects, the system iterates through these data entities, reading the 3D spatial coordinates associated with each entity. Subsequently, the system creates a global state management list data structure, inserting or updating each intelligent meta-object's instance pointer (or unique identifier) ​​and its current spatial coordinates as a record in an orderly manner. This list is continuously maintained during system runtime, dynamically reflecting the latest spatial positions of all objects. It achieves unified indexing and state tracking of all dispersed intelligent meta-object instances, providing an efficient management tool for the system to perceive the spatial relationships between users and any object in real time and quickly query interaction targets and their attributes. It forms the foundational architecture supporting real-time, large-scale multi-object interaction scenarios.

[0028] The system invokes the global state management list, traversing all registered smart meta-objects. Based on the thermodynamic coefficients and odor characteristic codes of each object in the list, combined with its spatial coordinates, the system runs an environmental field pre-calculation algorithm. This algorithm simulates the principle of physical diffusion, calculating the contribution of each object to the surrounding virtual space in terms of temperature and odor concentration. It then integrates the contributions of all objects on a three-dimensional spatial grid, generating a continuous baseline temperature distribution map and an environmental odor concentration background map covering the entire virtual environment. This pre-calculates and establishes the sensory background field of the entire virtual space, providing users with a basic thermal and olfactory atmosphere consistent with the scene setting before contact with specific objects. This enhances the overall realism and consistency of the environment and provides an initial reference benchmark for the dynamic update calculation of local sensory fields in subsequent real-time interactions.

[0029] Specifically, the real-time capture of users' multimodal interaction data and prediction of upcoming interaction events involves the following steps: The device uses built-in sensors to acquire the three-dimensional coordinates of the user's gaze point and pupil change data. The continuous position, motion velocity vector, and finger bending angle of each joint of both hands are obtained through a hand tracker. When no physical collision is detected, the baseline pressure micro-changes of the fingertip pressure sensors of the haptic glove are continuously read; The gaze point coordinates, motion velocity vector, and baseline pressure micro-changes are input into a trained temporal prediction model, which outputs one or more probabilistic event descriptions of specific types of contact between a user's limb and a specific intelligent meta-object within a predetermined future time window.

[0030] It should be noted that the system drives the integrated near-infrared eye-tracking module within the head-mounted display to begin operation. This module emits invisible infrared light towards the user's eyes and captures corneal reflection patterns via a miniature camera. The system's built-in dedicated image processing unit analyzes these patterns in real time, calculating the precise two-dimensional pixel position of the user's pupil center in the head-mounted display's screen coordinate system for each frame. Combining this with known user eye-tracking model parameters and device calibration parameters, the system uses triangulation to convert this position into world coordinates of the point where the user's gaze is focused in virtual three-dimensional space. Simultaneously, the system continuously monitors and records the continuous sequence of pupil diameter changes. By accurately acquiring the gaze point coordinates, the system can determine the potential object of the user's intended interaction. Combined with pupil change data, this provides auxiliary reference for assessing the user's cognitive load or emotional response, providing crucial visual intent input for subsequent prediction of interactive events.

[0031] The system utilizes multiple infrared cameras within the optical motion capture system, or the inertial measurement unit built into the handle and data glove, to continuously sample at a frequency of at least 90 Hz. For the optical system, it processes images of markers captured by the cameras and affixed to key joints of the user's hand, reconstructing the real-time posture of the hand skeleton model in three-dimensional space using multi-view vision algorithms. For the inertial or glove system, the system directly reads the raw data from the sensors at each joint. Subsequently, the system calculates the precise position of the palm and each finger joint in three-dimensional space, and derives its instantaneous velocity vector through differential calculation. Simultaneously, it directly reads or calculates the bending angle of each finger joint. The acquired continuous position, velocity, and posture angle data accurately describe the user's hand trajectory, the speed and direction of the intended action, and is the core kinematic input for analyzing and predicting upcoming physical interactions (such as grasping and touching).

[0032] Before the system's physics engine detects any collision or overlap between the user's virtual hand and any smart meta-objects, the system continuously reads the output signals of high-sensitivity thin-film pressure sensors or micro-strain gauges integrated under the pads of each fingertip of the haptic glove in parallel at a high sampling rate. The system pays particular attention to and records the baseline values ​​of these sensors' signals in the "non-contact" state and their subtle fluctuations, which may be caused by pre-contraction, minute tremors, or preparatory movements of the user's finger muscles. By monitoring the subtle changes in fingertip pressure before a collision, the system can obtain leading signals of the user's intended force and contact readiness. This supplements the interaction event prediction with the intentional cues about "force" that are missing in traditional kinematic data, improving the sensitivity and accuracy of the prediction model in judging the interaction force and contact timing.

[0033] The system performs time alignment and feature concatenation on the multi-channel time-series data (gaze coordinate sequence, hand joint kinematic sequence, and fingertip pressure micro-variation sequence) generated in real time from the first three steps, forming a comprehensive feature vector. This feature vector is then input into a recurrent neural network model pre-trained with a large amount of human-computer interaction data. Internally, this model processes these temporal features, analyzes their evolution patterns, and outputs a set of structured predictions. Each result includes: within a fixed short-term window (e.g., the next 300 milliseconds), the system anticipates which part of the user's hand, with what probability, which specific intelligent meta-object in the scene, in which region, and what type of contact event (e.g., touch, grasp) will occur. By integrating visual attention, kinematics, and pre-cognition information, and utilizing machine learning models to mine the temporal correlations, the system can predict the specific interaction event that is about to occur. This forward-looking prediction provides crucial lead time and contextual information for the subsequent generation of multi-sensory feedback strategies and high-precision temporal orchestration, which require a certain preparation time. It is a key intelligent decision-making step in enabling the system to shift from passive response to active collaboration.

[0034] Specifically, the step of generating a corresponding multi-sensory feedback strategy based on the predicted interaction events and their associated intelligent meta-object attributes involves the following steps: The probabilistic event description is parsed to extract the target intelligent element object identifier, predicted contact location, and contact mechanical parameters; Query the global state management list to obtain the surface material spectrum and odor feature encoding currently stored in the target intelligent meta-object; The surface material spectrum and the contact mechanical parameters are input into a tactile waveform generation network, which outputs a driving signal waveform that simulates the tactile sensation of the material under the contact parameters. Based on the odor feature code and the contact type, a cross-sensory modulation mapping table is queried to retrieve the corresponding modulating odor formula number and intensity modulation coefficient, which together constitute the multi-sensory feedback strategy.

[0035] It should be noted that the system receives a structured probabilistic event description data stream from the prediction module. By invoking a dedicated data parsing program, the system first identifies and extracts the target intelligent meta-object identifier string used to uniquely identify the virtual object in the description. Subsequently, the parsing program locates and reads the predicted user body part code that will come into contact (e.g., "the tip of the right index finger"). Finally, the program calculates the predicted set of contact mechanics parameters from the description. This typically includes the estimated normal velocity, tangential velocity vector, and predicted initial contact force direction at the moment of contact. Through precise parsing, the ambiguous probabilistic event is transformed into a clear interaction target, location, and mechanical quantity. This provides clear and specific input instructions for subsequent targeted data queries and signal generation, serving as the decision-making basis for achieving precise and customized feedback.

[0036] The system uses the target intelligent meta-object identifier extracted in the previous step as a query key to initiate a retrieval request to the global state management list that maintains the real-time state of all objects. The list management program quickly locates the corresponding data record based on this identifier and reads the target object's core sensory attribute fields from the record. These are primarily surface material spectra stored in digital spectrum format, describing its surface texture features, and odor feature codes represented in a specific encoding format, used to index its odor. By querying the central state list, the system ensures that the acquired material and odor attributes represent the object's latest and most accurate state. This provides an authoritative data source for generating multi-sensory feedback that conforms to the object's true physical characteristics at this moment, guaranteeing the authenticity and consistency of the feedback.

[0037] The system combines the retrieved surface material spectrum data with the analyzed contact mechanical parameters (such as normal velocity) into a feature vector, and inputs it into a pre-trained tactile waveform generation network (such as a deep learning model). This network performs calculations internally, simulating the physical vibration response of a given material under specific mechanical parameters, and ultimately generates a corresponding digital drive signal waveform that can be used to directly drive tactile actuators (such as linear resonant actuators). Utilizing a neural network model, it generates highly matched tactile waveforms in real-time based on the specific object material and interaction method, replacing the traditional method of retrieving preset patterns from a fixed waveform library. This allows the tactile sensation to change precisely with the interaction force and speed, significantly improving the realism and adaptability of the tactile simulation.

[0038] The system uses both the target object's odor feature encoding and the predicted contact type (such as "grasp" or "touch") as a composite index key to query an independently maintained cross-sensory modulation mapping table. This mapping table defines the moderating odor strategies that should be adopted for different objects in different interaction contexts to optimize the overall perceptual effect. The query result returns a corresponding odor formula number and a suggested intensity modulation coefficient, which are encapsulated together with the tactile drive waveform generated in the previous step to form the final multi-sensory feedback strategy package. This makes olfactory feedback no longer just the release of the object's inherent odor, but rather a strategic selection and intensity modulation based on the interaction type. The aim is to enhance or regulate the user's perception of primary senses such as touch through the olfactory channel, thereby proactively optimizing the overall effect of multi-sensory fusion at the system level and improving the coordination and immersion depth of the perceptual experience.

[0039] Specifically, based on the multi-sensory feedback strategy and the delay characteristics of each feedback channel, timing is arranged to form synchronous control commands. The specific steps are as follows: Receive the tactile drive signal waveform, the adjustable odor formula number, and the intensity modulation coefficient; Query the hardware latency database to obtain the current fixed response latency values ​​of the haptic actuator and odor release device; Based on the display time of the visual rendering frame, the start playback command of the tactile drive signal waveform is sent to the tactile driver by a first offset; the release command of the regulatory odor is sent to the odor controller by a second offset, wherein the second offset is greater than the first offset. All time-calibrated hardware control commands are packaged to generate the synchronization control instruction sequence.

[0040] It should be noted that the timing orchestration module receives structured data packets from the strategy generation module via an internal data bus. This data packet fully contains the digital waveform file of the haptic drive signal generated for this predicted interactive event, the specific odor formula number retrieved from the cross-sensory modulation map, and the calculated suggested odor release intensity modulation coefficient. The system temporarily stores this data in a cache to prepare input for subsequent delay compensation calculations, ensuring that the dynamically generated multi-sensory feedback parameters are completely and accurately transmitted to the synchronization control stage. This provides a clear operational object for subsequent precise timing scheduling based on hardware characteristics and serves as a key data interface connecting creative generation and physical execution.

[0041] The system accesses a locally stored hardware latency database. This database stores the calibration parameters of various feedback hardware currently connected to the system in tabular form. The timing orchestration module sends a query request to this database, retrieving and reading the fixed delay time (first delay value) from receiving a command to generating effective vibration for the currently used linear resonant haptic actuator, and the fixed delay time (second delay value) from receiving a command to the release of odor molecules for the microfluidic odor release device. By querying the preset latency database, the system obtains a precise quantitative understanding of the inherent response speed of hardware in different sensory channels. This is an absolute prerequisite for any effective timing compensation and synchronization calculation, transforming uncontrollable physical delays into predictable and manageable system parameters.

[0042] The system obtains the precise timestamp of the next frame that will be refreshed on the head-mounted display from the graphics rendering engine, using this as the "perceptual alignment target moment" for all sensory feedback. Subsequently, the orchestration module calculates: subtracting the first delay value of the tactile actuator from the target moment yields the "sending moment" of the tactile command; subtracting the second delay value of the odor device from the same target moment yields an earlier "sending moment" of the odor command. After calculation, the system inserts control commands with the corresponding sending timestamps into the command queues of the tactile driver and odor controller, respectively. By proactively sending commands based on vision and setting differentiated offsets for different levels of hardware latency, the system actively compensates for the time difference in hardware response at the command sending level. This ensures that the fast-responding touch and the slow-responding odor are as synchronized as possible with visual events at the brain's perception level, fundamentally solving the technical pain point of asynchronous multi-sensory responses.

[0043] After timing all individual instructions, the system creates an empty synchronization instruction sequence container. Then, following the order of instruction transmission, it systematically adds all hardware control commands, each pre-marked with a precise transmission timestamp, into this container. Finally, it packages these commands into a complete synchronization control instruction sequence data packet with timing information, integrating the scattered, independently scheduled channel instructions into a unified instruction package with a clear timeline and strict sequence. This instruction sequence serves as the sole authoritative basis for all subsequent hardware collaboration, ensuring that the distributed multi-sensory hardware can perform precisely in harmony, strictly adhering to a unified "score," like an orchestra. It is the ultimate carrier for achieving cross-channel sensing synchronization.

[0044] Specifically, the steps for executing synchronization control instructions, driving multi-sensory hardware to output feedback, and monitoring actual interaction data are as follows: The synchronization control command sequence is distributed to the corresponding hardware communication interface; The tactile actuator plays the drive signal waveform according to the instruction timestamp, and the odor controller activates the designated odor capsule and controls the release concentration according to the instruction. During the feedback output process, the pressure sensor and temperature sensor continuously collect data on the actual pressure distribution and temperature changes on the user's skin surface; The actual pressure distribution and temperature change data are compared with the predicted contact mechanical parameters in real time to generate a sensory consistency verification signal.

[0045] It should be noted that the central processing unit copies the generated synchronous control command sequence data packets with precise timestamps to the dedicated USB communication interface buffer connected to the tactile driver and the transmission queue of the low-latency wireless transmission module (such as Wi-Fi 6E or dedicated RF) connected to the odor controller via the system bus. This ensures that the command data accurately reaches the receiving end of each hardware subsystem, and ensures that the unified control plan, which has been precisely timed, is reliably and with low latency transmitted to each independent sensory feedback hardware terminal, providing a unified source of execution commands for the coordinated startup of all hardware devices.

[0046] The haptic actuator parses instructions locally. When the system clock reaches the timestamp specified in the instruction, it immediately reads the haptic drive signal waveform data from its buffer and converts it into an analog voltage signal to drive the linear resonant actuator to vibrate. Simultaneously, at an earlier specified timestamp, the odor controller locates the corresponding odor capsule valve on the microfluidic chip according to the formula number in the instruction and adjusts the power of the carrier gas pump according to the intensity coefficient in the instruction to control the release rate of odor molecules. The driving hardware strictly follows the timing instructions, transforming the digitized haptic waveform and odor strategy into a physical stimulus that the user can perceive. This is the final physical output link to achieve a synchronized "visual-tactile-olfactory" sensory experience.

[0047] During the operation of the haptic actuator, a thin-film pressure sensor array integrated into the fingertips of the glove measures and outputs a real-time pressure distribution map of the contact area at a high sampling rate. Simultaneously, temperature sensors (such as thermocouples) in close contact with the skin continuously monitor and record the temperature change curve at the contact point. This sensor data, after analog-to-digital conversion, is transmitted back to the system's main processor in real time. By objectively measuring the actual physical effects produced by multi-sensory feedback on the user, the subjective "user experience" is transformed into quantifiable and analyzable "interaction data," providing a real and objective input signal for the system to evaluate its own output effect.

[0048] The system compares the real-time received actual pressure distribution data (such as average pressure value) and temperature change curve with the predicted contact mechanical parameters (such as predicted pressure and predicted temperature change trends) output by the prediction module in step two for the same interactive event, performing point-by-point or feature-segment comparisons. Based on a preset algorithm (such as calculating root mean square error or trend correlation), the system outputs a quantified sensory consistency verification signal. This signal characterizes the degree of matching between predicted perception and actual perception. By comparing the "expected" and "actual" sensory data, the system can automatically diagnose the rendering accuracy of this multi-sensory mixed feedback. The generated verification signal serves as the direct decision-making basis for triggering subsequent adaptive optimization algorithms, giving the system the foundation for self-evaluation and continuous optimization capabilities.

[0049] Specifically, the steps for adaptively updating the attributes of the intelligent meta-object based on the monitored actual interaction data are as follows: If the sensory consistency verification signal is received, and the verification signal indicates that the actual interaction force continuously exceeds the predicted value by a threshold, then the interaction is determined to be an enhanced contact. Based on the mechanical data of the enhanced contact, calculate a property wear increment; Based on the attribute wear increment, the surface material spectrum of the target intelligent meta-object stored in the global state management list is attenuated and corrected. The corrected surface material spectrum is saved as a new attribute state of the smart meta-object for use in generating tactile waveforms for subsequent interactions.

[0050] It should be noted that the attribute update module continuously listens for and receives the sensor consistency verification signal stream from the sensor monitoring closed loop. Internally, this module maintains a state machine and an analysis window. When it detects that the actual measured pressure value represented by the verification signal is consistently higher than the predicted pressure value output by the prediction module for several consecutive analysis cycles, and the excess exceeds a preset fixed threshold, the state machine is triggered. This formally records and marks the complete interaction event as an "enhanced contact" event. By performing threshold judgment on the mechanical signal that continuously deviates from the prediction, the system can automatically identify "high-intensity" interaction events from ordinary interactions that may cause physical impact on virtual objects, providing clear event triggering conditions and classification criteria for subsequent attribute evolution calculations.

[0051] The system extracts key mechanical parameters, such as average pressure, peak pressure, and total duration, from the complete sensor data log of this event, which was marked as "intensified contact." Subsequently, the system calls a predefined attribute wear model, using the mechanical data of this event as input. Based on a simplified physical wear law (e.g., wear is positively correlated with the product of pressure and time), this model calculates a unitless value characterizing the degree of wear caused by this interaction, namely the "attribute wear increment." This transforms a specific, high-intensity physical interaction into a standardized "damage value" that can be applied to the digital attributes of an object, establishing a computable and repeatable mathematical model for the physical evolution of the virtual world.

[0052] Based on the calculated attribute wear increment, the system accesses the global state management list, locates the data record of the target intelligent meta-object, and positions its "Surface Material Spectrum" attribute field. The system applies a decay function to this spectrum data (usually a set of values ​​representing the intensity of different frequency components). This function proportionally reduces the intensity of the high-frequency components representing surface micro-roughness in the spectrum according to the magnitude of the wear increment, thus digitally simulating the physical process of an object's surface being "smoothed" by friction. This directly applies the quantified interaction effect (wear increment) to the object's essential physical properties (material spectrum), realizing a physically intuitive change in the virtual object's intrinsic state based on its real-world usage history, giving the object dynamic "life" characteristics.

[0053] After completing the attenuation correction calculation for the surface material spectrum, the system writes the updated spectrum data back to the original storage location of the intelligent meta-object in the global state management list, overwriting the old spectrum data. From then on, this attribute of the intelligent meta-object is permanently updated. When any user interacts with this object again in the future, the material basis used to synthesize tactile feedback, queried by the haptic waveform generation network in step three, will no longer be its original smooth state, but rather its state after this wear correction. This ensures that every change in the object's state is remembered and solidified by the system, making the evolution of the virtual environment continuous and consistent. This allows the interaction history to genuinely influence future sensory experiences, constructing a living virtual world with a time dimension and cumulative changes, fundamentally enhancing the depth and realism of immersion.

[0054] This embodiment also provides a multi-sensory hybrid interaction system for immersive virtual reality experiences, including: The meta-object loading module loads a virtual environment containing smart meta-objects, which encapsulate physical attributes and feedback rules. The interaction prediction module captures users' multimodal interaction data in real time and predicts upcoming interaction events; The strategy generation module generates corresponding multi-sensory feedback strategies based on the predicted interactive events and their associated intelligent meta-object attributes. The synchronization orchestration module performs timing orchestration based on the multi-sensory feedback strategy and the delay characteristics of each feedback channel to form synchronous control commands; The execution monitoring module executes synchronous control commands, drives multi-sensory hardware output feedback, and monitors actual interaction data; The attribute update module adaptively updates the attributes of smart meta objects based on the monitored actual interaction data.

[0055] This embodiment also provides a computer device applicable to a multi-sensory hybrid interaction method for immersive virtual reality experiences, comprising: a memory and a processor; the memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions to implement the multi-sensory hybrid interaction method for immersive virtual reality experiences as proposed in the above embodiment.

[0056] The computer device can be a terminal, comprising a processor, memory, communication interface, display screen, and input devices connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The communication interface is used for wired or wireless communication with external terminals; wireless communication can be achieved through Wi-Fi, carrier networks, NFC (Near Field Communication), or other technologies. The display screen can be an LCD screen or an e-ink screen. The input devices can be a touch layer covering the display screen, buttons, a trackball, or a touchpad on the computer device's casing, or an external keyboard, touchpad, or mouse.

[0057] This embodiment also provides a storage medium storing a computer program that, when executed by a processor, implements the multi-sensory hybrid interaction method for immersive virtual reality experiences as proposed in the above embodiments. The storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as Static Random Access Memory (SRAM), Electrically Erasable Programmable Read-Only Memory (EEPROM), Erasable Programmable Read Only Memory (EPROM), Programmable Red-Only Memory (PROM), Read-Only Memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.

[0058] In summary, this invention constructs a closed-loop interactive system integrating perception, decision-making, execution, and evolution by introducing intelligent meta-objects with encapsulated attributes, generating predictive dynamic feedback strategies, and implementing cross-modal temporal orchestration. This fundamentally changes the rigidity of the traditional "event-response" architecture. The system can dynamically synthesize and weave multi-sensory feedback signals based on the user's real-time interaction intentions and the embedded physical states of virtual objects, achieving a transformation from isolated sensory overlay to cross-modal perception fusion. Through forward-looking prediction and hardware latency compensation, the system effectively improves the temporal synchronization accuracy and perceptual consistency of multi-channel feedback, including visual, tactile, and olfactory feedback. Simultaneously, the system can continuously evolve and update the attributes of meta-objects based on actual interaction data, enabling the virtual environment to possess dynamic response and memory capabilities. This significantly enhances the naturalness of the interaction, the coherence of immersion, and the realism and adaptability of the virtual world.

[0059] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.

Claims

1. A multi-sensory mixed interaction method for immersive virtual reality experiences, characterized in that, Includes the following steps: Load a virtual environment containing smart meta-objects, which encapsulate physical properties and feedback rules; Capture users' multimodal interaction data in real time and predict upcoming interaction events; Based on the predicted interaction events and their associated intelligent meta-object attributes, generate corresponding multi-sensory feedback strategies. Based on the multi-sensory feedback strategy and the delay characteristics of each feedback channel, timing is arranged to form synchronous control commands; Execute synchronous control commands, drive multi-sensory hardware output feedback, and monitor actual interaction data; The attributes of the smart meta-object are adaptively updated based on the actual interaction data monitored.

2. The multi-sensory mixed interaction method for immersive virtual reality experiences of claim 1, wherein: The loading of the virtual environment includes intelligent meta-objects, which encapsulate physical attributes and feedback rules. The specific steps are as follows: Read the topology of the virtual environment and the identifiers of all objects from the scene database; Based on the identifier, a corresponding data entity is instantiated from the intelligent meta-object library. Each data entity contains at least a geometric model, surface material spectrum, thermodynamic coefficient, and odor feature encoding. Register all instantiated smart meta-objects to the global state management list according to their spatial coordinates; Based on the physical attributes of each object in the global state management list, a base sensory field covering the entire virtual space is initialized. The base sensory field includes a preliminary temperature gradient distribution and environmental odor background.

3. The multi-sensory hybrid interaction method for immersive virtual reality experiences of claim 2, wherein: The specific steps for capturing users' multimodal interaction data in real time and predicting upcoming interaction events are as follows: The device uses built-in sensors to acquire the three-dimensional coordinates of the user's gaze point and pupil change data. The continuous position, motion velocity vector, and finger bending angle of each joint of both hands are obtained through a hand tracker. When no physical collision is detected, the baseline pressure micro-changes of the fingertip pressure sensors of the haptic glove are continuously read; The gaze point coordinates, motion velocity vector, and baseline pressure micro-changes are input into a trained temporal prediction model, which outputs one or more probabilistic event descriptions of specific types of contact between a user's limb and a specific intelligent meta-object within a predetermined future time window.

4. The multi-sensory mixed interaction method for immersive virtual reality experiences of claim 3, wherein: The specific steps for generating corresponding multi-sensory feedback strategies based on predicted interaction events and their associated intelligent meta-object attributes are as follows: The probabilistic event description is parsed to extract the target intelligent element object identifier, predicted contact location, and contact mechanical parameters; Query the global state management list to obtain the surface material spectrum and odor feature encoding currently stored in the target intelligent meta-object; The surface material spectrum and the contact mechanical parameters are input into a tactile waveform generation network, which outputs a driving signal waveform that simulates the tactile sensation of the material under the contact parameters. Based on the odor feature code and the contact type, a cross-sensory modulation mapping table is queried to retrieve the corresponding modulating odor formula number and intensity modulation coefficient, which together constitute the multi-sensory feedback strategy.

5. The multi-sensory mixed interaction method for immersive virtual reality experiences of claim 4, wherein: Based on the multi-sensory feedback strategy and the delay characteristics of each feedback channel, timing is arranged to form synchronous control commands. The specific steps are as follows: Receive the tactile drive signal waveform, the adjustable odor formula number, and the intensity modulation coefficient; Query the hardware latency database to obtain the current fixed response latency values ​​of the haptic actuator and odor release device; Based on the display time of the visual rendering frame, the start playback command of the tactile drive signal waveform is sent to the tactile driver in advance by a first offset. The command to release the regulatory odor is sent to the odor controller ahead of time by a second offset, wherein the second offset is greater than the first offset; All time-calibrated hardware control commands are packaged to generate the synchronization control instruction sequence.

6. The multi-sensory hybrid interaction method for immersive virtual reality experiences as described in claim 5, characterized in that: The specific steps for executing synchronization control commands, driving multi-sensory hardware to output feedback, and monitoring actual interaction data are as follows: The synchronization control command sequence is distributed to the corresponding hardware communication interface; The tactile actuator plays the drive signal waveform according to the instruction timestamp, and the odor controller activates the designated odor capsule and controls the release concentration according to the instruction. During the feedback output process, the pressure sensor and temperature sensor continuously collect data on the actual pressure distribution and temperature changes on the user's skin surface; The actual pressure distribution and temperature change data are compared with the predicted contact mechanical parameters in real time to generate a sensory consistency verification signal.

7. The multi-sensory mixed interaction method for immersive virtual reality experiences of claim 6, wherein: The specific steps for adaptively updating the attributes of the intelligent meta-object based on the monitored actual interaction data are as follows: If the sensory consistency verification signal is received, and the verification signal indicates that the actual interaction force continuously exceeds the predicted value by a threshold, then the interaction is determined to be an enhanced contact. Based on the mechanical data of the enhanced contact, calculate a property wear increment; Based on the attribute wear increment, the surface material spectrum of the target intelligent meta-object stored in the global state management list is attenuated and corrected. The corrected surface material spectrum is saved as a new attribute state of the smart meta-object for use in generating tactile waveforms for subsequent interactions.

8. A multi-sensory mixed interaction system for immersive virtual reality experience, based on the multi-sensory mixed interaction method for immersive virtual reality experience according to any one of claims 1-7, characterized in that: include, The meta-object loading module loads a virtual environment containing smart meta-objects, which encapsulate physical attributes and feedback rules. The interaction prediction module captures users' multimodal interaction data in real time and predicts upcoming interaction events; The strategy generation module generates corresponding multi-sensory feedback strategies based on the predicted interactive events and their associated intelligent meta-object attributes. The synchronization orchestration module performs timing orchestration based on the multi-sensory feedback strategy and the delay characteristics of each feedback channel to form synchronous control commands; The monitoring module executes synchronous control commands, drives multi-sensory hardware output feedback, and monitors actual interaction data. The attribute update module adaptively updates the attributes of smart meta objects based on the monitored actual interaction data. 9.A computer device, comprising a memory and a processor, wherein the memory stores a computer program, and the computer device is characterized in that: When the processor executes the computer program, it implements the steps of the multi-sensory hybrid interaction method for immersive virtual reality experience as described in any one of claims 1 to 7.

10. A computer readable storage medium having stored thereon a computer program, characterized in that: When the computer program is executed by the processor, it implements the steps of the multi-sensory hybrid interaction method for immersive virtual reality experience as described in any one of claims 1 to 7.