Multi-mode sensing cooperation and real-time synchronization system

By using a multimodal perception collaboration and real-time synchronization system, the problem of independent control of multimodal perception channels in virtual reality systems has been solved, realizing the synchronous output of multi-sensory information, improving the realism and consistency of the immersive experience, and eliminating sensory conflicts.

CN121349296APending Publication Date: 2026-01-16BEIJING FANTASY PAI SHUSHI TECH CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202511417919.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-30
Publication Date
2026-01-16

AI Technical Summary

Technical Problem

In existing virtual reality systems, the independent control of multimodal perception channels leads to differences in time reference and a lack of spatial mapping, resulting in sensory fragmentation and motion sickness for users, and failing to provide realistic multisensory feedback and an immersive experience.

Method used

A multimodal perception collaboration and real-time synchronization system is adopted. Through a layered architecture of perception layer, control layer and execution layer, a unified perception quantization model is established to realize dynamic weight allocation and conflict resolution of multimodal data. A high-precision time synchronization mechanism is adopted to ensure that the signal delay across devices is in the millisecond level.

Benefits of technology

It achieves the synchronous output of multi-sensory information, enhances the realism and consistency of the immersive experience, eliminates sensory conflicts, reduces the risk of motion sickness, and provides a natural and integrated experience of multi-sensory feedback.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121349296A_ABST
    Figure CN121349296A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-mode perception collaboration and real-time synchronization system, which breaks through the limitation of vision / hearing in the prior art, covers vision, hearing, smell and movement, and realizes more accurate cooperation through a perception layer, a control layer and an execution layer. According to the invention, an immersive experience space full of shock and reality can be created, synchronous output of visual, auditory, tactile and olfactory signals is realized by means of high-precision equipment, and a series of cross-sensory signal superposition schemes are designed on the basis of response characteristics of human physiology and psychology to various sensory stimuli. Therefore, the overall perception of the experiencer to the space atmosphere is enhanced. Meanwhile, through weight distribution and conflict resolution, the system can flexibly adjust the sensing mode according to the environment and the user state. The multi-device cooperation error can be controlled at millisecond level and millimeter level, and the natural fusion of virtuality and reality is realized.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to a multi-modal perception coordination and real-time synchronization system. BACKGROUND

[0002] With the rapid development of virtual reality (VR), augmented reality (AR) and mixed reality (MR) technologies, immersive experience capsules have been widely used in travel and entertainment, popular science education, industrial simulation and other scenarios. The existing technology generally uses a single or a few modalities (vision, hearing) for information presentation, and outputs the content through a head-mounted display device or a fixed projection system. However, the above-mentioned existing solutions have the following obvious shortcomings in practical application:

[0003] Firstly, in terms of interaction and experience, the existing system relies heavily on wearable devices such as head-mounted displays (VR glasses) and data gloves, which brings significant burden and bondage to the user, limiting its application in long-time and large-scale activity scenarios. More importantly, this visual-dominated interaction mode cannot provide real tactile, force sense and other multi-sensory feedback, and the user experience stays at the "watching" level, making it difficult to have natural interaction with a real object (such as touching a virtual artifact). In addition, the immersive experience is isolated from the outside world, and the experimenter cannot easily capture and share the highlights in the virtual scene in the form of photos or videos, which weakens the technology's communicability and social attributes. At the same time, the existing virtual content lacks high-precision dynamic matching and registration capability with the real physical environment in which the user is located, resulting in a harsh fusion of virtual objects and real scenes, and a "out-of-play" feeling, making it difficult to achieve true immersion.

[0004] Secondly, in terms of system and technology, the more fundamental defect is that the existing multi-modal interaction system usually controls the visual, auditory, tactile, olfactory and even motion platform as independent units. Each channel often has independent processing threads and clocks, and lacks a unified, high-precision time synchronization reference and integrated space mapping and rendering engine. This architectural fragmentation leads to millisecond-level or even higher-level delay differences in the transmission, processing and feedback of different sensory information. For example, the moment the user sees the virtual impact and the moment the body feels the vibration feedback are not synchronized, or the sound field orientation update of the hearing after the head turns lags behind the visual picture update. This mismatch of multi-sensory information in space and time seriously damages the integrity and reality of the immersive experience, causes sensory conflicts for the user, and even leads to motion sickness, becoming a major technical bottleneck for improving user experience.

[0005] Therefore, it is particularly important to provide an improved multi-modal perception coordination and real-time synchronization system based on the deficiencies of the existing technology. SUMMARY

[0006] This invention proposes a multimodal perception collaboration and real-time synchronization system, which aims to solve problems such as perception distortion, feedback delay and multimodal conflict in the process of virtual and reality integration, and improve immersion and realism.

[0007] The system consists of a perception layer, a control layer, and an execution layer.

[0008] The perception layer collects multimodal data including vision, hearing, smell, and motion, and establishes a unified perception quantification model based on human physiological and psychological mechanisms. The model covers four major modules: vision, hearing, smell, and motion, and can accurately characterize the perceptual features of human senses regarding space, color, sound, smell, and motion.

[0009] The control layer uses a central control system and a Bayesian inference model to achieve dynamic weight allocation and conflict resolution of multimodal data; a high-precision time synchronization mechanism is used to ensure that cross-device signal delay is controlled at the millisecond level, thereby maintaining the coordination and consistency of virtual and real integration.

[0010] The execution layer, including projection displays, a six-degree-of-freedom motion platform, panoramic sound equipment, and odor generators, can generate corresponding multi-sensory feedback based on commands from the control layer. The system maintains a high degree of consistency in spatial positioning, sound field direction, color brightness, and odor diffusion, ensuring a seamless transition between the virtual and real environments.

[0011] In some embodiments, this application unifies and coordinates visual, auditory, olfactory, and motion perception in real time, and achieves a high degree of consistency between virtual content and the real environment in terms of geometric perspective, resolution, color, sound field, and motion through cross-modal synchronization and error control. The visual, auditory, olfactory, and motion perception modules all employ mechanisms based on a combination of physiological perception and psychological cognition.

[0012] In some embodiments, the visual perception modeling module is based on the physiological and psychological mechanisms of human spatial perception, including:

[0013] Visual physiological modeling, based on the geometric perspective laws of the human eye, minimum resolving angle (MAR), brightness perception, and color and color temperature matching principles, establishes a quantitative model to ensure that virtual objects are consistent with the real environment in terms of proportion, perspective, resolution, motion blur, brightness contrast, light source direction, shadows and colors.

[0014] Visual psychological modeling, combined with depth cues such as binocular parallax, motion parallax, and texture gradient, employs Gestalt psychology principles (continuity, closure, shared fate, etc.) to optimize scene contours and group perception, further enhancing immersion.

[0015] Through these mechanisms, virtual images achieve an effect that is difficult to distinguish from the real environment in terms of geometry, brightness, and color, ensuring visual consistency and a sense of naturalness.

[0016] In some embodiments, the auditory perception modeling module includes:

[0017] Auditory physiological modeling is employed to establish a spatial localization model based on binaural time difference (ITD) and binaural intensity difference (ILD). By combining critical bandwidth and equal loudness curves, the virtual sound source is made consistent with the real sound source in terms of direction, spectrum, and loudness. At the same time, a reverberation and sound field consistency model is introduced to ensure that the sound source matches the acoustic characteristics of the environment.

[0018] Auditory psychological modeling enhances the attentional discrimination of target sound sources by simulating the "cocktail party effect" through a sound source separation model; it ensures the stability of auditory experience in different environments through loudness constancy and timbre constancy; at the same time, it can adjust rhythm and frequency distribution to achieve psychological guidance of emotions and atmosphere.

[0019] Virtual sound is consistent with real sound in terms of direction, sound quality, loudness, and spatial sense, and can create a specific psychological atmosphere according to the needs of the scene.

[0020] In some embodiments, the olfactory perception modeling module includes a basic odor database and an odor generator, configured as follows:

[0021] Olfactory psychological modeling, through odor databases and perception matching mechanisms, combined with users' odor associations and contextual memories, enhances the user's sense of immersion in the environment.

[0022] Based on the physiological mechanism of "odor molecule-receptor-signal conversion", an olfactory physiological model was designed, and a multi-channel odor generator and a precise control module were designed. It can combine and release a variety of odor molecules under scene triggering to simulate environmental odors.

[0023] After testing, this application can realistically reproduce the atmosphere and smell in specific scenarios such as the deep sea and space, expanding the immersive experience from a single sense to a chemical sense.

[0024] In some embodiments, the motion-aware modeling module includes:

[0025] Motion physiology modeling, through a six-degree-of-freedom platform and a screen object displacement model, realizes the spatial mapping between the motion of virtual objects and the user's body motion, ensuring the natural correspondence between displacement, velocity and acceleration.

[0026] Motor psychological modeling, using motion constancy and body proprioception model, ensures that users do not experience distortion or dizziness during dynamic processes.

[0027] Provide feedback that is naturally coupled with body movements in virtual scenes to enhance the realism of motion interaction.

[0028] In some embodiments, the visual, auditory, olfactory, and kinesthetic models are coupled and integrated:

[0029] The various perception models are integrated within the Pandora central control system. The system dynamically assigns weights to multimodal data using a Bayesian inference model: in brightly lit scenes, visual weights are increased; in noisy environments, auditory weights are decreased while tactile and visual compensations are enhanced; and in specific scenarios, olfactory weights can be increased as context-triggered factors.

[0030] The control layer also implements a conflict resolution mechanism: when different sensory signals contradict each other (such as visual cues and motion feedback not matching), the system will automatically adjust the low-confidence signal to ensure the consistency of the multimodal experience.

[0031] A precise time protocol is used to ensure signal synchronization across devices, with a delay of ≤20ms, thus ensuring real-time performance.

[0032] The significant advancements of this invention lie in its multimodal perception collaboration and real-time synchronization system. This system overcomes the limitations of existing technologies that are solely visual / auditory, encompassing vision, hearing, smell, and motion. Through a perception layer, control layer, and execution layer, it achieves relatively precise coordination. In practical applications, users will no longer rely on a single sensory input to perceive their environment. Instead, they will comprehensively construct their perception of the environment through the synchronous interaction of multiple senses (such as vision, hearing, smell, and motion), significantly enhancing the realism and consistency of the immersive experience. This invention creates a stunning and realistic immersive experience space. It not only relies on high-precision equipment to achieve synchronous output of visual, auditory, tactile, and olfactory signals but also designs a series of cross-sensory signal superposition schemes based on the physiological and psychological response characteristics of humans to sensory stimuli, thereby enhancing the user's overall perception of the spatial atmosphere. Simultaneously, through weight allocation and conflict resolution, the system can flexibly adjust the perception mode according to the environment and user state. Multi-device collaboration errors are controllable at the millisecond and millimeter levels, achieving a natural fusion of virtual and reality. Attached Figure Description

[0033] Figure 1 This is a schematic diagram of the perception layer in this exemplary embodiment.

[0034] Figure 2 This is a schematic diagram of the control layer of this exemplary embodiment.

[0035] Figure 3 This is a schematic diagram of the execution layer of this exemplary embodiment. Detailed Implementation

[0036] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. The following description of at least one exemplary embodiment is merely illustrative and is in no way intended to limit the present invention or its application or use. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort are within the scope of protection of the present invention.

[0037] It should be noted that in the claims and specification of this patent, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one" does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.

[0038] In existing technologies, virtual reality, augmented reality, and mixed reality are driving the application of immersive experience cabins in multiple fields. Traditional systems primarily rely on single or a few sensory channels for information presentation, such as outputting visual content through head-mounted displays or fixed projection systems. However, these systems typically control visual, auditory, olfactory, and kinesthetic sensory channels independently, lacking a unified temporal reference and spatial mapping mechanism, resulting in significant delay differences between different modal signals. For example, in a virtual reality experience cabin, the projected image and the feedback from the motion platform may not be synchronized, and users may experience a mismatch between visual and vestibular information, leading to sensory disconnect and discomfort.

[0039] To address the aforementioned issues, the inventors first observed the time reference discrepancies and spatial mapping deficiencies caused by independent control of multimodal sensory channels. By analyzing the physiological mechanisms of multisensory integration in the human body, they discovered dynamic weight allocation and conflict resolution mechanisms at the neural processing level for different sensory modalities. Based on this, they proposed constructing a hierarchical architecture to achieve a closed-loop link for data acquisition, fusion control, and collaborative execution. Specifically, the approach is as follows: at the sensory layer, a multimodal quantification model conforming to physiological mechanisms is established to ensure the biological rationality of data acquisition; at the control layer, a probabilistic inference model is introduced to handle multi-source data conflicts, while simultaneously establishing a global clock synchronization mechanism; at the execution layer, multiple types of feedback devices are designed for collaborative response, forming a unified sensory output.

[0040] Therefore, this application proposes a multimodal perception collaboration and real-time synchronization system, including a perception layer, a control layer and an execution layer.

[0041] The perception layer is configured to collect visual, auditory, tactile, and motion perception data. Based on the physiological and psychological mechanisms of human perception, a multimodal perception quantification model is established. The model includes a visual perception modeling module, an auditory perception modeling module, an olfactory perception modeling module, and a motion perception modeling module.

[0042] The control layer includes a central control system configured to achieve dynamic weight allocation and conflict resolution of multimodal sensing data through a Bayesian inference model, and to achieve cross-device signal synchronization using a precise time protocol.

[0043] The execution layer is configured to drive the projection system, the six-degree-of-freedom platform, the immersive sound device, and the odor generator, and respond to control layer commands to generate collaborative feedback.

[0044] Specifically, the perception layer is primarily responsible for collecting raw data from different directions, providing the foundational information for subsequent processing and control. This layer includes visual perception, auditory perception, tactile and motion perception, and environmental perception. Next, we will introduce the devices used for each perception layer in this project and their main functions.

[0045] 1. Visual perception

[0046] Equipment configuration: 4K laser projector and optical waveguide holographic equipment.

[0047] Main functions: Using a 4K laser projector, high-precision reproduction of real scenes or virtual images can be achieved under high brightness and wide color gamut conditions, ensuring the clarity of details and color consistency of the projected image, meeting the demand for realistic presentation of complex textures and dynamic scenes. Simultaneously, the waveguide holographic device supports multi-angle holographic display and immersive visual interaction, allowing observers to obtain a stable and three-dimensional visual experience from different viewpoints. The combination of these two technologies is used for multi-dimensional, multi-scenario visual simulation and experiments, supporting the construction of an integrated sensory model at the visual level.

[0048] 2. Auditory perception

[0049] Equipment configuration: Equipped with a 64-channel speaker array and a Head Related Transfer Function (HRTF) microphone array. Main functions: Locating sound sources from all directions for precise capture of ambient sounds; eliminating environmental noise to ensure the system accurately recognizes and processes the user's voice commands.

[0050] Function Description: The auditory module ensures that the user's voice commands can be clearly recognized, and at the same time lays the foundation for immersive sound effects output, so that sound effects and visuals are synchronized to jointly enhance the user experience.

[0051] 3. Tactile and kinesthetic perception

[0052] Equipment configuration: Six-degree-of-freedom platform and Buttkicker LFE vibrator.

[0053] Main functions: Real-time monitoring of the platform's posture to capture the user's movement and position changes; transmitting motion data through vibration feedback to help users feel physical feedback during interaction.

[0054] Function Description: Through tactile and motion sensing, it not only provides users with a multi-sensory interactive experience, but also ensures that the system can quickly respond to the user's physical actions when the user operates, thus enhancing the natural interaction effect.

[0055] 4. Environmental perception

[0056] Equipment configuration: Equipped with a CNC odor generator, infrared temperature control module and humidity sensor.

[0057] Its main functions include monitoring odor concentration in the environment and adjusting odor release in a timely manner; controlling local temperature gradients to ensure smooth temperature changes; and providing feedback on changes in ambient humidity to support dynamic adjustment of environmental parameters. Function Description: The environmental perception module enables a more comprehensive immersive experience, responding in real-time to visual, auditory, and tactile sensations, and adjusting odor and temperature as needed to further enrich the user experience.

[0058] The control layer is responsible for integrating, processing, and making decisions on various information from the perception layer, and distributing the final instructions to the execution layer. Therefore, the control layer mainly has the following three functions: Rapid processing and response. The system uses a combination of Unreal Engine 5 (UE5) and an NVIDIA RTX 6000Ada GPU to achieve real-time image rendering at high resolutions (e.g., 4K) and high refresh rates (e.g., 120 frames per second). This ensures a smooth and realistic visual experience for the user during interaction. Furthermore, NVIDIA's PhysX technology is used to accurately calculate the motion of the six-degrees-of-freedom platform. This means that the acceleration error of the platform during movement can be controlled within 0.05G, ensuring smooth and accurate motion. Integration of multi-source information. First, data from different sensors is synchronized through a Precise Time Protocol (PTP) to ensure they are consistent in time. For example, visual and auditory data need to be processed at the same point in time to provide accurate environmental perception. Second, the ICP (Iterative Closest Point) algorithm is used to match the coordinates of the virtual environment with those of the real world. This facilitates the precise overlay of virtual objects onto the real environment in augmented reality (AR) or virtual reality (VR) applications. The system also implements efficient command transmission, using the EtherCAT bus protocol to transmit control signals to the six-degrees-of-freedom platform. EtherCAT is renowned for its high-efficiency communication capabilities, completing data transmission within a cycle as short as 1 millisecond, ensuring real-time platform response to commands. This enables motion control. Building upon this, the DMX512 protocol is used to control devices such as lights and scent generators for environmental control. DMX512 is a widely used protocol for stage lighting control, capable of updating data across 512 channels in approximately 23 milliseconds, equivalent to a refresh rate of approximately 44 times per second, ensuring instantaneous changes in environmental effects. For hard real-time communication, EtherCAT and PTP protocols are employed to ensure rapid transmission of critical data. For soft real-time communication, ROS2 and NDI protocols are used to achieve data fusion and video transmission, meeting the system's processing needs for different types of data.

[0059] Through the collaborative work of the aforementioned components, the control layer can efficiently process a large amount of information from the perception layer, make intelligent decisions, and quickly transmit instructions to the execution layer, ensuring the smooth operation of the entire system and improving the user experience.

[0060] The execution layer is responsible for translating the specific instructions issued by the control layer into actual operations and presenting the final effect to the user. This layer includes four execution units: visual, sound field, motion, and environment.

[0061] The visual execution unit utilizes the 7thSense Delta system to stitch and correct the images from multiple projectors, ensuring seamless image transitions without noticeable boundaries. Simultaneously, it receives volumetric video streams from the UE5 engine via a waveguide holographic device, displaying them at a 1920x1080@90Hz refresh rate to deliver a immersive, three-dimensional holographic effect. The sound execution unit leverages a Dolby Atmos processor to analyze sound field coordinates in real time, accurately distributing digital audio signals to a 64-channel speaker array, providing an immersive, three-dimensional auditory experience. The motion and environment execution unit uses a six-degrees-of-freedom platform to perform specific pose adjustments based on calculations from the physics engine, supporting tilt angles of ±20° and acceleration control within 0-1G to meet the precise motion requirements of complex dynamic scenes. Furthermore, the infrared temperature control module can adjust the local temperature according to preset scene requirements, maintaining a temperature range between 10-50℃ with an error within ±1℃, ensuring consistency between the environmental atmosphere and the visual and auditory effects. The execution layer, acting as the actual executor of the response, transforms the deeply processed instructions from the control layer into concrete visual, auditory, motion, and environmental changes, ensuring that the user can directly perceive the various responses made by the system. This top-down, progressive design achieves a complete closed loop from data acquisition and real-time analysis to final feedback display, ensuring the entire system operates collaboratively and efficiently.

[0062] Furthermore, this application utilizes data flow management as a key element to ensure smooth information flow and timely response between system layers. The entire data flow is divided into two directions: uplink data flow and downlink data flow. Both are ensured through dedicated communication protocols and buffering mechanisms to guarantee rapid information transmission and error handling in case of data loss.

[0063] Uplink data flow refers to the transmission of data from the perception layer to the control layer. The perception layer collects data from various sources in real time, such as participant location data obtained using a UWB (Ultra-Wideband) system, and environmental status data such as temperature and odor concentration in the field. This data is uploaded to the data fusion center via ROS2 nodes at a frequency of no less than 100 times per second to ensure real-time performance and continuity. A unified time synchronization protocol ensures that data from all sensors can be processed at the same time, avoiding information misalignment. In addition, we have used a preset caching mechanism and error checking algorithm to ensure that information is not lost as much as possible during transmission.

[0064] Downlink data flow refers to the transmission of information from the control layer to the execution layer. After data fusion and intelligent decision-making, the control layer transmits different types of operation commands to the execution layer, including real-time image data commands generated by the UE5 rendering engine, precise motion parameters transmitted to the six-degrees-of-freedom platform via the EtherCAT bus, and sound field coordinate commands for the Dolby Atmos audio system to adjust the sound field layout in real time. These commands are distributed according to their importance and urgency; emergency stop or fault response commands are given the highest priority to ensure the system can respond in the shortest possible time. All commands are transmitted by the control layer according to pre-defined rules through hard real-time and soft real-time channels to ensure response time and data accuracy.

[0065] Through the above technical solution, this application effectively solves the latency difference problem caused by independent control of multimodal perception channels, achieving millisecond-level synchronization of visual projection, motion feedback, sound field reconstruction, and odor release. This application can create a multimodal perception collaborative system that seamlessly integrates digital content and the physical environment, optimizing the synchronization and collaborative response of sensory feedback in immersive experiences. In practical applications, users will no longer rely on a single sensory input to perceive the environment, but will comprehensively construct their perception of the environment through multiple sensory interactions (such as vision, hearing, smell, and motion), significantly improving the realism and consistency of the immersive experience.

[0066] This application further proposes a visual perception modeling module based on the physiological and psychological mechanisms of human eye spatial perception. This model starts from multiple dimensions such as geometric perspective, resolution, lighting, motion, and color consistency, and combines physical modeling with the laws of human eye perception to ensure a high degree of consistency and immersion between virtual and reality.

[0067] To achieve a natural visual integration of virtual images and the real environment, this application establishes a visual sensory model. This model considers multiple dimensions, including geometric perspective, resolution, lighting, motion, and color consistency, and combines physical modeling with the principles of human visual perception to ensure a high degree of consistency and immersion between the virtual and real worlds.

[0068] For the geometric and perspective reconstruction of virtual parts in the visual representation, this application ensures that the size, proportion, and perspective of virtual objects remain consistent with real objects at different viewpoints through precise spatial mapping relationships. In order to calculate the perspective and motion performance of objects in a virtual scene, it is first necessary to establish the world coordinates (X, Y, Z) of the object relative to the observer's viewpoint, and use these coordinates in subsequent calculations.

[0069] The key point in analyzing and calculating perspective effects lies in matching the human eye's visual perspective, i.e., θ. real =θ screen This allows us to obtain the perspective ratio of real objects in a virtual scene:

[0070]

[0071] h is the height of the object displayed on the screen (this can also be applied to the object's length and width), D eye Z is the distance from the viewer's eye to the screen, Z is the distance from the viewer's eye to the object (including the object's depth in the image), and H is the object's actual height. px This refers to the pixel height of the object. When the object is tilted, this application uses a three-dimensional projection matrix to project the four corner points of the object onto a two-dimensional screen, calculates the reprojection error separately, and ensures that the error is less than a set threshold (such as 5px or 0.5% of the screen width), thereby eliminating perspective distortion.

[0072] For the resolution module, this invention uses the minimum resolvable angle (MAR) of the human eye as the criterion. Typically, MAR is 1 arcminute (1 / 60 degree). To avoid pixelation, it is essential to ensure that the angle subtended by each pixel in the human eye does not exceed this threshold. The minimum viewing distance formula is as follows:

[0073]

[0074] Where p represents the spacing between individual pixels, which can be calculated from the screen resolution and physical dimensions. During testing, this application displayed high-contrast resolution test images at different viewing distances. The final results showed that α ≤ 1.03' for static, high-brightness images and α ≤ 1.07' for dynamic, complex images. ′ When playing simple dynamic scenes, α≤1.22'.

[0075] For the object motion module, to ensure that the motion of virtual objects is consistent with that of real objects, this invention also establishes a screen object displacement model. First, the actual motion of the object needs to be converted into pixel velocity:

[0076]

[0077] Then we can calculate the pixel displacement for each frame:

[0078]

[0079] Among them, f r This is the video frame rate, and T is the duration of each frame. Based on testing, this application concludes that when Δ... pix When 1 ≤ Δ, the video only needs to be rendered normally and kept clear. When 1 < Δ pix When the value is ≤3, motion realism can be improved by adding motion trails. The blur length can be calculated using the formula: blur_pixels=V pix ·t e , where t eThis refers to the exposure time. When Δ pix When the value is greater than 3, frame interpolation and dynamic compensation algorithms need to be enabled, and stronger motion blur processing that matches the direction of motion and scaling needs to be applied to avoid distortions such as ghosting, frame skipping, or violations of perspective rules, so as to ensure that high-speed moving objects can still be presented smoothly and naturally when the depth changes.

[0080] Besides calculating the geometric motion of objects in the virtual image, the virtual image also needs to maintain a high degree of consistency with the real object in terms of brightness, color, and shadow direction. Therefore, it requires the mutual coupling of brightness module, light source angle and shadow module, and color and color temperature consistency module.

[0081] For the brightness module, at the experimental site, illuminance meters and luminance meters were used to measure the brightness of the screen under different inputs (full white, full black, grayscale) and the illuminance of the screen by ambient light when the projection was off.

[0082]

[0083] Where L screen ρ is the screen's brightness, E is the screen's diffuse reflection characteristic value. proj,avg It is the average illuminance of the system on the screen.

[0084] It's important to note that the screen brightness is composed of two parts, where L... display It refers to the brightness of the screen's own light emission / projection, L ambient It is the brightness of ambient light falling on the screen after diffuse reflection.

[0085] L screen =L display +L ambient

[0086] Therefore, it is necessary to measure the background brightness L of the screen area using a luminance meter with the projection / display off. ambient By adding or subtracting the extra brightness from diffuse reflection when comparing, an effective contrast value can be obtained by comparing "projected white brightness" and "black brightness + ambient brightness," where L... background Ambient light level:

[0087]

[0088] The application found through testing that when the contrast ratio falls within the range of 20–50, the human eye usually has difficulty distinguishing the brightness difference between the virtual image and the real environment; if it exceeds this range, it is necessary to adjust the projection brightness or reduce ambient light interference to achieve consistency between the virtual and real scene brightness.

[0089] In the light source direction and shadow module, a spectrophotometer can be used to measure the direction and relative intensity of light sources at several reference points in the scene, obtaining the direction and proportion of each main light source. Based on the obtained data, the same number of virtual light sources are established in the virtual scene, and the same direction and intensity ratio are set. In the comparative test, by projecting shadows under virtual and real objects, the consistency of shadow direction and length was measured. It was found that when the deviation is less than 2° and the length error is less than 10%, the deviation between the virtual and real scenes is difficult to be perceived by the naked eye.

[0090] For color matching, a colorimeter was used to measure the color of the physical environment (Lab value) and the color of the virtual image (Lab value). The color difference ΔE was calculated by subtracting the two values ​​using the CIEDE2000 color difference formula. In the tests conducted in this application, it was found that when ΔE≤2, most observers could hardly perceive the difference; when ΔE≤3, the color difference was acceptable; and when ΔE>5, the color difference between the virtual and real environments became quite noticeable and required correction.

[0091] To eliminate color temperature differences, a spectrophotometer is used to measure the ambient white point and the projector white point to obtain their color temperatures or xy chromaticity coordinates. If the color temperature difference is less than ±200K and ΔE≤3, it meets the consistency requirements. Otherwise, the white balance needs to be adjusted using the projector or rendering software to match the ambient white point. After calibration, the measurement is repeated until the consistency standard is met.

[0092] In addition to the physiological perception model, this application also constructs a psychological perception model of human vision for virtual and physical content. This method mainly includes four parts: binocular disparity modeling, static and dynamic cue modeling, application of Gestalt psychology principles, and visual constancy modeling, as detailed below:

[0093] The binocular parallax application module is configured to generate slightly different virtual images for the left and right eyes based on the natural distance between human eyes (approximately 6.5 cm), simulating the binocular parallax that occurs during actual observation. Using these differing images, the brain can calculate the spatial depth of objects, thereby generating a realistic stereoscopic visual experience in a virtual scene.

[0094] Static and dynamic cue modeling is based on the principle of linear perspective. In the virtual scene, a vanishing point is preset, and the scene's geometry is adjusted. For example, in an underwater tunnel scene, the tunnel's extension direction is aligned with the perspective of the observation window of the hull, thus enhancing the sense of depth. Furthermore, based on the principle of texture gradient, the texture density of the virtual object's surface is dynamically adjusted, reducing the texture density to 50% at a distance of 10 meters, simulating a realistic visual attenuation effect. Dynamic cue modeling is based on the principle of motion parallax, setting the displacement speed of near-field objects to be greater than that of distant objects. For example, in a scene combining a school of fish and distant mountains, the near-field fish move three times faster than the distant objects, thus enhancing the sense of spatial hierarchy and three-dimensionality.

[0095] The Gestalt psychology principles application module guides users to form a coherent whole when perceiving incomplete information. It is configured to use the principles of closure and continuity in a virtual underwater scene, designing cabin pipes and underwater relics with smooth, continuous lines, allowing users to automatically fill in the missing parts and form a complete structural perception. Through the figure-ground principle, it enhances the contrast and clarity of key objects (such as underwater relics and organisms), making them stand out in complex backgrounds and improving recognizability. Through the proximity principle, it groups similar organisms or coral groups, adjusting the spacing to create a visually dynamic natural community. Based on the principle of shared destiny, it controls schools of fish or whales to move along a unified trajectory in dynamic scenes, thereby enhancing the sense of motion integration and immersive experience.

[0096] Visual constancy modeling further incorporates constancy in size, shape, color, and brightness. Size and shape constancy allow for dynamic scaling of virtual creature models, ensuring they maintain their true size at different distances. Edge enhancement is applied to key structures like porthole frames when the viewpoint is tilted to prevent distortion due to perspective changes. Color constancy allows for pre-defined color correction models for different areas, ensuring stable and undistorted colors of key objects despite changes in ambient light intensity and color temperature; for example, the colors of creatures and relics remain consistent when transitioning from shallow to deep sea scenes. Brightness constancy allows for dynamic matching of virtual scene brightness with water depth and scattering, preventing any sense of incongruity.

[0097] To enhance immersion, this embodiment also incorporates optimizations based on peripheral visual characteristics. Motion blur effects (such as simulating rising bubbles or swimming fish) are introduced within a 30° to 130° field of view, leveraging the high sensitivity of peripheral vision to motion to improve immersion. Furthermore, by combining the 2-meter-wide forward porthole of the cabin with multiple laser projectors for seamless fusion, approximately 90° horizontal field of view is achieved, ensuring image continuity when the user turns their head, further enhancing the immersive experience of the grand scene.

[0098] Through the above technical solution, this application starts from the physiological mechanism of human eye spatial perception, and develops a basic visual model by scientifically utilizing the physiological mechanism of human eye spatial perception to simulate the entire process of human eye imaging. It transforms the physiological mechanism of human eye spatial perception into a complete set of digital visual models, and comprehensively constructs a realistic and seamless visual experience system from optical imaging, binocular parallax, static and dynamic cues to haptic feedback.

[0099] To achieve consistency between virtual sound sources and real sound fields in terms of directionality, spectral characteristics, loudness, and spatial perception, this application proposes a complete auditory physiological perception model. This model consists of a spatial positioning and direction modeling module, a frequency resolution and bandwidth matching module, a loudness and dynamic range modeling module, an environmental acoustic feature modeling module, and a real-time sound field adjustment submodule. Through multi-layered coupling, it achieves a highly realistic auditory immersion experience in the virtual environment.

[0100] In the physiological perception model, to achieve directional consistency between virtual sound sources and real-world sound sources, this application establishes a spatial localization and orientation perception model. This model is based on physical modeling using binaural time difference (ITD) and binaural intensity difference (ILD). By calculating the relative positional relationship between the virtual sound source and the user's ears, it generates binaural difference signals consistent with the real sound field. Specifically, based on the spatial coordinates (X, Y, Z) of the virtual sound source, the propagation path difference Δd between the left and right ears is first calculated, thereby obtaining the time delay. Where c is the speed of sound; and combining the directionality of sound source radiation and propagation attenuation, the sound intensity difference ΔL between the left and right ears is calculated. When the binaural signal generated by the above method is played through headphones or multi-channel speakers, a spatial localization sense consistent with the direction of the real sound source can be achieved. When the binaural difference signal error between the virtual sound source and the real sound source is less than ±5μs (ITD), more than 90% of the subjects cannot distinguish the directionality of the virtual and real sound sources.

[0101] To avoid discrepancies in spectral characteristics between virtual sound sources and the real environment, this application establishes a frequency resolution and bandwidth matching model. Based on the frequency range perceptible to the human ear (approximately 20Hz–20kHz) and critical bandwidth theory, this model performs frequency division filtering and reconstruction on the virtual sound source. Specifically, the spectral distribution of the virtual sound source is obtained through Fourier transform, and frequency bands are divided according to the Bark scale. The energy of each frequency band is weighted and adjusted to ensure consistency with the sound pressure level distribution of the real sound source within the critical band. Test results show that when the sound pressure level difference between each frequency band is less than ±1.5dB, most subjects find it difficult to perceive the frequency difference between virtual and real sounds.

[0102] To ensure that the virtual sound source maintains consistency with the real sound source in terms of loudness and dynamic range, this application establishes a modeling module based on the human ear's loudness perception principles. This module uses equal-loudness curves (ISO 226:2003) to correct the loudness of the virtual sound source. Specifically, a sound level meter is first used to measure the sound pressure level L of the real sound source in each frequency band. p(f), and then calculate its corresponding loudness level N(f). During virtual sound source synthesis, the gain of each frequency band is adjusted according to the loudness level curve to keep the total loudness consistent. In addition, to avoid distortion caused by excessive dynamic range, this application introduces an adaptive compression algorithm to control the instantaneous dynamics of the virtual sound source within the range of 40–60 dB, thereby ensuring that the sound is both natural and comfortable.

[0103] To ensure that the virtual sound source maintains consistency with the real sound field in terms of spatial perception and environmental uniformity, this application establishes a reverberation and sound field consistency model. This module, based on the Head-Related Transfer Function (HRTF) and the Sound Field Impulse Response Function (RIR), establishes a "digital acoustic map" of the virtual scene. The reverberation time (RT60) and early reflection parameters are obtained by collecting the impulse response of reference points within the experimental environment. In reality, sound propagation produces different reflection effects in different environments. For example, in a small, enclosed space, sound waves reflect frequently, producing a brief and rapidly decaying reverberation; while in an open space, sound waves reflect less, resulting in a longer reverberation time, and may be accompanied by low-frequency enhancement. Therefore, when modeling a small, enclosed space such as a submarine compartment, this application first collects sound wave reflection data from the real environment. It is determined that in this environment, sound wave reflection leads to a brief and rapidly decaying reverberation, with the reverberation time typically set between 0.3 and 0.5 seconds. Based on this data, this application developed a rapid decay reverberation algorithm, enabling virtual sound to decay rapidly in a very short time when simulating a confined space, thus reproducing this compact and clear sound effect. In contrast, in open spaces such as underwater canyons, sound wave reflections are much sparser, and the reverberation time is extended to 1.5 to 2 seconds. To address this, this application designed a multi-reflection simulation algorithm, which not only simulates the delay of sound waves reflecting off different surfaces but also appropriately enhances the low-frequency components, making the sound wider and more spatial. Subsequently, during virtual sound source rendering, the original signal is superimposed with the convolutional response to reproduce spatial reverberation consistent with the environment. Test results show that when the difference between the virtual sound source reverberation time and the real sound field is less than ±0.05s, subjects generally have difficulty distinguishing between virtual and real sounds. In the medium propagation modeling section, the sound propagation mechanism in underwater environments is drastically different from that in air. Water has a higher density, and sound travels at approximately 1500 meters per second, much faster than in air; simultaneously, high-frequency components are rapidly absorbed by water, while low-frequency sound waves dominate. In simulating underwater sound propagation, the speed of sound and frequency characteristics in the underwater environment are first determined using experimental data. Then, a low-pass filter technique is used to simulate underwater audio characteristics, filtering out frequencies above 500Hz to reproduce the dominance of low-frequency sound waves in the underwater environment. This processed sound more closely resembles the actual underwater auditory experience, allowing the listener to feel as if they are truly hearing deep, resonant sounds on the seabed.

[0104] To enhance the immersive experience, this application designs a real-time sound field adjustment submodule, configured to output sound sources directionally through a speaker array and convert impact sounds into low-frequency vibrations using a vibrator. The speaker array control achieves directionality enhancement through delay and sound pressure difference adjustment; for example, in a surround sound system, adjusting the left channel 0.1ms ahead of the right channel creates a sense of location from the left-side sound source. When a special event is triggered in the scene (such as a cabin impacting a reef), the system not only plays the corresponding impact sound but also outputs a 20–50Hz low-frequency signal through the vibrator, converting the sound into perceptible vibrations. Combined with a visual shockwave effect, this achieves consistency in cross-modal perception.

[0105] In the psychological perception model, this application establishes a sound source separation model based on human auditory attention. This model dynamically enhances the target sound source signal when multiple sound sources coexist by simulating the cocktail party effect. Specifically, the system first extracts the masking features of the target sound source based on the sound source direction and characteristic frequency band, and then uses a masking inversion algorithm to suppress non-target sound sources. Results show that when the signal-to-noise ratio is improved by more than 6 dB, subjects can quickly focus their attention on the target sound source, thus obtaining a selective auditory experience consistent with real-world scenarios.

[0106] To address the issue of perceived stability of virtual sound sources under varying environmental conditions, this application establishes an auditory constancy model. Based on loudness and timbre constancy theories, this model adaptively corrects the virtual sound source. Specifically, under different background noise and environmental reverberation conditions, the system maintains the timbre characteristics of the virtual sound source through normalization processing while dynamically adjusting the sound pressure level to ensure the stability of loudness perception. Experiments show that when the ambient noise varies from 30–60 dBA, the subjective loudness fluctuation of the corrected virtual sound source is less than ±5%.

[0107] In some embodiments, this application further establishes an emotion and atmosphere regulation module based on auditory psychological effects. This module guides the user to generate specific emotional responses by adjusting the rhythm, frequency distribution, and loudness changes of virtual sound sources. For example, when a tense atmosphere needs to be created, the system increases high-frequency energy and adds irregular rhythms; when a soothing atmosphere needs to be created, the system reduces high-frequency components and enhances stable low-frequency sounds. Psychoacoustic experiments have verified that this module can effectively improve the immersion and emotional guidance effect of the virtual environment.

[0108] After establishing a basic sound field model, this application further realizes the dynamic adjustment of sound response in the scene to meet the real-time interactive needs of immersive experience.

[0109] For example, within the immersive experience cabin, this application installed surround speakers on both sides of the cabin. By precisely controlling the sound playback delay (e.g., making the left channel 0.1 milliseconds ahead of the right channel), the effect of sound coming from a specific direction is simulated. Simultaneously, when playing high-frequency sound effects through the top speakers, the volume in the right ear can be appropriately reduced to create a sound field from the upper left. This approach helps to construct an accurate three-dimensional sound field and, through preset sound effect templates, automatically adjusts the intensity, direction, and frequency of sound in complex sound mixing environments, making the final sound effect closer to the real environment and meeting the experiencer's expectations for spatial positioning. Secondly, the system also incorporates various user behavior responses and event triggering mechanisms. When a special event occurs in the scene, such as the cabin colliding with a reef, the system's preset script will be automatically activated. At this time, not only will the corresponding impact sound be played, but low-frequency vibrations will also be triggered simultaneously, typically between 20 and 50 Hz. Combined with visual shockwave effects, this allows the experiencer to both hear a powerful sound and feel the vibrations, truly experiencing the power and urgency of the event.

[0110] Through the above technical solution, this application establishes a high-fidelity, dynamic-response digital sound field model to realistically reproduce the acoustic characteristics and motion feedback in the virtual environment, thereby providing accurate and coherent auditory information for the immersive experience cabin, enabling key sound field effects in the multimodal system to be synchronized with other sensory outputs, and ultimately improving the realism and consistency of the immersive experience.

[0111] This embodiment provides a method for constructing and implementing an olfactory physiological perception model. The model's design is based on the physiological mechanism of the olfactory system, where odor molecules are converted into neural signals via binding to olfactory epithelial receptors, and then associated with the emotional and memory systems in the cerebral cortex, thereby forming specific olfactory perceptions and psychological responses. Based on this, this embodiment uses technical means to realize the complete link from odor release to psychological mapping.

[0112] Specifically, this embodiment uses an FPGA as the central control unit to achieve precise scheduling of odor release through pre-stored odor-scene mapping parameters. The odor generating device includes a multi-channel odor generation module, preferably a 16-channel structure, with each channel containing a replaceable odor reservoir, capable of releasing different odor molecules independently or in combination according to scene requirements. Simultaneously, the device is also equipped with an auxiliary airflow delivery device, using a micro-pump to drive the airflow, making the diffusion speed and direction of the odor in space controllable, thereby simulating the natural diffusion process of gases in a real environment.

[0113] Regarding the establishment of the odor library, this embodiment selected approximately thirty basic odors, including common scents such as seawater, metallic rust, algae, and pine. These odors were subjectively evaluated by real people during the experimental phase and categorized with different psychological feelings such as calmness, tension, and mystery, forming a correspondence between "odor and psychological label." In actual operation, when the system receives a scene trigger signal, the FPGA control unit will invoke the corresponding odor combination according to the pre-set mapping logic, and strengthen or weaken the corresponding psychological effect by adjusting the concentration and release duration. For example, when a user enters a deep-sea scene, the system automatically triggers the mixed release of salt mist, algae, and humid odors to create a calm and mysterious atmosphere; when entering a space capsule environment, it releases a combination of metallic and enclosed odors, combined with low-flow airflow control, to create a feeling of isolation and tension; when simulating a deep-sea hydrothermal vent area, the system also triggers a sulfur odor, combined with ambient lighting effects and airflow temperature adjustment, thereby enhancing the immersive experience of the high-temperature, high-pressure environment.

[0114] To ensure the continuity and realism of the experience, this embodiment further incorporates a concentration gradient control and air circulation mechanism. During odor release, the system gradually increases or decreases the odor concentration based on the user's spatial location or the progression of the scene, achieving a gradual effect that conforms to realistic diffusion patterns and avoiding abrupt sensory transitions. Simultaneously, the air circulation device activates after each odor release, essentially eliminating residual odors within thirty seconds, providing a fresh environment for the next odor trigger. Furthermore, this embodiment considers olfactory adaptation; when the same odor is continuously released for more than two minutes, the system automatically switches to a neutral odor, such as clean air, maintaining this for approximately thirty seconds to help the user's olfactory receptors regain sensitivity, thereby preventing sensory desensitization caused by prolonged exposure.

[0115] The olfactory physiological perception model proposed in this embodiment achieves a high degree of fit between precise odor release and perceptual effect through the organic combination of hardware control, odor database, psychological label mapping, and dynamic triggering mechanism. This model not only provides a realistic immersive experience in complex virtual environments such as the deep sea and space, but also boasts performance advantages such as low latency (less than 0.5 seconds) and concentration error control within ±5%, ensuring the reliability and applicability of the system in scientific experiments and entertainment applications. As a result, users can obtain a multimodal immersive experience highly integrated with visual, auditory, and tactile signals, thereby significantly enhancing the realism and interactivity of the virtual environment.

[0116] In this embodiment, a motion perception model for a six-degree-of-freedom motion platform (including longitudinal movement X, lateral movement Y, vertical movement Z, and roll, pitch, and yaw) is proposed. This model consists of a physiological perception layer, a psychological perception layer, and a motion control mapping layer, and is primarily used for immersive experiences in virtual-real fusion scenarios such as deep sea and space.

[0117] Regarding the physiological perception layer, this embodiment first sets the motion perception threshold range based on the characteristics of the human vestibular system and proprioception. Studies have shown that the minimum perception threshold for linear acceleration in humans is approximately 8.5 cm / s² in the anterior-posterior direction. 2 The horizontal speed is approximately 6.5 cm / s. 2 The minimum angle at which a person can perceive rotation and tilt is approximately 1–3°. Therefore, in this embodiment, the output linear acceleration of the control platform is preferably in the range of 6–9 cm / s². 2 The tilt angle is no less than 2° to ensure that users can clearly perceive key motion events, rather than being in a vague or imperceptible range. To achieve individualized parameter matching, the system is equipped with a perception threshold calibration unit. Through balance platform testing and high-precision accelerometers, gyroscopes, and other sensors, it collects motion sensitivity data of users under acceleration conditions of 0.1–2G and angular velocity conditions of 1–10° / s, and establishes an individualized perception threshold database, thereby providing a scientific basis for subsequent motion control and experience adjustment.

[0118] In terms of motion control, this embodiment uses a six-degree-of-freedom electric platform as the main motion output hardware. This platform has a high load-bearing capacity (preferably above 2.5 tons), can simulate a maximum acceleration of 1G, and achieves a repeatability accuracy of ±0.1mm. For the control strategy, an inverse kinematics solver is used for pose calculation and motion planning, ensuring high-precision control and repeatability in complex motion sequences. The platform employs an electric cylinder drive system to reduce noise and avoid the risk of hydraulic leakage; simultaneously, a low-frequency vibrator (frequency range 5–50Hz) is integrated into the floor structure to provide additional vibration feedback in specific scenarios. Through the above design, the platform can achieve multi-dimensional motion such as lifting, translation, pitch, roll, and yaw, while controlling the response latency to within 30 milliseconds, thereby ensuring the synchronization between the virtual image and physical motion.

[0119] To address the issue of distorted experience caused by mechanical vibration or platform drift during motion, this embodiment incorporates a dynamic compensation unit. When the speed of an object in the virtual scene exceeds 3 m / s, the system automatically triggers reverse micro-vibration feedback (preferably with an amplitude of approximately 2 cm and a frequency of 5 Hz). It also monitors platform vibration in real time using an integrated IMU sensor, compensating for the projected image's displacement at a rate of 3.2 m / s. This effectively eliminates visual jitter and achieves dynamic coupling between the virtual image and the platform's motion.

[0120] At the psychological perception level, this embodiment establishes a correspondence between motion perception and psychological effects. Standardized testing determines the user's sensitivity under different motion conditions, and this is combined with the psychological concepts of "controllability" and "intention." Figure 1 The "consistency" theory ensures that the platform's motion direction and amplitude are highly consistent with the motion in the virtual scene, thereby reducing perceptual conflict. Simultaneously, this embodiment introduces an "adaptation / motion sickness adjustment mechanism." When the system detects a potential tendency for the user to experience dizziness, it adjusts the motion amplitude or adds transition buffers to reduce motion sickness. Furthermore, this embodiment follows the cross-modal consistency principle, ensuring that visual cues (such as screen tilt and acceleration) and motion feedback are strictly synchronized in time, avoiding a sense of disjointed immersion caused by inconsistencies in signals from different channels.

[0121] In terms of motion control mapping layer, this embodiment adopts a motion mapping method driven by virtual images, which converts motion parameters such as speed, acceleration, and direction in the virtual scene into platform pose control signals in real time. The signals are decomposed by a time-domain filtering algorithm. Short-term rapid changes (such as explosion impact, rapid turning) are directly mapped to the platform's instantaneous response, while long-term slow drift (such as submersible descent, spacecraft continuous acceleration) are simulated by combining platform tilt angle and displacement washing to ensure long-term stable operation within a limited physical space.

[0122] For example, in deep-sea scenes, when the submersible descends, the platform slowly sinks and generates a pitch angle to enhance the sense of depth; when encountering underwater currents, the platform triggers rapid lateral shaking and low-frequency vibrations to enhance realism. In space scenes, when the spacecraft accelerates into orbit, the platform tilts forward and maintains a sense of thrust; when encountering micrometeorite impacts, it triggers instantaneous vibrations and yaw, combined with a dynamic compensation mechanism to reduce screen shake. Furthermore, this embodiment also designs a feedback mechanism based on interactive operation. For example, when the user operates the virtual joystick, the platform will simultaneously generate a slight tilt of approximately ±5° to simulate the feeling of real resistance.

[0123] Through the above methods, this embodiment achieves a high degree of consistency between the virtual scene and the physical platform in motion perception, which not only satisfies the realism of physiological perception, but also takes into account the comfort and safety of psychological experience, thereby constructing a complete motion perception integrated model.

[0124] This application further proposes a dynamic weight allocation system including an environment-driven priority module and a conflict detection and compensation module. The environment-driven priority module allocates 60% weight to visual feedback, 25% to motion feedback, and 15% to sound effects when ambient light is bright; and allocates 40% weight to auditory feedback, 40% to tactile feedback, and 20% to visual feedback when ambient light is insufficient. The conflict detection and compensation module automatically reduces the platform's motion amplitude and increases the visual compensation tilt angle to 20° when the difference between visual and vestibular feedback acceleration exceeds 0.2G.

[0125] Specifically, in multimodal systems, the importance of various sensory signals varies under different environments. To ensure the best experience for users, this application employs a dynamic weighting strategy, automatically adjusting the output proportion of each sense based on scene characteristics and user behavior. Through dynamic weighting and conflict resolution mechanisms, it ensures that sensory information such as vision, hearing, touch, and motion remain synchronized and coordinated under different environments and user behaviors. This strategy not only reduces discomfort caused by perceptual conflicts but also significantly enhances the overall realism and stability of the immersive experience.

[0126] The human eye automatically adjusts the size of its pupils according to the intensity of ambient light. When the ambient light is strong, the pupils constrict; when the light is dim, the pupils dilate, thereby regulating the amount of light entering the retina to achieve the best visual effect.

[0127] Therefore, visual information is typically most abundant when ambient light is bright. In this case, visual information is weighted at 60%, presenting a richly detailed underwater spectacle through high-resolution projection. Simultaneously, motion feedback (such as slight tilting of the platform) accounts for 25%, and sound effects (such as ambient sound) account for 15%. This allocation ensures that users primarily rely on vision to obtain information in bright environments, while other senses play a supporting role. For example, the underwater scenery outside the cabin, presented through 4K high-resolution projection, combined with slight platform movements, greatly enhances the sense of spatial realism.

[0128] In low-light environments, such as deep-sea caves, visual details become weaker, making auditory and tactile information particularly crucial. In this case, the weighting of auditory and tactile senses is increased to 40%, while visual information is reduced to 20%. For example, when exploring a dimly lit deep-sea cave, this application enhances sonar echoes and platform vibrations, allowing users to rely on these auxiliary information to perceive their environment, thus compensating for the lack of visual information. This characteristic can also be utilized when designing dinosaur activities; by adjusting ambient lighting to control the user's pupil size, localized areas (such as dinosaur images) can be highlighted and clearly visible on the same display screen, while the background appears deep black, expanding the sense of visual space and achieving the desired immersive effect. Besides adjusting the weighting of corresponding senses under different lighting conditions, responsiveness to user behavior is also very important. When users actively operate (e.g., using a virtual periscope), tactile feedback becomes even more important. In this case, increasing the weighting of tactile feedback to 50% allows users to feel more realistic resistance and tactile sensations during operation, enhancing the realism of the operational feedback. When the user is merely observing the environment, the default weighting of each sense remains, ensuring a balanced overall experience. This dynamic weighting mechanism allows the system to automatically optimize the output of each sensory signal based on the real-time environment and the user's state, achieving the best immersive effect.

[0129] In multimodal systems, inconsistencies can sometimes arise between different sensory information, leading to disjointed experiences or motion sickness for the user. To address this, this application also designs conflict detection and adaptive compensation mechanisms to ensure that all sensory signals remain consistent and coordinated.

[0130] The conflict detection mechanism in the system continuously monitors signal differences between various sensory modules, paying particular attention to the matching of visual and vestibular signals. For example, when the platform is stationary but the visual display shows high-speed motion, or when the acceleration difference between visual and vestibular feedback exceeds 0.2G, this application can automatically identify potential perceptual conflicts. Through this real-time monitoring, situations that may trigger motion sickness can be captured and identified in a timely manner, providing trigger signals for subsequent compensation.

[0131] When the system detects a sensory conflict, it promptly activates corresponding compensation strategies based on the specific situation. These compensation strategies are divided into visual-dominant compensation and vestibular-dominant compensation. If the system detects a significant discrepancy between visual information and vestibular signals (e.g., the platform is actually tilted by 10°, but the visual effect shows excessive movement), this application automatically reduces the platform's movement amplitude while increasing visual compensation, such as adjusting the visually displayed tilt to 20°. This method reduces the conflict caused by the inconsistency between visual and actual movement, allowing the user to achieve a more stable motion perception. When the vestibular signal is detected as more reliable, this application can weaken the effect of visual motion blur and help the user recalibrate orientation and motion perception by displaying fixed reference objects (e.g., the cabin instrument panel). This method stabilizes the user's spatial positioning and reduces the risk of motion sickness.

[0132] This application further proposes a precise time protocol to achieve synchronization in the following ways: a global clock alignment module, configured to unify the clock reference of the projector system, the six-degree-of-freedom platform, the immersive sound device, and the scent generator; and a dynamic delay compensation module, configured to adjust the output of each module in advance by pre-setting the next plot design.

[0133] Specifically, in a multi-sensory system, different devices (such as projectors, sound systems, odor generators, and motion platforms) each have their own independent operating clocks. If these devices are out of sync, delays will occur, leading to incoordination between visual, auditory, olfactory, and motor feedback. To address this issue, this application employs a Precise Time Protocol (PTP) to unify the clock reference of all devices. This ensures that all devices operate strictly according to a unified timeframe during startup, data transmission, and signal output, guaranteeing that the delay of each modal signal is controlled within 15 milliseconds. For example, in an explosion scenario, the light effects, loud noise, sulfurous smell, and platform vibration must be strictly synchronized. Actual testing shows that when the time error of all devices is controlled within 15 milliseconds, the perceptual error of the experiencer is less than 20 milliseconds—a value close to the minimum threshold of human neural response, thus resulting in a high degree of coordination and consistency in the experience of multimodal stimuli.

[0134] In multimodal experiences, besides time synchronization, it is crucial to ensure that the outputs of each sense correspond to their spatial locations; that is, sound, smell, and visual scenes should all precisely correspond to the same virtual spatial location. Based on this, this application employs Head Related Transfer Function (HRTF) technology to achieve spatial binding between sound and vision. HRTF can calculate the direction in which a virtual sound source should appear in a specific scene based on the shape of a person's head and ears. When a virtual sound source (such as a deep-sea probe) is set up in the system, its coordinates in the virtual scene are bound to the visual projection coordinates in real time, ensuring that the direction of the sound heard by the user is consistent with the trajectory of the object's movement on the screen. Experimental data shows that when the error in the sound direction is controlled within 3°, the user's accuracy in perceiving the direction of the sound source can be improved to 92%, while if the error reaches 5°, the accuracy drops to only 68%.

[0135] Besides sight and hearing, smell is also an important component of the multimodal experience. This application also designs a dynamic control system that synchronizes smell release with virtual scene events. For example, when the user approaches the virtual volcano, the system automatically adjusts the release intensity and direction of the sulfur odor nozzles. If the user is near the left-hand observation window, the system will amplify the odor output on the left and adjust the odor concentration to match visual effects such as heat waves and red light, thereby enhancing the overall spatial realism.

[0136] This application further proposes a dual-layer projection integrated device in the execution layer, consisting of a main projection screen and a holographic screen.

[0137] The main projection screen is the display layer responsible for projecting the basic visual content. It can be implemented using high-resolution laser projection equipment, with precise calibration ensuring high-fidelity presentation of core visual elements. The holographic gauze screen is an enhanced display layer that generates stereoscopic images using the dynamic scattering characteristics of microparticle media. It can be implemented using a suspended microparticle screen combined with short-throw projection technology, simulating physical depth through light field distribution. The main projection screen and the holographic gauze screen are complementary in their physical structure; the former provides a stable two-dimensional image substrate, while the latter superimposes three-dimensional light field information with spatial distribution characteristics.

[0138] Specifically, the main projection screen, as the basic display layer, is responsible for outputting high-precision color and shape information, ensuring the clarity and positioning accuracy of key visual elements. The holographic gauze screen generates interactive light field information at different depths in space by adjusting the density distribution and optical reflection path of suspended particles, such as simulating the spatial diffusion effect of dynamic environmental elements like smoke and dust in a virtual scene. The integrated device achieves decoupling and fusion of visual information through a layered projection mechanism: the main projection screen maintains the stability of the main structure of the image, while the holographic gauze screen dynamically adjusts the light field distribution in the depth dimension according to scene requirements, so that the depth cues output by the visual system and the acceleration feedback of the motion perception module and the sound field positioning of the auditory perception module achieve physiological consistency matching.

[0139] Compared to existing technologies, traditional single-projection systems can only present visual content on a single plane, resulting in a lack of scene depth and insufficient stereoscopic visual cues. In contrast, dual-layer projection integrated devices, through physically superimposed light field distribution, enable the visual system to simultaneously output two-dimensional planar information and three-dimensional spatial light field information without the need for wearing stereoscopic display devices, significantly enhancing the physiological matching degree of multimodal perception coordination.

[0140] Through the above technical solution, this application solves the problem of the single visual level caused by a single projection system. By working together with the main projection screen and the holographic gauze screen, the depth dimension of the visual presentation is expanded while maintaining the accuracy of the main subject of the image. This makes the visual depth cues, motion feedback and sound field positioning form a spatial consistency that is more in line with the human perception mechanism, thereby improving the realism and immersion of multimodal perception collaboration.

[0141] This application further proposes a technical solution for integrating the ROS2 framework into the Pandora central control system, integrating three communication protocols: DMX512, EtherCAT, and OSC.

[0142] Specifically, the Pandora central control system is built on ROS2 and can integrate the communication protocols of various devices (such as DMX512, EtherCAT, OSC), enabling unified distribution and collaborative work of cross-device and cross-modal data.

[0143] Through the above technical solution, the instruction transmission latency in this application is as low as 5 milliseconds. The system can support the concurrent control of thousands of devices at the same time and is the central hub of the entire system. It has extremely strong compatibility and scalability.

[0144] This application further proposes a fault-tolerant redundancy mechanism, which switches to the backup bus when the main communication link fails, and ensures automatic degradation of output when the projection system fails through redundancy mode.

[0145] Specifically, in order to ensure the stability of data transmission and the continuity of system operation, this application also introduces a multi-layered fault tolerance and redundancy mechanism.

[0146] First, this application employs dual-link communication. When the primary EtherCAT network fails, the system automatically switches to the backup RS485 bus within 500 milliseconds, thus preventing interruption of command transmission. Second, the system incorporates a pre-defined data caching mechanism, with real-time computing nodes reserving a 10-millisecond data buffer. This buffer replenishes data promptly in the event of momentary data loss, preventing command interruption or erroneous transmission. Finally, for critical equipment, such as the projection system, an N+1 redundancy mode is used: when one device fails, other devices automatically take over its work, ensuring the system continues to operate normally, although the image quality may be reduced to a low-resolution mode, without interrupting the overall display. In summary, data flow management ensures fast and accurate information flow from the perception layer to the control layer, and then from the control layer to the execution layer, while fault tolerance and redundancy design improve the system's stability in the face of network failures or momentary data loss. The entire process incorporates strict data transmission rules, real-time buffering, and backup networks, guaranteeing a smooth user experience during interaction.

[0147] Through the above technical solution, this application effectively solves the problem of synchronization failure caused by communication interruption or equipment failure in multimodal collaborative systems, ensuring that the system can still maintain a basic operating state when some components are abnormal, and avoiding the collapse of immersive experience caused by service interruption.

[0148] This application creates a stunning and realistic immersive experience space. It not only relies on high-precision equipment to achieve simultaneous output of visual, auditory, tactile, and olfactory signals, but also designs a series of cross-sensory signal overlay schemes based on the physiological and psychological response characteristics of humans to sensory stimuli, thereby enhancing the user's overall perception of the spatial atmosphere. Firstly, the visual system plays a crucial role in the immersive experience, not only presenting concrete images but also directly influencing the user's evaluation of spatial perception. In this application, dual-layer projection technology is used to create a striking contrast between the visually prominent areas and the background. For example, in the dinosaur cabin or deep-sea experience cabin, by presenting the main area of ​​the observation window as a high-resolution, richly detailed image, while using an 18% gray holographic screen as the foreground layer, the background appears as a deep, uniform black. This design fully utilizes the human eye's perception of contrast and figure-ground separation, allowing the user to not only clearly capture the main subject information but also feel the expanded spatial depth and layering, thus constructing a strong visual impact and an immersive space. Furthermore, when showcasing underwater scenes, the flowing schools of fish, swaying corals, and constantly changing light and shadow effects in the virtual visuals enhance visual continuity. This visual design fully utilizes the principle of the human eye's zoning processing of the central high-resolution area and the surrounding dynamic information, allowing users to capture rich details when observing up close, while quickly perceiving the dynamic flow of movement and light in the surrounding areas, thus creating a strong sense of spatial depth and an immersive atmosphere.

[0149] Secondly, the auditory system not only plays a crucial role in playing environmental sound effects but also in conveying spatial positioning and dynamic emotions. This application will utilize Dolby Atmos technology, employing a multi-channel speaker system to precisely position sound within the virtual space. Simultaneously, in critical scenarios (such as explosions or emergencies), the system utilizes high-precision timestamp synchronization (PTP protocol) to ensure that the latency of visual, auditory, and other sensory feedback is controlled within 20 milliseconds, allowing each sound to precisely match the screen image, olfactory feedback, and motion state. This design not only allows the user to "hear" precise spatial dynamics but also enhances the overall spatial perception of the environment through the directionality of sound.

[0150] Regarding the olfactory system, given the unique advantage of smell in directly evoking emotional memories, this application designs a context-linked odor release scheme. For example, in simulated lava flow or hydrothermal vent scenarios, when the user sees the lava in motion, the system simultaneously releases a burnt smoke odor and adjusts the color temperature of the external color-changing lights to approximately 1800K, creating a high-temperature, scorching atmosphere. This design fully considers the psychological response of humans to the association between smell and context. By combining low-color-temperature light effects with the odor, it strengthens the linkage between vision and smell, making the user feel as if they are in a real environment.

[0151] To ensure high coordination among the various sensory modules, this application employs multimodal latency optimization technology. Through real-time monitoring and compensation, it ensures that the motion trajectories in the virtual scene (such as rapidly passing schools of fish or dynamic underwater scenes) are synchronized with the actual movements of the user, with latency errors controlled within 20 milliseconds. Throughout the multimodal overlay process, this application fully considers the physiological and psychological response patterns of the human body to different sensory stimuli, such as the sensitivity of the human eye to high contrast, the ear's directional perception of low-frequency sound waves, and the prominent role of smell in emotional memory. Through this multimodal collaborative model based on human perception mechanisms, not only is precise control achieved at the device level, but a unified and coordinated sensory impression is also formed at the user's psychological level. Based on cross-sensory signal overlay design, this application strives to construct an immersive experience system full of spatial atmosphere and emotional resonance by precisely controlling the output of various modalities such as vision, hearing, smell, and touch, leveraging the physiological and psychological characteristics of human perception.

[0152] Furthermore, since different users have different sensitivities to sensory stimulation, this application adopts a personalized adaptation strategy, adjusting the modal parameters in real time based on the user's physiological and subjective feedback to ensure that each user can obtain the optimal experience.

[0153] For users prone to motion sickness, the acceleration of the motion platform can be reduced, for example, from 1G to 0.3G, and the display of visually stable reference lines (such as a fixed dashboard) can be enhanced to help users better calibrate their sense of direction. For users with a sensitive sense of smell, the system will appropriately reduce the concentration of irritating odors (for example, from 50ppm to 30ppm) and lengthen the interval between odor releases to avoid discomfort caused by excessively strong odors.

[0154] This personalized optimization strategy ensures that the system can automatically adjust various parameters according to the needs of different users, providing a realistic and comfortable multimodal immersive experience.

[0155] In this application, each functional module (such as visual control, motion feedback, environmental management, and sound field control) is designed as an independent unit, and all communication between modules adopts a unified interface standard to ensure consistent data format and transmission protocol. This allows devices from different manufacturers or technology platforms to seamlessly integrate into the system, reducing compatibility issues. If new sensing or control technologies (such as adding olfactory or tactile modules) need to be added in the future, the new module can be integrated simply through a standard API or unified communication protocol without affecting the core functionality of the entire system.

[0156] Modular design allows for the addition of new functional modules in the future without altering the overall system architecture. For example, if wind simulation or haptic gloves are needed in the future, they can be directly added to the system through reserved interfaces. This design pattern ensures that the system can be continuously expanded according to actual needs without facing difficulties in modification due to an overly fixed structure. Overall, the modular design and expansion strategy build a flexible and efficient "building block" platform for the system. Each functional unit can operate independently and also collaborate with other modules through standard interfaces, collectively forming a stable and easily expandable cross-modal intelligent control system.

[0157] To comprehensively verify the various modules in the system and their integration effects, this application designed the following multiple test methods.

[0158] 1. Projection synchronization test

[0159] Testers used a high-speed camera (e.g., a 1000fps camera) to capture the output images from the dual projectors, focusing on analyzing the frame refresh time difference and image stitching error between the two projectors during image stitching. The aim was to ensure that the projection system could achieve seamless image stitching at high frame rates, guaranteeing a smooth overall visual experience.

[0160] 2. Color-changing lamp matching test

[0161] Testers used a spectroradiometer to measure the color temperature of the projected area, while simultaneously checking the color consistency of the color-changing lamp output. The purpose of this was to verify whether the color temperature difference between the real-time rendered image and the color-changing lamp output was within an acceptable range, ensuring that the color effect seen by the user was consistent and uniform.

[0162] 3. Lighting scene switching test

[0163] Testers recorded the response time required from triggering the switching command to the new lighting scene reaching the design standard. The test was repeated 100 times and the average value was taken. The purpose of this was to verify whether the entire system could respond quickly (<0.5 seconds) during scene switching, ensuring the continuity of the immersive experience and maintaining a relatively stable visual effect even when the plot scene changes.

[0164] 4. Data Flow and Communication Verification

[0165] Testers monitored uplink data (user location, physiological data, environmental conditions, etc.) and downlink data (rendering commands, motion parameters, sound field coordinates, etc.) in real time to analyze whether the data transmission frequency and latency met design requirements. This was done to ensure efficient and stable information flow between modules during data transmission, while prioritizing the transmission of critical commands in a multi-protocol environment.

[0166] 5. Continuous stability test

[0167] Testers operated the system under extreme conditions, including high temperature (45°C) and high humidity (80%), conducting continuous 24 / 7 testing. The aim was to verify the system's long-term stability and fault tolerance, ensuring a failure rate of less than 0.1%.

[0168] Test results of this application

[0169] Indicator Target value Measured value Compliance status Project synchronization delay ≤ 16 ms 15.2 ms √ Color temperature error of colored light ≤5% 3.8% √ Light scene switching time ≤ 0.5 seconds 0.47 seconds √ Multi-protocol instruction conflict rate ≤0.1% 0.08% √

[0170] This application integrates various advanced devices and communication protocols, adopting a layered and modular design to achieve fully closed-loop control from data acquisition and real-time fusion to precise command transmission. Practice has proven that the system, while ensuring low latency and high precision, also possesses the advantages of stable operation and flexible expansion, providing strong support for intelligent control in complex scenarios.

[0171] Through practical deployment in cultural and tourism venues and science exhibitions, this application has not only received high praise from visitors, but its economic benefits have also been preliminarily verified. The high participation rate in cultural and tourism experiences and science exhibitions, as well as the optimization of film and television production processes, all indicate that this technology has broad market application prospects and significant commercial promotion potential.

Claims

1. A multimodal sensing, collaborative, and real-time synchronization system, characterized in that, include: The perception layer collects multimodal data of vision, hearing, touch and motion, and establishes a unified perception quantification model based on human physiological and psychological mechanisms. The perception quantification model includes a visual perception modeling module, an auditory perception modeling module, an olfactory perception modeling module and a motion perception modeling module. The control layer uses a central control system and a Bayesian inference model to achieve dynamic weight allocation and conflict resolution of multimodal data. A high-precision time synchronization mechanism is employed to ensure that cross-device signal delay is controlled within milliseconds; The execution layer includes a projection display, a six-degree-of-freedom motion platform, a panoramic sound device, and an odor generating device. The execution layer generates corresponding multi-sensory feedback according to the instructions of the control layer to ensure that the system maintains a high degree of consistency in spatial positioning, sound field direction, color brightness, and odor diffusion, thus ensuring a seamless connection between the virtual and real environments.

2. The multimodal sensing, collaborative, and real-time synchronization system as described in claim 1, characterized in that, The visual perception modeling module is based on the physiological and psychological mechanisms of human spatial perception, including: Visual physiological modeling, based on the geometric perspective rules, minimum resolution angle, brightness perception, color and color temperature matching principles of the human eye, establishes a quantitative model to ensure that virtual objects are consistent with the real environment in terms of proportion, perspective, resolution, motion blur, brightness contrast, light source direction, shadow and color. Visual psychological modeling, combined with depth cues including binocular parallax, motion parallax, and texture gradient, optimizes scene contours and group perception using Gestalt psychology principles.

3. The multimodal sensing collaboration and real-time synchronization system as described in claim 2, characterized in that, The auditory perception modeling module includes: Auditory physiological modeling is used to establish a spatial localization model based on binaural time difference and binaural intensity difference. By combining critical bandwidth and equal loudness curves, the virtual sound source is made consistent with the real sound source in terms of direction, spectrum and loudness. At the same time, a reverberation and sound field consistency model is introduced to ensure that the sound source matches the acoustic characteristics of the environment. Auditory psychological modeling enhances the attentional discrimination of target sound sources through a sound source separation model; it ensures the stability of auditory experience in different environments through loudness constancy and timbre constancy; and it adjusts rhythm and frequency distribution to achieve psychological guidance of emotions and atmosphere.

4. The multimodal sensing, collaborative, and real-time synchronization system as described in claim 3, characterized in that, The olfactory perception modeling module includes a basic odor database and an odor generator, configured as follows: Olfactory psychological modeling, through odor database and perception matching mechanism, combined with users' odor association and contextual memory, enhances the user's sense of immersion in the environment; Olfactory physiological modeling, based on the physiological mechanism of "odor molecule-receptor-signal conversion", uses a multi-channel odor generator and a precise control module to combine and release various odor molecules under scene triggering to simulate environmental odors.

5. The multimodal sensing, collaborative, and real-time synchronization system as described in claim 4, characterized in that, The motion perception modeling module includes: Motion physiology modeling, through a six-degree-of-freedom platform and screen object displacement model, realizes the spatial mapping between the motion of virtual objects and the user's body motion, ensuring the natural correspondence between displacement, velocity and acceleration; Motor psychological modeling, using motion constancy and body proprioception model, ensures that users do not experience distortion or dizziness during dynamic processes.

6. The multimodal sensing, coordination, and real-time synchronization system as described in claim 5, characterized in that, The visual perception modeling module, auditory perception modeling module, olfactory perception modeling module, and motion perception modeling module are coupled and integrated: The various perception models are integrated in the Pandora central control system, which uses a Bayesian inference model to dynamically assign weights to multimodal data: in scenes with strong lighting, the visual weight is increased; In noisy environments, auditory weight decreases while tactile and visual compensation increases; in specific scenarios, olfactory weight can be enhanced as a contextual trigger factor. The control layer also executes a conflict resolution mechanism: when different sensory signals conflict, the system will automatically adjust the low-confidence signal to ensure the consistency of the multimodal experience; The system uses a precise time protocol to ensure signal synchronization across devices, with a delay of ≤20ms, thus ensuring real-time performance.

Citation Information

Cited By

  • Abnormal co-debt association mode graph learning identification system for multi-source time series data

    CN121767089A

  • Abnormal co-debt correlation pattern graph learning and identification system for multi-source time series data

    CN121767089B