Virtual reality interaction system integrating spatial positioning and multi-modal interaction for movie playing
By combining visual SLAM, inertial measurement unit and semantic segmentation into a three-element fusion localization algorithm, and integrating voice, gesture and gaze multimodal recognition, high-precision spatial localization and natural interaction are achieved. This solves the problem of localization and interaction fusion in immersive movie playback systems in complex environments and improves the user experience.
Patent Information
- Application Number
- CN202510950077.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-10
- Publication Date
- 2025-10-28
AI Technical Summary
Existing technologies cannot achieve high-precision spatial positioning in complex dynamic environments and lack deep integration of multimodal interaction, resulting in an unnatural user experience and insufficient immersion.
A three-element fusion localization algorithm combining visual SLAM, inertial measurement unit, and semantic segmentation is adopted. Global pose is generated by confidence weighting, and multimodal recognition of voice, gesture, and gaze is combined to dynamically schedule interaction weights, realize intelligent switching of system state and adaptive rendering, and support multi-user collaborative management.
It significantly improves spatial positioning accuracy and stability, enhances the recognition accuracy and naturalness of multimodal interactions, provides a smooth mode switching experience, reduces latency in multi-user scenarios, and improves system robustness.
Smart Images

Figure CN120848727A_ABST
Abstract
Description
Technical Field
[0001] This relates to the field of virtual reality interaction technology, specifically to an immersive movie playback system that integrates spatial positioning and multimodal interaction. Background Technology
[0002] With the continuous development of virtual reality (VR) technology, immersive movie playback and interactive systems based on VR platforms have become a research hotspot in academia and industry. On the one hand, spatial positioning technology is increasingly widely used in VR interaction: traditional methods mostly rely on single visual SLAM (such as ORB-SLAM) or pure inertial measurement units (IMU) to estimate the user's helmet position; there are also hybrid positioning schemes that use ultra-wideband (UWB) or external base stations to improve accuracy. For example, some advanced VR headsets enhance positioning robustness through tight coupling of vision and IMU, but drift is still prone to occur in low light or dynamic environments. On the other hand, research on the integration of multimodal interaction methods (voice, gestures, gaze, body posture) in virtual scenes is also progressing rapidly: existing systems can recognize gesture commands through controllers or cameras and achieve voice control through microphone arrays; some headsets also support eye trackers to track the user's gaze point for menu selection or gaze guidance. However, existing multimodal solutions are often independent of each other, lack dynamic weight scheduling of different interaction signals, and have low correlation with the user's real-time spatial position, making it difficult to form a closed-loop mapping between "user position – interaction intention – content scene". While some commercial products offer basic scene roaming and simple interaction, when high-precision positioning and natural interaction are combined, users often need to manually switch modes or rely on additional hardware support.
[0003] In summary, existing technologies have shortcomings such as the inability to achieve high-precision spatial positioning in complex dynamic environments and the lack of deep fusion of multimodal interactions based on positioning. Summary of the Invention
[0004] To address the shortcomings of existing technologies, such as the inability to achieve high-precision spatial positioning in complex dynamic environments and the lack of deep fusion of multimodal interactions based on positioning, the technical solution provided by this invention is as follows:
[0005] A virtual reality interaction method for movie playback that integrates spatial positioning and multimodal interaction includes:
[0006] Step 1, System Initialization: The software loads the visual SLAM, inertial measurement unit calibration module, scene semantic segmentation model, and speech, gesture, and gaze multimodal recognition model, and outputs the ready modules;
[0007] Step 2, Spatial localization fusion: The software acquires visual SLAM pose, IMU pose and semantic-assisted pose in parallel, and fuses them according to confidence level to generate global pose and localization confidence vector.
[0008] Step 3, Multimodal Interaction Fusion: The software receives the positioning confidence vector output from Step 2 and the voice, gesture, and gaze features collected in parallel. It dynamically fuses these features based on the recognition confidence of each modality and the spatial context weights, and outputs a unified interaction vector.
[0009] Step 4, Joint State Scheduling: The software constructs the system state based on the fused pose from Step 2 and the interaction vector from Step 3, calculates the spatial confidence and interaction intent strength, intelligently switches between three modes: idle, roaming, and interactive, and outputs the current mode and control commands.
[0010] Step 5, Adaptive Rendering and Feedback: The software receives the mode and control commands output in Step 4, combines the fixed position confidence dynamic weighted user perspective with the director's preset shot, renders the virtual camera perspective, adjusts the spatial sound effects in real time, and generates haptic feedback commands.
[0011] Step 6, Multi-user Collaborative Management: The software collects the fusion status and confidence level of each user, synthesizes a weighted group view, resolves interactive resource conflicts based on confidence level, and supports smooth recovery after disconnection and reconnection;
[0012] Step 7, Expanding Interfaces and Logs: The software dynamically activates script plugins and third-party AI assistants based on group status and session information, adjusts the log collection frequency as needed, and outputs a list of activated modules and log data.
[0013] Furthermore, the confidence-weighted fusion uses the ratio of the localization covariance of visual, inertial, and semantic approaches to generate weights.
[0014] Furthermore, the spatial context weights are obtained by linearly mapping and normalizing the localization confidence vector output in step two, and are used to adjust the attention allocation for each modality.
[0015] Furthermore, the mode switching dynamically calculates the mode priority based on the ratio of spatial confidence to interaction intent intensity, and determines the current execution mode through normalized Soft-max.
[0016] Furthermore, adaptive rendering uses positional confidence as interpolation coefficients to dynamically and smoothly switch between the user's free viewpoint and the director's preset viewpoint, and attenuates spatial sound effects in real time according to the distance between the user and the virtual sound source.
[0017] A virtual reality interactive system integrating spatial positioning and multimodal interaction for movie playback is also provided, comprising:
[0018] Module 1, System Initialization: The software loads the visual SLAM, inertial measurement unit calibration module, scene semantic segmentation model, and speech, gesture, and gaze multimodal recognition model, and outputs the ready modules;
[0019] Module 2, Spatial Localization Fusion: The software acquires visual SLAM pose, IMU pose, and semantic-assisted pose in parallel, and then fuses them according to confidence level to generate a global pose and localization confidence vector.
[0020] Module 3, Multimodal Interaction Fusion: The software receives the positioning confidence vector output from step 2 and the voice, gesture, and gaze features collected in parallel. It dynamically fuses these features based on the recognition confidence of each modality and the spatial context weights, and outputs a unified interaction vector.
[0021] Module 4, Joint State Scheduling: The software constructs the system state based on the fused pose from step 2 and the interaction vector from step 3, calculates the spatial confidence and interaction intent strength, intelligently switches between three modes: idle, roaming, and interactive, and outputs the current mode and control commands.
[0022] Module 5, Adaptive Rendering and Feedback: The software receives the mode and control commands output from step 4, combines the fixed-position confidence dynamic weighted user perspective with the director's preset shots, renders the virtual camera perspective, adjusts spatial sound effects in real time, and generates haptic feedback commands.
[0023] Module Six, Multi-User Collaborative Management: The software collects the fusion status and confidence level of each user, weights and synthesizes a group view, prioritizes the resolution of interactive resource conflicts based on confidence level, and supports smooth recovery after disconnection and reconnection;
[0024] Module 7, Extended Interfaces and Logs: The software dynamically activates script plugins and third-party AI assistants based on group status and session information, adjusts the log collection frequency as needed, and outputs a list of activated modules and log data.
[0025] A computer storage medium for storing a computer program, which, when read by a computer, is used by the computer to execute the method described thereon.
[0026] A computer, including a processor and a storage medium, executes the method when the processor reads a computer program stored in the storage medium.
[0027] A computer program product, which, as a computer program, implements the method when the computer program is executed.
[0028] Compared with the prior art, the advantages of the technical solution provided by the present invention are as follows:
[0029] The visual-inertial-semantic ternary fusion localization algorithm in this scheme significantly improves spatial localization accuracy and stability. Compared with methods that rely solely on single visual SLAM or pure IMU, it can maintain sub-centimeter-level error in low-light and dynamic scenes, and effectively suppress drift and jitter.
[0030] The attention-based multimodal interaction fusion engine introduced in this solution can dynamically allocate the weights of voice, gesture, and gaze inputs according to the reliability of real-time interaction signals. Compared with traditional isolated recognition of each modality, it achieves a more natural and coherent interpretation of user commands, improving recognition accuracy by more than 15%.
[0031] The joint state scheduler in this solution achieves smooth switching between idle, roaming, and interactive modes through dual metrics of spatial confidence and interaction intent. Compared to existing fixed threshold triggering strategies, it avoids sudden changes in perspective and provides users with a seamless transition experience.
[0032] The adaptive rendering module in this solution dynamically weights and fuses the virtual camera's perspective and spatial sound effects based on positional confidence. This ensures that the perspective better matches the user's roaming needs at high confidence levels and adheres more closely to the director's presets at low confidence levels. Simultaneously, the spatial sound effects are precisely attenuated according to the user's orientation and distance. Overall, the immersive experience is significantly improved compared to traditional fixed perspectives and static sound effects.
[0033] The multi-user collaborative synchronization mechanism in this scheme utilizes spatial confidence weighted synthesis of group views and prioritizes resource allocation based on confidence when conflicts occur. Compared with common master-slave replication or unified broadcast mechanisms, it reduces latency by nearly 30% in large-scale user scenarios and exhibits better fairness and robustness.
[0034] Suitable for virtual reality movie playback and experience systems that require a high degree of immersion and natural interaction. Attached Figure Description
[0035] Figure 1 This is a flowchart illustrating a virtual reality interaction method that integrates spatial positioning and multimodal interaction for movie playback. Detailed Implementation
[0036] To make the advantages and benefits of the technical solution provided by the present invention clearer, the technical solution provided by the present invention will now be further described in conjunction with the accompanying drawings, specifically:
[0037] Implementation Method 1: This implementation method provides a virtual reality interaction method for movie playback that integrates spatial positioning and multimodal interaction, including:
[0038] Step 1, System Initialization: The software loads the visual SLAM, inertial measurement unit calibration module, scene semantic segmentation model, and speech, gesture, and gaze multimodal recognition model, and outputs the ready modules;
[0039] Step 2, Spatial localization fusion: The software acquires visual SLAM pose, IMU pose and semantic-assisted pose in parallel, and fuses them according to confidence level to generate global pose and localization confidence vector.
[0040] Step 3, Multimodal Interaction Fusion: The software receives the positioning confidence vector output from Step 2 and the voice, gesture, and gaze features collected in parallel. It dynamically fuses these features based on the recognition confidence of each modality and the spatial context weights, and outputs a unified interaction vector.
[0041] Step 4, Joint State Scheduling: The software constructs the system state based on the fused pose from Step 2 and the interaction vector from Step 3, calculates the spatial confidence and interaction intent strength, intelligently switches between three modes: idle, roaming, and interactive, and outputs the current mode and control commands.
[0042] Step 5, Adaptive Rendering and Feedback: The software receives the mode and control commands output in Step 4, combines the fixed position confidence dynamic weighted user perspective with the director's preset shot, renders the virtual camera perspective, adjusts the spatial sound effects in real time, and generates haptic feedback commands.
[0043] Step 6, Multi-user Collaborative Management: The software collects the fusion status and confidence level of each user, synthesizes a weighted group view, resolves interactive resource conflicts based on confidence level, and supports smooth recovery after disconnection and reconnection;
[0044] Step 7, Expanding Interfaces and Logs: The software dynamically activates script plugins and third-party AI assistants based on group status and session information, adjusts the log collection frequency as needed, and outputs a list of activated modules and log data.
[0045] The confidence-weighted fusion uses the ratio of the localization covariance of visual, inertial and semantic approaches to generate weights.
[0046] The spatial context weights are obtained by linearly mapping and normalizing the localization confidence vector output in step two, and are used to adjust the attention allocation for each modality.
[0047] The mode switching dynamically calculates the mode priority based on the ratio of spatial confidence to interaction intent intensity, and determines the current execution mode through normalized Soft-max.
[0048] Adaptive rendering uses positional confidence as interpolation coefficients to dynamically and smoothly switch between the user's free viewpoint and the director's preset viewpoint, and attenuates spatial sound effects in real time according to the distance between the user and the virtual sound source.
[0049] Implementation Method 2, see below Figure 1 This embodiment is a further description of the technical solution provided in Embodiment 1, specifically:
[0050] Step 1: System Startup and Module Loading
[0051] When a user puts on a VR headset and launches the application, the system loads the positioning, semantic, and multimodal recognition modules.
[0052] After the user puts on the helmet and launches the "Immersive Viewing" application, the system automatically enters the initialization interface and begins loading: visual SLAM and IMU calibration model, semantic segmentation model, and voice / gesture / eye gaze recognition model.
[0053] The helmet is left to stand still for 3–5 seconds to complete the IMU zero-bias convergence, and the user can see the loading progress indicator.
[0054] Once the application has finished loading, it prompts the user to select a viewing room or movie scene. After the user confirms, they proceed to the next step.
[0055] The output modules are now ready, and the user can move around freely and begin interacting.
[0056] Step 2 Spatial Positioning Fusion
[0057] Users move their heads or controllers in a virtual scene, and the system integrates multi-source positioning in real time to generate high-precision pose.
[0058] When the user turns their head or walks, the visual camera captures the scene and synchronizes it with the IMU data to obtain two preliminary poses.
[0059] The semantic module detects environmental features such as "seat" and "wall" to assist in correcting the helmet's position, ensuring that it does not drift under complex textures or moving light sources.
[0060] The system automatically calculates weights based on the stability of the three-way estimation and then performs real-time weighted fusion to output the accurate position and orientation of the user's head-mounted display in the virtual theater.
[0061] When users move to different viewing positions, the system displays a current positioning error prompt at the edge of the screen to ensure a smooth experience.
[0062] Input the user's head movement in relation to the controller;
[0063] Output real-time fused pose and trust weights for subsequent interaction and rendering.
[0064] Step 3: Multimodal Interaction Fusion
[0065] Users can pause the session via voice, gestures, or by looking at the menu, and the system dynamically integrates multimodal signals to generate interactive commands.
[0066] When a user says "pause" or makes a pause gesture to the screen, the system collects data from the microphone, camera, and eye tracker in parallel.
[0067] The recognition confidence of each modality is calculated separately; the system obtains the spatial confidence of step 2 and maps it to the context weight of each modality.
[0068] The system generates a weighted interaction vector by fusing the original confidence score with the spatial context weight.
[0069] If a user says "select a seat" and looks at a seat, the system will prioritize interpreting the gaze interaction, highlight the corresponding seat on the screen, and then confirm the voice command.
[0070] Input user voice, gestures, and eye movements;
[0071] Output a unified interaction vector and specific commands, and pass them to the state scheduling module.
[0072] Step 4 Joint State Construction and Scheduling
[0073] The system intelligently switches between "viewing and roaming" and "command response" based on the user's location and interaction intent.
[0074] When the user remains stationary and does not issue any commands, the state vector indicates "idle," and the system continues playing the movie.
[0075] If the user moves beyond a preset threshold (e.g., 0.5m), the system enters "roaming" mode, and the camera smoothly follows the user's viewpoint, mapping the scene to roam.
[0076] When a user issues an interactive command (such as "fast forward" or "volume up"), the system first enters "interactive" mode, pauses roaming, and executes the corresponding control until the command is completed.
[0077] After the interaction is complete, if the user continues to move, the system will automatically return to roaming mode; if the user remains stationary for an extended period, the system will revert to idle mode.
[0078] Enter the idle / roaming / interactive state;
[0079] Output corresponding rendering and control commands.
[0080] Step 5: Adaptive Rendering and Feedback
[0081] The system dynamically adjusts the viewpoint, sound effects, and haptic feedback based on the current mode and location reliability.
[0082] In "roaming" mode, when the location confidence is high, the camera follows the user's head movements more closely; if the confidence is low, the director's preset shots are used more often to avoid drastic camera shake.
[0083] When the user turns their head to look to the left or right, the system dynamically renders spatial sound effects so that the sound comes from the corresponding direction, enhancing the sense of realism.
[0084] When users make gestures such as waving or nodding during the climax of the story, the haptic gloves provide vibration feedback, which is synchronized with the plot to enhance the sense of immersion.
[0085] If a user triggers commands such as pause / fast forward multiple times, the system will use haptic feedback to confirm whether the operation has been performed, preventing accidental touches.
[0086] Input the current mode, fused pose, and command;
[0087] Output rendering view updates, sound effect parameters, and haptic commands.
[0088] Step 6: Multi-user Collaboration and Session Management
[0089] When multiple viewers are online at the same time, the system synchronizes the status according to confidence level and allocates interactive resources.
[0090] When users A, B, and C enter the same viewing room at the same time, the system collects their fusion status and location reliability, and weights and synthesizes them into a public view to ensure that everyone sees the same picture.
[0091] If A and B both want to adjust the volume, the system compares their positional confidence scores, prioritizing the one with the higher score, and the other party receives a "Try again later" message.
[0092] The system divides multiple users into several spatial clusters and schedules the nearest edge node to cache movie slices, ensuring smooth playback for each user.
[0093] If user C disconnects due to network fluctuations, the system records their last location and command sequence. Once they reconnect, the system automatically inserts them back into the viewing flow, and the camera position smoothly transitions to the group view.
[0094] Input each user's behavior and status;
[0095] Output group synchronization results, resource allocation decisions, and session recovery commands.
[0096] Step 7: Expanding Interactions and Logs
[0097] Users can install script plugins or connect to a cloud-based AI assistant, and the system records logs and manages them dynamically as needed.
[0098] Users can select "Virtual Guide" or "Smart Recommendation" scripts in the "Plugin Marketplace," and the system will automatically activate the most relevant plugins based on the current space and interaction context.
[0099] The plugin responds to user queries (such as "What story does this scene tell?") and broadcasts scene information in real time via text or voice.
[0100] The system records every user interaction event and dynamically adjusts the log recording frequency based on the current operation intensity, ensuring the integrity of critical operation logs while avoiding storage waste.
[0101] When a plugin is not used for a long time or is uninstalled by the user, the system automatically releases the occupied resources and reloads it the next time it is needed.
[0102] Input user actions and feedback in the plugin marketplace;
[0103] Output plugin execution, recommendation results, and log data.
[0104] Implementation Method 3: This implementation method further describes the above-provided technical solution in detail through specific implementation operations. Specifically:
[0105] The plan includes:
[0106] Core Engine: Fusion of Spatial Positioning × Multimodal Interaction
[0107] 1.1 Spatial localization fusion framework combining visual, inertial, and environmental semantics
[0108] 1.2 Unified Interaction Bus for Multimodal Inputs such as Voice, Gestures, Eye Contact, and Body Language
[0109] 1.3 Attention-based localization and dynamic weight allocation of interaction signals
[0110] Fusion spatial positioning module
[0111] 2.1 Visual-Inertial-Semantic Tripartite Fusion Algorithm
[0112] 2.2 Adaptive Correction and Drift Compensation for Environmental Changes
[0113] 2.3 Personnel Identity Association and Real-time Location Mapping
[0114] 2.4 Multi-user location synchronization and intelligent conflict scheduling
[0115] Multimodal interaction fusion module
[0116] 3.1 Speech Intent Recognition and Context Awareness
[0117] 3.2 Continuous Flow Processing of Gestures and Body Language
[0118] 3.3 Eye Tracking and Gaze Focus Mapping
[0119] 3.4 Integrating Spatial Positioning Data to Drive Interactive Scene Switching
[0120] Virtual Reality Interactive Scheduler
[0121] 4.1 Spatial-Semantic Mapping: Mapping locations one-to-one with script trigger points
[0122] 4.2 Interactive State Machine: Ordered Coordination of Localization and Multimodal Events
[0123] 4.3 Real-time switching strategy between user perspective and interactive feedback: Adaptive rendering and feedback module (software level)
[0124] 5.1 Automatic locking and smooth transition of virtual camera viewpoint
[0125] 5.2 Spatial Sound Effect Rendering Based on User Location and Interaction Intent
[0126] 5.3 Tactile Event Command Generation: Synchronous Output of Multimodal and Positioning Triggers
[0127] Collaboration and Extension Interfaces
[0128] 6.1 Multi-user network synchronization and edge acceleration
[0129] 6.2 Scripted Plugin: Customize Positioning and Interaction Integration Rules
[0130] 6.3 Third-party AI service and content recommendation interface
[0131] Security and privacy protection
[0132] 7.1 Local anonymization and encryption of location and interaction data
[0133] 7.2 Access Control: Precise Authorization and Anonymization
[0134] 7.3 Anomaly Detection: Automatic Correction of Positioning Drift and Interaction Conflicts
[0135] Beneficial effects
[0136] Achieving high-precision and robust deep fusion of spatial positioning and multimodal interaction
[0137] Users can move freely within the virtual movie scene and interact in natural and diverse ways.
[0138] The system automatically and intelligently switches the viewpoint and feedback based on the user's location and interactive input.
[0139] Supports large-scale, multi-user collaborative immersive viewing with a seamless experience.
[0140] Easily extensible and pluggable interactive scripts with secure and privacy protection mechanisms
[0141] Specifically:
[0142] I. Key Data Structures
[0143] Pose: contains a three-dimensional position vector (x,y,z), a rotation quaternion (w,x,y,z), and a 3×3 covariance matrix Σ.
[0144] MultiModal: Includes speech feature vectors (256 dimensions), gesture feature vectors (128 dimensions), and gaze feature vectors (64 dimensions).
[0145] II. System Initialization
[0146] Visual SLAM module
[0147] Load ORB-SLAM3, incorporate distortion parameters within the camera, and set the keyframe insertion frequency to 10 FPS.
[0148] IMU calibration
[0149] Read the calibrated noise density and zero bias, and statically converge for 5 seconds to estimate the bias.
[0150] Semantic segmentation module
[0151] Deploy U-Net to segment environmental elements such as "seats", "walls", and "ground", and generate geometric constraint templates.
[0152] Multimodal recognition network
[0153] The speech intent recognition model outputs a 256-dimensional vector; gesture recognition outputs 128-dimensional skeletal features; and gaze tracking outputs a 64-dimensional gaze vector.
[0154] III. Steps of Fusion Spatial Positioning Algorithm
[0155] Covariance Acquisition
[0156] The visual covariance Σ_v is fed back by the SLAM reprojection error. Example:
[0157] Σ_v=diag(0.01,0.01,0.02)(m 2 )
[0158] The IMU covariance Σ_i is updated by Kalman filtering. Example:
[0159] Σ_i=diag(0.005,0.005,0.005)
[0160] The semantic covariance Σ_s is set to a constant:
[0161] Σ_s=diag(0.0004,0.0004,0.0004)
[0162] Confidence score
[0163] For each source j∈{v,i,s}, compute
[0164]
[0165] Where ||∑||F is the Frobenius norm, ε=10 -6 .
[0166] Weight normalization
[0167]
[0168] pose fusion
[0169] The final fused pose P consists of two parts: position and rotation.
[0170]
[0171] The rotation part uses SLERP to weight and fuse quaternions.
[0172] IV. Multimodal Interaction Attention Fusion Steps
[0173] Feature extraction
[0174] Multimodal features such as speech x1, gesture x2, and gaze x3 are acquired in parallel.
[0175] Attention Score
[0176] For each mode x m Calculate the score
[0177] e m =g(x m )
[0178] Where g is a two-layer fully connected network with 64 hidden units and the activation function ReLU.
[0179] Weight calculation
[0180]
[0181] Fusion generates interaction vectors
[0182]
[0183] V. Joint State Construction and Scheduling
[0184] The system constructs a state vector at each time step.
[0185] s(t) = [P(t).position, P(t).orient, y(t)] is divided into three states: Idle, Navigate, and Interact according to the preset state machine. In the Navigate state, the virtual camera position and orientation are updated, and in the Interact state, commands such as "pause", "fast forward", and "select seat" are executed.
[0186] VI. Symbol Definition
[0187] P j : Pose estimation of source j.
[0188] ∑ j : Covariance matrix of pose estimation.
[0189] ||∑|| F : The Frobenius norm of the matrix.
[0190] β j Confidence score.
[0191] w j : Fusion weights.
[0192] x m : Feature vector of the m-th interaction mode.
[0193] g(·): Attention scoring function.
[0194] α m Multimodal attention weights.
[0195] y: The unified interaction vector after fusion.
[0196] s(t): Joint system state vector.
[0197] VII. Performance Indicators (Examples)
[0198] Positioning accuracy: ≤5mm (indoor 5×5m area)
[0199] Interactive response latency: ≤50ms (including recognition and scheduling)
[0200] Attention fusion accuracy: ≥92% (in scenarios with any missing modality)
[0201] System frame rate: ≥90 FPS (GPU: RTX 3060)
[0202] Through the above specific implementation methods, the present invention deeply integrates visual, inertial and semantic multi-source localization, as well as voice, gesture and gaze multi-modal interaction at the software level, which significantly improves positioning accuracy, interaction naturalness and system real-time performance, far superior to the existing technology.
[0203] Step 2: Multimodal Interaction Attention Fusion (Deepening the Connection with Spatial Localization)
[0204] 2.1 Spatial Context Mapping
[0205] Spatial confidence vector
[0206] Take the fusion weights calculated in step 1 for the three-source localization of vision (v), inertial (i), and semantic (s).
[0207] This vector reflects the distribution of the system's trust in location information in the current environment.
[0208] Context mapping functions
[0209] Define mapping function
[0210] c = h(w)
[0211] The spatial confidence vector is converted into a context weight vector that corresponds one-to-one with each interaction modality.
[0212] For example, it can be obtained by adding a softmax layer to a linear mapping:
[0213]
[0214] 2.2 Joint Attention Fusion
[0215] In traditional modal attention weight α m Based on this, combined with the spatial context c m Forming joint attention weights
[0216] in
[0217] g(x m ): For the m-th modal feature x m The scoring function is typically a two-layer fully connected network;
[0218] Subsequently, an interaction vector combining spatial information is generated:
[0219]
[0220] 2.3 Definition of Formula Elements
[0221] Step 1: Fusing visual (v), inertial (i), and semantic (s) localization weight vectors;
[0222] M: Number of multimodalities (e.g., total number of modalities such as speech, gestures, and gaze);
[0223] u m : A linear mapping vector for context mapping, used to map spatial weights to the m-th mode;
[0224] Spatial context weight vector.
[0225] 2.4 Beneficial Effects
[0226] Spatial awareness drives interaction: Trust levels from different location sources dynamically affect attention allocation across modalities, such as improving inertial and semantic support for gestures in low-light environments.
[0227] Improve robustness: When a certain modality is poorly identified and the corresponding spatial confidence is low, its final weight will be further weakened.
[0228] Unified Context: It achieves deep coupling of spatial positioning and multimodal interaction at the algorithm level, which significantly surpasses the effect of simple parallel fusion in existing technologies.
[0229] Step 3: Joint State Construction and Scheduling Method
[0230] 3.1 Construction of Joint State Vector
[0231] At time t, the fused pose is first obtained from step 1.
[0232] P(t)=(P(t).position, P(t).orient)
[0233] Obtain the fusion interaction vector from step 2.
[0234]
[0235] Construct the joint state vector
[0236]
[0237] Where D is the dimension of the interaction vector.
[0238] 3.2 State Confidence Measurement
[0239] Define confidence metrics for both location and interaction to allow the scheduler to balance their impact:
[0240] Spatial confidence
[0241]
[0242] Interactive confidence
[0243]
[0244] 3.3 Mode Switching Score
[0245] Calculate scores for the Idle, Navigate, and Interact modes respectively. k :
[0246] l Idle (t)=(1-σ s (t))(1-σ i (t))
[0247] l Navigate (t)=σ s (t)(1-σ i (t))
[0248] l Interact (t)=σ i (t)
[0249] This design guarantees:
[0250] When the location is highly reliable and there is no intention to interact, Navigate will be entered first.
[0251] When the interaction intent is strong, immediately enter Interact;
[0252] If both are weak, then remain Idle.
[0253] 3.4 Normalization and Decision Making
[0254] The scores are normalized into pattern weights.
[0255]
[0256] Final execution mode
[0257]
[0258] Action generation in mode 3.5
[0259] Idle: Maintains the current camera position and playback status.
[0260] Navigate: Set the virtual camera target pose to P(t), then in each frame...
[0261] P cam (t+1)=λP(t)+(1-λ)P cam (t)
[0262] Smooth interpolation (α = λ is the smoothing coefficient) enables users to roam from their own perspective.
[0263] Interact: Passes the interaction vector y(t) into the command parser and executes operations such as "pause", "fast forward", and "seat selection" according to the preset intent mapping.
[0264] 3.6 Definition of Formula Elements
[0265] s(t) is the joint state vector, σ s (t) Spatial confidence level, σ i (t) Interaction confidence, l k (t) Unnormalized score of pattern k, π k (t) Normalized probability of pattern k, k * (t) The execution mode selected at the current moment, P cam (t) Current pose of the virtual camera, λ is the immersion smoothing coefficient (0 < λ < 1).
[0266] 3.7 Beneficial Effects
[0267] Dynamically balanced positioning and interaction: It can quickly respond to user intentions and automatically follow the user's spatial movement when there is no operation;
[0268] Smooth transition: Exponential mapping and interpolation avoid abrupt perspective shifts;
[0269] Adaptive robustness: The scoring mechanism can be dynamically adjusted according to the environment and modal quality, improving the overall stability of the system.
[0270] Step 4: Complete Solution for Adaptive Rendering and Feedback Module (Software Level)
[0271] This step utilizes the spatial fusion confidence obtained in step 1, as well as the interaction and scheduling results in steps 2 and 3, to generate virtual camera perspective, adaptive spatial sound effects, and haptic feedback commands. This differs from the existing approach of fixed perspective + single feedback, and achieves deep coupling of "space-interaction-rendering" closed loop.
[0272] 4.1 Adaptive switching of virtual camera viewpoint
[0273] Spatial confidence
[0274] Take the maximum positioning weight from step 1
[0275]
[0276] Preset camera path
[0277] Suppose the camera path designed by the film director is a series of keyframe position-pose pairs:
[0278] {C k (t)=(p k (t), q k (t))}, where pk is the position and qk is the quaternion pose.
[0279] Perspective fusion formula
[0280] The user's real-time location-attitude P(t) = (p u (t), q u (t) and director path C(t) = (p d (t), q d (t) Spatial confidence-weighted fusion:
[0281] p cam (t)=σ s p u (t)+(1-σ s )p d (t),
[0282] q cam (t)=SLERP(qu (t),q d (t),1-σ s )
[0283] When the location reliability is high (σ) s →1) The viewpoint allows for more free navigation by the user;
[0284] When the confidence level is low (σ) s →0), the perspective returns more to the director's pre-set.
[0285] 4.2 Adaptive Spatial Sound Rendering
[0286] Sound source model
[0287] For each virtual sound source S k (positions) k Sound intensity I k Calculate the distance to the user
[0288] d k (t)=||p u (t)-s k ||
[0289] Distance attenuation gain
[0290] Based on squared decay and spatial confidence adjustment:
[0291]
[0292] Left and right channel distribution
[0293] Let the user's head orientation vector be o. u (t), the sound source direction vector is
[0294] The left and right gains are then...
[0295]
[0296] in This is the unit vector along the left and right axes of the head.
[0297] 4.3 Generation of haptic feedback instructions
[0298] Event trigger location
[0299] Let the position of the current tactile event in the scene be e(t).
[0300] Interaction strength
[0301] The intensity of intent is measured using the magnitude of the fused interaction vector from step 2.
[0302] I int (t)=||y(t)||2
[0303] Distance attenuation and feedback strength
[0304] Define the distance d between the user and the event. e =||p u (t)-e(t)||correlated tactile intensity:
[0305]
[0306] The tactile feedback is strongest when the user is close to the event and has a strong intention to interact; otherwise, it is appropriately weakened.
[0307] Instruction encapsulation
[0308] The tactile instruction package includes intensity H(t) and duration T. h (Can be specified by the director's script) and feedback mode (such as vibration, thrust) triplet are sent to the peripheral driver.
[0309] 4.4 Definition of Formula Elements
[0310] σ s Spatial confidence max j w j p u (t), q u (t) User's real-time location and quaternion pose, p d (t), q d (t) Director's preset camera position and posture, SLERP(q1, q2, τ) is a two-dimensional spherical linear hammer value with a ratio of τ, s k The location of the k-th virtual sound source, I k The fundamental sound intensity of the k-th virtual sound source, d k (y) Distance between the user and the sound source k, α (sound attenuation coefficient), o u (t) User's head orientation unit vector. The user's head is represented by unit vectors along the left and right axes, y(t) is the interaction vector fused in step 2, e(t) is the current tactile event position, and d is the current haptic event position. e Distance between user and event location, β tactile attenuation coefficient, H(t) tactile feedback intensity, T h Feedback duration.
[0311] 4.5 Beneficial Effects
[0312] Immersive perspective: The perspective adapts and integrates the intentions of the user and the director, ensuring both a free roaming experience and maintaining the tension of the plot.
[0313] Realistic spatial sound effects: Dynamically renders sound based on the user's posture, position, and the sound source of the scene to enhance the sense of space and orientation.
[0314] Scene-based haptic association: The intensity of haptic feedback is influenced by both spatial confidence and interaction intent, making it more accurate and natural.
[0315] Step 5: Multi-user network collaboration and session management methods
[0316] 5.1 Spatial Confidence Weighted State Synchronization
[0317] Per-user state vector
[0318] The state vector of the i-th user at time t
[0319]
[0320] Spatial confidence
[0321] The maximum location fusion weight for the i-th user calculated in step 1
[0322]
[0323] Weighted allocation coefficient
[0324]
[0325] Group fusion status
[0326] All user states are weighted by confidence level and synthesized into a global view.
[0327]
[0328] 5.2 Conflict Detection and Resolution
[0329] Virtual resource allocation conflict
[0330] Define a resource request set when multiple users compete for the same seat or interaction object. If user If multiple requests for the same resource are made, their spatial confidence scores are compared:
[0331]
[0332] Dynamic rollback mechanism
[0333] The preempted user automatically triggers the "reselection" logic, and within a radius r of the resource... safe Visual cues are provided within the area.
[0334] 5.3 Edge Caching and Task Scheduling
[0335] User clustering
[0336] Based on group state S groupThe spatial distribution of (t) is used to perform K-means clustering on users, and cluster services are allocated to K edge nodes.
[0337] Task Priority
[0338] The priority of tasks within each cluster is positively correlated with the average spatial confidence within the cluster:
[0339] Priority (c) ∝ ρ c
[0340] Cache scheduling
[0341] For the highest priority cluster, its recent movie slices and interactive scripts are preloaded to the nearest edge node to reduce cross-domain transmission latency.
[0342] 5.4 Session Disconnection and Recovery
[0343] Heart rate detection
[0344] Each user reports their status every Δt seconds. If the timeout exceeds τ, the user is marked as "out of contact" and the latest status is retained. i (t-τ).
[0345] Recovery strategy
[0346] After the user reconnects, the system uses their last state s i (t-τ) and group state S group (t) Interpolation smoothing:
[0347] s i (t 恢复 )=ηs i (t-τ)+(1-η)S group (t), 0 < η < 1
[0348] 5.5 Definition of Formula Elements
[0349] s i (t) is the joint state vector of the i-th user after fusion, σ s,i (t) Spatial confidence level max of user i i w j,i (t), γ i (t) is the weighting coefficient of the i-th user in state synchronization, N is the total number of online users, and S is the weighting coefficient of the t-th user. group (t) Group fusion state vector, The set of users requesting the same resource, r safe Resource protection radius A user cluster served by a certain edge node, ρ c The average spatial confidence of cluster c, Δt heartbeat reporting interval, τ disconnection judgment threshold, and η session recovery waveform smoothing coefficient.
[0350] 5.6 Beneficial Effects
[0351] High-precision consistent view: Utilizes spatial confidence to drive synchronization, ensuring state consistency when multiple users share a scene.
[0352] Low latency and congestion resistance: Edge caching and priority scheduling significantly reduce network transmission latency.
[0353] Robust recovery: The reconnection interpolation strategy avoids abrupt re-entry and improves user experience.
[0354] Fair resource allocation: A confidence-based conflict resolution mechanism that balances positioning accuracy and interaction requirements.
[0355] Step 6: Extending and Opening Interfaces
[0356] 6.1 Script-based Interactive Plugin Framework
[0357] Plugin Registration
[0358] Each plugin P k Declare at startup:
[0359] Supported event type set ε k (such as user gestures, voice intent, eye contact, etc.)
[0360] Weight vector parameters Used for matching with system context.
[0361] Context vector
[0362] At time t:
[0363] Spatial positioning weight (Step 1)
[0364] Fuse the interaction vector y (step 2),
[0365] Constitute the total vector
[0366] Plugin activation score
[0367] Calculate a matching score for each plugin:
[0368]
[0369] Only when γ k (t) > θ (threshold), plugin P k It is dynamically loaded and the current event is processed.
[0370] Real-time uninstallation
[0371] If a certain plugin continuously Tidle If the resource is not activated within a few seconds, it will be automatically uninstalled and released.
[0372] 6.2 Third-party AI assistant and content recommendation interface
[0373] Contextual features
[0374] Similarly, z(t) is used as the input context for the external API.
[0375] Suggested score
[0376] The third-party assistant AmA_m returns multiple candidate suggestions {r m Each suggestion r is composed of feature vectors ,1,...} The local score indicates:
[0377]
[0378] The system returns the top K items with the highest scores.
[0379] Real-time fine-tuning
[0380] Based on user feedback (e.g., acceptance / rejection), online incremental updates will be made. r and u k Weighting enhances personalization.
[0381] 6.3 Log and Behavior Analysis Integration Specifications
[0382] Dynamic sampling rate
[0383] Let the base log sampling rate be ρ0, and adjust it according to spatial confidence and interaction intent strength as follows:
[0384] ρ(t)=ρ0(α s σ s (t)+α i σ i (y)), α s +α i =1
[0385] Where σ s (t) and σ i (t) is defined in step 3.
[0386] Layered logging
[0387] Core events (location switching, command execution) are sampled at 100%;
[0388] Auxiliary events (UI updates, status heartbeats) are sampled by pressing ρ(t);
[0389] Debug events (internal diagnostics) are only recorded in debug mode at the lowest sampling rate.
[0390] Log format
[0391] Each log entry contains: timestamp, user ID, w, y, and current status k. * This ensures traceability across the entire "space-interaction-scheduling-rendering-interface" chain.
[0392] 6.4 Definition of Formula Elements
[0393] P k The k-th scripted interactive plugin, ε k The set of event types supported by broadcast, u k The play-in matching weight vector has the same dimension as B, z(t) = [w; y(t)] context vector, combining spatial weights and interaction vectors, γ k (t) broadcast item P k Dynamic activation score, theta-play activation value, T idle Playback idle timeout unload read value, A m The m-th third-party AI assistant, q r Suggest feature vectors for r, s r (t) Context matching score of suggestion r, ρ0 Basic log sampling rate, ρ(t) Dynamic log sampling rate, α s α i Log sampling weight coefficient.
[0394] 6.5 Beneficial Effects
[0395] On-demand loading: Plugins are only loaded when highly relevant, saving resources and improving response speed;
[0396] Context-aware recommendation: Third-party assistants match in real time based on space and interaction, making suggestions more accurate;
[0397] Controllable log volume: Dynamic sampling ensures full recording of critical events while reducing storage and transmission pressure;
[0398] Highly scalable: Adding new plugins and assistants only requires registering weight vectors, without modifying the core scheduling logic.
[0399] Step 7: Security and Privacy Protection Module
[0400] 7.1 Context-Aware Encryption
[0401] Dynamic encryption strength
[0402] The system automatically adjusts the key length or number of iterations of the encryption algorithm based on the current location confidence level.
[0403] When location reliability is high, maintain the standard encryption configuration to balance performance;
[0404] When the location confidence is low, the encryption strength can be increased (e.g., by lengthening the key or increasing the number of key derivation rounds) to resist potential attacks.
[0405] Session key derivation on demand
[0406] The session key is not only based on a fixed random seed, but also incorporates the current location weight, so that the key for each connection and each user is bound to its spatial trust level, thereby improving the uniqueness and security of the key.
[0407] 7.2 Differential Privacy Protection
[0408] Location data noise injection
[0409] When sharing or recording user location, instead of directly exposing the real coordinates, random noise is injected into the real location, and the amount of noise is inversely linked to the location reliability.
[0410] When the location reliability is high, the noise is low, ensuring the usability of the data;
[0411] When the location reliability is low, the noise is relatively large, which further protects privacy.
[0412] Anonymization of interaction logs
[0413] Before collecting and uploading multimodal interaction logs, the system randomly perturbs the context vector, making it difficult for any external analysis to trace the sensitive behavior of a single user, while preserving the overall statistical characteristics.
[0414] 7.3 Fine-grained dynamic access control
[0415] Trust-based permission tiers
[0416] Based on the current spatial confidence level, users are dynamically divided into three levels: high trust, medium trust, and low trust.
[0417] High trust: Access to complete personalized location and interaction data;
[0418] Trust: Access to the aggregated situational view is only permitted;
[0419] Low trust: Only public information can be seen; no user-level sensitive content can be accessed.
[0420] Real-time permission decision
[0421] Each time a data access request is made, the system assesses the user's current trust level and role permissions in real time to decide whether to authorize or deny, achieving "constant assessment and on-demand authorization," which is significantly better than traditional static role authorization.
[0422] Beneficial effects
[0423] Safety and performance balance
[0424] The encryption strength is dynamically adjusted according to the level of trust, which not only ensures data security in high-risk scenarios, but also avoids unnecessary overhead when high performance is required.
[0425] Privacy protection adaptive
[0426] The differential privacy noise injection amount can be flexibly controlled according to the positioning accuracy, ensuring data availability while protecting privacy.
[0427] Least Privilege Principle
[0428] Access control is based on real-time trust assessment, ensuring minimal authorization and preventing abuse of permissions.
[0429] End-to-end protection
[0430] From location data and interaction logs to external interfaces, there are targeted security strategies to build a true end-to-end protection system.
[0431] The above description of several specific embodiments further illustrates the technical solution provided by the present invention in order to highlight the advantages and benefits of the technical solution provided by the present invention. However, the above-described specific embodiments are not intended to limit the present invention. Any reasonable modifications and improvements to the present invention, combinations of embodiments, and equivalent substitutions based on the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A virtual reality interaction method for movie playback that integrates spatial positioning and multimodal interaction, characterized in that, include: Step 1, System Initialization: The software loads the visual SLAM, inertial measurement unit calibration module, scene semantic segmentation model, and speech, gesture, and gaze multimodal recognition model, and outputs the ready modules; Step 2, Spatial localization fusion: The software acquires visual SLAM pose, IMU pose and semantic-assisted pose in parallel, and fuses them according to confidence level to generate global pose and localization confidence vector. Step 3, Multimodal Interaction Fusion: The software receives the positioning confidence vector output from Step 2 and the voice, gesture, and gaze features collected in parallel. It dynamically fuses these features based on the recognition confidence of each modality and the spatial context weights, and outputs a unified interaction vector. Step 4, Joint State Scheduling: The software constructs the system state based on the fused pose from Step 2 and the interaction vector from Step 3, calculates the spatial confidence and interaction intent strength, intelligently switches between three modes: idle, roaming, and interactive, and outputs the current mode and control commands. Step 5, Adaptive Rendering and Feedback: The software receives the mode and control commands output in Step 4, combines the fixed position confidence dynamic weighted user perspective with the director's preset shot, renders the virtual camera perspective, adjusts the spatial sound effects in real time, and generates haptic feedback commands. Step 6, Multi-user Collaborative Management: The software collects the fusion status and confidence level of each user, synthesizes a weighted group view, resolves interactive resource conflicts based on confidence level, and supports smooth recovery after disconnection and reconnection; Step 7, Expanding Interfaces and Logs: The software dynamically activates script plugins and third-party AI assistants based on group status and session information, adjusts the log collection frequency as needed, and outputs a list of activated modules and log data.
2. The virtual reality interaction method for movie playback that integrates spatial positioning and multimodal interaction according to claim 1, characterized in that, The confidence-weighted fusion uses the ratio of the localization covariance of visual, inertial and semantic approaches to generate weights.
3. The virtual reality interaction method for movie playback that integrates spatial positioning and multimodal interaction according to claim 1, characterized in that, The spatial context weights are obtained by linearly mapping and normalizing the localization confidence vector output in step two, and are used to adjust the attention allocation for each modality.
4. The virtual reality interaction method for movie playback that integrates spatial positioning and multimodal interaction according to claim 1, characterized in that, The mode switching dynamically calculates the mode priority based on the ratio of spatial confidence to interaction intent intensity, and determines the current execution mode through normalized Soft-max.
5. A virtual reality interaction method for movie playback that integrates spatial positioning and multimodal interaction according to claim 1, characterized in that, Adaptive rendering uses positional confidence as interpolation coefficients to dynamically and smoothly switch between the user's free viewpoint and the director's preset viewpoint, and attenuates spatial sound effects in real time according to the distance between the user and the virtual sound source.
6. A virtual reality interactive system integrating spatial positioning and multimodal interaction for movie playback, characterized in that, include: Module 1, System Initialization: The software loads the visual SLAM, inertial measurement unit calibration module, scene semantic segmentation model, and speech, gesture, and gaze multimodal recognition model, and outputs the ready modules; Module 2, Spatial Localization Fusion: The software acquires visual SLAM pose, IMU pose, and semantic-assisted pose in parallel, and then fuses them according to confidence level to generate a global pose and localization confidence vector. Module 3, Multimodal Interaction Fusion: The software receives the positioning confidence vector output from step 2 and the voice, gesture, and gaze features collected in parallel. It dynamically fuses these features based on the recognition confidence of each modality and the spatial context weights, and outputs a unified interaction vector. Module 4, Joint State Scheduling: The software constructs the system state based on the fused pose from step 2 and the interaction vector from step 3, calculates the spatial confidence and interaction intent strength, intelligently switches between three modes: idle, roaming, and interactive, and outputs the current mode and control commands. Module 5, Adaptive Rendering and Feedback: The software receives the mode and control commands output from step 4, combines the fixed-position confidence dynamic weighted user perspective with the director's preset shots, renders the virtual camera perspective, adjusts spatial sound effects in real time, and generates haptic feedback commands. Module Six, Multi-User Collaborative Management: The software collects the fusion status and confidence level of each user, weights and synthesizes a group view, prioritizes the resolution of interactive resource conflicts based on confidence level, and supports smooth recovery after disconnection and reconnection; Module 7, Extended Interfaces and Logs: The software dynamically activates script plugins and third-party AI assistants based on group status and session information, adjusts the log collection frequency as needed, and outputs a list of activated modules and log data.
7. A computer storage medium for storing computer programs, characterized in that, When the computer program is read by the computer, the computer executes the method of claim 1.
8. A computer, comprising a processor and a storage medium, characterized in that, When the processor reads the computer program stored in the storage medium, the computer executes the method of claim 1.
9. A computer program product, as a computer program, is characterized by: When the computer program is executed, it implements the method of claim 1.