Augmented reality system and method for immersive experience

By collecting and analyzing physical environment and user interaction data in real time, generating rendering priorities and interaction event sequences, and dynamically compensating for virtual-reality synchronization errors, the problems of insufficient dynamic environment adaptation and user interaction personalization in augmented reality technology are solved, and the synchronization and coherence of the immersive experience are improved.

CN120747433AActive Publication Date: 2025-10-03GUANGZHOU INTEREST ISLAND INFORMATION TECH CO LTD

Patent Information

Application Number
CN202511269803.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-08
Publication Date
2025-10-03
Estimated Expiration
2045-09-08

AI Technical Summary

Technical Problem

Existing augmented reality technology has shortcomings in dynamic environment adaptation and user interaction personalization, which leads to deviations in the superposition of virtual content and real scenes, affecting immersion and user experience.

Method used

By collecting dynamic spatial data of the physical environment and user interaction behavior data in real time, we conduct environmental semantic modeling and user behavior analysis, generate rendering priority sequences and virtual-reality interaction event sequences, and dynamically compensate for virtual-reality synchronization errors.

Benefits of technology

It improves the synchronization between virtual content and the physical environment and the personalized interactive experience, enhances the user's immersion and experience continuity, and reduces discomfort reactions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120747433A_ABST
    Figure CN120747433A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of augmented reality application, and discloses an augmented reality system and method for immersive experience. The method comprises the following steps: collecting dynamic spatial data of a physical environment in real time, and synchronously obtaining an interactive behavior data stream of a user terminal; performing environment semantic modeling on the dynamic spatial data to generate multi-level feature description of a physical environment, and determining a rendering priority sequence of virtual-real superposition content according to the multi-level feature description; extracting behavior pattern characteristics in historical interaction data of the user, performing immersion correlation analysis on the behavior pattern characteristics to obtain environment cognitive preference parameters of the current user, and generating a virtual-real interaction event sequence in combination with the environment cognitive preference parameters and the interaction behavior data flow; and dynamically compensating the virtual-real synchronization error in the immersive experience process according to the rendering priority sequence and the virtual-real interaction event sequence. According to the method, the accuracy and naturalness of virtual-real fusion can be improved, the personalized requirements of the user are met, and the overall immersive experience effect is enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of augmented reality application technology, and in particular to an augmented reality system and method for immersive experience. Background Art

[0002] With the rapid development of information technology, augmented reality technology has gradually moved from the laboratory to practical applications, showing great potential in various fields such as education, medical care, entertainment, and industry. However, current augmented reality technology still faces many challenges in achieving an immersive experience. When it comes to dynamic environmental adaptation, existing technologies mostly rely on pre-set static environmental models. This makes it difficult to quickly update environmental data when the physical environment changes in real time, leading to discrepancies between the virtual content and the real scene. For example, in indoor scenes, if an object moves or the lighting suddenly changes, the virtual content often fails to respond in time, disrupting the user's immersion. The lack of personalization in the user interaction experience is also a prominent issue. Traditional methods often use a unified interaction logic, ignoring the differences in user behavior and cognitive patterns. Some users may focus more on the details of a scene, while others prefer to perceive the overall environment. These differentiated needs are difficult to effectively meet with existing technologies. Virtual-reality synchronization errors are a key factor affecting the immersive experience. Due to delays in the acquisition of dynamic spatial data and real-time changes in user interactions, temporal or spatial misalignment often occurs between virtual content and the physical environment. This misalignment directly reduces user immersion and can even cause discomfort such as dizziness. Dynamically compensating for these errors and improving the naturalness and accuracy of virtual-reality integration are key challenges in the current development of augmented reality technology. Summary of the Invention

[0003] The purpose of the present invention is to provide an augmented reality system and method for immersive experience to solve the problems raised in the above background technology.

[0004] To achieve the above objectives, the present invention provides an augmented reality method for immersive experience, the method comprising: Collect dynamic spatial data of the physical environment in real time and simultaneously obtain interactive behavior data streams from user terminals; Performing environmental semantic modeling on the dynamic spatial data to generate a multi-level feature description of the physical environment, and determining a rendering priority sequence of virtual and real overlay content based on the multi-level feature description; Extracting behavioral pattern features from historical user interaction data, performing immersion correlation analysis on the behavioral pattern features to obtain the current user's environmental cognition preference parameters, and combining the environmental cognition preference parameters with the interaction behavior data stream to generate a virtual-reality interaction event sequence; Dynamically compensate for virtual-reality synchronization errors during the immersive experience process according to the rendering priority sequence and the virtual-reality interaction event sequence.

[0005] Preferably, the environmental semantic modeling includes: constructing a three-dimensional topological structure of a physical environment according to the dynamic spatial data; The semantic anchor points of the interactive objects are marked in the three-dimensional topological structure by using a spatiotemporal association algorithm.

[0006] Preferably, determining the rendering priority sequence of virtual-real overlay content according to the multi-level feature description includes: Analyzing the ambient light intensity change curve and spatial motion trajectory in the multi-level feature description; Calculating the visual penetration coefficient of the virtual and real superimposed content based on the ambient light intensity change curve; A rendering priority sequence is generated according to the spatial motion trajectory and the visual penetration coefficient.

[0007] Preferably, extracting behavioral pattern features from user historical interaction data includes: Call the gaze dwell time distribution and gesture trigger frequency matrix in the user behavior knowledge base; The behavioral feature extraction model outputs behavioral pattern features from the gaze dwell duration distribution and gesture trigger frequency matrix.

[0008] Preferably, generating a virtual-reality interaction event sequence includes: Calculate the remaining difference between the current experience duration and the preset immersion threshold; Adjusting the timestamp alignment strategy of the virtual-real interaction event according to the remaining difference; An event stream reorganization algorithm is used to merge the environmental cognition preference parameters and the interactive behavior data stream into a virtual-reality interactive event sequence.

[0009] Preferably, dynamically compensating for the virtual-reality synchronization error during the immersive experience includes: generating a spatial position compensation amount according to the rendering priority sequence; Calculating a time delay compensation amount based on the virtual-reality interaction event sequence; The spatial position compensation amount and the time delay compensation amount are input into the error compensation engine to perform virtual-real synchronization calibration.

[0010] Preferably, the process of generating the virtual-reality interaction event sequence includes: Identify multimodal input features in interactive behavior data streams; Performing weighted superposition of the multimodal input features and the environmental cognition preference parameters through a feature fusion gateway; Based on the superposition results, a sequence of virtual-reality interaction events carrying spatiotemporal markers is generated.

[0011] Preferably, the workflow of the feature fusion gateway includes: Establish a real-time feedback channel between environmental cognitive preference parameters and multimodal input features; When an abnormal jump in multimodal input features is detected, a dynamic correction mechanism for environmental cognitive preference parameters is triggered; The corrected environmental cognitive preference parameters are re-weighted and superimposed with the multimodal input features.

[0012] Preferably, the method further comprises: Generate a compensation intensity adjustment factor according to the accumulated value of the virtual and real synchronization errors; The virtual-real synchronization compensation map is constructed by using the spatial position compensation amount, time delay compensation amount and compensation intensity adjustment factor; The execution frequency of the error compensation engine is controlled based on the virtual-real synchronization compensation map.

[0013] Preferably, the present invention further includes an augmented reality system for immersive experience, configured to execute the above-mentioned augmented reality method for immersive experience, the system comprising: The environmental perception module is used to collect dynamic spatial data of the physical environment from multiple environmental sensors and synchronously obtain the interactive behavior data stream of the user terminal; a semantic modeling module for performing environmental semantic modeling on the dynamic spatial data in a virtual-reality fusion mode, generating a multi-level feature description of the physical environment, and determining a rendering priority sequence of virtual-reality overlay content based on the multi-level feature description; An interaction analysis module is used to extract behavioral pattern features from historical user interaction data, perform immersion correlation analysis on the behavioral pattern features to obtain environmental cognition preference parameters, and generate a sequence of virtual-reality interaction events; A synchronization compensation module is used to perform dynamic compensation operations on virtual-reality synchronization errors during the immersive experience process according to the rendering priority sequence and the virtual-reality interaction event sequence.

[0014] Compared with the prior art, the present invention has the following beneficial effects: Real-time collection of dynamic spatial data from the physical environment and interactive behavior data streams from user terminals lays the foundation for high-precision virtual-reality fusion. Environmental semantic modeling of dynamic spatial data and the generation of multi-level feature descriptions enable a more comprehensive analysis of the physical environment's attributes and structure. The resulting rendering priority sequence for virtual-reality overlays ensures that the presentation of virtual content more closely matches the actual characteristics of the environment, avoiding interference from irrelevant information. This allows users to focus on key content and enhance the coherence of the experience. Extracting behavioral pattern features from historical user interaction data and conducting immersion correlation analysis can deeply explore users' personalized needs and cognitive habits. The obtained environmental cognitive preference parameters combined with the virtual-reality interaction event sequence generated by the real-time interactive behavior data stream can make the interaction process more in line with user expectations, reduce the sense of operational discomfort, and make users interact with virtual content more natural and smooth, thereby enhancing users' sense of identification with the experience. Dynamic compensation for virtual-reality synchronization errors is performed based on the rendering priority sequence and virtual-reality interaction event sequence. This can adjust the presentation status of virtual content in real time and promptly correct deviations in time or space, so that virtual content and the physical environment always maintain a high degree of consistency, avoiding immersion interruptions caused by asynchrony. This allows users to more easily feel immersed in the experience, reduces the occurrence of discomfort reactions, and enhances the overall immersive experience from multiple dimensions. This method works synergistically from three levels: environmental perception, user cognition, and error compensation, forming a closed-loop optimization system that can adapt to complex and changing physical environments and diverse user needs, enabling augmented reality technology to demonstrate better performance in actual applications and bringing users a higher-quality, more immersive experience. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] Figure 1 This is a timing diagram of the augmented reality method for immersive experience according to the present invention; Figure 2 Flowchart generated for the rendering priority sequence; Figure 3 A flowchart generated for the virtual-reality interaction event sequence; Figure 4 This is a flow chart of dynamic compensation of virtual and real synchronization errors; Figure 5 Flowchart of how the feature fusion gateway works. DETAILED DESCRIPTION

[0016] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0017] See also Figure 1 The present invention provides an augmented reality method for immersive experience, the method comprising: The system collects dynamic spatial data of the physical environment in real time, obtains the environmental point cloud sequence and inertial motion trajectory through the depth camera array, inertial measurement unit and lidar sensor deployed on the user terminal, and simultaneously captures the interactive behavior data stream through the user terminal's touch sensor, eye tracking unit and voice acquisition module. The interactive behavior data stream includes gesture coordinate sequence, eye movement vector and voice command feature vector.

[0018] Perform environmental semantic modeling on the dynamic spatial data: input the point cloud sequence into a three-dimensional convolutional neural network, extract the spatial geometric topological features and texture semantic features, and generate a multi-level feature description including a macroscopic scene structure layer, a mesoscopic object semantic layer, and a microscopic surface attribute layer; based on the multi-level feature description, analyze the changing gradient of the ambient light intensity within a continuous time window and the motion acceleration of the interactive objects to generate a rendering priority sequence of virtual and real superimposed content.

[0019] Extract behavioral pattern features from historical user interaction data: call the historical gaze dwell time distribution matrix and gesture trigger frequency matrix from the user behavior knowledge base, extract behavioral association rules in the time dimension through the long short-term memory neural network, and output environmental cognitive preference parameters; input the environmental cognitive preference parameters and the real-time interaction behavior data stream into the event sequence generator, and output the virtual-real interaction event sequence through the timestamp alignment strategy.

[0020] Dynamic compensation of virtual-reality synchronization errors is performed based on the rendering priority sequence and the virtual-reality interaction event sequence: a spatial compensation vector is generated based on the spatial position offset in the rendering priority sequence, and a time compensation coefficient is calculated based on the timestamp delay of the virtual-reality interaction event sequence; the spatial compensation vector and the time compensation coefficient are input into the error compensation engine, and the calibration operation is completed through spatial transformation matrix reprojection and timestamp interpolation algorithm.

[0021] Example 1: See Figure 2The system analyzes the ambient light intensity change curves and spatial motion trajectories in the multi-level feature descriptions. The visual penetration coefficient of the virtual and real superimposed content is calculated based on the ambient light intensity change curves. A rendering priority sequence is generated based on the spatial motion trajectories and visual penetration coefficients. In the environmental semantic modeling phase, the dynamic spatial data is used to construct a 3D topological structure of the physical environment using a point cloud registration algorithm. Starting from point cloud data in consecutive frames, the iterative closest point algorithm is used to register the discrete frame point clouds to a unified world coordinate system, generating an environmental topological mesh. This mesh consists of a set of vertices and triangular facets. Vertex coordinates contain spatial position information, while triangular facet indices describe the geometric continuity of the environmental surface. Within the generated 3D topological structure, semantic anchor labeling is performed: an object detection model is used to identify the categories of interactive objects in the point cloud, and a spatiotemporal correlation algorithm is used to calculate the spatial continuity index of each object in consecutive frames. Specifically, the change in the Euclidean distance between the object's center of mass between adjacent frames is calculated. When this change in multiple consecutive frames is below a preset spatial displacement threshold, the object is deemed to have stable spatial properties and a semantic anchor is labeled at the corresponding mesh vertex. The semantic anchor data includes object category encoding, spatial pose matrix and dynamically updated identifier. The pose matrix consists of rotation vector and translation vector, which is used to describe the precise orientation of the object in three-dimensional space.

[0022] When determining the rendering priority sequence based on multi-level feature descriptions, the ambient light intensity change curve is first analyzed. The light intensity values ​​within a continuous time window are collected by a photometric sensor array deployed on the user terminal to generate a time-light intensity function curve. A temporal differential operation is performed on this function curve to calculate the gradient value of the light intensity change per unit time. This gradient value is input into a nonlinear mapping function to output the visual penetration coefficient. This coefficient is negatively correlated with the transparency control parameter of the virtual and real superimposed content, that is, when the light gradient increases, the penetration coefficient decreases, and the opacity of the corresponding virtual and real superimposed content needs to be increased to maintain visual consistency. At the same time, the instantaneous dynamic parameters of the moving object are extracted from the spatial motion trajectory: the bounding box of the moving object is identified by differential operations on the point clouds of adjacent frames, and the instantaneous velocity vector is calculated based on the displacement of the center of mass. The visual penetration coefficient and the instantaneous velocity vector are input into the priority decision module to generate a rendering priority sequence. The decision logic includes dual conditional judgments: when the instantaneous velocity modulus of a moving object exceeds the dynamic threshold and the visual penetration coefficient is lower than the critical range, the virtual and real superimposed content associated with the object is given the highest rendering priority; when the object is in a low-speed motion state and the visual penetration coefficient is in the medium range, it is given secondary rendering priority; when the scene elements remain static and the penetration coefficient is in the high value range, it is given basic rendering priority.

[0023] The process of analyzing spatial motion trajectories involves a trajectory prediction algorithm for moving objects. Within a 3D topological grid, trajectory modeling is performed for objects with stable semantic anchor points within consecutive frames. A Kalman filter is used to predict the object's spatial coordinates at the next moment, and a polynomial function representing the trajectory is generated based on historical position data. The derivative of this function is used to calculate instantaneous acceleration, which serves as an auxiliary parameter in rendering priority decisions. For example, for high-speed moving objects, if their acceleration continues to increase, their rendering priority is further increased.

[0024] The processing of the ambient light intensity curve includes an anti-interference mechanism. Because the ambient light sensor is susceptible to interference from local shadows or transient reflections, a sliding window mean filter is used to smooth the raw light data and eliminate impulse noise. The smoothed light curve is then subjected to gradient calculation to prevent outliers from causing sudden changes in the visual penetration coefficient. The mapping function for the light gradient value uses a piecewise linear model: a gentle mapping curve is used in low-gradient ranges to prevent frequent fluctuations in the penetration coefficient caused by weak light fluctuations; a steep mapping curve is used in high-gradient ranges to ensure a rapid response to strong light changes.

[0025] The final generation of the rendering priority sequence needs to integrate the spatial position correlation. In the three-dimensional topological structure, the scene area is divided according to the spatial distribution density of the semantic anchor points. For areas with dense anchor points, even if the motion parameters of a single object do not reach the threshold, the overall rendering priority of the area is increased due to the concentrated interaction risk; for areas with sparse anchor points, only the priority of independent objects that reach the threshold is increased. At the same time, the priority sequence adopts a dynamic refresh mechanism: the sequence content is updated every time a frame of point cloud data is processed, and each entry in the sequence contains a unique identifier of the virtual and real superimposed content, a priority weight value, and an associated spatial grid area code. The sequence is transmitted to the rendering pipeline, driving the GPU resources to allocate rendering computing power according to the priority weight.

[0026] During the projection phase of virtual and real overlay content, the rendering priority sequence directly controls shader parameters. The highest-priority content uses a high-quality shading model, employing physical lighting calculations and anti-aliasing; the next-highest-priority content uses a simplified shading model; and the base-priority content uses pre-baked lightmaps. The visual penetration coefficient is converted into an alpha channel value, and the alpha blending parameters of the overlay content are dynamically adjusted through the fragment shader. As the penetration coefficient decreases, the opacity weight is increased to maintain visibility of the overlay content in bright light environments; as the penetration coefficient increases, the opacity weight is reduced to prevent the overlay content from obscuring real-world details in low-light environments.

[0027] The dynamic update mechanism for semantic anchors ensures spatial consistency. When an annotated object is detected to have moved beyond the tolerance range, spatial continuity checks are re-evaluated. If the new position still meets the consecutive frame displacement threshold, the semantic anchor pose matrix is ​​updated. If the displacement exceeds the limit, resulting in a break in continuity, the original anchor is deleted and a new object detection process is triggered. Anchor update events trigger the reconstruction of the rendering priority sequence, ensuring that changes in the state of moving objects are promptly reflected in rendering decisions.

[0028] The optimization of the three-dimensional topology structure adopts an incremental update strategy. Newly acquired point cloud frames are aligned with the existing mesh through a local registration algorithm. The mesh expansion operation is only performed on the newly added environmental areas to avoid the computational load caused by global reconstruction. The mesh vertex attributes contain timestamps. When the vertex is not updated for a set period of time, it will automatically become invalid, preventing outdated spatial data from interfering with semantic analysis. The spatial resolution of the topological mesh is adaptively adjusted according to the complexity of the scene: the mesh subdivision density is automatically increased in areas with dense semantic anchor points, and the subdivision density is reduced in open areas to save computing resources.

[0029] The entire process forms a closed-loop processing flow: point cloud data input drives 3D topology construction and semantic anchor annotation; anchor information is combined with lighting and motion parameters to generate a rendering priority sequence; this sequence controls the visual presentation of virtual and real overlays; and user interaction is fed back to the point cloud acquisition end, influencing the processing focus of subsequent frames. This closed-loop mechanism enables real-time response to dynamic changes in the physical environment, ensuring that augmented reality content remains visually and interactively synchronized with the physical space.

[0030] Example 2: See Figure 3The residual difference between the current experience duration and the preset immersion threshold is calculated. The timestamp alignment strategy for virtual-reality interaction events is adjusted based on this residual difference. An event stream reassembly algorithm is used to fuse environmental cognitive preference parameters with the interaction behavior data stream into a virtual-reality interaction event sequence. The behavioral pattern feature extraction process is based on a structured dataset stored in the user behavior knowledge base. The gaze dwell duration distribution matrix uses a three-dimensional tensor data structure: the first dimension records the spatial region grid code, the second dimension represents the time slot division, and the third dimension stores the user's cumulative gaze duration within the corresponding spatiotemporal unit. The gesture trigger frequency matrix is ​​a two-dimensional matrix: the row vectors index predefined gesture type codes, and the column vectors record the number of times each gesture is triggered in a specific spatial region per unit time. When processing this input data, a two-channel convolutional neural network model is configured with a one-dimensional convolution kernel group in the first channel, performing a sliding convolution operation along the time dimension of the gaze matrix. The convolution kernel width is set to the number of consecutive time slots. Through multi-layer convolution, gaze pattern features along the time dimension are extracted, including implicit patterns such as periodic gaze behavior and attention shift intervals. The second channel uses a square convolution kernel to perform two-dimensional convolution on the space-gesture type plane of the gesture frequency matrix to capture the association pattern between gesture type and spatial position, such as the tendency of zoom gestures or rotation gestures to be triggered frequently in a specific area.

[0031] The fusion mechanism for the dual-channel output feature vectors includes a feature alignment operation. Due to the dimensionality differences between the gaze feature vector and the gesture feature vector, the dual-channel outputs are projected into a latent space of uniform dimensionality via a fully connected layer. After feature concatenation within the latent space, a nonlinear transformation is performed through two fully connected layers to output a behavioral pattern feature vector. This vector contains compressed information about the user's interaction habits, such as visual preferences for specific spatial dimensions and high-frequency interaction gesture combinations.

[0032] Environmental cognitive preference parameters are generated through immersion correlation analysis. After receiving the behavioral pattern feature vector, the analysis module first decomposes it into a spatial attention component and an interactive behavior component. The spatial attention component is input into a cluster comparison unit, where it is matched against attention distribution templates from typical historical scenarios. This outputs a user's cognitive propensity score for open spaces, enclosed spaces, or transitional spaces. The interactive behavior component is input into a temporal pattern parser, which detects the temporal regularity of periodic interactive behaviors and outputs a user's preference index for continuous versus discrete interactions. Finally, the spatial propensity score and interactive preference index are combined to generate a multi-dimensional set of environmental cognitive preference parameters.

[0033] When the generation process of the virtual-reality interaction event sequence is initiated, the remaining difference between the current experience duration and the preset immersion threshold is first calculated. The immersion threshold is set based on the user's historical experience data: samples of the duration of previous effective immersion experiences are extracted, and the ideal experience duration benchmark is obtained through Gaussian distribution fitting. The remaining difference calculation module continuously monitors the progress of the current experience duration. When the remaining time is lower than a certain proportion of the benchmark value, the dynamic adjustment of the timestamp alignment strategy is triggered. The adjustment mechanism uses a timeline compression algorithm: the timestamp sequence of the original interaction behavior data stream is nonlinearly remapped to proportionally reduce the event time interval. The compression ratio is dynamically calculated based on the decreasing amplitude of the remaining difference, and the maximum compression ratio is used when the remaining time approaches the critical threshold.

[0034] The event stream reassembly algorithm performs multi-source data fusion operations. Environmental cognition preference parameters are first encoded into a weight matrix, where the row vectors correspond to different interaction modalities (such as eye movements, gestures, and voice), and the column vectors represent the weighted distribution values ​​of sub-features within each modality. The multimodal input in the interactive behavior data stream is parsed into a feature tensor: eye movement data is converted into a sequence of gaze focus coordinates, gesture data is parsed into tuples of type codes and spatial coordinates, and voice commands are converted into intent vectors through a semantic parsing engine. In the reassembly layer, the Hadamard product operation is performed on the features of each modal data and the weight matrix to achieve preference-weighted feature strengthening or weakening. For example, when the environmental cognition preference parameters indicate that the user pays close attention to visual details, the corresponding coefficient of the eye movement feature in the weight matrix is ​​significantly improved.

[0035] The weighted multimodal features are integrated through the spatiotemporal alignment module. This module establishes a spatiotemporal coordinate system mapping relationship: the eye focus coordinates are associated with the world coordinate system anchor point, the gesture space coordinates are converted to the world coordinate system through the hand tracking model, and the speech intent vector is bound to the spatial direction according to the current user's orientation. The aligned feature vectors are sorted by timestamp to generate discrete event units with spatiotemporal markers. Each event unit contains three elements: a generation timestamp accurate to the millisecond level (based on atomic clock synchronization), a three-dimensional position coordinate in the world coordinate system, an event type code, and a feature vector.

[0036] Dynamic optimization of event sequences includes a mechanism for filtering out anomalous events. During event stream reassembly, the spatiotemporal continuity of event units is monitored in real time. By calculating the ratio of the spatial distance change to the time interval between adjacent events, jump events, potentially caused by sensor noise, are detected. When anomalous patterns such as spatial teleportation or temporal reversal are detected, an event interpolation correction program is triggered. This program uses historical event sequences to construct a spatiotemporal trajectory prediction model, replacing the coordinates of anomalous events with predicted values ​​to maintain the physical plausibility of the event sequence.

[0037] The output structure of the event sequence utilizes a hierarchical indexing design. The base layer stores raw event units in strict timestamp order. The middle layer establishes a spatial region index, categorizing events by coordinates into pre-set spatial grid cells. The top layer constructs an event type hash table, enabling fast retrieval by interaction type. When subsequent modules request event data within a specific spatiotemporal range, the multi-layer indexing allows for rapid location of relevant event subsets, reducing computational latency for real-time queries.

[0038] The entire implementation process forms a data-driven closed-loop optimization. After executing a newly generated virtual-reality interaction event sequence, its triggering effects are recorded in the user behavior knowledge base, which is used to update the gaze retention matrix and gesture frequency matrix. When significant changes in user behavior patterns are detected, the environmental cognition preference parameters are recalculated, in turn influencing the generation logic of subsequent event sequences. This mechanism enables the system to adapt to the gradual evolution of user interaction habits, maintaining the effectiveness of the immersive experience during extended use.

[0039] The timestamp alignment strategy adjustment process includes a smooth transition mechanism. When the compression ratio needs to change, a gradual switching algorithm is employed: the compression factor is gradually adjusted over multiple consecutive processing cycles to avoid experience interruptions caused by sudden changes in event time intervals. Information about compression parameter changes is added to the event sequence metadata area for reference by subsequent compensation modules. For example, if the event sequence is detected to have undergone time compression, the time delay compensation module will adjust its reference clock synchronization strategy accordingly.

[0040] The conflict resolution mechanism for multimodal input ensures the inherent consistency of event sequences. When modalities such as eye movement, gesture, and speech generate semantically conflicting inputs at the same time (e.g., eye movement focused on area A while gestures manipulated area B), the reconstructive algorithm initiates modal confidence assessment: It assigns decision weights based on the recognition accuracy of each modality in that scenario from historical data, prioritizes feature data from high-confidence modalities, and marks potential conflicting events for special processing by subsequent modules. The feature vector in the event unit includes a conflict flag, which triggers additional spatial position verification during the subsequent virtual-reality synchronization calibration process.

[0041] The entire implementation architecture is designed to balance computational efficiency and accuracy. The behavioral feature extraction model utilizes a lightweight convolution kernel configuration, optimized for fixed-point arithmetic when deployed on mobile devices. The event stream reassembly algorithm supports pipeline parallel processing, allowing feature extraction and weighting operations for different modalities, such as eye movements, gestures, and voice, to be executed in independent threads. Finally, synchronization barriers are used for spatiotemporal alignment and integration. This design effectively utilizes multi-core processor resources and meets the stringent real-time responsiveness requirements of augmented reality applications.

[0042] Example 3: See Figure 4, generate spatial position compensation based on the rendering priority sequence, calculate time delay compensation based on the virtual-real interaction event sequence, and input the spatial position compensation and time delay compensation into the error compensation engine to perform virtual-real synchronization calibration. When the virtual-real synchronization error dynamic compensation operation is started, the generation of spatial position compensation is based on the parsing result of the rendering priority sequence. Each entry in the sequence contains the spatial offset information of the virtual-real superposition content, and the offset comes from the difference between the three-dimensional topological structure in the environment semantic modeling stage and the actual sensor detection position. The process of the pose solution algorithm processing this difference value includes the separate calculation of the rotation component and the translation component: the direction offset is converted into a rotation compensation vector through quaternion interpolation, and the translation compensation vector is generated through Euclidean distance decomposition, and finally combined into a six-degree-of-freedom compensation parameter. The parameter is formatted as a homogeneous transformation matrix, and the matrix elements contain three axial rotation angles and three-dimensional translations.

[0043] The calculation of time delay compensation depends on the timestamp alignment of the virtual-reality interaction event sequence. The millisecond timestamp carried by each entry in the event sequence is compared with the high-precision clock of the rendering system to generate a time difference sequence. The compensation coefficient matrix is ​​generated through statistical analysis of the time difference sequence: the standard deviation and mean of the time difference within a fixed time window are calculated. The standard deviation reflects the degree of delay fluctuation, and the mean reflects the baseline delay. The compensation coefficient is composed of a linear combination of the baseline delay and the degree of fluctuation, and its mathematical expression is: ; in: Indicates the Class events in The compensation coefficient for the time period, is the mean time difference, is the standard deviation of the time difference, and is a preset weight factor. This coefficient matrix is ​​organized by event type and time block, supporting refined delay compensation for specific interaction scenarios.

[0044] The error compensation engine consists of two main execution units: the spatial transformation module and the frame buffer controller. The spatial transformation module receives the six-degree-of-freedom compensation parameters and converts them into projection matrix corrections. This implementation utilizes a chain of affine transformations: first, rotation compensation is applied to the model matrix of the augmented reality content, updating the object orientation through quaternion multiplication. Secondly, translation compensation is applied to the view matrix, modifying the origin of the observation coordinate system. Finally, the projection matrix is ​​fine-tuned within the frustum to correct for perspective distortion caused by positional offsets. Each transformation step utilizes an incremental update mode to avoid the computational burden of global matrix reconstruction.

[0045] The frame buffer controller adjusts the rendering timing according to the compensation coefficient matrix. The controller maintains a triple buffer architecture: the front buffer receives the newly rendered frame, the middle buffer performs time compensation processing, and the back buffer outputs to the display device. When the value exceeds the threshold, the dynamic frame scheduling mechanism is started: the interpolation algorithm is used to generate transition frames between the front-end buffer and the middle buffer. The number of transition frames is proportional to the value of the buffer. At the same time, the triggering timing of the vertical synchronization signal of the back-end buffer is adjusted to keep the timestamp of the final output frame synchronized with the changes in the physical environment.

[0046] The generation of virtual-reality interaction event sequences begins with the recognition of multimodal input features. The interaction behavior data stream is deconstructed into three parallel processing channels: the gesture recognition channel uses a skeletal tracking algorithm to extract the spatial coordinates of hand joints, generating gesture trajectory vectors; the eye tracking channel calculates the angle between the pupil center and the corneal reflection point, outputting the projection of the gaze focus onto the screen coordinate system; and the speech processing channel uses voiceprint separation technology to extract user commands, which are then converted into semantic action codes using an intent recognition model. Each channel's output is accompanied by a confidence score and a timestamp.

[0047] The feature fusion gateway implements a three-layer weighted superposition architecture. The first layer performs spatial association operations: the gesture spatial coordinates are converted from the screen coordinate system to the world coordinate system, and the spatial proximity with the eye focus coordinates is calculated. The proximity calculation uses an adaptive radius judgment method: when the Euclidean distance between the gesture operation point and the eye focus is less than the dynamic threshold, it is judged as effective collaborative interaction, and a collaborative interaction flag is generated. The second layer performs semantic integration: the semantic vector of the voice command and the collaborative interaction flag are spliced ​​into a mixed feature vector, and the dimension is compressed through a fully connected network. The third layer implements environmental cognition fusion: the compressed feature vector is matrix multiplied with the environmental cognition preference parameters. The environmental cognition preference parameters are organized as a weight matrix, the row vectors of the matrix correspond to different interaction scenarios, and the column vectors represent the weight distribution scheme of the feature elements in each scenario.

[0048] The resulting virtual-reality interaction event sequence utilizes a dual spatiotemporal tagging mechanism. The spatial tag contains the 3D position coordinates and coverage radius in the world coordinate system, calculated using the weighted centroid of the gesture control point and the eye focus. The temporal tag includes the precise timestamp of the event generation moment (synchronized to an atomic clock) and the valid duration interval. Each event entry is accompanied by a confidence value output by the feature fusion gateway. This value is calculated by combining the intermediate results of each layer's processing and is used to assess the reliability of subsequent compensation operations.

[0049] The spatial position compensation process includes a dynamic verification phase. After the compensation engine applies the six-degree-of-freedom parameters, the environment semantic modeling module verifies the post-compensation virtual-real alignment in real time. Key semantic anchor points are selected within the 3D topology and the positional deviation between the projected virtual content and the actual anchor points is calculated. If the deviation exceeds the pre-compensation baseline, a rollback mechanism for the compensation parameters is triggered, restarting the pose calculation process. This feedback mechanism prevents overcompensation caused by sensor noise.

[0050] Special case processing of time delay compensation involves event sequence prediction. When the value shows a monotonically increasing trend, a pre-compensation mechanism is activated: an autoregressive prediction model is established based on the historical time difference series to generate compensation coefficients for future time slices in advance. These pre-compensation coefficients are injected into the forward scheduling module of the frame buffer controller, enabling the rendering pipeline to prepare frame processing plans for high-latency periods in advance. The prediction model uses a sliding window training mechanism, retaining only sample data from the most recent time period, ensuring rapid response to changes in user interaction patterns.

[0051] Multimodal input conflicts are resolved in the third layer of the feature fusion gateway. When semantic conflicts arise between gesture, eye movement, and voice input features (e.g., eye movement gazing at location A while gestures are directed at location B), the environmental cognitive preference parameters initiate conflict resolution. Based on the reliability of each modality in similar scenarios as documented in the user's historical behavioral data, the weight matrix is ​​dynamically adjusted. For example, if historical data indicates a user prefers gesture-driven operations, the coefficient for gesture features in the weight matrix is ​​automatically increased, while the coefficient for eye movement features is correspondingly decreased. The conflict resolution log is appended to the event sequence entry for subsequent analysis modules.

[0052] Throughout the compensation system's operation, the spatial transformation module and the frame buffer controller maintain state synchronization. A distributed transaction mechanism ensures the atomicity of both types of compensation operations. When spatial compensation parameters and temporal compensation coefficients need to be updated simultaneously, a version control marker is established. Output operations are executed only when the version markers of the two types of compensation data match, avoiding spatial-temporal compensation misalignment caused by transmission delays. The version marker, which combines a timestamp hash value with a spatial grid code, ensures accurate state synchronization.

[0053] The event sequences for virtual-reality interaction are stored using a spatiotemporal partitioned database. Event entries are partitioned by spatial grid index in the world coordinate system, and within each partition, they are sorted by millisecond timestamps. This structure supports efficient spatial range queries and temporal range retrieval. When the synchronization compensation module requests recent events in a specific area, it can quickly return results through a parallel query mechanism. The database implements an incremental backup strategy, retaining only event data from active interaction areas and transferring data from inactive areas to cold storage to conserve resources.

[0054] Example 4: See Figure 5 , establish a real-time feedback channel between environmental cognitive preference parameters and multimodal input features. When an abnormal jump in the multimodal input features is detected, the dynamic correction mechanism of the environmental cognitive preference parameters is triggered, and the corrected environmental cognitive preference parameters are re-weighted and superimposed with the multimodal input features. The real-time feedback channel of the feature fusion gateway adopts a circular buffer structure. The buffer contains continuous storage units, each of which records the complete state snapshot of the most recent feature fusion. The state snapshot contains four types of core data: environmental cognitive preference parameter version number, multimodal input feature vector, weighted superposition result value, and timestamp mark. The buffer adopts a first-in-first-out management strategy. The new data overwrites the earliest record to maintain the traceability of the historical state within a fixed time window.

[0055] When a multimodal input feature undergoes an abnormal jump, the jump detection algorithm is activated based on a dynamic threshold. The detection criteria include two main criteria: first, calculating the rate of change of the Euclidean distance between the current frame's feature vector and its historical mean, where the historical mean is the average of the most recent records in the ring buffer; and second, analyzing the standard deviation mutation of each dimension of the feature vector. The anomaly flag is triggered when the following logical relationship is met: the rate of change of the distance exceeds a specific multiple of the historical fluctuation range, and the standard deviation increments of at least two dimensions exceed independent thresholds. For key data fragments from a particular detection process, see Table 1.

[0056] Table 1: Key data fragments during a certain detection process.

[0057]

[0058] The triggering process of the dynamic correction mechanism includes four stages: the first stage immediately suspends the current weighted overlay pipeline and freezes the feature input port; the second stage starts the historical scene matching retrieval, uses the spatial topological features of the current environmental semantic model as the retrieval key, and searches for the scene record with the highest similarity in the historical behavior pattern feature library; the third stage calls the environmental cognitive preference parameters corresponding to the historical scene to replace the current parameters. If the match fails, the default parameter template is enabled; the fourth stage reinitializes the weighted overlay matrix.

[0059] The adaptive weight allocation algorithm is executed during feature re-integration. The confidence scores of each modal input feature are derived from the real-time output of the front-end recognition module: gesture recognition confidence is calculated based on the stability of skeletal joint tracking, eye movement confidence depends on the accuracy of pupil contour recognition, and speech confidence is derived from the signal-to-noise ratio of speech endpoint detection. The weight calculation uses a nonlinear mapping function to achieve exponential weight gain for high-confidence features. The specific allocation process includes a three-step normalization operation: first, the raw confidence of the sub-features within each modality is converted into relative weights; second, the weights are balanced between the modalities to prevent over-dominance of a single modality; and finally, the weight sum constraint is applied to ensure the reversibility of the superposition matrix.

[0060] Take the scenario where user gestures suddenly drift as an example: when the hand tracking sensor is disturbed by strong light and causes coordinates to jump, the confidence level of the gesture features drops sharply to a low value range. At this time, the feature fusion gateway detects that the rate of change of the X / Y coordinates deviates significantly from the historical pattern and immediately freezes the current fusion process. The system searches the historical database and finds that users rely more on eye movement interaction in past scenes under similar lighting conditions, and then calls the corresponding historical parameter template. In the new round of weighted superposition, the weight of the gesture feature is compressed to a smaller proportion of the baseline value, and the weight of the eye movement and voice features is increased accordingly, so that the final output virtual-reality interaction event sequence is free from contamination by abnormal gesture data.

[0061] The ring buffer depth is optimized using a scene-aware adjustment strategy. In physically stable indoor environments, the buffer depth is set to a high value to capture long-term behavioral patterns. When users enter dynamic outdoor environments, the buffer depth is automatically reduced to increase sensitivity to transient changes. Depth adjustment is based on the complexity of the 3D topology: the density change rate of semantic anchor points and the entropy of the spatial distribution are calculated. When the entropy value exceeds a threshold, the buffer depth is reduced.

[0062] Historical scene matching retrieval utilizes a two-level index structure. The first-level index is based on the topological fingerprint of the environment semantic model, which encodes the three-dimensional grid structure into a fixed-length feature code using a hash function. The second-level index is based on temporal context labels, marking the user's activity type within the current time period. The retrieval process is performed in the intersection space of the two-level indexes, prioritizing historical records with a topological fingerprint match greater than a critical value and consistent time tags. The matching degree is calculated using the Hamming distance to measure the difference in topological fingerprints, and the behavioral pattern feature vector of the historical record must be in the same cluster as the current vector.

[0063] The reconstruction process of the weighted overlay matrix includes a feature dimension adaptation mechanism. When the historical environmental cognitive preference parameters used in the replacement are inconsistent with the current multimodal input feature dimensions, a dimensional projection transformation is initiated: using a pre-trained autoencoder network, the historical parameters are embedded into the current feature space, maintaining parameter semantic invariance. The projection network uses a fully connected architecture: the input layer receives the historical parameter vector, the hidden layer performs nonlinear transformations, and the output layer dimensionality is aligned with the current feature space. This process ensures that parameters from different sources are effectively integrated into a unified space.

[0064] The exception handling logging mechanism provides traceability and analysis capabilities. Each dynamic correction event generates a detailed log entry, including the exception trigger timestamp, transition signature, source of the replacement parameters, weight distribution scheme, and a summary of the final output event sequence. Log entries are stored using binary compression and appended to the corresponding virtual-reality interaction event sequence metadata area. When multiple consecutive correction events occur in the same spatial area, an offline optimization task for environmental perception preference parameters is triggered, updating the user behavior model through batch training.

[0065] The entire implementation process establishes a fault isolation boundary. The feature fusion gateway's exception handling unit runs in an independent sandbox environment, exchanging data with the main interaction pipeline through a secure channel. If the correction mechanism fails continuously for more than a threshold number of times, a degraded processing mode is initiated: bypassing the weighted fusion process, directly outputting the original feature sequence of the multimodal input, and simultaneously sending a notification of experience degradation to the user terminal. The sandbox environment is automatically reset after each successful correction, clearing the temporary state to ensure the purity of subsequent processing.

[0066] The final output of the virtual-reality interaction event sequence is appended with a correction flag. This flag consists of two status bits: the high bit indicates whether it has undergone dynamic correction, and the low bit records the type of correction (historical parameter replacement / default template activation). The subsequent synchronization compensation module adjusts its processing strategy based on the flag bit status: for corrected event sequences, the sensitivity of spatial position compensation is reduced; for events using the default template, the safety margin for time delay compensation is increased. This flag transmission mechanism forms a cross-module collaborative fault-tolerant chain.

[0067] Example 5: The cumulative value statistics of the virtual-real synchronization error adopt a sliding window counter mechanism. The counter records the triggering events of the spatial position compensation amount and the time delay compensation amount within a fixed time period. Each event contains a compensation type identifier, a trigger timestamp and a spatial area code. When the density of compensation events in a specific spatial area within the window exceeds the dynamic threshold, the compensation intensity adjustment factor generation process is triggered. The adjustment factor calculation introduces a nonlinear response curve: when the density of compensation events is in a low range, a gentle growth mode is adopted, and when the density enters a high range, it switches to an exponential growth mode, so that the factor value matches the actual demand. The generated adjustment factor is attached with a regional label and stored in conjunction with the compensation data of the corresponding spatial position.

[0068] A hierarchical mapping strategy is implemented to convert spatial position compensations into a three-dimensional vector field. The translational component in the original compensation data is directly used as the base value of the vector field, while the rotational component is converted into a direction vector. In the three-dimensional spatial grid, each voxel cell stores the average compensation vector at that location, and spatial continuity is achieved through interpolation of adjacent cell vectors. Vector field updates are performed using incremental optimization: newly generated compensations are first compared with the historical mean. When the deviation exceeds the tolerance, local reconstruction of the vector field is initiated; otherwise, a weighted average smoothing method is used for updating. Field intensity distribution maps are visualized through color coding for system status monitoring.

[0069] Time delay compensation is processed as a time gradient field. The time delay value is mapped to the gradient change rate along the time axis, and a delay distribution model is established in the four-dimensional space-time coordinate system. The time gradient field is partitioned by spatial region, and each partition maintains an independent time delay change curve. The curve sampling points contain timestamps, delay baseline values, and fluctuation amplitudes, and a continuous gradient function is generated through curve fitting. The gradient field update mechanism includes outlier filtering: when the deviation between the new delay value and the predicted value of the fitted curve exceeds the historical fluctuation range, it is marked as a temporary outlier and not temporarily stored.

[0070] The compensation intensity adjustment factor participates in the tensor product operation as a scalar field. The outer product operation of the scalar field and the three-dimensional vector field generates a second-order tensor field, which describes the intensity modulation relationship of the spatial position compensation. Simultaneously, the scalar field and the time gradient field undergo a tensor contraction operation to generate the intensity correction coefficient in the time dimension. The result of the tensor product operation constitutes the core data structure of the virtual-real synchronization compensation map. Each spatiotemporal coordinate point in the map stores a six-tuple parameter: the spatial compensation vector, the time delay gradient, the adjustment factor intensity value, and their combined effect coefficient.

[0071] The execution frequency control of the error compensation engine is based on real-time analysis of the compensation map. In the three-dimensional spatial grid, the product of the modulus of the compensation vector and the strength of the adjustment factor is calculated for each unit. When the product value exceeds the critical threshold, the unit is marked as a high-frequency compensation area. For high-frequency compensation areas, the engine execution frequency is increased to an integer multiple of the base value, and the multiple is dynamically calculated based on the composite effect coefficient. At the same time, the slope change of the time gradient field is detected: when the delay gradient of a specific spatial area continues to increase, a step-by-step increase strategy is implemented for the time compensation frequency of that area.

[0072] A strength adjustment mechanism is introduced during the execution of spatial position compensation. Before applying the six-degree-of-freedom compensation parameters, the translation vector modulus and rotation angle values ​​are scaled by the composite effect coefficient. The scaling ratio is positively correlated with the strength of the adjustment factor, but an upper threshold is set to prevent overcompensation. When high-intensity compensation is triggered repeatedly in the same spatial region, a compensation effect evaluation loop is initiated: the rate of change of the virtual-real alignment deviation before and after compensation is compared. If the deviation does not decrease by the expected percentage, the subsequent compensation intensity is increased by a step value.

[0073] The frequency control of time delay compensation is linked to the frame buffer depth. When the compensation map indicates a persistently high temporal gradient in a certain area, the frame buffer controller dynamically increases the number of intermediate buffer layers. These additional layers are used to store pre-rendered frame sequences, with interpolation algorithms generating transition frames to fill the delay gaps. The buffer layer depth is proportional to the temporal gradient value, but is soft-limited due to hardware memory constraints. Furthermore, the trigger period of the vertical sync signal is fine-tuned based on the gradient change rate, shortening it to improve response speed when latency fluctuates significantly.

[0074] The compensation map update mechanism establishes a closed-loop feedback loop. After each error compensation engine execution, new spatial position deviation and time delay data are collected as feedback input. This data is compared with the map prediction value to calculate the residual. When the residual exceeds a threshold, a local correction of the map is triggered: the spatial vector field adjusts the vector direction through backpropagation, the temporal gradient field corrects the curve fitting parameters, and the adjustment factor updates the intensity value according to the residual ratio. The version number of the corrected map is incremented to ensure that subsequent operations are based on the latest state.

[0075] Resource scheduling for high-frequency compensation areas implements priority management. Spatial grid cells are divided into different compensation levels based on a comprehensive score of the three-dimensional vector field modulus and the strength of the adjustment factor. The engine execution thread prioritizes the highest-level areas, and areas of equal level are sorted in descending order of temporal gradient value. When system resources are limited, a compensation result reuse strategy is implemented for lower-level areas: compensation parameters from historically similar states are reused after linear transformation to reduce real-time computational load. Reuse decisions must ensure that spatial position deviations are within an acceptable range.

[0076] The execution frequency control module includes a safety protection mechanism. A maximum frequency threshold for a single region is set to prevent system overload. When the frequency of a region approaches the threshold, collaborative compensation for adjacent regions is triggered: compensation vectors for adjacent grids are interpolated, and part of the compensation task is distributed to surrounding areas. The engine core temperature and memory usage are also monitored. When hardware indicators exceed the limit, global frequency degradation is initiated: the execution frequency of all regions is reduced by a fixed ratio until the system returns to normal. Degradation status information is added to the compensation map metadata area for subsequent analysis.

[0077] The entire implementation process forms an adaptive dynamic compensation system. Newly generated synchronization error data is continuously fed into the statistics module to update the cumulative value; the updated adjustment factor drives the reconstruction of the compensation map; the optimized map controls the engine's execution strategy; and the execution results are fed back to the error acquisition terminal, forming a closed loop. This system achieves refined control of virtual-reality synchronization errors within system resource constraints by synergistically adjusting compensation intensity and frequency. The parameter adaptability during long-term operation enables the system to respond to changes in user behavior patterns and fluctuating environmental conditions.

[0078] Compensation effect evaluation data is stored in a historical knowledge base. Key metrics such as residual values, correction parameters, and resource utilization generated by each closed-loop feedback loop are archived by scenario. When a similar environment semantic model is activated again, the historically optimal compensation configuration is prioritized to initialize the graph parameters. The knowledge base creates an index linking environmental topology fingerprints with compensation parameter sets, enabling rapid matching of historical solutions through semantic similarity retrieval. This mechanism accelerates system convergence in new scenarios and improves user experience consistency.

[0079] The compensation graph's persistent storage uses a differential incremental strategy. Full storage is performed only when the graph structure undergoes substantial changes; regular updates only record parameter differential increments. The storage format includes binary data blocks and metadata description files, supporting fast loading and version rollbacks. A lightweight compression algorithm is enabled for mobile device deployment to reduce storage space while ensuring real-time access performance. The graph version management service ensures data consistency across multiple terminals in the distributed system.

[0080] It should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that includes a list of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus.

[0081] While embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions, and variations may be made to these embodiments without departing from the principles and spirit of the invention, and that the scope of the invention is defined by the appended claims and their equivalents.

Claims

1. An augmented reality method for immersive experience, characterized in that: The method comprises: Collect dynamic spatial data of the physical environment in real time and simultaneously obtain interactive behavior data streams from user terminals; Performing environmental semantic modeling on the dynamic spatial data to generate a multi-level feature description of the physical environment, and determining a rendering priority sequence of virtual and real overlay content based on the multi-level feature description; Extracting behavioral pattern features from historical user interaction data, performing immersion correlation analysis on the behavioral pattern features to obtain the current user's environmental cognition preference parameters, and combining the environmental cognition preference parameters with the interaction behavior data stream to generate a virtual-reality interaction event sequence; Dynamically compensate for virtual-reality synchronization errors during the immersive experience process according to the rendering priority sequence and the virtual-reality interaction event sequence.

2. The augmented reality method for immersive experience according to claim 1, wherein: The environmental semantic modeling includes: constructing a three-dimensional topological structure of a physical environment according to the dynamic spatial data; The semantic anchor points of the interactive objects are marked in the three-dimensional topological structure by using a spatiotemporal association algorithm.

3. The augmented reality method for immersive experience according to claim 1, wherein: Determining the rendering priority sequence of virtual and real overlay content according to the multi-level feature description includes: Analyzing the ambient light intensity change curve and spatial motion trajectory in the multi-level feature description; Calculating the visual penetration coefficient of the virtual and real superimposed content based on the ambient light intensity change curve; A rendering priority sequence is generated according to the spatial motion trajectory and the visual penetration coefficient.

4. The augmented reality method for immersive experience according to claim 1, wherein: Extracting behavioral pattern features from historical user interaction data includes: Call the gaze dwell time distribution and gesture trigger frequency matrix in the user behavior knowledge base; The behavioral feature extraction model outputs behavioral pattern features from the gaze dwell duration distribution and gesture trigger frequency matrix.

5. The augmented reality method for immersive experience according to claim 1, wherein: Generating a virtual-reality interaction event sequence includes: Calculate the remaining difference between the current experience duration and the preset immersion threshold; Adjusting the timestamp alignment strategy of the virtual-real interaction event according to the remaining difference; An event stream reorganization algorithm is used to merge the environmental cognition preference parameters and the interactive behavior data stream into a virtual-reality interactive event sequence.

6. The augmented reality method for immersive experience according to claim 1, wherein: Dynamic compensation for virtual-reality synchronization errors during immersive experience includes: generating a spatial position compensation amount according to the rendering priority sequence; Calculating a time delay compensation amount based on the virtual-reality interaction event sequence; The spatial position compensation amount and the time delay compensation amount are input into the error compensation engine to perform virtual-real synchronization calibration.

7. The augmented reality method for immersive experience according to claim 1, wherein: The generation process of the virtual-reality interaction event sequence includes: Identify multimodal input features in interactive behavior data streams; Performing weighted superposition of the multimodal input features and the environmental cognition preference parameters through a feature fusion gateway; Based on the superposition results, a sequence of virtual-reality interaction events carrying spatiotemporal markers is generated.

8. The augmented reality method for immersive experience according to claim 7, wherein: The workflow of the feature fusion gateway includes: Establish a real-time feedback channel between environmental cognitive preference parameters and multimodal input features; When an abnormal jump in multimodal input features is detected, a dynamic correction mechanism for environmental cognitive preference parameters is triggered; The corrected environmental cognitive preference parameters are re-weighted and superimposed with the multimodal input features.

9. The augmented reality method for immersive experience according to claim 1, wherein: Also includes: Generate a compensation intensity adjustment factor according to the accumulated value of the virtual and real synchronization errors; The virtual-real synchronization compensation map is constructed by using the spatial position compensation amount, time delay compensation amount and compensation intensity adjustment factor; The execution frequency of the error compensation engine is controlled based on the virtual-real synchronization compensation map.

10. An augmented reality system for immersive experience, configured to execute the augmented reality method for immersive experience according to any one of claims 1 to 9, characterized in that: The system comprises: The environmental perception module is used to collect dynamic spatial data of the physical environment from multiple environmental sensors and synchronously obtain the interactive behavior data stream of the user terminal; a semantic modeling module for performing environmental semantic modeling on the dynamic spatial data in a virtual-reality fusion mode, generating a multi-level feature description of the physical environment, and determining a rendering priority sequence of virtual-reality overlay content based on the multi-level feature description; An interaction analysis module is used to extract behavioral pattern features from historical user interaction data, perform immersion correlation analysis on the behavioral pattern features to obtain environmental cognition preference parameters, and generate a virtual-reality interaction event sequence; A synchronization compensation module is used to perform dynamic compensation operations on virtual-reality synchronization errors during the immersive experience process according to the rendering priority sequence and the virtual-reality interaction event sequence.

Citation Information

Patent Citations

  • Multi-vehicle parallel intelligent cooperative search and rescue system based on digital twinning and construction method thereof

    CN118092437A

  • Immersive space virtual-real interaction method and system based on meta universe

    CN119311124A

  • Cognitive rendering of inputs in virtual reality environments

    US20200104580A1

Cited By

  • Interaction system based on immersion degree

    CN121523553A

  • An immersion-based interactive system

    CN121523553B

  • Digital technology grabbed graph enhancement method and system

    CN121810894A

  • Digital technology grab pattern enhancement method and system

    CN121810894B