An augmented reality system and method for immersive experiences
By collecting and analyzing physical environment and user interaction data of augmented reality systems in real time, rendering priorities and interaction event sequences are generated, and virtual-real synchronization errors are dynamically compensated. This solves the problems of dynamic environment adaptation and personalized user interaction in augmented reality technology, and improves the coherence and immersion of the immersive experience.
Patent Information
- Application Number
- CN202511269803.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-08
- Publication Date
- 2025-11-07
- Estimated Expiration
- 2045-09-08
AI Technical Summary
Existing augmented reality technologies have shortcomings in adapting to dynamic environments and personalizing user interactions, resulting in discrepancies between the overlay of virtual content and real-world scenes and a reduction in user immersion.
By collecting dynamic spatial data of the physical environment and user interaction data in real time, environmental semantic modeling and user behavior analysis are performed to generate rendering priority sequences and virtual-real interaction event sequences for overlaying virtual and real content, and to dynamically compensate for virtual-real synchronization errors.
It achieves high-precision synchronization between virtual content and the physical environment, enhancing the user's immersion and experience continuity, and adapting to complex and ever-changing physical environments and diverse user needs.
Smart Images

Figure CN120747433B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of augmented reality application, in particular to an augmented reality system and method for immersive experience. BACKGROUND
[0002] With the rapid development of information technology, augmented reality technology gradually moves from the laboratory to practical application, showing great potential in education, medical treatment, entertainment, industry and other fields. However, the current augmented reality technology still faces many challenges in realizing immersive experience.
[0003] In terms of dynamic environment adaptation, most existing technologies rely on pre-set static environment models. When the physical environment changes in real time, it is difficult to update the environment data quickly, resulting in deviation in the superposition of virtual content and real scene. For example, in an indoor scene, if the position of an object moves or the light suddenly changes, the virtual content often cannot respond in time, destroying the user's sense of immersion.
[0004] Insufficient personalization of user interaction experience is also a major problem. Traditional methods mostly use unified interaction logic, ignoring the differences in behavior habits and cognitive patterns of different users. Some users may pay more attention to detailed information in the scene, while others tend to perceive the overall environment. This differentiated demand is difficult to meet effectively in existing technologies.
[0005] Virtual-real synchronization error is always a key factor affecting immersive experience. Due to the delay in collecting dynamic spatial data, real-time changes in user interaction behavior, and other reasons, there is often a time or spatial mismatch between virtual content and physical environment. This error directly reduces the user's sense of immersion and even causes users to feel dizzy and other discomforts. How to dynamically compensate for these errors and improve the naturalness and accuracy of virtual-real fusion is a problem that needs to be solved in the development of current augmented reality technology. SUMMARY
[0006] The purpose of the present application is to provide an augmented reality system and method for immersive experience to solve the problems raised in the background art.
[0007] To achieve the above purpose, the present application provides an augmented reality method for immersive experience, which comprises:
[0008] Real-time collection of dynamic spatial data of the physical environment and synchronous acquisition of interaction behavior data stream of the user terminal;
[0009] Environment semantic modeling of the dynamic spatial data to generate a multi-level feature description of the physical environment, and determining a rendering priority sequence of virtual-real superposition content according to the multi-level feature description;
[0010] extracting behavior pattern features in user historical interaction data, performing immersion degree correlation analysis on the behavior pattern features to obtain an environmental cognitive preference parameter of a current user, and generating a virtual-real interaction event sequence in combination of the environmental cognitive preference parameter and the interaction behavior data stream;
[0011] dynamically compensating for virtual-real synchronization errors in the immersive experience process according to the rendering priority sequence and the virtual-real interaction event sequence.
[0012] Preferably, the environmental semantic modeling comprises:
[0013] constructing a three-dimensional topological structure of a physical environment according to the dynamic spatial data;
[0014] marking semantic anchor points of an interactable object in the three-dimensional topological structure through a space-time correlation algorithm.
[0015] Preferably, determining a rendering priority sequence of virtual-real superimposed content according to the multi-level feature description comprises:
[0016] analyzing an environmental light intensity variation curve and a spatial motion trajectory in the multi-level feature description;
[0017] calculating a visual penetration coefficient of virtual-real superimposed content based on the environmental light intensity variation curve;
[0018] generating a rendering priority sequence according to the spatial motion trajectory and the visual penetration coefficient.
[0019] Preferably, extracting behavior pattern features in user historical interaction data comprises:
[0020] calling a gaze dwell duration distribution and a gesture trigger frequency matrix in a user behavior knowledge base;
[0021] outputting behavior pattern features from the gaze dwell duration distribution and the gesture trigger frequency matrix through a behavior feature extraction model.
[0022] Preferably, generating a virtual-real interaction event sequence comprises:
[0023] calculating a remaining difference between a current experience duration and a preset immersion threshold;
[0024] adjusting a timestamp alignment strategy of virtual-real interaction events according to the remaining difference;
[0025] fusing the environmental cognitive preference parameter and the interaction behavior data stream into a virtual-real interaction event sequence through an event stream reorganization algorithm.
[0026] Preferably, dynamically compensating for virtual-real synchronization errors in the immersive experience process comprises:
[0027] generate a spatial position compensation amount according to the rendering priority sequence;
[0028] calculate a time delay compensation amount based on the virtual-real interaction event sequence;
[0029] input the spatial position compensation amount and the time delay compensation amount into an error compensation engine to perform virtual-real synchronization calibration.
[0030] Preferably, the generation process of the virtual-real interaction event sequence comprises:
[0031] identify multi-modal input features in the interaction behavior data stream;
[0032] weight and superimpose the multi-modal input features and environment cognition preference parameters through a feature fusion gateway;
[0033] generate a virtual-real interaction event sequence carrying a space-time label according to the superimposition result.
[0034] Preferably, the workflow of the feature fusion gateway comprises:
[0035] establish a real-time feedback channel for the environment cognition preference parameters and the multi-modal input features;
[0036] when an abnormal jump of the multi-modal input features is detected, trigger a dynamic correction mechanism for the environment cognition preference parameters;
[0037] reweight and superimpose the corrected environment cognition preference parameters and the multi-modal input features.
[0038] Preferably, the method further comprises:
[0039] generate a compensation intensity adjustment factor according to the cumulative value of the virtual-real synchronization error;
[0040] construct a virtual-real synchronization compensation atlas through the spatial position compensation amount, the time delay compensation amount and the compensation intensity adjustment factor;
[0041] control the execution frequency of the error compensation engine based on the virtual-real synchronization compensation atlas.
[0042] Preferably, the present application further comprises an augmented reality system for immersive experience, which is used to execute an augmented reality method for immersive experience as described above, and the system comprises:
[0043] an environment perception module for collecting dynamic space data of a physical environment from a plurality of environment sensors, and synchronously acquiring an interaction behavior data stream of a user terminal;
[0044] a semantic modeling module, configured to perform semantic modeling on the dynamic space data in the virtual-real fusion mode, generate a multi-level feature description of the physical environment, and determine a rendering priority sequence of the virtual-real superimposed content according to the multi-level feature description;
[0045] an interaction analysis module, configured to extract behavior pattern features from the user historical interaction data, perform immersion degree correlation analysis on the behavior pattern features to obtain an environment cognitive preference parameter, and generate a virtual-real interaction event sequence;
[0046] a synchronization compensation module, configured to perform dynamic compensation on the virtual-real synchronization error in the immersive experience process according to the rendering priority sequence and the virtual-real interaction event sequence.
[0047] Compared with the prior art, the present application has the following advantages:
[0048] By collecting the dynamic space data of the physical environment and the interaction behavior data stream of the user terminal in real time, a foundation is laid for realizing high-precision virtual-real fusion. The semantic modeling on the dynamic space data and the generation of the multi-level feature description can more comprehensively analyze the attributes and structure of the physical environment, and the rendering priority sequence of the virtual-real superimposed content determined on this basis can make the presentation of the virtual content more in line with the actual features of the environment, avoid the interference of irrelevant information, enable the user to focus more on the key content, and enhance the coherence of the experience.
[0049] The extraction of the behavior pattern features from the user historical interaction data and the immersion degree correlation analysis can deeply mine the personalized needs and cognitive habits of the user, and the virtual-real interaction event sequence generated by combining the environment cognitive preference parameter with the real-time interaction behavior data stream can make the interaction process more in line with the user's expectations, reduce the sense of discomfort in operation, enable the user to interact with the virtual content more naturally and smoothly, and improve the user's sense of identity to the experience.
[0050] The dynamic compensation on the virtual-real synchronization error according to the rendering priority sequence and the virtual-real interaction event sequence can adjust the presentation state of the virtual content in real time, correct the deviation in time or space in a timely manner, make the virtual content always highly consistent with the physical environment, avoid the interruption of the sense of immersion caused by asynchronization, enable the user to more easily produce the feeling of being in the scene during the experience process, reduce the occurrence of uncomfortable reactions, and improve the overall immersive experience effect from multiple dimensions.
[0051] This method synergistically acts from three aspects of environment perception, user cognition and error compensation, forms a closed-loop optimization system, can adapt to complex and changeable physical environments and diversified user needs, makes the augmented reality technology exhibit better performance in actual application, and brings better quality and more immersive experience to the user. BRIEF DESCRIPTION OF DRAWINGS
[0052] Figure 1 a timing diagram for the augmented reality method for immersive experience of the present application;
[0053] Figure 2 a flowchart for generating the rendering priority sequence;
[0054] Figure 3 a flowchart for generating the virtual-real interaction event sequence;
[0055] Figure 4 a flowchart for dynamic compensation of virtual-real synchronization error;
[0056] Figure 5 a flowchart for the feature fusion gateway working. DETAILED DESCRIPTION
[0057] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative work fall within the scope of protection of the present application.
[0058] Referring to Figure 1 The present application provides an augmented reality method for immersive experience, which comprises:
[0059] Real-time acquisition of dynamic spatial data of a physical environment, acquisition of environment point cloud sequence and inertial motion trajectory through a depth camera array, an inertial measurement unit and a laser radar sensor deployed on a user terminal, and synchronous capture of interaction behavior data stream through a touch sensor, an eye tracking unit and a voice acquisition module of the user terminal, the interaction behavior data stream containing gesture coordinate sequence, eye movement vector and voice instruction feature vector.
[0060] Environment semantic modeling on the dynamic spatial data: input of the point cloud sequence into a three-dimensional convolutional neural network, extraction of spatial geometric topology features and texture semantic features, generation of multi-level feature description containing macroscopic scene structure layer, mesoscopic object semantic layer and microscopic surface attribute layer; based on the multi-level feature description, analysis of the change gradient of environment light intensity within a continuous time window and the motion acceleration of interactive objects, and generation of a rendering priority sequence of virtual-real superimposed content.
[0061] Extracting behavior pattern features in user historical interaction data: calling historical gaze residence duration distribution matrix and gesture trigger frequency matrix from user behavior knowledge base, extracting behavior association rules in time dimension through long short-term memory neural network, outputting environment cognitive preference parameters; inputting the environment cognitive preference parameters and real-time interaction behavior data stream into event sequence generator, outputting virtual-real interaction event sequence through timestamp alignment strategy.
[0062] Performing virtual-real synchronization error dynamic compensation according to the rendering priority sequence and the virtual-real interaction event sequence: generating a spatial compensation vector according to the spatial position offset in the rendering priority sequence, and calculating a time compensation coefficient based on the timestamp delay of the virtual-real interaction event sequence; inputting the spatial compensation vector and the time compensation coefficient into the error compensation engine, and completing the calibration operation through spatial transformation matrix reprojection and timestamp interpolation algorithm.
[0063] Embodiment 1: refer to Figure 2 , analyze the environmental light intensity variation curve and the spatial motion trajectory in the multi-level feature description, calculate the visual penetration coefficient of the virtual-real superimposed content based on the environmental light intensity variation curve, and generate the rendering priority sequence according to the spatial motion trajectory and the visual penetration coefficient. In the environment semantic modeling link, the dynamic spatial data constructs the three-dimensional topological structure of the physical environment through point cloud registration algorithm. This process starts from the continuous frame point cloud data, uses the iterative closest point algorithm to register the discrete frame point cloud to the unified world coordinate system, and generates the environment topological grid. The grid is composed of vertex set and triangular patch, and the vertex coordinates contain spatial position information and the triangular patch index describes the geometric continuity of the environment surface. Perform semantic anchor point labeling operation in the generated three-dimensional topological structure: identify the interactive object categories in the point cloud through the target detection model, and calculate the spatial continuity index of each object in the continuous frames through the space-time correlation algorithm. Specifically, calculate the Euclidean distance change of the object centroid between adjacent frames, and when the change is less than the preset spatial displacement threshold in continuous multiple frames, it is determined that the object has stable spatial properties, and the semantic anchor point is labeled at the corresponding grid vertex. The semantic anchor point data includes object category code, spatial pose matrix and dynamic update identifier, wherein the pose matrix is composed of rotation vector and translation vector, which is used to describe the accurate orientation of the object in three-dimensional space.
[0064] When determining the rendering priority sequence based on the multi-level feature description, first, the ambient light intensity change curve is analyzed. The light intensity value in the continuous time window is collected by the photometric sensor array deployed in the user terminal to generate the time-light intensity function curve. The time domain differential operation is performed on the function curve to calculate the light intensity change gradient value per unit time. The gradient value is input into the nonlinear mapping function to output the visual penetration coefficient. The coefficient is negatively correlated with the transparency control parameter of the virtual-real superimposed content, that is, when the light gradient increases, the penetration coefficient decreases, and the virtual-real superimposed content needs to increase the opacity to maintain visual consistency. At the same time, the instantaneous dynamic parameters of the moving object are extracted from the spatial motion trajectory: the motion object bounding box is identified by the point cloud difference operation of adjacent frames, and the instantaneous speed vector is calculated based on the centroid displacement. The visual penetration coefficient and the instantaneous speed vector are input into the priority decision module to generate the rendering priority sequence. The decision logic includes double condition judgment: when the instantaneous speed module length of the moving object exceeds the dynamic threshold and the visual penetration coefficient is lower than the critical range, the virtual-real superimposed content associated with the object is given the highest rendering priority; when the object is in a low-speed motion state and the visual penetration coefficient is in the medium interval, the secondary rendering priority is given; when the scene elements remain static and the penetration coefficient is in the high value interval, the basic rendering priority is given.
[0065] The analysis process of the spatial motion trajectory involves a motion object trajectory prediction algorithm. In the three-dimensional topological grid, the object with stable semantic anchor points in the continuous frame is subjected to motion trajectory modeling: the spatial coordinates of the object at the next time are predicted by the Kalman filter, and the motion trajectory polynomial function is generated in combination with the historical position data. The derivative of the function is used to calculate the instantaneous acceleration, and the acceleration value is used as an auxiliary parameter to participate in the rendering priority decision. For example, for a high-speed moving object, if its acceleration continues to increase, its rendering priority level is further strengthened.
[0066] The processing of the ambient light intensity change curve includes an anti-interference mechanism. Since the ambient light sensor is easily disturbed by local shadows or instantaneous reflections, a sliding window mean filter is used to smooth the original light data to eliminate impulse noise. The smoothed light curve is then subjected to gradient calculation to avoid abnormal values causing the visual penetration coefficient to jump. The mapping function of the light gradient value uses a piecewise linear model: a gentle mapping curve is used in the low gradient interval to prevent frequent fluctuations in the penetration coefficient caused by weak light fluctuations; an abrupt mapping curve is used in the high gradient interval to ensure that the penetration coefficient responds quickly to strong light changes.
[0067] The final generation of the rendering priority sequence needs to integrate the spatial position correlation. In the three-dimensional topology, the scene area is divided according to the spatial distribution density of the semantic anchor points. For the anchor point dense area, even if the single object motion parameter does not reach the threshold, the rendering priority of the whole area is improved due to the concentration of interaction risk; for the anchor point sparse area, only the independent object reaching the threshold is improved in priority. At the same time, the priority sequence adopts a dynamic refreshing mechanism: the sequence content is updated after completing one frame of point cloud data processing, and each entry in the sequence contains a unique identifier of virtual-real superimposed content, a priority weight value and an associated spatial grid area code. The sequence is transmitted to the rendering pipeline to drive the GPU resource to allocate rendering computing power according to the priority weight.
[0068] In the virtual-real superimposed content projection stage, the rendering priority sequence directly controls the shader parameters. The highest priority content enables a high-quality shading model, adopts physical light calculation and anti-aliasing processing; the secondary priority content enables a simplified shading model; the basic priority content adopts a pre-baked light map. The visual penetration coefficient is converted into a transparency channel value to dynamically adjust the Alpha blending parameters of the superimposed content through the fragment shader. When the penetration coefficient decreases, the opacity weight is increased to maintain the visibility of the superimposed content in a strong light environment; when the penetration coefficient increases, the opacity weight is reduced to avoid the superimposed content blocking the real environment details in a weak light environment.
[0069] The dynamic updating mechanism of the semantic anchor points guarantees the spatial consistency. When it is detected that the labeled object displacement exceeds the tolerance range, the spatial continuity determination is re-executed: if the new position still satisfies the continuous frame displacement threshold condition, the semantic anchor point pose matrix is updated; if the displacement is out of limit and the continuity is interrupted, the original anchor point is deleted and the new object detection process is triggered. The anchor point updating event will trigger the reconstruction of the rendering priority sequence, ensuring that the motion object state changes are timely reflected in the rendering decision.
[0070] The optimization of the three-dimensional topology adopts an incremental updating strategy. The newly collected point cloud frame is aligned with the existing grid through a local registration algorithm, and only the grid expansion operation is performed on the newly added environment area to avoid the calculation load caused by global reconstruction. The grid vertex attributes contain a timestamp marker, which is automatically invalidated when the vertex is not updated for a set time limit, preventing outdated spatial data from interfering with semantic analysis. The spatial resolution of the topology grid is adaptively adjusted according to the scene complexity: the grid subdivision density is automatically increased in the semantic anchor point dense area, and the subdivision density is reduced in the empty area to save computing resources.
[0071] The whole process forms a closed-loop processing flow: point cloud data input drives three-dimensional topology construction and semantic anchor labeling; anchor information combined with lighting and motion parameters generates rendering priority sequence; sequence controls visual performance of virtual-real superimposed content; user interaction behavior feedback to point cloud acquisition end, affecting the processing focus of subsequent frames. This closed-loop mechanism realizes real-time response to dynamic changes in the physical environment, making the augmented reality content always synchronized with the physical space in visual and interactive aspects.
[0072] Embodiment 2: refer to Figure 3 The remaining difference between the current experience duration and the preset immersion threshold is calculated, and the timestamp alignment strategy of virtual-real interaction events is adjusted according to the remaining difference. The environment cognition preference parameters and interaction behavior data stream are fused into a virtual-real interaction event sequence using an event stream reorganization algorithm. The extraction process of behavior pattern features is based on the structured data set stored in the user behavior knowledge base. The gaze residence duration distribution matrix uses a three-dimensional tensor data structure: the first dimension records the spatial region grid code, the second dimension represents the time slot division, and the third dimension stores the cumulative gaze duration of the user in the corresponding space-time unit. The gesture trigger frequency matrix is a two-dimensional matrix: the row vector indexes the pre-defined gesture type code, and the column vector records the trigger times of each gesture in a specific spatial region within a unit time. When the dual-channel convolutional neural network model processes the above input data, the first channel is configured with a one-dimensional convolution kernel group, which performs sliding convolution operation along the time dimension of the gaze matrix. The convolution kernel width is set to the number of continuous time slots, and the gaze pattern features in the time dimension are extracted through multiple convolution layers, including periodic gaze behavior, attention shift interval, and other implicit rules. The second channel uses a square convolution kernel to perform two-dimensional convolution in the space-gesture type plane of the gesture frequency matrix, capturing the association patterns of gesture types and spatial positions, such as the tendency of high-frequency trigger zoom gestures or rotation gestures in specific regions.
[0073] The fusion mechanism of the dual-channel output feature vector includes feature alignment operation. Since there is a dimensional difference between the gaze feature vector and the gesture feature vector, the dual-channel output is projected to a unified dimensional hidden space through a fully connected layer. After feature splicing in the hidden space, two layers of fully connected network are used for non-linear transformation, and the behavior pattern feature vector is output. This vector contains compressed user interaction habit information, such as visual preference for specific spatial dimensions, high-frequency interaction gesture combination patterns, and other abstract expressions.
[0074] The generation of the environmental cognitive preference parameter is achieved through immersion degree correlation analysis. After the analysis module receives the behavior pattern feature vector, it is first decomposed into a spatial attention component and an interaction behavior component. The spatial attention component is input into the clustering comparison unit to match the attention distribution template in the historical typical scene, and the output is the user's cognitive tendency score for open space, closed space or transition space. The interaction behavior component is input into the time sequence pattern analyzer to detect the time regularity of periodic interaction behavior, and output the user's preference index for continuous interaction and discrete interaction. Finally, the spatial tendency score and the interaction preference index are fused to generate the environmental cognitive preference parameter set containing multiple parameters.
[0075] When the generation process of the virtual-real interaction event sequence starts, first calculate the remaining difference between the current experience duration and the preset immersion threshold. The setting of the immersion threshold is based on the user's historical experience data: extract the duration samples of the previous effective immersion experience, and obtain the ideal experience duration reference value through Gaussian distribution fitting. The remaining difference calculation module continuously monitors the current experience duration progress, and when the remaining time is below a certain proportion of the reference value, the dynamic adjustment of the timestamp alignment strategy is triggered. The adjustment mechanism uses a time axis compression algorithm: the timestamp sequence of the original interaction behavior data stream is nonlinearly remapped, so that the event time interval is proportionally reduced. The compression ratio is dynamically calculated according to the decreasing amplitude of the remaining difference, and the maximum compression ratio is used when the remaining time is close to the critical threshold.
[0076] The event stream reorganization algorithm performs multi-source data fusion operation. The environmental cognitive preference parameter is first encoded into a weight matrix, and the row vector of the matrix corresponds to different interaction modalities (such as eye movement, gesture, voice), and the column vector represents the weight distribution value of each sub-feature in the modality. The multi-modal input in the interaction behavior data stream is parsed into a feature tensor: eye movement data is converted into a sequence of gaze focus coordinates, gesture data is parsed into a tuple of type encoding and spatial coordinates, and voice instructions are converted into an intent vector through a semantic analysis engine. In the reorganization layer, each modality data feature and the weight matrix are subjected to Hadamard product operation to realize feature enhancement or weakening with preference weighting. For example, when the environmental cognitive preference parameter shows that the user pays high attention to visual details, the corresponding coefficient of the eye movement feature in the weight matrix is significantly improved.
[0077] The weighted multi-modal features are integrated through the space-time alignment module. The module establishes a space-time coordinate system mapping relationship: the eye movement focus coordinates are associated with the anchor points of the world coordinate system, the gesture spatial coordinates are converted to the world coordinate system through a hand tracking model, and the voice intent vector is bound to the spatial direction according to the current user orientation. The sorted feature vector is sorted by timestamp to generate discrete event units carrying space-time labels. Each event unit contains three elements: a millisecond-level generation timestamp (based on atomic clock synchronization), a three-dimensional position coordinate in the world coordinate system, an event type code and a feature vector.
[0078] The dynamic optimization of the event sequence includes an abnormal event filtering mechanism. The spatiotemporal continuity of the event units is monitored in real time during the event stream reorganization process: by calculating the ratio of the spatial distance change rate of adjacent events to the time interval, jump events that may be caused by sensor noise are detected. When abnormal patterns such as spatial transient or time reverse order are detected, an event interpolation correction program is triggered. The program uses the historical event sequence to build a spatiotemporal trajectory prediction model, and uses the predicted value to replace the abnormal event coordinates to maintain the physical rationality of the event sequence.
[0079] The output structure of the event sequence adopts a hierarchical index design. The basic layer stores the original event units in strict chronological order; the intermediate layer establishes a spatial region index, and the events are divided into preset spatial grid units according to the coordinates; the top layer constructs an event type hash table to support fast retrieval by interaction type. When the subsequent module requests event data within a specific spatiotemporal range, the relevant event subset can be quickly located through multi-layer indexing, reducing the computational delay of real-time queries.
[0080] The entire implementation process forms a data-driven closed-loop optimization. The newly generated virtual-real interaction event sequence is recorded to the user behavior knowledge base after execution, which is used to update the gaze residence matrix and gesture frequency matrix. When a significant change in user behavior pattern is detected, the environmental cognitive preference parameters are recalculated, which in turn affects the generation logic of subsequent event sequences. This mechanism enables the system to adapt to the gradual evolution of user interaction habits and maintain the effectiveness of the immersive experience over a long period of use.
[0081] The adjustment process of the timestamp alignment strategy includes a smooth transition mechanism. When the compression ratio needs to be changed, a gradual switching algorithm is used: the compression coefficient is adjusted gradually over multiple consecutive processing periods to avoid experience discontinuity caused by sudden changes in event time intervals. The compression parameter change information is added to the metadata area of the event sequence for reference by subsequent compensation modules. For example, when it is detected that the event sequence has been time-compressed, the time delay compensation module adjusts its reference clock synchronization strategy accordingly.
[0082] The conflict resolution mechanism of multi-modal input ensures the internal consistency of the event sequence. When eye movement, gesture, speech and other modalities produce semantic conflict inputs in the same period (such as eye movement fixation area A while gesture operates area B), the reorganization algorithm starts the modal confidence evaluation: according to the recognition accuracy of each modality in the scene in the historical data, the decision weight is allocated, and the feature data of the high-confidence modality is preferentially adopted, while the potential conflict events are marked for special processing by subsequent modules. The feature vector in the event unit contains a conflict marker, which triggers an additional spatial position verification process during the subsequent virtual-real synchronization calibration process.
[0083] The entire implementation architecture design takes into account both computational efficiency and accuracy requirements. The behavior feature extraction model uses a lightweight convolution kernel configuration, enabling fixed-point number operation optimization when deployed on mobile terminals. The event stream reorganization algorithm supports pipeline parallel processing, allowing features extraction and weighting operations of different modalities such as eye movement, gestures, and speech to be executed in independent threads, and finally integrating the space-time alignment through a synchronization barrier mechanism. This design effectively utilizes multi-core processor resources, meeting the stringent real-time response requirements of augmented reality applications.
[0084] Embodiment 3: Refer to Figure 4 , according to the rendering priority sequence to generate the spatial position compensation amount, based on the virtual-real interaction event sequence to calculate the time delay compensation amount, input the spatial position compensation amount and the time delay compensation amount into the error compensation engine to perform virtual-real synchronization calibration. When the virtual-real synchronization error dynamic compensation operation is started, the generation of the spatial position compensation amount is based on the analysis result of the rendering priority sequence. Each entry in the sequence contains spatial offset information of virtual-real superimposed content, and the offset is derived from the difference value between the three-dimensional topological structure in the environmental semantic modeling stage and the actual sensor detection position. The process of the pose solution algorithm processing the difference value includes the separation calculation of the rotation component and the translation component: the direction offset is converted into a rotation compensation vector through the quaternion interpolation method, and the translation compensation vector is generated through the Euclidean distance decomposition, and finally combined into a six-degree-of-freedom compensation parameter. The parameter is formatted as a homogeneous transformation matrix, and the matrix elements include three axis rotation angles and three dimensional translation amounts.
[0085] The calculation of the time delay compensation amount depends on the timestamp alignment state of the virtual-real interaction event sequence. The millisecond-level timestamp carried by each entry in the event sequence is compared with the high-precision clock of the rendering system, producing a time difference sequence. The compensation coefficient matrix is generated through statistical analysis of the time difference sequence: the standard deviation and the mean of the time difference within a fixed time window are calculated, the standard deviation reflects the delay fluctuation degree, and the mean reflects the reference delay amount. The compensation coefficient is composed of the linear combination of the reference delay amount and the fluctuation degree, which is mathematically expressed as:
[0086] ;
[0087] Wherein: represents the compensation coefficient of the type event in the time period, is the mean of the time difference, is the standard deviation of the time difference, and are preset weight factors. The coefficient matrix is organized by event type and time block, supporting fine delay compensation for specific interaction scenarios.
[0088] The error compensation engine includes two execution units, a spatial transformation module and a frame buffer controller. The spatial transformation module receives the six-degree-of-freedom compensation parameters and converts them into a projection matrix correction amount. The specific implementation adopts an affine transformation chain operation: first, a rotation compensation is applied to the model matrix of the augmented reality content, and the object orientation is updated through quaternion multiplication; second, a translation compensation is applied to the view matrix, and the observation coordinate system origin position is modified; finally, a perspective distortion caused by position offset is corrected by fine-tuning the view frustum of the projection matrix. Each step of transformation operation adopts an incremental update mode to avoid the computational burden of global matrix reconstruction.
[0089] The frame buffer controller adjusts the rendering timing according to the compensation coefficient matrix. The controller maintains a triple buffer architecture: the front-end buffer receives a new rendering frame, the middle-end buffer performs time compensation processing, and the back-end buffer outputs to the display device. When the value exceeds the threshold value, a dynamic frame scheduling mechanism is started: transition frames are generated between the front-end buffer and the middle-end buffer through an interpolation algorithm, and the number of transition frames is in a positive proportional relationship with the value; at the same time, the vertical synchronization signal triggering time of the back-end buffer is adjusted to make the timestamp of the final output frame synchronized with the physical environment change.
[0090] The generation of the virtual-real interaction event sequence starts from the recognition of multi-modal input features. The interaction behavior data stream is deconstructed into three parallel processing channels: the gesture recognition channel extracts the spatial coordinates of the hand joint points through a skeleton tracking algorithm to form a gesture trajectory vector; the eye tracking channel calculates the vector angle between the pupil center and the corneal reflection point to output the projection position of the gaze focus in the screen coordinate system; the speech processing channel extracts user instructions through voiceprint separation technology and converts them into semantic action codes through an intent recognition model. The output results of each channel are attached with confidence scores and timestamp labels.
[0091] The feature fusion gateway implements a three-layer weighted superposition architecture. The first layer performs spatial correlation operations: the spatial coordinates of the gesture are converted from the screen coordinate system to the world coordinate system, and the spatial proximity calculation is performed with the eye focus coordinates. The proximity calculation adopts an adaptive radius judgment method: when the Euclidean distance between the gesture operation point and the eye focus is less than the dynamic threshold, it is determined as an effective cooperative interaction, and a cooperative interaction flag is generated. The second layer performs semantic integration: the semantic vector of the voice instruction is spliced with the cooperative interaction flag to form a mixed feature vector, and the dimension is compressed through a fully connected network. The third layer implements environment cognition fusion: the compressed feature vector and the environment cognition preference parameter are subjected to matrix multiplication operation. The environment cognition preference parameter is organized as a weight matrix, and the row vectors of the matrix correspond to different interaction situations, and the column vectors represent the weight distribution scheme of the feature elements under each situation.
[0092] The final generated virtual-real interaction event sequence adopts a space-time dual labeling mechanism. The spatial label includes a three-dimensional position coordinate in the world coordinate system and a coverage radius, which is calculated by the weighted centroid of the gesture operation point and the eye gaze focus; the time label includes the precise timestamp (synchronized to the atomic clock) of the event generation time and the effective duration interval. Each event entry is attached with the confidence value output by the feature fusion gateway, which is calculated by the intermediate results of each layer processing, and is used for reliability evaluation of subsequent compensation operations.
[0093] The execution process of spatial position compensation includes a dynamic verification link. After the compensation engine applies the six-degree-of-freedom parameters, the real-time detection of the virtual-real alignment degree after compensation is performed by the environmental semantic modeling module: in the three-dimensional topological structure, key semantic anchor points are selected, and the position deviation between the virtual content projection position and the actual anchor point is calculated. If the deviation value is higher than the baseline value before compensation, the rollback mechanism of the compensation parameters is triggered, and the pose solving process is restarted. This feedback mechanism prevents the phenomenon of excessive compensation caused by sensor noise.
[0094] The special scene processing of time delay compensation involves event sequence prediction. When it is detected that the value of the time difference between the current time window and the previous time window is in a monotonically increasing trend for consecutive multiple time windows, the pre-compensation mechanism is started: based on the historical time difference sequence, an autoregressive prediction model is established, and the compensation coefficient of the future time slice is generated in advance. The pre-compensation coefficient is injected into the forward scheduling module of the frame buffer controller, so that the rendering pipeline can prepare the frame processing scheme in advance for the high delay period. The prediction model adopts a sliding window training mechanism, only retains the sample data of the recent time period, and ensures the rapid response to the change of user interaction mode.
[0095] The solution to the conflict of multi-modal input is implemented in the third layer of the feature fusion gateway. When the input features of gestures, eye movements, and voices are contradictory in the semantic level (such as the eye gaze position A and the gesture pointing position B), the environmental cognitive preference parameters start the conflict resolution mode: according to the reliability records of each modality in similar scenarios in the user historical behavior data, the distribution proportion of the weight matrix is dynamically adjusted. For example, when the historical data shows that the user is more inclined to gesture dominant operation, the coefficient of the gesture feature in the weight matrix is automatically increased, and the coefficient of the eye movement feature is correspondingly decreased. The conflict resolution log is attached to the event sequence entry for use by the subsequent analysis module.
[0096] During the operation of the entire compensation system, the spatial transformation module and the frame buffer controller maintain state synchronization. Through the distributed transaction mechanism, the atomicity of the two types of compensation operations is ensured: when the spatial compensation parameters and the time compensation coefficients need to be updated synchronously, version control markers are established; only when the version identification of the two types of compensation data matches, the output operation is performed, avoiding the misplacement of space-time compensation due to transmission delay. The version identification includes the combination of the timestamp hash value and the space grid code, ensuring the accuracy of state synchronization.
[0097] The storage of virtual-real interaction event sequences adopts a space-time partition database. Event entries are stored in partitions indexed by a spatial grid of the world coordinate system, and each partition is sorted by millisecond-level timestamps. This structure supports efficient spatial range queries and time range retrieval. When the synchronization compensation module requests recent events in a specific area, the results can be quickly returned through a parallel query mechanism. The database implements an incremental backup strategy, retaining only event data from active interaction areas, and converting non-active area data to cold storage to save resources.
[0098] Embodiment 4: Refer to Figure 5 , a real-time feedback channel is established between the environmental cognitive preference parameters and the multi-modal input features. When an abnormal jump in the multi-modal input features is detected, a dynamic correction mechanism for the environmental cognitive preference parameters is triggered, and the corrected environmental cognitive preference parameters are re-weighted and superimposed with the multi-modal input features. The real-time feedback channel of the feature fusion gateway adopts a ring buffer structure. The buffer contains continuous storage units, each of which records a complete state snapshot of the latest feature fusion. The state snapshot includes four types of core data: environmental cognitive preference parameter version number, multi-modal input feature vector, weighted superposition result value, and timestamp marker. The buffer adopts a first-in, first-out management strategy, with new data covering the oldest records, maintaining historical state traceability within a fixed time window.
[0099] When the multi-modal input features experience an abnormal jump, the jump detection algorithm is activated based on a dynamic threshold. The detection criteria include two criteria: first, calculate the rate of change of the Euclidean distance between the current frame feature vector and the historical mean, which is taken from the average of the last several records in the ring buffer; second, analyze the standard deviation mutation of each dimension of the feature vector. When the following logical relationship is met, the abnormal flag is triggered: the distance change rate exceeds a certain multiple of the historical fluctuation range, and the standard deviation increment of at least two dimensions breaks through the independent threshold. Key data segments during a detection process are shown in Table 1.
[0100] Table 1: Key data segments during a detection process.
[0101]
[0102] The trigger process of the dynamic correction mechanism includes four stages: the first stage immediately suspends the current weighted superposition pipeline and freezes the feature input port; the second stage starts historical scene matching retrieval, using the spatial topological features of the current environmental semantic model as the retrieval key to find the most similar scene record in the historical behavior pattern feature library; the third stage calls the environmental cognitive preference parameters corresponding to the historical scene to replace the current parameters, and if the matching fails, the default parameter template is enabled; the fourth stage reinitializes the weighted superposition matrix.
[0103] The adaptive weight assignment algorithm is executed at the feature re-fusion stage. The confidence scores of each modality input feature are derived from the real-time outputs of the front-end recognition modules: the gesture recognition confidence is based on the stability of the tracked skeletal joints, the eye movement confidence depends on the accuracy of the identified pupil contour, and the speech confidence is sourced from the signal-to-noise ratio of the speech endpoint detection. The weight calculation employs a non-linear mapping function to give exponential weight gain to high-confidence features. The specific assignment process involves three normalization operations: first, the original confidence of each sub-feature within a modality is converted into a relative weight; second, the weights are balanced across modalities to prevent a single modality from dominating; and finally, the weight sum constraint is applied to ensure the invertibility of the superposition matrix.
[0104] Take the scenario of a user's gesture burst drifting as an example: when the hand tracking sensor is disturbed by strong light, causing the coordinates to jump, the gesture feature confidence drops sharply to a low value range. At this time, the feature fusion gateway detects that the X / Y coordinate change rate deviates significantly from the historical pattern, and immediately freezes the current fusion process. The system searches the historical database and finds that the user relies more on eye movement interaction in similar lighting conditions in the past, and then calls the corresponding historical parameter template. In the new round of weighted superposition, the gesture feature weight is compressed to a small proportion of the baseline value, and the eye movement and speech feature weights are correspondingly increased, so that the final output of the virtual-real interaction event sequence is immune to abnormal gesture data pollution.
[0105] The depth optimization of the ring buffer adopts a scene-aware adjustment strategy. In indoor scenes with stable physical environments, the buffer depth is set to a larger value to capture long-period behavior patterns; when the user enters a dynamic outdoor environment, the buffer depth is automatically reduced to increase the sensitivity to transient changes. The depth adjustment is based on the complexity index of the three-dimensional topological structure: the semantic anchor point density change rate and the spatial distribution entropy value are calculated, and when the entropy value increases beyond the threshold, the buffer depth is reduced.
[0106] The historical scene matching retrieval adopts a two-level index structure. The first-level index is established on the topological fingerprint of the environmental semantic model, which encodes the three-dimensional grid structure into a fixed-length feature code through a hash function; the second-level index is based on the time context label, which marks the user's activity type in the current time period. The retrieval process is carried out in the intersection space of the two-level indexes, and the historical records with a topological fingerprint matching degree greater than the critical value and consistent time labels are preferentially returned. The matching degree calculation uses the Hamming distance to measure the topological fingerprint difference, and at the same time requires that the behavior pattern feature vector of the historical record and the current vector are in the same clustering cluster.
[0107] The reconstruction process of the weighted superposition matrix contains a feature dimension adaptation mechanism. When the historical environment cognition bias parameter used as a substitute is inconsistent with the current multi-modal input feature dimension, a dimension projection conversion is started: through a pre-trained auto-encoder network, the historical parameter is embedded into the current feature space, maintaining the semantic invariance of the parameter. The projection network adopts a fully connected architecture, the input layer receives the historical parameter vector, the hidden layer performs nonlinear transformation, and the output layer dimension is aligned with the current feature space. This process ensures that parameters from different sources can be effectively fused in a unified space.
[0108] The abnormal handling log recording mechanism provides traceability analysis capability. Each dynamic correction event generates a detailed log entry, including: abnormal trigger timestamp, jump feature identifier, substitute parameter source, weight allocation scheme, and final output event sequence summary. The log entry is stored by binary encoding compression and attached to the corresponding virtual-real interaction event sequence metadata area. When multiple correction events occur in the same spatial region, an offline optimization task of the environment cognition bias parameter is triggered, and the user behavior model is updated through batch training.
[0109] The entire implementation process establishes a fault isolation boundary. The abnormal handling unit of the feature fusion gateway runs in an independent sandbox environment and exchanges data with the main interaction pipeline through a secure channel. When the correction mechanism fails continuously more than a threshold number of times, a degradation processing mode is started: bypass the weighted fusion link and directly output the original feature sequence of the multi-modal input, while sending an experience degradation notification to the user terminal. The reset operation of the sandbox environment is automatically executed after each successful correction, clearing the temporary state to ensure the purity of subsequent processing.
[0110] The final output virtual-real interaction event sequence is attached with a correction flag bit. The flag bit contains two status bits: the high bit indicates whether dynamic correction is experienced, and the low bit records the correction type (historical parameter substitution / default template enabled). The subsequent synchronous compensation module adjusts the processing strategy according to the state of the flag bit: for the corrected event sequence, the sensitivity of spatial position compensation is reduced; for events using the default template, the safety margin of time delay compensation is increased. This marking delivery mechanism forms a cross-module collaborative fault tolerance chain.
[0111] Embodiment 5: Cumulative value statistics of virtual-real synchronization error adopts a sliding window counter mechanism. The counter records the triggering events of spatial position compensation and time delay compensation within a fixed time period, each event containing compensation type identification, triggering timestamp and spatial region code. When the compensation event density of a specific spatial region in the window exceeds the dynamic threshold, the compensation intensity adjustment factor generation process is triggered. The adjustment factor calculation introduces a nonlinear response curve: when the compensation event density is in the low interval, a gentle growth mode is adopted, and when the density enters the high interval, an exponential growth mode is switched, so that the factor value matches the actual demand. The generated adjustment factor is attached with a region label and stored together with the compensation data of the corresponding spatial position.
[0112] The process of converting spatial position compensation to three-dimensional vector field implements a hierarchical mapping strategy. The translation component in the original compensation data is directly used as the basis value of the vector field, and the rotation component is converted into a direction vector. In the three-dimensional space grid, each voxel unit stores the average compensation vector at that position, and spatial continuity is achieved through vector interpolation of adjacent units. Vector field update adopts incremental optimization: the newly generated compensation is first compared with the historical average, and when the deviation exceeds the tolerance, local reconstruction of the vector field is started; otherwise, it is updated by weighted average smoothing. The field intensity distribution map is visualized by color coding for system state monitoring.
[0113] The time delay compensation quantity is processed to construct a time gradient field. The time delay value is mapped to the gradient change rate along the time axis, and a delay distribution model is established in the four-dimensional space-time coordinate system. The time gradient field is managed by spatial region, and each subregion maintains an independent time delay change curve. The curve sampling points include timestamp, delay reference value and fluctuation amplitude, and a continuous gradient function is generated by curve fitting. The gradient field update mechanism includes abnormal point filtering: when the deviation between the new delay value and the predicted value of the fitted curve exceeds the historical fluctuation range, it is marked as a temporary abnormal point and not stored.
[0114] The compensation intensity adjustment factor is a scalar field participating in tensor product operation. The scalar field and the three-dimensional vector field perform outer product operation to generate a second-order tensor field, which describes the intensity modulation relationship of spatial position compensation. At the same time, the scalar field and the time gradient field perform tensor contraction operation to generate intensity correction coefficients in the time dimension. The tensor product operation result constitutes the core data structure of the virtual-real synchronization compensation atlas, and each space-time coordinate point in the atlas stores six-tuple parameters: spatial compensation vector, time delay gradient, adjustment factor intensity value and their synthesis effect coefficient.
[0115] The execution frequency of the error compensation engine is controlled based on real-time analysis of the compensation atlas. In the three-dimensional spatial grid, the product of the modulus of the compensation vector and the adjustment factor intensity is calculated for each cell. When the product value exceeds the critical threshold, the cell is marked as a high-frequency compensation area. For high-frequency compensation areas, the engine execution frequency is increased to an integer multiple of the base value, and the multiple is dynamically calculated based on the composite effect coefficient. At the same time, the slope change of the time gradient field is detected: when the delay gradient in a specific spatial region continues to increase, a stepwise incremental strategy is implemented for the time compensation frequency of that region.
[0116] An intensity adjustment mechanism is introduced in the execution process of spatial position compensation. Before applying the six-degree-of-freedom compensation parameters, the translation vector modulus and rotation angle values are scaled by the composite effect coefficient. The scaling ratio is positively correlated with the adjustment factor intensity, but an upper threshold is set to prevent overcompensation. When the same spatial region triggers high-intensity compensation multiple times in a row, a compensation effect evaluation loop is started: compare the rate of change of the virtual-real alignment deviation before and after compensation, and if the deviation decreases by less than the expected proportion, increase the subsequent compensation intensity by a step value.
[0117] The frequency control of time delay compensation is linked with the frame buffer depth. When the compensation atlas shows that a certain region's time gradient field is running at a high level continuously, the frame buffer controller dynamically increases the number of intermediate buffer layers. The new buffer layers are used to store pre-rendered frame sequences, and transition frames are generated to fill the delay gap through interpolation algorithms. The buffer layer depth is proportional to the time gradient value, but it is soft-limited by the hardware memory capacity. At the same time, the triggering period of the vertical synchronization signal is fine-tuned according to the gradient change rate, and when the delay fluctuates dramatically, the period is shortened to improve the response speed.
[0118] The update mechanism of the compensation atlas builds a closed-loop feedback. After each execution of the error compensation engine, new spatial position deviation and time delay data are collected as feedback input. The data is compared with the atlas prediction value to calculate the residual error, and when the residual error exceeds the threshold, local correction of the atlas is triggered: the spatial vector field adjusts the vector direction through backpropagation, the time gradient field corrects the curve fitting parameters, and the adjustment factor updates the intensity value in proportion to the residual error. The version number of the corrected atlas is incremented to ensure that subsequent operations are based on the latest state.
[0119] Resource scheduling for high-frequency compensation areas implements priority management. According to the comprehensive score of the three-dimensional vector field modulus and the adjustment factor intensity, the spatial grid cells are divided into different compensation levels. The engine execution thread prioritizes processing the highest level areas, and areas with the same level are arranged in descending order of time gradient value. When system resources are tight, a compensation result reuse strategy is used for low-level areas: the compensation parameters of similar historical states are reused after linear transformation, reducing real-time calculation load. The reuse decision must satisfy the condition that the spatial position deviation is within an acceptable range.
[0120] The execution frequency control module contains a safety protection mechanism. A single-area maximum frequency threshold is set to prevent system overload, and when the frequency of a certain area approaches the threshold, the neighboring area is triggered to compensate collaboratively: the compensation vector of the adjacent grid is calculated and interpolated, and part of the compensation task is allocated to the surrounding area. At the same time, the engine core temperature and memory occupancy are monitored, and when the hardware indicators exceed the limit, the global frequency is degraded: the execution frequency of all areas is reduced by a fixed proportion until the system returns to normal. The degradation state information is added to the compensation graph metadata area for subsequent analysis.
[0121] The entire implementation process forms an adaptive dynamic compensation system. Newly generated synchronization error data is continuously input into the statistical module to update the cumulative value; the updated adjustment factor drives the reconstruction of the compensation graph; the optimized graph controls the execution strategy; and the execution result is fed back to the error collection end to form a closed loop. This system adjusts the compensation strength and frequency collaboratively to achieve fine control of virtual and real synchronization errors under system resource constraints. The parameter self-adaptation capability in long-term operation enables the system to respond to changes in user behavior patterns and fluctuations in environmental conditions.
[0122] The compensation effect evaluation data is stored in the historical knowledge base. Key indicators such as residual error values, correction parameters, and resource occupancy generated by each closed-loop feedback are archived by scene classification. When a similar environmental semantic model is activated again, the historical optimal compensation configuration is preferentially called to initialize the graph parameters. The knowledge base establishes an index to associate environmental topology fingerprints with compensation parameter sets, supporting fast matching of historical solutions through semantic similarity retrieval. This mechanism accelerates the convergence speed of the system in new scenarios and improves user experience consistency.
[0123] The persistent storage of the compensation graph adopts a differential incremental strategy. Full storage is only performed when the graph structure changes substantially, and regular updates only record parameter differential increments. The storage format includes binary data blocks and metadata description files, supporting fast loading and version rollback. When deployed on mobile terminals, lightweight compression algorithms are enabled to reduce storage space occupation while ensuring real-time access performance. The graph version management service ensures data consistency among multiple terminals in a distributed system.
[0124] It should be noted that, in this text, relational terms such as first and second are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "contain" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device that includes a series of elements not only includes those elements, but also includes other elements not explicitly listed or inherent to such a process, method, article or device.
[0125] While embodiments of the application have been shown and described, it is to be understood that the embodiments described are merely exemplary of the principles and application of the present application. Numerous modifications and adaptions can be effected without departing from the spirit and scope of the present application, which is not limited to the exact construction and arrangement described. It is intended, therefore, to cover all modifications and adaptions that fall within the scope of the claims and their equivalents.
Claims
1. An augmented reality method for immersive experiences, characterized by, The method comprises: Real-time collection of dynamic space data of a physical environment, and synchronous acquisition of an interactive behavior data stream of a user terminal; Environment semantic modeling of the dynamic space data, generation of a multi-level feature description of the physical environment, and determination of a rendering priority sequence of virtual-real superimposed content according to the multi-level feature description; Extraction of behavior pattern features in user historical interaction data, immersion degree correlation analysis of the behavior pattern features, obtaining of an environment cognitive preference parameter of a current user, and generation of a virtual-real interactive event sequence combining the environment cognitive preference parameter and the interactive behavior data stream; Dynamic compensation of virtual-real synchronization errors in an immersive experience process according to the rendering priority sequence and the virtual-real interactive event sequence; The determination of the rendering priority sequence of virtual-real superimposed content according to the multi-level feature description comprises: Analyzing the environmental light intensity change curve and the space motion trajectory in the multi-level feature description; Calculating the visual penetration coefficient of virtual-real superimposed content based on the environmental light intensity change curve; Generating a rendering priority sequence according to the space motion trajectory and the visual penetration coefficient; The dynamic compensation of virtual-real synchronization errors in the immersive experience process comprises: Generating a space position compensation amount according to the rendering priority sequence; Calculating a time delay compensation amount based on the virtual-real interactive event sequence; Inputting the space position compensation amount and the time delay compensation amount into an error compensation engine to perform virtual-real synchronization calibration.
2. The augmented reality method for immersive experience of claim 1, wherein, The environment semantic modeling comprises: Constructing a three-dimensional topological structure of the physical environment according to the dynamic space data; Marking semantic anchor points of interactive objects in the three-dimensional topological structure through a space-time correlation algorithm.
3. The augmented reality method for immersive experience of claim 1, wherein, The extraction of behavior pattern features in user historical interaction data comprises: Calling the gaze residence duration distribution and gesture trigger frequency matrix in the user behavior knowledge base; Outputting behavior pattern features from the gaze residence duration distribution and gesture trigger frequency matrix through a behavior feature extraction model.
4. The augmented reality method for immersive experience of claim 1, wherein, The generation of the virtual-real interactive event sequence comprises: Calculating the remaining difference between the current experience duration and a preset immersion threshold value; Adjusting the timestamp alignment strategy of virtual-real interactive events according to the remaining difference; Fusing the environment cognitive preference parameter and the interactive behavior data stream into a virtual-real interactive event sequence through an event stream reorganization algorithm.
5. The augmented reality method for immersive experience of claim 1, wherein, The generation process of the virtual-real interactive event sequence comprises: Identifying multi-modal input features in the interactive behavior data stream; Weighted superimposition of the multi-modal input features and the environment cognitive preference parameter through a feature fusion gateway; Generation of a virtual-real interactive event sequence carrying space-time labels according to the superimposition result.
6. The augmented reality method for immersive experience of claim 5, wherein, The workflow of the feature fusion gateway comprises: Establishing a real-time feedback channel for the environment cognitive preference parameter and the multi-modal input features; Triggering a dynamic modification mechanism for the environment cognitive preference parameter when an abnormal jump of the multi-modal input features is detected; Re-weighted superimposition of the modified environment cognitive preference parameter and the multi-modal input features.
7. The augmented reality method for immersive experiences of claim 1, wherein, Further comprising: Generation of a compensation intensity adjustment factor according to the cumulative value of the virtual-real synchronization error; Construction of a virtual-real synchronization compensation map through the space position compensation amount, the time delay compensation amount, and the compensation intensity adjustment factor; Control the execution frequency of the error compensation engine based on the virtual-real synchronous compensation atlas.
8. An augmented reality system for immersive experience, configured to perform an augmented reality method for immersive experience according to any one of claims 1 to 7, characterized in that, The system comprises: An environment perception module for collecting dynamic space data of a physical environment from a plurality of environment sensors and synchronously acquiring an interactive behavior data stream of a user terminal; A semantic modeling module for performing environment semantic modeling on the dynamic space data in a virtual-real fusion mode, generating a multi-level feature description of the physical environment, and determining a rendering priority sequence of virtual-real superimposed content according to the multi-level feature description; An interaction analysis module for extracting behavior pattern features from user historical interaction data, performing immersion degree correlation analysis on the behavior pattern features to obtain environment cognitive preference parameters, and generating a virtual-real interaction event sequence; A synchronous compensation module for performing dynamic compensation operations on virtual-real synchronization errors in an immersive experience process according to the rendering priority sequence and the virtual-real interaction event sequence.
Citation Information
Patent Citations
Multi-vehicle parallel intelligent cooperative search and rescue system based on digital twinning and construction method thereof
CN118092437A
Cognitive rendering of inputs in virtual reality environments
US20200104580A1