Digital exhibition hall panoramic roaming construction method based on virtual-real fusion

By using multi-sensor dynamic calibration and modular management, the problems of virtual content offset and time-consuming updates in digital exhibition halls have been solved, enabling personalized immersive experiences and rapid updates to meet the application needs of various scenarios.

CN120876791AInactive Publication Date: 2025-10-31NANJING YIXUAN INTERNET TECHNOLOGY CO LTD

Patent Information

Application Number
CN202510983707.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-17
Publication Date
2025-10-31
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

In existing technologies, the offset between virtual content and physical space in digital exhibition halls accumulates over time in complex environments, and content updates are time-consuming and prone to failure, resulting in homogenized user experiences and difficulty in meeting personalized needs.

Method used

Through multi-sensor dynamic calibration and algorithm collaboration, combined with modular management and personalized path planning, it enables real-time capture and correction of micro-deformations in physical space, supports modular configuration and rapid updates of virtual content, and personalized path generation.

Benefits of technology

Maintaining precise alignment between virtual content and physical space in complex environments provides a personalized, immersive experience, supports rapid content updates and multi-terminal adaptation, and enhances the practical value and user appeal of digital exhibition halls.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120876791A_ABST
    Figure CN120876791A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of computer graphics and virtual reality, and discloses a virtual-real fusion-based panoramic roaming construction method for a digital exhibition hall, which effectively solves the problems of easy drift and offset in a virtual environment in a traditional system through a multi-sensor fusion algorithm and a closed-loop optimization mechanism. Under the complex environment of people flow change, illumination fluctuation and the like in an exhibition hall, accurate alignment of virtual content and physical space can still be kept, it is ensured that a user obtains coherent and stable visual experience in the roaming process, and a reliable space reference is provided for virtual-real interaction; the light and shadow effect and the material performance of the virtual content are naturally fused with the physical environment through the physical rendering and multi-sensory interaction technology, and meanwhile, a multi-dimensional immersion sense is constructed in combination with the spatial sound effect and interaction feedback; when the user interacts with the virtual exhibit, sensory experience close to reality can be obtained, and the sense of substitution and attraction of the digital exhibition hall are remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the fields of computer graphics and virtual reality technology, specifically a method for constructing panoramic digital exhibition hall tours based on virtual-real fusion. Background Technology

[0002] With the integration of digital technology and the exhibition field, digital exhibition halls, as a new form of display, provide users with an immersive experience through the combination of virtual content and physical space, which has become a research hotspot. In existing technologies, relevant solutions have realized the digital replication of exhibition halls and basic roaming functions, but there are still significant limitations in core issues such as the stability of virtual-real integration in complex environments, the efficiency of dynamic content management, and the personalization of user experience.

[0003] Comparative document 1 (publication number CN119494916A, "A Digital Exhibition Hall Server Based on 3D Scene Editing and Its Application Method") realizes the online replication of offline exhibition halls through a 3D scene editor and supports VR roaming function. However, it relies on a single 3D model data and static reference points for virtual-real matching and lacks a dynamic calibration mechanism based on multi-source data. In scenarios such as disturbances caused by the flow of people in the exhibition hall and temperature changes leading to slight deformation of the physical space, the static reference points cannot perceive and correct deviations in real time, resulting in the accumulation of offset between virtual content and physical space over time. At the same time, the content adopts an overall storage architecture, and updates require a full release, which is not only time-consuming but also prone to update failure due to network fluctuations, making it difficult to adapt to the high-frequency iteration needs of exhibition hall content.

[0004] Comparative document 2 (publication number CN118689314A, "Method, Device and Equipment for Exhibition Hall Construction Based on Metaverse") proposes a metaverse exhibition hall framework that supports dynamic loading of exhibit models, but it does not build a real-time coordinate system matching mechanism between physical space and virtual space. The anchoring accuracy is easily affected by factors such as ambient light intensity and object occlusion, resulting in spatial misalignment of virtual content when the viewpoint is switched. In addition, the guide path relies entirely on manual preset and cannot be dynamically adjusted based on user interest tags (such as science and technology, art preferences) or real-time traffic data, resulting in homogenized roaming experiences for different users and making it difficult to meet personalized needs. Summary of the Invention

[0005] The purpose of this invention is to provide a method for constructing a panoramic digital exhibition hall based on virtual-real fusion, so as to solve the problems mentioned in the background art.

[0006] This invention addresses the issues of static benchmarks in Reference Document 1 being unable to handle minor deformations in physical space and easy shifts within virtual environments, as well as the shortcomings of Reference Document 2, such as anchoring accuracy being greatly affected by the environment and lacking personalized paths. It achieves real-time capture and correction of minor deformations in physical space through multi-sensor dynamic calibration and algorithm collaboration (extended Kalman filter short-term estimation and graph optimization long-term correction). At the same time, it combines modular management, personalized path planning, and cross-terminal immersion enhancement to form a systematic solution.

[0007] To achieve the above objectives, the present invention provides the following technical solution: a method for constructing a panoramic digital exhibition hall based on virtual-real fusion, the specific steps of which are as follows:

[0008] S1. Panoramic Image and Spatial Data Acquisition: By deploying high-parameter panoramic cameras, depth cameras, and inertial navigation sensors, multi-view images, depth, and pose data are acquired. The data is then aligned with timestamps and calibrated using the IEEE 1588 protocol to ensure uniformity and high precision, providing reliable input for step S2.

[0009] S2. Matching of physical and virtual coordinate systems: Based on the multi-source data obtained in step S1, a multi-sensor fusion algorithm is adopted, including extended Kalman filtering for short-term estimation of real-time pose and graph optimization for long-term correction of global pose. The two work together to make the modeling error of the physical coordinate system ≤2cm. Combined with the virtual space reference point, the rigid body transformation is achieved through the iterative nearest point algorithm, and the matching error is ≤3cm.

[0010] S3. Setting and binding virtual content anchor points: Based on the coordinate system established in step S2, identify anchor points such as QR codes and SIFT feature points, bind virtual content using SLAM technology, and ensure that the virtual content drifts ≤1cm when the viewing angle changes through dynamic optimization, so as to provide stable spatial anchoring for step S4.

[0011] S4. Modular content configuration and management: After anchor point binding is completed in step S3, the virtual content is divided into modules such as 3D models. The CMS system enables modular configuration, incremental hot updates and version management, flexibly expands the content, and provides scalable support for step S5.

[0012] S5. Personalized Roaming Path Generation: Through the modular content management in step S4, combined with data such as exhibition hall layout and content distribution, an initial path is generated using the A* algorithm. After optimization by machine learning, roaming routes that match the interests of different users are recommended, providing a path basis for step S6.

[0013] S6. Multi-terminal interaction adaptation: Based on the personalized path generated in step S5, for terminals such as virtual reality headsets, adaptation is carried out through distortion correction, touch mapping and other mechanisms to ensure consistent and smooth multi-terminal interaction, laying the terminal foundation for step S7.

[0014] S7. Rendering Optimization and Immersive Enhancement: After completing multi-terminal adaptation in step S6, rendering performance is optimized by using region culling and LOD management, and immersiveness is enhanced by combining PBR lighting and spatial sound effects to improve the user's immersive interactive experience in different scenarios.

[0015] Preferably, the specific steps for panoramic image and spatial data acquisition in step S1 are as follows:

[0016] S11. Multi-type sensor collaborative data acquisition: Three types of high-parameter equipment are scientifically deployed in the exhibition hall: a 12K resolution, 30fps panoramic camera, with a 360-degree horizontal field of view and a 190-degree vertical field of view, can clearly capture details such as exhibit textures and wall decorations in every corner of the exhibition hall; a depth camera with a ranging range of 0.5-10 meters and an accuracy of ±1%, can accurately measure the spatial distance between exhibits and the environment, ensuring data reliability; and an inertial navigation sensor with a 100Hz sampling rate, with a heading accuracy of 0.1° and a pitch / roll accuracy of 0.05°, outputs spatial pose information in real time. The three types of equipment work together to acquire images, depth, and spatial data of the exhibition hall from all angles.

[0017] S12. Data calibration ensures consistent accuracy: During data acquisition, the IEEE 1588 protocol is used to align the timestamps of various data types, strictly controlling the synchronization error to within 1ms to avoid data misalignment due to time deviations. Simultaneously, feature point matching is used to correlate multi-sensor data, extracting common features to establish associations, keeping the calibration error within 5cm. This provides unified and high-quality basic data support for accurate matching of the physical and virtual spatial coordinate systems.

[0018] Preferably, the specific steps for matching the physical space and virtual space coordinate systems in step S2 are as follows:

[0019] S21. Precise Modeling of Physical Space Coordinate System: Based on the multi-source data obtained in step S1, a multi-sensor fusion algorithm is used to construct a physical space coordinate system. The extended Kalman filter state update formula is used to treat position, velocity, and attitude as state variables. Equations are constructed by combining sensor noise characteristics, and state estimation is optimized through prediction and update. The graph optimization algorithm uses keyframes as nodes and relative pose constraints between nodes as edges. Accumulated errors are eliminated by minimizing the global energy function, thereby improving the modeling accuracy. The keyframe selection interval for the graph optimization algorithm is 500ms or the moving distance is ≥1m, and the number of iterations for minimizing the global energy function is ≥10.

[0020] Extended Kalman Filter State Update Formula Expression:

[0021]

[0022] In the formula: Let K be the updated state vector at time k (containing position, velocity, attitude, etc.); k For Kalman gain; z k P is the sensor measurement value; k|k-1 To predict the state covariance; R k To measure the noise covariance; H k The observation matrix at time k is used to represent the state vector. Mapping to the observation space, such as converting position, attitude and other states into distance observations from the depth camera; h(·) is the observation model, specifically a nonlinear function based on the physical characteristics of the sensor;

[0023] For example, the distance observation model for a depth camera is as follows:

[0024]

[0025] Where (x) p ,y p ,z p (x) represents the coordinates of a point in physical space. c ,y c ,z c () represents the camera coordinates;

[0026] This formula is the core of multi-sensor fusion and is used to dynamically optimize the state estimation of the physical space coordinate system. By fusing information from multiple sources such as panoramic images and depth data, it reduces the error caused by sensor noise and achieves accurate modeling of the physical space coordinate system.

[0027] This formula originates from an extended form of the Kalman filter, which was systematically explained by Grewal et al. in "Global Positioning Systems, Inertial Navigation, and Integration". It is widely used in the field of multi-sensor fusion and is suitable for state estimation scenarios of nonlinear systems.

[0028] Graph optimization algorithm expression:

[0029]

[0030] In the formula: x is the set of poses for all keyframes; e i (x i The first term () represents the single-node prior error, originating from sensor noise during keyframe acquisition (such as the heading angle error of an inertial navigation sensor). The calculation formula is as follows: ( (Initial sensor measurement); e ij (x i ,x j The term ) represents the inter-node constraint error, indicating the relative pose deviation between keyframes i and j. The calculation formula is e.ij (x i ,x j )=x j -T ij x i (T ij (where i is the theoretical transformation matrix from j);

[0031] This formula is used to optimize the physical space coordinate system model by minimizing the global energy function, eliminating sensor cumulative errors, improving the consistency of coordinate system modeling, and providing a stable coordinate reference for subsequent anchor point binding.

[0032] The formula originates from probabilistic graphical model theory and was introduced into the SLAM field by Kaess et al. in "iSAM: Incremental Smoothing and Mapping". It is suitable for high-precision pose optimization in large-scale spaces.

[0033] S22. Precise matching of virtual and real space coordinates: Combining the preset reference points in the virtual space (whose coordinate information is accurate and evenly distributed), the spatial transformation relationship between physical feature points and reference points is calculated. The iterative nearest point algorithm is used to repeatedly solve the rotation matrix and translation vector of rigid body transformation, so that the matching error is ≤3cm. This lays a solid coordinate foundation for the spatial anchoring of virtual content and smoothly connects to the subsequent anchor point setting work.

[0034] Formula for solving the transformation matrix using the Iterative Closest Point (ICP) algorithm:

[0035]

[0036] In the formula: R is the rotation matrix; t is the translation vector; p i For physical space feature points; q i is the virtual reference point; n is the number of feature points participating in the matching, with a selection standard of 250 (to ensure matching stability), and the physical space feature points and the virtual reference point must have a one-to-one correspondence (e.g., associated through SIFT feature matching);

[0037] This formula is used to solve for the rigid body transformation parameters (rotation matrix R and translation vector t) between physical and virtual spaces. By minimizing the sum of squared Euclidean distances between physical feature points and virtual reference points, it achieves high-precision coordinate system matching. The formula guarantees the minimization of the spatial position error of corresponding points in both physical and virtual spaces.

[0038] This formula was proposed by Besl and McKay in 1992 in "A Method for Registration of 3-DShapes". It is a classic algorithm for point cloud registration and coordinate system calibration and is widely used in fields such as virtual-real fusion and robot localization.

[0039] Preferably, the specific steps for setting and binding virtual content anchor points in step S3 are as follows:

[0040] S31. Diverse Anchor Point Recognition and Positioning: Based on the coordinate system established in step S2, multiple anchor points are set in physical space. The QR code anchor point adopts an anti-distortion algorithm, which can achieve fast recognition at 10 frames / second even if it is tilted or blurred. The SIFT feature point anchor point is accurately positioned due to its rotation and scaling invariance, and is suitable for areas without obvious markings. The infrared marker anchor point uses an infrared camera to capture the position, which is not affected by ambient light and has a positioning accuracy of 2cm, providing a variety of options for binding.

[0041] S32, SLAM technology achieves stable binding: By using SLAM technology to track the user's position and posture in real time, virtual exhibits and other content are bound to physical locations. Through closed-loop detection and beam adjustment optimization, continuous dynamic correction is made to ensure that the virtual content drifts ≤1cm when the viewing angle changes, eliminating deviation and providing stable anchoring support for modular management, thus ensuring the accuracy of virtual content display.

[0042] Preferably, the specific steps for modular content configuration and management in step S4 are as follows:

[0043] S41. Modular Division of Virtual Content: After binding is completed in step S3, the virtual content is divided into independent modules. The 3D model adopts the glTF format, which includes geometric and PBR material parameters and animation information, and supports adaptive fusion with physical environment lighting. Interactive elements define click, drag and other operations and response logic. Audio and video support multiple formats and can adapt to the network. Metadata covers exhibit name, background and other information to achieve orderly splitting and standardized management of content.

[0044] S42. Efficient Content Management Mechanism: Through the backend CMS system, users can perform operations such as adding and modifying content via a visual configuration module. Incremental update technology is used to reduce bandwidth and time consumption, and version history is recorded to support rollback, thereby improving operation and maintenance efficiency and providing a flexible and scalable content foundation for personalized roaming path generation.

[0045] Preferably, the specific steps for generating the personalized roaming path in step S5 are as follows:

[0046] S51. Multi-dimensional data supports path generation: Based on the modular management in step S4, multi-dimensional data is integrated, and passage areas and obstacles are marked according to the exhibition hall grid map; importance weights are assigned according to content distribution; historical user behavior and real-time traffic data are collected, and the initial path is generated by combining the A* algorithm with a weighted graph. The weights cover factors such as length and importance, providing a reference for optimization.

[0047] S52. Personalized Path Intelligent Optimization: Utilizing machine learning models to analyze user interests and build user profiles. First-time visitors are recommended a general path containing key content based on their registration information; repeat visitors are recommended a path containing relevant new content based on their historical behavior, maximizing the fit with their needs and providing precise roaming guidance for multi-terminal interaction adaptation.

[0048] Preferably, the specific steps of multi-terminal interaction adaptation in step S6 are as follows:

[0049] S61. Adaptation to mainstream terminal interaction features: Based on the path generated in step S5, it adapts to different terminals. VR headsets ensure clear display through distortion correction, and posture information mapping enables natural switching of viewpoints with a latency of ≤20ms. Mobile terminals automatically adjust the content layout and map touch events to virtual operations, such as clicking to view details and swiping to switch viewpoints.

[0050] S62. Cross-device consistent experience guarantee: The web browser uses WebGL rendering and adjusts the video stream quality according to network conditions through an adaptive bitrate push mechanism to ensure smooth playback. All devices provide a consistent interactive experience with unified operation logic, laying a solid terminal foundation for rendering optimization and enhanced immersion.

[0051] Preferably, the specific steps for rendering optimization and immersion enhancement in step S7 are as follows:

[0052] S71, Precise Rendering Performance Optimization: Rendering is optimized based on the adaptation in step S6. The frustum culling algorithm reduces the rendering load; streaming loading technology preloads content to avoid delays; LOD grading uses different precision models and textures for objects at different distances to ensure visual effects while maintaining a frame rate of ≥30fps on low-configuration devices, ensuring smooth system operation.

[0053] S72. Enhanced Immersion in Multiple Dimensions: By combining PBR lighting technology, the lighting parameters of virtual content are adjusted according to the physical exhibition hall light source to achieve light and shadow fusion; during interaction, tactile feedback, visual flashing and auditory prompts enhance the sense of realism; spatial sound effects generate directional sound effects based on the location of the sound source and the user's orientation, and multiple technologies work together to enhance the immersive experience.

[0054] Preferably, in step S31, the infrared marker anchor point maintains a positioning accuracy of ≤2cm and an identification response time of ≤100ms under ambient light intensity of 0-10000 lux. When the infrared marker anchor point has an occlusion area of ​​≤30%, the identification response time is extended by no more than 50ms, while still maintaining a positioning accuracy of ≤2cm.

[0055] The beneficial effects of this invention are as follows:

[0056] 1. This invention effectively solves the problem of virtual content drifting and offset in traditional systems by using a multi-sensor fusion algorithm and closed-loop optimization mechanism. Even in complex environments such as changes in visitor flow and lighting fluctuations in exhibition halls, it can maintain precise alignment between virtual content and physical space, ensuring a consistent and stable visual experience for users during navigation and providing a reliable spatial benchmark for virtual-real interaction.

[0057] 2. This invention uses physical rendering and multi-sensory interaction technology to naturally integrate the lighting effects and material representation of virtual content with the physical environment. At the same time, it combines spatial sound effects and interactive feedback to create a multi-dimensional sense of immersion. When users interact with virtual exhibits, they can obtain a sensory experience that is close to reality, which significantly enhances the sense of immersion and attractiveness of digital exhibition halls.

[0058] 3. By adopting a modular content management and multi-terminal adaptation architecture, this invention supports the rapid updating, expansion and cross-device synchronization of virtual content, and can adapt to various scenarios such as museums, corporate exhibition halls and temporary exhibitions without large-scale adjustments; at the same time, the personalized path generation mechanism can meet the interests and needs of different users, taking into account both efficient operation and maintenance and diversified applications, greatly improving the practical value and promotion potential of the system. Attached Figure Description

[0059] Figure 1 This is a flowchart of the method for constructing a panoramic digital exhibition hall based on virtual-real fusion according to the present invention. Detailed Implementation

[0060] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0061] like Figure 1 As shown in the figure, this invention provides a method for constructing a panoramic digital exhibition hall based on virtual-real fusion. The specific steps of this method are as follows:

[0062] S1. Panoramic Image and Spatial Data Acquisition: By deploying high-parameter panoramic cameras, depth cameras, and inertial navigation sensors, multi-view images, depth, and pose data are acquired. The data is then aligned with timestamps and calibrated using the IEEE 1588 protocol to ensure uniformity and high precision, providing reliable input for step S2.

[0063] S2. Matching of physical and virtual coordinate systems: Based on the multi-source data obtained in step S1, a multi-sensor fusion algorithm is adopted, including extended Kalman filtering for short-term estimation of real-time pose and graph optimization for long-term correction of global pose. The two work together to make the modeling error of the physical coordinate system ≤2cm. Combined with the virtual space reference point, the rigid body transformation is achieved through the iterative nearest point algorithm, and the matching error is ≤3cm.

[0064] S3. Setting and binding virtual content anchor points: Based on the coordinate system established in step S2, identify anchor points such as QR codes and SIFT feature points, bind virtual content using SLAM technology, and ensure that the virtual content drifts ≤1cm when the viewing angle changes through dynamic optimization, so as to provide stable spatial anchoring for step S4.

[0065] S4. Modular content configuration and management: After anchor point binding is completed in step S3, the virtual content is divided into modules such as 3D models. The CMS system enables modular configuration, incremental hot updates and version management, flexibly expands the content, and provides scalable support for step S5.

[0066] S5. Personalized Roaming Path Generation: Through the modular content management in step S4, combined with data such as exhibition hall layout and content distribution, an initial path is generated using the A* algorithm. After optimization by machine learning, roaming routes that match the interests of different users are recommended, providing a path basis for step S6.

[0067] S6. Multi-terminal interaction adaptation: Based on the personalized path generated in step S5, for terminals such as virtual reality headsets, adaptation is carried out through distortion correction, touch mapping and other mechanisms to ensure consistent and smooth multi-terminal interaction, laying the terminal foundation for step S7.

[0068] S7. Rendering Optimization and Immersive Enhancement: After completing multi-terminal adaptation in step S6, rendering performance is optimized by using region culling and LOD management, and immersiveness is enhanced by combining PBR lighting and spatial sound effects to improve the user's immersive interactive experience in different scenarios.

[0069] This invention effectively solves the problem of virtual content drifting and offset in traditional systems through multi-sensor fusion algorithms and closed-loop optimization mechanisms. Even in complex environments such as changing visitor flow and fluctuating lighting in exhibition halls, it maintains precise alignment between virtual content and physical space, ensuring a consistent and stable visual experience for users and providing a reliable spatial benchmark for virtual-real interaction. Simultaneously, through physically based rendering and multi-sensory interaction technologies, the lighting effects and material representation of virtual content are naturally integrated with the physical environment. Combined with spatial sound effects and interactive feedback, a multi-dimensional immersive experience is constructed, allowing users to obtain a near-realistic sensory experience when interacting with virtual exhibits, significantly enhancing the immersion and attractiveness of digital exhibition halls. Furthermore, the modular content management and multi-terminal adaptation architecture support rapid updates, expansions, and cross-device synchronization of virtual content. It can adapt to various scenarios such as museums, corporate exhibition halls, and temporary exhibitions without large-scale adjustments, and the personalized path generation mechanism can meet the interests of different users, balancing efficient operation and maintenance with diverse applications, greatly enhancing the system's practical value and promotional potential.

[0070] The specific steps for panoramic image and spatial data acquisition in step S1 are as follows:

[0071] S11. Multi-type sensor collaborative data acquisition: Three types of high-parameter equipment are scientifically deployed in the exhibition hall: a 12K resolution, 30fps panoramic camera, with a 360-degree horizontal field of view and a 190-degree vertical field of view, can clearly capture details such as exhibit textures and wall decorations in every corner of the exhibition hall; a depth camera with a ranging range of 0.5-10 meters and an accuracy of ±1%, can accurately measure the spatial distance between exhibits and the environment, ensuring data reliability; and an inertial navigation sensor with a 100Hz sampling rate, with a heading accuracy of 0.1° and a pitch / roll accuracy of 0.05°, outputs spatial pose information in real time. The three types of equipment work together to acquire images, depth, and spatial data of the exhibition hall from all angles.

[0072] S12. Data calibration ensures consistent accuracy: During data acquisition, the IEEE 1588 protocol is used to align the timestamps of various data types, strictly controlling the synchronization error to within 1ms to avoid data misalignment due to time deviations. Simultaneously, feature point matching is used to correlate multi-sensor data, extracting common features to establish associations, keeping the calibration error within 5cm. This provides unified and high-quality basic data support for accurate matching of the physical and virtual spatial coordinate systems.

[0073] The specific steps for matching the physical space and virtual space coordinate systems in step S2 are as follows:

[0074] S21. Precise Modeling of Physical Space Coordinate System: Based on the multi-source data obtained in step S1, a multi-sensor fusion algorithm is used to construct a physical space coordinate system. The extended Kalman filter state update formula is used to treat position, velocity, and attitude as state variables. Equations are constructed by combining sensor noise characteristics, and state estimation is optimized through prediction and update. The graph optimization algorithm uses keyframes as nodes. The keyframe selection criteria are: one frame is collected every 500ms or when the user moves a distance ≥1m to ensure a balance between the density of pose constraints and computational efficiency. The relative pose constraints between nodes are used as edges. Accumulated errors are eliminated by minimizing the global energy function, thereby improving the modeling accuracy.

[0075] Extended Kalman filter updates the real-time pose every 10ms, and graph optimization performs cumulative error correction on the global pose every 500ms. The two are run alternately to achieve high-precision modeling.

[0076] Extended Kalman Filter State Update Formula Expression:

[0077]

[0078] In the formula: Let K be the updated state vector at time k (containing position, velocity, attitude, etc.); k For Kalman gain; z k P is the sensor measurement value; k|k-1 To predict the state covariance; R k To measure the noise covariance; H k The observation matrix at time k is used to represent the state vector. Mapping to the observation space, such as converting position, attitude and other states into distance observations from the depth camera; h(·) is the observation model, specifically a nonlinear function based on the physical characteristics of the sensor;

[0079] For example, the distance observation model for a depth camera is as follows:

[0080]

[0081] Where (x) p ,y p ,z p (x) represents the coordinates of a point in physical space. c ,y c ,z c () represents the camera coordinates;

[0082] Graph optimization algorithm expression:

[0083]

[0084] In the formula: x is the set of poses for all keyframes; e i (x iThe first term () represents the single-node prior error, originating from sensor noise during keyframe acquisition (such as the heading angle error of an inertial navigation sensor). The calculation formula is as follows: ( (Initial sensor measurement); e ij (x i ,x j The term ) represents the inter-node constraint error, indicating the relative pose deviation between keyframes i and j. The calculation formula is e. ij (x i ,x j )=x j -T ij x i (T ij (where i is the theoretical transformation matrix from j);

[0085] S22. Precise matching of virtual and real space coordinates: Combining the preset reference points in the virtual space (whose coordinate information is accurate and evenly distributed), the spatial transformation relationship between physical feature points and reference points is calculated. The iterative nearest point algorithm is used to repeatedly solve the rotation matrix and translation vector of rigid body transformation, so that the matching error is ≤3cm. This lays a solid coordinate foundation for the spatial anchoring of virtual content and smoothly connects to the subsequent anchor point setting work.

[0086] Formula for solving the transformation matrix using the Iterative Closest Point (ICP) algorithm:

[0087]

[0088] In the formula: R is the rotation matrix; t is the translation vector; p i For physical space feature points; q i The virtual reference point is denoted as n; n is the number of feature points participating in the matching, with a selection standard of 250 (to ensure matching stability), and the physical space feature points and the virtual reference point must have a one-to-one correspondence (e.g., associated through SIFT feature matching).

[0089] The specific steps for setting and binding virtual content anchor points in step S3 are as follows:

[0090] S31. Diverse Anchor Point Recognition and Positioning: Based on the coordinate system established in step S2, multiple anchor points are set in physical space. The QR code anchor point adopts an anti-distortion algorithm, which can achieve fast recognition at 10 frames / second even if it is tilted or blurred. The SIFT feature point anchor point is accurately positioned due to its rotation and scaling invariance, and is suitable for areas without obvious markings. The infrared marker anchor point uses an infrared camera to capture the position, which is not affected by ambient light and has a positioning accuracy of 2cm, providing a variety of options for binding.

[0091] S32. SLAM technology achieves stable binding: SLAM technology tracks the user's position and posture in real time, binding virtual exhibits and other content to their physical locations. Through closed-loop detection (triggered when the feature point overlap of 5 consecutive key frames is ≥75%) and bundle adjustment optimization (using the Levenberg-Marquardt iterative algorithm, with a single iteration time ≤10ms), the iteration terminates when the error change between two consecutive iterations is ≤0.1cm, ensuring that virtual content drift is controlled within ≤1cm. Continuous dynamic correction ensures that virtual content drift is ≤1cm when the viewing angle changes, eliminating offset and providing stable anchoring support for modular management, guaranteeing the accuracy of virtual content display.

[0092] The specific steps for modular content configuration and management in step S4 are as follows:

[0093] S41. Modular Division of Virtual Content: After binding is completed in step S3, the virtual content is divided into independent modules. The 3D model adopts the glTF format, which includes geometric and PBR material parameters and animation information, and supports adaptive fusion with physical environment lighting. Interactive elements define click, drag and other operations and response logic. Audio and video support multiple formats and can adapt to the network. Metadata covers exhibit name, background and other information to achieve orderly splitting and standardized management of content.

[0094] S42. Efficient Content Management Mechanism: Through the backend CMS system, users can perform operations such as adding and modifying content via a visual configuration module. Incremental update technology is used to reduce bandwidth and time consumption, and version history is recorded to support rollback, thereby improving operation and maintenance efficiency and providing a flexible and scalable content foundation for personalized roaming path generation.

[0095] The specific steps for generating the personalized roaming path in step S5 are as follows:

[0096] S51. Multi-dimensional data supports path generation: Based on the modular management in step S4, multi-dimensional data is integrated, and passage areas and obstacles are marked according to the exhibition hall grid map; importance weights are assigned according to content distribution; historical user behavior and real-time traffic data are collected, and the initial path is generated using the A* algorithm combined with a weighted graph. The weights cover factors such as length and importance, with length accounting for 60% (the shorter the path, the higher the weight) and importance accounting for 40% (assigned based on exhibit popularity and user interest tags), providing a reference for optimization;

[0097] S52. Personalized Path Intelligent Optimization: Utilizing machine learning models to analyze user interests, the model employs a collaborative filtering algorithm. It calculates the similarity of user interest vectors using cosine similarity, with a value range of [0,1]. Users with a similarity ≥ 0.6 are considered similar. Interest tags are generated based on user history and similar user preferences to construct a user profile. Model input features include the number of times a user clicks on exhibits, dwell time, and content type preferences in historical paths. Path weight parameters are optimized using a gradient descent algorithm. First-time visitors are recommended a general path containing key content based on registration information; repeat visitors are recommended paths with relevant new content based on historical behavior, maximizing alignment with user needs and providing precise roaming guidance for multi-terminal interaction adaptation.

[0098] The specific steps for multi-terminal interaction adaptation in step S6 are as follows:

[0099] S61. Adaptation to mainstream terminal interaction features: Based on the path generated in step S5, it adapts to different terminals. VR headsets ensure clear display through distortion correction, and posture information mapping enables natural switching of viewpoints with a latency of ≤20ms. Mobile terminals automatically adjust the content layout and map touch events to virtual operations, such as clicking to view details and swiping to switch viewpoints.

[0100] S62. Cross-device consistent experience guarantee: The web browser uses WebGL rendering and adjusts the video stream quality according to network conditions through an adaptive bitrate push mechanism to ensure smooth playback. All devices provide a consistent interactive experience with unified operation logic, laying a solid terminal foundation for rendering optimization and enhanced immersion.

[0101] The specific steps for rendering optimization and immersion enhancement in step S7 are as follows:

[0102] S71. Precise Rendering Performance Optimization: Rendering is optimized based on the adaptation in step S6. The frustum culling algorithm reduces the rendering load; streaming loading technology preloads content to avoid delays; LOD grading uses different precision models and textures for objects at different distances. Specifically: high precision models (≥10,000 triangle faces) are used within 5 meters of the user, medium precision models (3,000-10,000 triangle faces) are used between 5 and 10 meters, and low precision models (≤3,000 triangle faces) are used above 10 meters. This ensures visual quality while maintaining a frame rate of ≥30fps on low-configuration devices, ensuring smooth system operation.

[0103] S72. Enhanced Immersion in Multiple Dimensions: Combining PBR lighting technology, the lighting parameters of virtual content are adjusted according to the physical exhibition hall's light source. Specifically, the intensity and color temperature parameters of the physical exhibition hall's light source are collected in real time through sensors, and the lighting parameters of the virtual content are mapped and adjusted at a 1:1 ratio to ensure natural light and shadow transitions and achieve light and shadow fusion. During interaction, tactile feedback, visual flashing, and auditory cues enhance the sense of realism. Spatial sound effects generate directional sound effects based on the location of the sound source and the user's orientation. Multiple technologies work together to enhance the immersive experience.

[0104] In step S31, the infrared marker anchor point maintains a positioning accuracy of ≤2cm and an identification response time of ≤100ms under ambient light intensity of 0-10000 lux. When the infrared marker anchor point has an occlusion area of ​​≤30%, the identification response time is extended by no more than 50ms, while still maintaining a positioning accuracy of ≤2cm.

[0105] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus.

[0106] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.

Claims

1. A method for constructing a panoramic digital exhibition hall roaming experience based on virtual-real fusion, characterized in that: The specific steps of this method are as follows: S1. Panoramic Image and Spatial Data Acquisition: By deploying high-parameter panoramic cameras, depth cameras, and inertial navigation sensors, multi-view images, depth, and pose data are acquired. The data is then aligned with timestamps and calibrated using the IEEE 1588 protocol to ensure uniformity and high precision. S2. Matching of physical and virtual coordinate systems: Based on the multi-source data obtained in step S1, a multi-sensor fusion algorithm is adopted, including extended Kalman filtering for short-term estimation of real-time pose and graph optimization for long-term correction of global pose. The two work together to make the modeling error of the physical coordinate system ≤2cm. Combined with the virtual space reference point, the rigid body transformation is achieved through the iterative nearest point algorithm, and the matching error is ≤3cm. S3. Setting and binding virtual content anchor points: Based on the coordinate system established in step S2, identify QR codes and SIFT feature point anchor points, use SLAM technology to bind virtual content, and ensure that the virtual content drifts ≤1cm when the viewing angle changes through dynamic optimization. S4. Modular Content Configuration and Management: After anchor point binding is completed in step S3, the virtual content is divided into 3D model modules. The CMS system enables modular configuration, incremental hot updates, and version management, allowing for flexible content expansion. S5. Personalized Roaming Path Generation: Through the modular content management in step S4, combined with the exhibition hall layout and content distribution data, an initial path is generated using the A* algorithm. After optimization by machine learning, roaming routes that match the interests of different users are recommended. S6. Multi-terminal interaction adaptation: Based on the personalized path generated in step S5, for virtual reality headset terminals, through distortion correction and touch mapping mechanism adaptation, the interaction of multiple terminals is consistent and smooth. S7. Rendering Optimization and Immersive Enhancement: After completing multi-terminal adaptation in step S6, region culling and LOD management are used to optimize rendering performance, and PBR lighting and spatial sound effects are combined to enhance immersiveness.

2. The method for constructing a panoramic digital exhibition hall roaming experience based on virtual-real fusion according to claim 1, characterized in that: The specific steps for panoramic image and spatial data acquisition in step S1 are as follows: S11. Multi-type sensor collaborative acquisition: Three types of high-parameter equipment are deployed in the exhibition hall: a panoramic camera with 12K resolution and 30fps frame rate; a depth camera with a ranging range of 0.5-10 meters and an accuracy of ±1%; and an inertial navigation sensor with a sampling rate of 100Hz. The three types of equipment work together to acquire images, depth and spatial data of the exhibition hall from all directions. S12. Data calibration ensures consistent accuracy: During the acquisition process, the IEEE1588 protocol is used to align the timestamps of various types of data, and the synchronization error is strictly controlled within 1ms. The feature point matching method is used to associate multi-sensor data, extract common features to establish association, and control the calibration error within 5cm.

3. The method for constructing a panoramic digital exhibition hall roaming experience based on virtual-real fusion according to claim 1, characterized in that: The specific steps for matching the physical space and virtual space coordinate systems in step S2 are as follows: S21. Precise Modeling of Physical Space Coordinate System: Based on the multi-source data acquired in S1, a multi-sensor fusion algorithm is used to construct a physical space coordinate system. The extended Kalman filter state update formula is used to treat position, velocity, and attitude as state variables. Equations are constructed by combining sensor noise characteristics, and state estimation is optimized through prediction and update. The graph optimization algorithm uses keyframes as nodes and relative pose constraints between nodes as edges. Accumulated errors are eliminated by minimizing the global energy function. The keyframe selection interval for the graph optimization algorithm is 500ms or the moving distance is ≥1m, and the number of iterations for minimizing the global energy function is ≥10. S22. Precise matching of virtual and real space coordinates: Combining the preset reference point in virtual space, the spatial transformation relationship between physical feature points and reference points is calculated. The iterative nearest point algorithm is used to repeatedly solve the rotation matrix and translation vector of rigid body transformation, so that the matching error is ≤3cm.

4. The method for constructing a panoramic digital exhibition hall roaming based on virtual-real fusion according to claim 1, characterized in that: The specific steps for setting and binding virtual content anchor points in step S3 are as follows: S31. Diverse Anchor Point Recognition and Positioning: Based on the coordinate system established in step S2, various anchor points are set in the physical space. The QR code anchor points adopt an anti-distortion algorithm; SIFT feature point anchor points are suitable for areas without obvious markings; infrared marker anchor points use infrared cameras to capture positions, with a positioning accuracy of up to 2cm. S32. SLAM technology achieves stable binding: SLAM technology tracks the user's position and posture in real time, binding the virtual exhibit content to the physical position. Through closed-loop detection and beam adjustment optimization, it continuously and dynamically corrects the virtual content drift of ≤1cm when the viewing angle changes. After the closed-loop detection of SLAM technology is triggered, the iteration termination condition of beam adjustment optimization is the error change amount ≤0.1cm.

5. The method for constructing a panoramic digital exhibition hall roaming experience based on virtual-real fusion according to claim 1, characterized in that: The specific steps for modular content configuration and management in step S4 are as follows: S41. Modular division of virtual content: After binding is completed in step S3, the virtual content is divided into independent modules. The 3D model adopts glTF format, which includes geometric and PBR material parameters and animation information, and supports adaptive fusion with physical environment lighting. Interactive elements define click, drag operations, and response logic; audio and video support multiple formats and can adapt to the network. Metadata includes exhibit names and background information; S42. Efficient content management mechanism: Through the backend CMS system, users can add and modify content via a visual configuration module. Incremental update technology is used to reduce bandwidth and time consumption, and version history is recorded to support rollback.

6. The method for constructing a panoramic digital exhibition hall roaming experience based on virtual-real fusion according to claim 1, characterized in that: The specific steps for generating the personalized roaming path in step S5 are as follows: S51. Multi-dimensional data supports path generation: Based on the modular management of step S4, multi-dimensional data is integrated, and passage areas and obstacles are marked according to the exhibition hall grid map; importance weights are assigned according to content distribution. Collect user history and real-time traffic data, and use the A* algorithm combined with a weighted graph to generate an initial path. The weights cover length and importance factors. In the path weights of the A* algorithm, the length factor accounts for 60% and the content importance factor accounts for 40%. S52. Personalized Path Intelligent Optimization: Utilize machine learning models to analyze user interests and build profiles. For first-time visitors, a general path containing key content is recommended based on registration information; for repeat visitors, a path containing relevant new content is recommended based on historical behavior.

7. The method for constructing a panoramic digital exhibition hall roaming based on virtual-real fusion according to claim 1, characterized in that: The specific steps for multi-terminal interaction adaptation in step S6 are as follows: S61. Adaptation to mainstream terminal interaction features: Based on the path generated in step S5, it adapts to different terminals. VR headsets ensure clear display through distortion correction, and posture information mapping enables natural switching of viewpoints with a latency of ≤20ms. Mobile terminals automatically adjust the content layout and map touch events to virtual operations. S62, Cross-device consistent experience guarantee: The web browser uses WebGL rendering and adjusts the video stream quality according to network conditions through an adaptive bitrate push mechanism to ensure smooth playback. All types of devices provide a consistent interactive experience and unified operation logic.

8. The method for constructing a panoramic digital exhibition hall roaming experience based on virtual-real fusion according to claim 1, characterized in that: The specific steps for rendering optimization and immersion enhancement in step S7 are as follows: S71. Precise Rendering Performance Optimization: Rendering is optimized based on the adaptation in step S6. The frustum culling algorithm reduces the rendering load. Streaming loading technology preloads content to avoid delays. LOD grading uses different precision models and textures for objects at different distances to ensure visual effects while enabling low-configuration devices to achieve a frame rate of ≥30fps. S72. Enhanced Immersion in Multiple Dimensions: Combining PBR lighting technology, the lighting parameters of virtual content are adjusted according to the physical exhibition hall light source to achieve light and shadow fusion; during interaction, tactile feedback, visual flashing and auditory prompts enhance the sense of realism; spatial sound effects generate directional sound effects based on the location of the sound source and the user's orientation.

9. The method for constructing a panoramic digital exhibition hall roaming based on virtual-real fusion according to claim 4, characterized in that: In step S31, the infrared marker anchor point maintains a positioning accuracy of ≤2cm and an identification response time of ≤100ms under ambient light intensity of 0-10000 lux. When the infrared marker anchor point has an occlusion area of ​​≤30%, the identification response time is extended by no more than 50ms, while still maintaining a positioning accuracy of ≤2cm.

Citation Information

Patent Citations

  • Exhibition hall construction method, device and equipment based on element universe

    CN118689314A

  • Digital exhibition hall server based on 3D scene editing and application method thereof

    CN119494916A

Cited By

  • Digital media interactive display and immersive experience generation system and method

    CN121070506A

  • A system and method for generating interactive digital media displays and immersive experiences.

    CN121070506B

  • Digital display method and system for tourism products

    CN121280678A