An outdoor immersive travel and culture mixed positioning and dynamic rendering optimization method
By using a hybrid positioning model and multi-dimensional virtual-real fusion technology, the problem of low positioning accuracy and insufficient dynamic adjustment capability of traditional cultural and tourism navigation in complex outdoor scenarios has been solved, realizing a high-precision, personalized, and immersive cultural and tourism experience.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- JIANGSU XUNGAO INTELLIGENT TECH CO LTD
- Filing Date
- 2025-05-28
- Publication Date
- 2026-05-12
AI Technical Summary
Traditional tourism navigation methods rely heavily on positioning in complex outdoor scenarios, have low navigation accuracy, and lack dynamic adjustment capabilities. They are unable to provide users with real-time, adaptive, and context-sensitive navigation in open scenic areas with multiple targets and paths, thus affecting user experience.
A hybrid positioning model is adopted, which combines ultra-wideband ranging signals, visual inertial odometry signals, real-time carrier phase differential signals and inertial navigation signals. Spatial coordinates are generated through a multi-layer signal processing filtering model and weighted least squares method. Combined with a user mobility situation prediction model and a multi-dimensional virtual-real fusion model, augmented reality spatial scenes and interactive content are generated.
It improves the stability and accuracy of outdoor positioning, achieves deep integration of virtual content and the real environment, enhances users' immersion and cultural interaction participation, and provides personalized cultural interactive content.
Smart Images

Figure CN120576758B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of high-precision positioning technology, specifically to an outdoor immersive cultural tourism hybrid positioning and dynamic rendering optimization method. Background Technology
[0002] As cultural tourism scenarios continue to develop towards immersion and intelligence, tourists' navigation needs in outdoor environments are becoming increasingly diversified and real-time. Traditional cultural tourism navigation methods, such as paper maps, QR code signs, and handheld audio guides, suffer from weak interactivity, strong reliance on location, and low navigation accuracy, making it difficult to meet the comprehensive needs of contemporary tourists for immersive, self-guided tours and route guidance.
[0003] In complex outdoor scenarios, due to the varied terrain and wide activity areas of tourists, navigation systems face multiple challenges, including how to guide users to complete path identification, pose adjustment, and point of interest access in real time. Most existing systems rely on fixed paths or preset points for navigation information prompts, lacking the ability to dynamically adjust based on the user's current state. This is especially problematic in open scenic areas with multiple goals and paths, where the real-time nature, adaptability, and contextual relevance of navigation information are difficult to guarantee, impacting the navigation experience.
[0004] Furthermore, current navigation devices often fail to effectively integrate information such as user movement status, orientation changes, and spatial reference points for comprehensive navigation control. This results in navigation commands easily becoming detached from the actual location relationships within the scene, making continuous and stable guidance difficult. While some systems have attempted to use multiple sensors for assisted navigation (such as geomagnetism, accelerometers, and gyroscopes), a complete multi-source fusion navigation method suitable for outdoor cultural and tourism scenarios is still lacking. This makes it difficult to achieve synchronous management of the user's dynamic location relationships within the scene, the navigation path, and environmental elements.
[0005] To address this, a hybrid positioning and dynamic rendering optimization method for outdoor immersive cultural tourism is proposed. Summary of the Invention
[0006] The purpose of this invention is to provide an optimized method for hybrid positioning and dynamic rendering in outdoor immersive cultural tourism scenarios. This method includes acquiring the positioning signal of a mobile terminal in a target service area through a hybrid positioning model. The positioning signal includes ultra-wideband ranging signals, visual-inertial odometry signals, real-time carrier phase differential signals, and inertial navigation signals. The positioning signal is processed using a multi-layer signal processing filtering model to obtain coordinate data. The coordinate data is then solved using a weighted least squares method to generate spatial coordinates. Geographical environmental parameters of the target service area are acquired, and the spatial coordinates, geographic environmental parameters, and virtual objects are integrated using a multi-dimensional virtual-real fusion model to generate an augmented reality spatial scene. Finally, a multi-modal driving model analyzes the user's movement trajectory data, spatial coordinates, and the augmented reality spatial scene to generate interactive content. This makes outdoor immersive cultural tourism scenarios more intelligent.
[0007] To achieve the above objectives, the present invention provides the following technical solution:
[0008] An outdoor immersive cultural tourism hybrid positioning and dynamic rendering optimization method includes:
[0009] The positioning signal of the mobile terminal in the target service area is obtained through a hybrid positioning model; the positioning signal includes ultra-wideband ranging signal, visual inertial odometry signal, real-time carrier phase differential signal and inertial navigation signal; the positioning signal is processed through a multi-layer signal processing filtering model to obtain coordinate data, and the coordinate data is calculated by weighted least squares method to generate spatial coordinates;
[0010] A user mobility situation prediction model is constructed based on spatial coordinates to analyze the user's movement trajectory, generate navigation prediction trajectory and navigation confidence; virtual objects are pre-rendered based on the navigation prediction trajectory to generate a preloading strategy.
[0011] The geographic environment parameters of the target service area are obtained, and the spatial coordinates, geographic environment parameters, preloading strategies and virtual objects are integrated through a multi-dimensional virtual-real fusion model to generate an augmented reality spatial scene.
[0012] Interactive content is generated by analyzing user movement trajectory data, spatial coordinates, and augmented reality spatial scenes using a multimodal driven model.
[0013] Preferably, the hybrid positioning model includes: an ultra-wideband positioning layer, which determines the distance relationship between the mobile terminal and each base station based on a multilateral positioning algorithm and generates an ultra-wideband ranging signal;
[0014] The visual inertial odometry layer collects visual feature points of the environment, combines them with motion data from the inertial measurement unit to obtain position trajectory information in the local coordinate system, and outputs visual inertial odometry signals.
[0015] The global navigation and positioning layer obtains positioning signals, and obtains a global position reference through differential processing of carrier phase observations between the base station and the rover station, generating a real-time carrier phase differential signal;
[0016] The inertial navigation layer measures the linear acceleration, angular velocity, and magnetic field strength of the mobile terminal, provides high-frequency attitude and motion state data, and generates inertial navigation signals.
[0017] Preferably, the multi-layer signal processing filtering model includes: a first filtering layer, which uses the Kalman filtering algorithm to preprocess the ultra-wideband ranging signal, the real-time carrier phase differential signal, and the inertial navigation signal, removes signal outliers through state prediction and observation update steps, and outputs preliminary position estimation data;
[0018] The second filtering layer uses a complementary filtering algorithm to fuse the local map data constructed from the visual inertial odometry signal with the preliminary position estimation data, and outputs multi-source fused position data.
[0019] The weight allocation mechanism calculates the confidence weight coefficient of each positioning source in real time based on signal strength, environmental occlusion, and sensor confidence, and outputs dynamic weight parameters.
[0020] The coordinate calculation module uses dynamic weight parameters to perform weighted least squares calculation on multi-source fused location data to generate three-dimensional spatial coordinate data.
[0021] Preferably, the user mobility situation prediction model includes:
[0022] The trajectory feature extraction module analyzes historical spatial coordinate sequences based on a sliding time window to extract the trajectory features of user movement, including speed change patterns, direction change frequency, and stop point distribution features.
[0023] The motion state recognition module identifies the user's motion state by combining the current motion state parameters, including four basic states: walking, standing and watching, moving quickly, and exploring directions.
[0024] The path prediction algorithm integrates trajectory features and motion state to predict the user's movement path and generate a navigation prediction trajectory that includes location coordinates. The navigation confidence is calculated by measuring the degree of deviation between the navigation prediction trajectory and the actual trajectory. When the navigation confidence is not lower than a preset threshold, the virtual object is pre-rendered based on the navigation prediction trajectory to generate a preloading strategy.
[0025] Preferably, the multidimensional virtual-real fusion model includes: an environmental perception layer, which collects information on light intensity, weather conditions, terrain elevation, and object surface material of the target service area in real time to obtain geographic environmental parameters;
[0026] The digital asset matching layer selects virtual buildings, characters, and scene elements corresponding to the current spatial coordinates based on a 3D digital model library, and adjusts the lighting, material reflection, and physical properties of virtual objects according to geographical environment parameters.
[0027] The spatial registration layer uses coordinate data to locate virtual objects that have been adapted to the environment to their corresponding positions in the real scene, and uses coordinate transformation algorithms to align virtual content with the physical environment.
[0028] The dynamic synchronization layer monitors changes in geographic environment parameters and preloading strategies in real time, updates the visual appearance of virtual objects, and generates augmented reality spatial scenes.
[0029] Preferably, the multimodal driving model includes: a user behavior analysis layer, which performs trajectory segmentation and pattern recognition on user movement trajectory data, and extracts user behavior feature parameters at different spatial locations, including dwell time, movement speed and interaction frequency;
[0030] The content matching layer retrieves historical and cultural information, architectural backgrounds, and personal stories related to the current location from a pre-built knowledge graph based on the user's current coordinate data and augmented reality spatial scene information.
[0031] The multimodal fusion layer comprehensively analyzes text descriptions, image resources, and audio materials to generate dialogue text that conforms to the current scene context and creates visual elements that match the historical background.
[0032] The personalized generation layer dynamically adjusts the information weights retrieved by the content matching layer based on personal preference features extracted from behavioral feature parameters, generating customized interactive content, including virtual character dialogues, historical scene reproduction, and interactive task design.
[0033] Preferably, multiple ultra-wideband base stations are deployed in the target service area, and the base stations achieve time synchronization wirelessly. The mobile terminal calculates the three-dimensional spatial coordinate data of the mobile terminal by measuring the time difference of arrival of signals from different base stations and using the hyperbolic positioning principle.
[0034] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0035] 1. This invention proposes a multi-source hybrid positioning model that integrates ultra-wideband ranging, visual inertial odometry, real-time carrier phase differential, and inertial navigation, and performs data fusion and dynamic weighted calculation through a multi-layer signal filtering mechanism. This invention effectively overcomes positioning errors caused by building obstruction, signal attenuation, and complex terrain, improving positioning stability and robustness. Especially in densely populated and structurally complex scenarios such as cultural tourism, traditional technologies often suffer from severe positioning jumps, causing virtual content misalignment and affecting user immersion. This invention, by combining Kalman and complementary filtering, significantly improves the stability and accuracy of spatial coordinates, providing a solid foundation for subsequent augmented reality rendering and trajectory prediction, thereby achieving a smoother and more natural human-computer interaction experience.
[0036] 2. This invention designs a user movement status prediction model based on spatial trajectory feature extraction and motion state recognition. It can dynamically predict the user's next movement path and introduces a navigation confidence mechanism to evaluate the consistency between the predicted path and actual behavior. Especially in application scenarios with limited mobile computing power, this method can detect changes in user behavior in advance, preprocess and cache key content, allowing augmented reality content to "get ahead," improving the system's response speed to user behavior and the coherence of content interaction, thereby enhancing user immersion and cultural interaction participation.
[0037] 3. This invention constructs a multi-dimensional virtual-real fusion model comprising four layers: environmental perception, digital asset matching, spatial registration, and dynamic synchronization. This model can perceive geographical environmental parameters such as lighting, weather, and terrain within the service area in real time, and match corresponding virtual assets based on the current coordinates. It then adjusts lighting effects, physical properties, and material reflections to ensure visual consistency between virtual objects and the real environment. This invention can "embed" virtual content into real-world scenes, achieving a seamless transition in visual perception through spatial coordinate alignment and visual effect synchronization. This not only enhances the realism of content presentation but also strengthens the reproduction of historical scenes and user immersion in cultural and tourism scenarios.
[0038] 4. This invention introduces a four-layer driving mechanism: user behavior analysis, content matching, multimodal fusion, and personalized generation. Based on users' behavioral characteristics in space, it mines individual preferences and combines augmented reality spatial scene information to intelligently retrieve cultural content related to the current location from a knowledge graph. This content is then integrated with text, images, and audio materials to generate context-appropriate interactive content. This invention can dynamically generate personalized cultural interactive content based on user interests, such as virtual character guides and historical dialogue replays, providing users with a unique immersive experience while on the move. Attached Figure Description
[0039] Figure 1This invention provides a schematic diagram of a method for optimizing outdoor immersive cultural tourism hybrid positioning and dynamic rendering.
[0040] Figure 2 This is a schematic diagram of the UWB+SLAM+GPS hybrid positioning technology architecture provided in an embodiment of the present invention;
[0041] Figure 3 This is a schematic diagram illustrating the working principle of UWB positioning technology provided in an embodiment of the present invention;
[0042] Figure 4 This is a schematic diagram illustrating the working principle of the hybrid positioning technology provided in an embodiment of the present invention;
[0043] Figure 5 This is a schematic diagram illustrating the data processing and fusion of hybrid positioning technology provided in an embodiment of the present invention. Detailed Implementation
[0044] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0045] This invention aims to construct an outdoor immersive cultural tourism system based on a hybrid positioning technology of UWB (Ultra-Wideband) and SLAM (Simultaneous Localization and Mapping), and employs a multimodal large-model-driven approach to provide tourists with an unprecedented cultural tourism experience. The system deeply integrates virtual reality (VR), augmented reality (AR), and mixed reality (MR) technologies, seamlessly blending the digital restoration of historical sites with the real environment through high-precision positioning, real-time rendering, and intelligent interaction.
[0046] Example 1
[0047] Please see Figure 1 This invention provides a method for optimizing hybrid positioning and dynamic rendering in outdoor immersive cultural tourism, the technical solution of which is as follows:
[0048] The positioning signal of the mobile terminal in the target service area is obtained through a hybrid positioning model; the positioning signal includes ultra-wideband ranging signal, visual inertial odometry signal, real-time carrier phase differential signal and inertial navigation signal; the positioning signal is processed through a multi-layer signal processing filtering model to obtain coordinate data, and the coordinate data is calculated by weighted least squares method to generate spatial coordinates;
[0049] A user mobility situation prediction model is constructed based on spatial coordinates to analyze the user's movement trajectory, generate navigation prediction trajectory and navigation confidence; virtual objects are pre-rendered based on the navigation prediction trajectory to generate a preloading strategy.
[0050] The geographic environment parameters of the target service area are obtained, and the spatial coordinates, geographic environment parameters, preloading strategies and virtual objects are integrated through a multi-dimensional virtual-real fusion model to generate an augmented reality spatial scene.
[0051] Interactive content is generated by analyzing user movement trajectory data, spatial coordinates, and augmented reality spatial scenes using a multimodal driven model.
[0052] The hybrid positioning model includes: an ultra-wideband positioning layer, which determines the distance relationship between the mobile terminal and each base station based on a multilateral positioning algorithm and generates an ultra-wideband ranging signal;
[0053] The visual inertial odometry layer collects visual feature points of the environment, combines them with motion data from the inertial measurement unit to obtain position trajectory information in the local coordinate system, and outputs visual inertial odometry signals.
[0054] The global navigation and positioning layer obtains positioning signals, and obtains a global position reference through differential processing of carrier phase observations between the base station and the rover station, generating a real-time carrier phase differential signal;
[0055] The inertial navigation layer measures the linear acceleration, angular velocity, and magnetic field strength of the mobile terminal, provides high-frequency attitude and motion state data, and generates inertial navigation signals.
[0056] In this embodiment, a hybrid positioning model consisting of an ultra-wideband positioning layer, a visual-inertial odometry layer, a global navigation positioning layer, and an inertial navigation layer is constructed. This effectively integrates the advantages of multiple positioning methods, significantly improving the positioning accuracy and robustness of mobile terminals in complex outdoor environments. The ultra-wideband positioning layer enhances the system's stability in scenarios with signal obstruction or multipath interference; the visual-inertial odometry layer improves the system's self-awareness when external signals are lacking; the global navigation positioning layer strengthens the overall accuracy of positioning data through a global position reference with centimeter-level precision; and the inertial navigation layer provides the system with high-frequency attitude and motion state data, ensuring continuous positioning capabilities in highly dynamic scenarios. Through this four-layer synergistic and complementary approach, the hybrid positioning model possesses stronger environmental adaptability and real-time response capabilities, providing a solid positioning support foundation for subsequent dynamic rendering and immersive interaction of augmented reality content.
[0057] Multiple ultra-wideband base stations are deployed in the target service area. The base stations synchronize their time wirelessly. The mobile terminal calculates its three-dimensional spatial coordinates by measuring the time difference of arrival of signals from different base stations and using the hyperbolic positioning principle.
[0058] In this embodiment, by deploying multiple ultra-wideband base stations within the target service area according to a preset spatial geometry and achieving wireless time synchronization between the base stations, high-precision and high-stability three-dimensional spatial positioning can be achieved while ensuring flexible system deployment. This effectively improves the real-time performance and reliability of the positioning system in complex outdoor environments. This method is particularly suitable for cultural and tourism scenarios where GPS signals are limited or inaccurate. It can build high-performance positioning capabilities without relying on ground signal infrastructure, providing a precise location information foundation for subsequent user behavior prediction, virtual-real fusion rendering, and dynamic loading of immersive interactive content, thereby significantly enhancing the intelligence and immersion of the cultural and tourism experience.
[0059] The multi-layer signal processing filtering model includes: a first filtering layer, which uses the Kalman filtering algorithm to preprocess the ultra-wideband ranging signal, real-time carrier phase differential signal and inertial navigation signal, removes signal outliers through state prediction and observation update steps, and outputs preliminary position estimation data;
[0060] The second filtering layer uses a complementary filtering algorithm to fuse the local map data constructed from the visual inertial odometry signal with the preliminary position estimation data, and outputs multi-source fused position data.
[0061] The weight allocation mechanism calculates the confidence weight coefficient of each positioning source in real time based on signal strength, environmental occlusion, and sensor confidence, and outputs dynamic weight parameters.
[0062] The coordinate calculation module uses dynamic weight parameters to perform weighted least squares calculation on multi-source fused location data to generate three-dimensional spatial coordinate data.
[0063] The weight allocation mechanism determines dynamic weights by evaluating the signal quality indicators of each positioning source in real time. When vegetation obstruction is detected, causing global navigation signal attenuation, the fusion weight of ultra-wideband positioning and visual inertial odometry is automatically increased. When a narrow channel environment is identified as limiting the propagation of ultra-wideband signals, the weight ratio of visual inertial odometry in constructing local maps is increased.
[0064] In this embodiment, a multi-layer signal processing filtering model is constructed to effectively improve the fusion accuracy and anti-interference capability of multi-source positioning signals in complex outdoor environments. The first filtering layer uses the Kalman filter algorithm to effectively eliminate abnormal fluctuations and noise interference, ensuring the stability and continuity of the initial position estimation. The second filtering layer is based on a complementary filtering algorithm, which enhances the local spatial awareness capability and further improves the fusion consistency and spatial accuracy of multi-source data. A weight allocation mechanism is introduced, thereby enabling the fusion process to have high adaptability and environmental awareness. Finally, the coordinate calculation module performs weighted least squares calculation based on dynamic weight parameters to generate accurate three-dimensional spatial coordinate data, providing a solid and reliable spatial positioning foundation for subsequent navigation prediction, virtual object rendering, and immersive interaction, significantly improving the overall accuracy, stability, and application breadth of the system.
[0065] The user mobility situation prediction model includes:
[0066] The trajectory feature extraction module analyzes historical spatial coordinate sequences based on a sliding time window to extract the trajectory features of user movement, including speed change patterns, direction change frequency, and stop point distribution features.
[0067] The motion state recognition module identifies the user's motion state by combining the current motion state parameters, including four basic states: walking, standing and watching, moving quickly, and exploring directions.
[0068] The path prediction algorithm integrates trajectory features and motion state to predict the user's movement path and generate a navigation prediction trajectory that includes location coordinates. The navigation confidence is calculated by measuring the degree of deviation between the navigation prediction trajectory and the actual trajectory. When the navigation confidence is not lower than a preset threshold, the virtual object is pre-rendered based on the navigation prediction trajectory to generate a preloading strategy.
[0069] In this embodiment, by introducing a user movement status prediction model, intelligent perception and path prediction of user behavior patterns in outdoor cultural and tourism scenarios are achieved, significantly enhancing the system's foresight and interactive response capabilities. The trajectory feature extraction module can accurately identify key features such as speed, direction changes, and dwell time, comprehensively depicting the user's movement behavior patterns; the motion state recognition module helps to better understand the user's current behavioral context. The path prediction algorithm dynamically predicts the user's future location, not only generating specific navigation trajectories but also calculating navigation confidence through deviation analysis from the actual trajectory. When navigation confidence decreases, the system can pre-render virtual objects within the target area and generate corresponding pre-loading strategies, thereby effectively reducing content loading delays, improving the real-time performance and continuity of virtual-real fusion, and providing users with a smoother and more natural immersive interactive experience.
[0070] The multidimensional virtual-real fusion model includes: an environmental perception layer, which collects information on light intensity, weather conditions, terrain elevation, and object surface material in the target service area in real time to obtain geographic environmental parameters;
[0071] The digital asset matching layer selects virtual buildings, characters, and scene elements corresponding to the current spatial coordinates based on a 3D digital model library, and adjusts the lighting, material reflection, and physical properties of virtual objects according to geographical environment parameters.
[0072] The spatial registration layer uses coordinate data to locate virtual objects that have been adapted to the environment to their corresponding positions in the real scene, and uses coordinate transformation algorithms to align virtual content with the physical environment.
[0073] The dynamic synchronization layer monitors changes in geographic environment parameters and preloading strategies in real time, updates the visual appearance of virtual objects, and generates augmented reality spatial scenes.
[0074] The dynamic synchronization layer establishes a mapping relationship between changes in geographic environment parameters and preloading strategies and the attributes of virtual objects. Under rainy or snowy weather conditions, it automatically adds wet or snow texture effects to the surface of virtual buildings. During the day-night cycle, it adjusts the shadow projection direction and brightness intensity of virtual objects according to the real-time lighting angle.
[0075] In this embodiment, by constructing a multi-dimensional virtual-real fusion model, deep integration of virtual content and the real environment is achieved across multiple dimensions, including space, lighting, materials, and dynamic changes, significantly enhancing the realism and immersion of augmented reality scenes. The environment perception layer provides a precise environmental foundation for the adaptation of virtual objects; the digital asset matching layer selects matching virtual buildings, characters, and scene elements from a 3D model library based on the current coordinates and environmental parameters, and intelligently adjusts their lighting, material reflections, and physical properties, allowing virtual content to naturally blend into the real scene. The spatial registration layer ensures spatial consistency and positional accuracy from the user's perspective; the dynamic synchronization layer endows the virtual scene with a high degree of dynamic adaptability and environmental responsiveness. This model greatly enhances the visual coherence, interactive realism, and user immersion experience of augmented reality systems, providing strong technical support for the digital innovation of cultural and tourism scenarios.
[0076] The multimodal driving model includes: a user behavior analysis layer, which performs trajectory segmentation and pattern recognition on user movement trajectory data, and extracts user behavior feature parameters at different spatial locations, including dwell time, movement speed and interaction frequency;
[0077] The content matching layer retrieves historical and cultural information, architectural backgrounds, and personal stories related to the current location from a pre-built knowledge graph based on the user's current coordinate data and augmented reality spatial scene information.
[0078] The multimodal fusion layer comprehensively analyzes text descriptions, image resources, and audio materials to generate dialogue text that conforms to the current scene context and creates visual elements that match the historical background.
[0079] The personalized generation layer dynamically adjusts the information weights retrieved by the content matching layer based on personal preference features extracted from behavioral feature parameters, generating customized interactive content, including virtual character dialogues, historical scene reproduction, and interactive task design.
[0080] The personalized generation layer analyzes users' interactive behavior preferences during historical tours, identifies users' attention to different themes such as architectural art, historical figures, or cultural legends, and adjusts the theme focus and presentation of content generation accordingly. It generates architectural craftsmanship analysis content for users who prefer architectural details and biographical interactive scenarios for users who are interested in historical stories.
[0081] In this embodiment, by constructing a multimodal driven model and fully integrating user behavior data with augmented reality content resources, personalized and contextualized content generation and dynamic interaction are achieved in immersive cultural tourism scenarios, effectively enhancing the depth and engagement of the user experience. The user behavior analysis layer provides data support for understanding user interests and participation methods; the content matching layer ensures the historical relevance and contextual consistency of the content. The multimodal fusion layer enriches the user's perceptual layers in the virtual scene; and the personalized generation layer provides biographical sketches and interactive scenario recreations for users who are passionate about historical stories. This multidimensional perception and intelligent adaptation mechanism not only enhances the relevance and attractiveness of the content but also encourages users to build emotional resonance during interaction, enhancing the effectiveness and depth of cultural dissemination.
[0082] Table 1 Technical Feature Comparison and Verification Table
[0083]
[0084] This invention significantly improves positioning accuracy, stability, and real-time response capabilities in complex outdoor environments by constructing a multi-source fusion hybrid positioning model and optimizing signal processing and user movement prediction mechanisms. This provides a solid positioning foundation for subsequent augmented reality dynamic rendering and immersive interaction, enabling virtual content to deeply integrate with the real environment in terms of space, lighting, materials, and dynamic changes, greatly enhancing the realism and immersion of the scene. Simultaneously, the system can intelligently perceive user behavior and predict trajectories, combining multimodal driving to achieve personalized and contextualized content preloading and interaction, effectively reducing rendering latency. This provides users with a smooth, natural, highly engaging, and emotionally resonant intelligent cultural tourism experience, enhancing the effect of cultural dissemination, as detailed in Table 1.
[0085] Example 2
[0086] This invention integrates ultra-wideband (UWB) positioning, simultaneous localization and mapping (SLAM) technology, a real-time dynamic rendering engine, and a multimodal large-model-driven strategy to construct an immersive interactive model covering the entire domain and multiple scenarios. It achieves centimeter-level positioning accuracy, dynamic synchronization of virtual and real scenes, and high-degree-of-freedom human-computer interaction in complex outdoor environments, significantly improving the realism and security of the user experience. It also provides flexible content iteration and hardware compatibility support for later operation. The technical solution involves obtaining the mobile terminal's positioning signal in the target service area through a hybrid positioning model. The positioning signal includes ultra-wideband ranging signals, visual-inertial odometry signals, real-time carrier phase differential signals, and inertial navigation signals. The positioning signal is processed through a multi-layer signal processing filtering model to obtain coordinate data, and spatial coordinates are generated by calculating the coordinate data using a weighted least squares method.
[0087] A user mobility situation prediction model is constructed based on spatial coordinates to analyze the user's movement trajectory, generate navigation prediction trajectory and navigation confidence; virtual objects are pre-rendered based on the navigation prediction trajectory to generate a preloading strategy.
[0088] The geographic environment parameters of the target service area are obtained, and the spatial coordinates, geographic environment parameters, preloading strategies and virtual objects are integrated through a multi-dimensional virtual-real fusion model to generate an augmented reality spatial scene.
[0089] Interactive content is generated by analyzing user movement trajectory data, spatial coordinates, and augmented reality spatial scenes using a multimodal driven model.
[0090] To achieve the above objectives, this invention first constructs a hybrid positioning model and a multi-layer signal processing and filtering model. It employs a hybrid positioning technology combining UWB and SLAM, and integrates GPS RTK (Real-Time Kinematic) technology with an IMU (Inertial Measurement Unit) module to achieve multi-source data fusion. Based on the UWB+SLAM+GPS hybrid positioning technology architecture, a complete multi-source positioning fusion model is constructed. This architecture demonstrates how GPS, UWB, and SLAM—three core positioning technologies—work collaboratively to achieve different levels of immersive experience through mobile devices and head-mounted displays. The architectural principles are described in [link to architecture description]. Figure 2 .
[0091] This demonstration showcases the multi-module collaborative working principle of an outdoor mixed reality positioning system. It is based on three major positioning technologies: GPS, UWB, and SLAM, and constructs a complete working chain through hierarchical data flow. On the left side of the diagram, the GPS module provides a global geographic coordinate framework via satellite signals, serving as the spatial reference for the entire system. The UWB module communicates with the head-mounted display (HUD) through a deployed base station network, utilizing ultra-wideband pulse signals to achieve centimeter-level local spatial calibration. The SLAM module, through real-time data fusion from visual sensors and inertial measurement units, autonomously constructs an environmental map and maintains continuous positioning in areas lacking GPS signals. These three underlying technology modules perform heterogeneous computing and fusion of positioning data on the mobile device—the mobile device integrates multi-source positioning information through a dynamic weighted algorithm, eliminating error drift from individual technologies, and transmits the processed high-precision coordinates to the HUD via arrows. The HUD aligns virtual and real spaces based on the fused positioning data. The AR+MR arrow on the left represents lightweight augmented reality applications, while the LBE VR and LBS MR+VR arrows on the right correspond to immersive mixed reality scenarios supported by the HUD: in open areas, the combination of GPS and SLAM enables road-level navigation and spatial semantic interaction. Data transmission and fusion between modules enable the system to have dynamic fault tolerance. For example, when the GPS signal is interrupted, SLAM and UWB maintain positioning continuity through coordinate system alignment, while the predictive rendering technology at the head-mounted display compensates for data transmission delays, ensuring the stability of the virtual-real fusion image. Through complementary collaboration of multiple technology layers, the overall architecture achieves multi-precision positioning coverage from centimeter to meter level in complex outdoor environments, providing a reliable spatial computing foundation for mixed reality applications.
[0092] In the specific implementation process, the first step is to construct a point cloud map. A high-precision 3D scan of the entire archaeological site is performed using LiDAR equipment. LiDAR determines the distance to object surfaces by emitting laser beams and measuring the time it takes for them to reflect back, thus generating detailed 3D point cloud data. During the scanning process, the scanning frequency and angle of the LiDAR are controlled to ensure that sufficiently dense point cloud data is acquired to accurately reflect the detailed features of the site. For example, for a palace foundation within the site, the LiDAR scans at a rate of millions of points per second, covering every corner of the foundation, including remaining walls, column bases, steps, etc. After the scan is completed, the raw point cloud data is obtained.
[0093] Raw point cloud data typically contains a large amount of noise and redundant information, requiring preprocessing. First, a statistical filtering method is used to remove outliers. This method calculates the average distance of each point to its neighbors based on the distance distribution between each point and its neighbors, and sets a distance threshold. If the average distance between a point and its neighbors exceeds this threshold, the point is considered an outlier and is removed. For example, for a tree in an archaeological site, due to foliage and wind disturbance, the lidar may collect some suspended points that are not part of the tree trunk; these points can be removed using statistical filtering.
[0094] Voxel filtering is used to downsample point cloud data, dividing the point cloud space into a series of small cubes (voxels). Only one point is retained within each voxel, typically chosen as the centroid of all points within the voxel. The density of the point cloud data can be controlled by adjusting the voxel size. For example, for a flat area of ground at an archaeological site, a larger voxel size can be used to reduce the amount of point cloud data while preserving the basic shape of the ground.
[0095] After preprocessing the point cloud data, the point cloud map needs to be optimized. An octree-based data structure is used to organize the point cloud map. An octree is a recursive spatial partitioning structure that continuously divides a 3D space into eight subspaces until the amount of point cloud data in each subspace is less than a set threshold. The octree structure allows for rapid retrieval and access to point cloud data, improving the efficiency of subsequent localization algorithms. For example, during localization, the corresponding leaf node in the octree can be quickly located based on the current position, and only the point cloud data within that leaf node can be processed, thus reducing computational load.
[0096] After the point cloud map is constructed, the UWB positioning system is used. Specifically, UWB positioning consists of multiple base stations and mobile tags. Base stations are fixedly installed around the site area, forming a positioning network covering the entire area. Mobile tags are integrated into the MR headsets worn by visitors. UWB base stations calculate the distance between the mobile tag and the base station by emitting ultra-wideband pulse signals and measuring the time difference of arrival (TDOA) or time of flight (TOF) of the signal. Figure 3 A schematic diagram illustrating the working principle of UWB positioning technology.
[0097] Figure 3The left side features a power supply network centered around the first and second power supply modules. A 5V power supply provides basic power to each unit. This dual-power design ensures system redundancy and stability while achieving balanced power distribution through parallel power supply paths indicated by arrows. The UWB signal generation and receiving unit in the center serves as the core signal source, generating nanosecond-level ultra-wideband signals via an internal pulse generator. The generated raw signal is transmitted to the UWB signal adjustment module to the right via arrows. The signal adjustment module on the right constitutes a complete RF processing chain: the differential coupling unit first converts the single-ended signal into a more interference-resistant differential signal; the first microwave switching unit switches the signal transmission path according to the operating mode, achieving dynamic isolation between the transmitting and receiving channels; the RF power amplification unit boosts the gain of the modulated signal to ensure sufficient spatial coverage; the second microwave switching unit further optimizes channel isolation after signal amplification, and finally, the transmitting antenna module radiates the optimized UWB pulse signal into space. The ranging processor unit on the left interacts bidirectionally with the signal generation unit via arrows, real-time analyzing the time-of-flight data returned by the receiver, and communicating ranging data with external devices via serial port and USB modules. The entire architecture clearly embodies the dual flow logic of signal and energy—vertical power supply and horizontal signal processing form a three-dimensional synergy. The signal is gradually enhanced in terms of anti-interference and radiation efficiency through multi-level calibration from the generation end, and finally achieves accurate ranging and positioning functions through spatial propagation.
[0098] Multilateral localization (MLU) algorithms can be used to calculate the position of a mobile tag in three-dimensional space using these distance differences. The basic principle of MLU is to determine the mobile tag's position by knowing the locations of multiple base stations and the distance differences between the mobile tag and each base station. This is a problem involving solving a system of nonlinear equations, typically solved using iterative algorithms (such as Newton's method or gradient descent).
[0099] To improve the accuracy and stability of UWB positioning, precise calibration of the base stations is required. The calibration process includes measuring the precise location and time synchronization of the base stations. The location of the base stations can be measured using high-precision GPS RTK equipment. Time synchronization can be achieved via wired or wireless methods. For example, an architecture with one master base station and multiple slave base stations can be used. The master base station obtains precise time information through a GPS receiver and synchronizes this time information to each slave base station via wired or wireless means.
[0100] In addition to UWB positioning, this invention also incorporates GPS RTK technology. GPS RTK technology utilizes real-time differential processing of carrier phase observations between the base station and the rover, achieving positioning accuracy at the centimeter or even millimeter level. In open outdoor environments, GPS RTK can provide a global position reference, as detailed in [reference needed]. Figure 4 A schematic diagram illustrating the working principle of hybrid positioning technology.
[0101] In practical implementation, the dynamic reference system and the rover system achieve a collaborative process of high-precision positioning through multi-source data fusion. In the diagram, the GNSS satellite at the top simultaneously transmits signals to both systems: the dynamic reference system on the left acquires satellite positioning data through its GNSS receiver, while its inertial measurement unit continuously collects the vehicle's motion status. The two systems are deeply fused using a GNSS / INS tight-combination algorithm—satellite signals correct accumulated errors in inertial navigation, and inertial data compensates for blind spots caused by brief interruptions in satellite signals, ultimately outputting a high-precision reference position. This correction information is transmitted in real-time to the rover system on the right via a radio and data link. While receiving raw satellite observation data, the rover system acquires differential correction information from the dynamic reference system via a radio. It first uses RTK (Real-Time Kinematic) technology combined with DGNSS (Differential Global Navigation Satellite System) algorithms to eliminate common errors such as ionospheric delay and orbital errors. Then, it tightly combines its own inertial measurement unit's real-time motion data with GNSS signals. Finally, a synchronous relative position synthesis module spatiotemporally aligns the differential correction results, inertial navigation calculation results, and dynamic reference coordinates, outputting a continuous positioning result with centimeter-level accuracy. The diagram clearly shows the bidirectional data flow characteristics: the dynamic reference system provides differential reference to the rover station, which in turn achieves dynamic positioning by fusing its own multi-source sensing data. The two form a closed-loop correction network through a radio station. In situations where satellite signals are blocked or in dynamic scenarios, the data stream from the inertial measurement unit continuously fills the positioning gap, ensuring the stability and reliability of the overall system output.
[0102] For example, a GPS RTK base station was set up at the highest point in the center of the archaeological site park. The base station receives signals from both the BeiDou and GPS systems and provides centimeter-level positioning references for the mobile station through carrier phase differential technology. Considering the impact of ancient buildings on satellite signals, the GPS weight was reduced in areas with severe obstruction, and the location was mainly calculated using UWB and visual positioning.
[0103] By employing a GNSS / INS tight combination algorithm, satellite positioning data and inertial navigation data are deeply fused to construct a user mobility situation prediction model, achieving continuous and stable position output. The synchronous relative position synthesis module ensures time synchronization and spatial consistency between the moving reference system and the rover system.
[0104] To further improve the accuracy and robustness of positioning, this invention also introduces an IMU module. The IMU module consists of an accelerometer, a gyroscope, and a magnetometer. The accelerometer measures the object's acceleration, the gyroscope measures the angular velocity, and the magnetometer measures the magnetic field strength. By integrating the data from these sensors, the object's attitude (such as pitch angle, roll angle, and yaw angle) and relative displacement information can be obtained.
[0105] IMU modules typically output high frequencies (hundreds or even thousands of hertz), providing high-frequency attitude and motion information to compensate for the limitations of GPS RTK and UWB in data update frequency. However, the integration operations of the IMU generate accumulated errors, causing positioning accuracy to drift over time. Therefore, it is necessary to utilize UWB and GPS RTK position information to correct for IMU drift.
[0106] This invention employs a Kalman filter algorithm to fuse data from UWB, GPS RTK, and IMU. The Kalman filter is a recursive estimation algorithm capable of estimating the positioning state from a series of noisy measurement data.
[0107] The basic idea of the Kalman filter algorithm is to predict the current state based on the state equation and the observation equation, and then correct the predicted state based on the observations to obtain the optimal estimate. The Kalman filter algorithm can fuse data from UWB, GPS RTK, and IMU to obtain high-precision and highly stable positioning results. After the positioning module is built, it mainly includes a positioning algorithm module, a data processing module, an MR content rendering module, and a user interaction module. The positioning algorithm module is responsible for receiving data from UWB, GPS RTK, and IMU, running the Kalman filter algorithm, and outputting high-precision positioning results.
[0108] Figure 5The schematic diagram of hybrid positioning technology data processing and fusion clearly illustrates the complete processing chain of multi-source sensor data fusion and error correction. Its core modules form a progressive processing structure according to the data flow direction. The four major data acquisition modules on the left constitute the system input layer: the SINS data acquisition unit outputs the three-dimensional attitude, position, and velocity information of the carrier in real time through accelerometers and gyroscopes, but its inertial navigation characteristics will accumulate errors over time; the CAN / OBD bus module extracts velocity and position data from the target body, providing kinematic constraints for dead reckoning; the UWB short-range positioning module provides centimeter-level relative position reference through ultra-wideband pulse ranging; and the BDS+GPS dual-mode satellite module outputs absolute geographic coordinates, but is susceptible to signal blockage. The raw data from the four modules converges to the particle filter processing layer in the middle via arrows—each positioning source is equipped with an independent sub-filter, which eliminates the specific errors of each sensor through probability distribution simulation, and calculates the credibility weight of each data source. After initial correction, the position, velocity, and confidence parameters flow to the main filter on the right via arrows. A Kalman filter framework is used for spatiotemporal alignment and multi-source fusion: the absolute coordinates of satellite positioning are mutually calibrated with the relative position of UWB data, and the dynamic trajectory of inertial navigation is constrained by target dead reckoning data to form a continuous and stable positioning output. The Sage-Husa adaptive filter at the bottom constitutes a closed-loop correction mechanism, analyzing the statistical characteristics of the system output error in real time, dynamically adjusting the parameter compensation values of the dead reckoning model, and feeding the corrections back to the forward processing nodes. The data flow between modules reveals the reverse flow logic of error compensation data. For example, when satellite signals are lost, the system automatically increases the fusion weight of inertial navigation and UWB data, while the target body steering angle data corrects the attitude reckoning error of the inertial navigation. Finally, through the cascading action of multiple filters, adaptive positioning with centimeter- to meter-level accuracy is achieved.
[0109] The user interaction module is responsible for processing user input (such as gestures, voice, etc.) and providing corresponding feedback.
[0110] For example, for complex terrain within an archaeological site, the localization algorithm needs to be adjusted to adapt to changes in the terrain. For obstructions within the site, the MR content rendering module needs to be optimized to prevent virtual content from penetrating the obstructions.
[0111] In addition to the positioning module, this invention also employs a multi-dimensional virtual-real fusion model and a multi-modal driven model to provide tourists with an intelligent interactive experience. A multi-modal large model refers to a model capable of processing multiple modalities of data (such as text, images, and voice). The multi-modal large model used in this invention is based on the Transformer architecture and has undergone targeted optimization.
[0112] The training process of a multimodal large model includes three stages: data preprocessing, model training, and model evaluation.
[0113] The data preprocessing stage requires collecting a large amount of text, image, and audio data, and then cleaning, labeling, and formatting the data. For example, for text data, operations such as word segmentation, stop word removal, and stemming are needed. For image data, operations such as cropping, scaling, and normalization are required. For audio data, operations such as feature extraction and noise reduction are needed.
[0114] During the model training phase, a combination of pre-training and fine-tuning is employed. First, the model is pre-trained using large-scale unlabeled data, enabling it to learn general language and visual knowledge. Then, the model is fine-tuned using project-specific labeled data, adapting it to specific tasks and scenarios.
[0115] During model training, a novel loss function was employed that comprehensively considers the correlation between text, images, and speech. The image loss function uses a contrastive loss function, while the speech loss function uses a CTC (Connectionist Temporal Classification) loss function.
[0116] During the model evaluation phase, various metrics are used to assess the model's performance. For example, for text generation tasks, metrics such as BLEU, ROUGE, and METEOR can be used. For image generation tasks, metrics such as Inception Score and FID can be used. For speech recognition tasks, metrics such as WER (Word Error Rate) can be used.
[0117] Multimodal big data models enable a variety of intelligent interactive functions. For example, visitors can ask questions via voice to obtain information about the site. The multimodal big data model can understand the visitor's intent based on the voice input, retrieve relevant information from the knowledge base, and answer the visitor's questions in the form of voice or text.
[0118] This embodiment implements a user behavior prediction algorithm based on machine learning, which combines Long Short-Term Memory (LSTM) networks and Hidden Markov Models (HMMs). The trajectory feature extraction module uses a sliding time window to analyze historical location data and extract feature parameters such as speed change patterns, turning angle distribution, and stop point clustering.
[0119] The motion state recognition module identifies six basic motion states by analyzing the patterns of acceleration and angular velocity changes: stationary, slow walking, fast movement, standing and watching, directional exploration, and taking photos. Each state has its specific motion characteristics and duration range.
[0120] The path prediction algorithm integrates historical trajectory patterns and current movement status to predict possible paths within a future period. A graph-based path network model was established, dividing the archaeological park into multiple regions of interest. By analyzing a large amount of visitor behavior data, a transition probability matrix between regions was established. When the prediction confidence level reaches a high level, the virtual content of the target area is pre-rendered.
[0121] The generated preloading strategy includes: 3D model geometry preloading, preloading the building model and character model of the target area; texture resource pre-caching, pre-rendering texture maps according to the current lighting conditions; audio material preparation, preloading background music and environmental sound effects; and interactive script initialization, preparing virtual character dialogue and interaction logic.
[0122] The environmental perception layer collects environmental parameters in real time through multi-sensor fusion. Light intensity is measured by an ambient light sensor, covering the entire range from low indoor light to strong outdoor light; weather conditions are determined by temperature and humidity sensors and barometric pressure sensors; terrain elevation is obtained through lidar and binocular cameras to acquire depth information; and object surface material is identified by analyzing spectral reflectance characteristics to identify different materials such as stone, wood, and metal.
[0123] The digital asset matching layer has established a digital asset library containing thousands of high-precision 3D models, covering multiple categories such as ancient architectural components, human figures, and artifacts. It automatically matches the corresponding historical scene based on the current location coordinates; for example, at the location of a palace ruin, it automatically loads the corresponding palace reconstruction model; at the location of a garden ruin, it loads the corresponding garden landscape model.
[0124] The real-time adjustment of lighting and shadow effects is achieved using a physically based rendering engine. It calculates the sun's position and angle in real time, dynamically adjusting the direction and intensity of shadows cast by virtual objects. Material reflection effects adaptively adjust based on ambient light intensity, ensuring that the brightness and darkness of virtual objects remain consistent with the real environment.
[0125] The spatial registration layer uses a 3D map constructed with SLAM technology as a spatial reference, and employs coordinate transformation algorithms to accurately locate virtual content in real space. An absolute positioning mechanism based on marker points is established, and QR code markers are placed at key locations to correct for accumulated errors.
[0126] The dynamic synchronization layer enables real-time synchronization between virtual content and environmental changes. For example, when the weather changes from sunny to cloudy, the lighting intensity of the virtual scene is automatically adjusted, and the shadow contrast is reduced; when rain is detected, rain effects and wet reflections are added to the surface of virtual buildings.
[0127] The user behavior analysis layer uses deep learning algorithms to analyze tourists' behavioral preferences. It records data such as the duration of tourists' stays in different areas, their movement trajectories, and the frequency of their interactions to build individual interest models. For example, if it detects that a tourist spends a long time in front of architectural ruins and frequently observes them up close, it can be determined that they have a strong interest in architectural art.
[0128] The content matching layer constructs a semantic network containing massive amounts of cultural knowledge entries based on knowledge graph technology. The knowledge graph covers multiple dimensions, including historical figures, architectural techniques, cultural events, and the uses of artifacts, storing complex relationships between entities through a graph database. Relevant content is intelligently retrieved from the knowledge graph based on the visitor's current location and interests.
[0129] The multimodal fusion layer integrates natural language processing, computer vision, and speech synthesis technologies. The text generation module is based on a large language model architecture and can generate vivid narrative text based on historical context; the image generation module uses advanced image synthesis technology to generate visual elements that conform to historical context based on text descriptions; the speech synthesis module supports multiple dialects and ancient speech styles to enhance cultural immersion.
[0130] The personalized generation layer dynamically adjusts the content generation strategy based on user profiles: for architecture enthusiasts, it focuses on introducing the architectural structure and craftsmanship features, generating interactive displays of architectural details; for history enthusiasts, it emphasizes telling historical events and stories of people, creating time-traveling narrative scenes; for art enthusiasts, it highlights decorative art and cultural relic value, providing interactive experiences for art appreciation; and for families with children, it designs fun game tasks to spread cultural knowledge in an entertaining and educational way.
[0131] Safety has also been fully considered in this invention. The positioning module monitors the user's location in real time and provides timely alerts when the user approaches the protected area or dangerous area of the site, ensuring the safety of tourists and the site. Based on point cloud maps, it provides users with the best travel route, avoiding densely populated or construction areas, thus improving the visitor experience and safety.
[0132] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.
Claims
1. A method for optimizing hybrid positioning and dynamic rendering in outdoor immersive cultural tourism, characterized in that, include: The positioning signal of the mobile terminal in the target service area is obtained through a hybrid positioning model; the positioning signal includes ultra-wideband ranging signal, visual inertial odometry signal, real-time carrier phase differential signal and inertial navigation signal; the positioning signal is processed through a multi-layer signal processing filtering model to obtain coordinate data, and the coordinate data is calculated by weighted least squares method to generate spatial coordinates; A user mobility situation prediction model is constructed based on spatial coordinates to analyze the user's movement trajectory and generate navigation prediction trajectory and navigation confidence. The virtual objects are pre-rendered based on the navigation prediction trajectory to generate a preloading strategy; The user mobility situation prediction model includes: The trajectory feature extraction module analyzes historical spatial coordinate sequences based on a sliding time window to extract the trajectory features of user movement, including speed change patterns, direction change frequency, and stop point distribution features. The motion state recognition module identifies the user's motion state by combining the current motion state parameters, including four states: walking, standing and watching, moving quickly, and exploring directions. The path prediction algorithm integrates trajectory features and motion state to predict the user's movement path and generate a navigation prediction trajectory that includes location coordinates. The navigation confidence is calculated by measuring the degree of deviation between the navigation prediction trajectory and the actual trajectory. When the navigation confidence is not lower than a preset threshold, the virtual object is pre-rendered based on the navigation prediction trajectory to generate a preloading strategy. The geographic environment parameters of the target service area are obtained, and the spatial coordinates, geographic environment parameters, preloading strategies and virtual objects are integrated through a multi-dimensional virtual-real fusion model to generate an augmented reality spatial scene. Interactive content is generated by analyzing user movement trajectory data, spatial coordinates, and augmented reality spatial scenes using a multimodal driven model.
2. The outdoor immersive cultural tourism hybrid positioning and dynamic rendering optimization method according to claim 1, characterized in that: The hybrid positioning model includes: an ultra-wideband positioning layer, which determines the distance relationship between the mobile terminal and each base station based on a multilateral positioning algorithm and generates an ultra-wideband ranging signal; The visual inertial odometry layer collects visual feature points of the environment, combines them with motion data from the inertial measurement unit to obtain position trajectory information in the local coordinate system, and outputs visual inertial odometry signals. The global navigation and positioning layer obtains positioning signals, and obtains a global position reference through differential processing of carrier phase observations between the base station and the rover station, generating a real-time carrier phase differential signal; The inertial navigation layer measures the linear acceleration, angular velocity, and magnetic field strength of the mobile terminal, provides high-frequency attitude and motion state data, and generates inertial navigation signals.
3. The outdoor immersive cultural tourism hybrid positioning and dynamic rendering optimization method according to claim 1, characterized in that: The multi-layer signal processing filtering model includes: a first filtering layer, which uses the Kalman filtering algorithm to preprocess the ultra-wideband ranging signal, real-time carrier phase differential signal and inertial navigation signal, removes signal outliers through state prediction and observation update steps, and outputs preliminary position estimation data; The second filtering layer uses a complementary filtering algorithm to fuse the local map data constructed from the visual inertial odometry signal with the preliminary position estimation data, and outputs multi-source fused position data. The weight allocation mechanism calculates the confidence weight coefficient of each positioning source in real time based on signal strength, environmental occlusion, and sensor confidence, and outputs dynamic weight parameters. The coordinate calculation module uses dynamic weight parameters to perform weighted least squares calculation on multi-source fused location data to generate three-dimensional spatial coordinate data.
4. The outdoor immersive cultural tourism hybrid positioning and dynamic rendering optimization method according to claim 1, characterized in that: The multidimensional virtual-real fusion model includes: an environmental perception layer, which collects information on light intensity, weather conditions, terrain elevation, and object surface material in the target service area in real time to obtain geographic environmental parameters; The digital asset matching layer selects virtual buildings, characters, and scene elements corresponding to the current spatial coordinates based on a 3D digital model library, and adjusts the lighting, material reflection, and physical properties of virtual objects according to geographical environment parameters. The spatial registration layer uses coordinate data to locate virtual objects that have been adapted to the environment to their corresponding positions in the real scene, and uses coordinate transformation algorithms to align virtual content with the physical environment. The dynamic synchronization layer monitors changes in geographic environment parameters and preloading strategies in real time, updates the visual appearance of virtual objects, and generates augmented reality spatial scenes.
5. The outdoor immersive cultural tourism hybrid positioning and dynamic rendering optimization method according to claim 1, characterized in that: The multimodal driving model includes: a user behavior analysis layer, which performs trajectory segmentation and pattern recognition on user movement trajectory data, and extracts user behavior feature parameters at different spatial locations, including dwell time, movement speed and interaction frequency; The content matching layer retrieves historical and cultural information, architectural backgrounds, and personal stories related to the current location from a pre-built knowledge graph based on the user's current coordinate data and augmented reality spatial scene information. The multimodal fusion layer comprehensively analyzes text descriptions, image resources, and audio materials to generate dialogue text that conforms to the current scene context and creates visual elements that match the historical background. The personalized generation layer dynamically adjusts the information weights retrieved by the content matching layer based on personal preference features extracted from behavioral feature parameters, generating customized interactive content, including virtual character dialogues, historical scene reproduction, and interactive task design.
6. The outdoor immersive cultural tourism hybrid positioning and dynamic rendering optimization method according to claim 2, characterized in that: Multiple ultra-wideband base stations are deployed in the target service area. The base stations synchronize their time wirelessly. The mobile terminal calculates its three-dimensional spatial coordinates by measuring the time difference of arrival of signals from different base stations and using the hyperbolic positioning principle.