Historical space-time perspective system for context awareness and multi-dimensional context awareness calculation method
By using an embodied intelligent agent system with a cloud-edge-device architecture, combined with multi-dimensional context-aware computing and spatiotemporal anchor-driven dynamic registration, the problem of spatiotemporal misalignment and cognitive gap in the display of cultural heritage by AR technology has been solved, achieving high-precision virtual-real fusion and personalized interaction, thus improving the user experience.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- 李忠孝
- Filing Date
- 2025-12-26
- Publication Date
- 2026-04-21
AI Technical Summary
Existing AR technology suffers from problems such as spatiotemporal misalignment, cognitive barriers, and static interaction in cultural heritage displays, making it impossible to achieve high-precision virtual-real fusion, personalized interaction, and a deeply immersive experience.
The embodied intelligent agent system, which adopts a cloud-edge-device architecture, achieves high-precision positioning and personalized content presentation through multi-dimensional context-aware computing and spatiotemporal anchor-driven dynamic registration, combined with multimodal perception data.
It achieves a unified understanding of high-precision spatiotemporal perception and user intent, providing an immersive and personalized historical spatiotemporal interactive experience, and enhancing the depth of understanding and interactive quality of cultural heritage displays.
Smart Images

Figure CN121900618A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of artificial intelligence, augmented reality and digital cultural heritage, and in particular to a context-aware historical spatiotemporal perspective system and a multi-dimensional context-aware computing method. Background Technology
[0002] Currently, augmented reality (AR) technology has been widely used in the digital display of cultural heritage. Existing technological solutions can be mainly divided into the following three categories:
[0003] (1) Static AR systems based on markers or image recognition
[0004] Static AR systems based on markers or image recognition rely on pre-set physical markers or 2D images to trigger fixed, pre-rendered 2D / 3D content, such as animations and models. Their core drawback lies in:
[0005] a) The content is completely disconnected from the spatiotemporal context of the real scene, and cannot reflect the historical accumulation and changes;
[0006] b) The interaction method is rigid, and users passively watch without personalized guidance.
[0007] (2) A simple AR navigation system based on location services
[0008] A simple location-based AR navigation system triggers media content at specific geographic coordinates via GNSS or Bluetooth beacons. Its core drawback is:
[0009] a) Low positioning accuracy (meter level), unable to achieve centimeter-level virtual-real fusion;
[0010] b) The content presentation consists of discrete information silos, making it impossible to construct a continuous historical narrative flow;
[0011] c) Unable to perceive changes in user status and environment and make dynamic adjustments accordingly.
[0012] (3) General AR platform based on SLAM
[0013] While general-purpose AR platforms based on SLAM can achieve basic spatial perception and virtual content placement, their core shortcomings lie in:
[0014] a) Lacking knowledge of historical and cultural fields, the virtual content is unrelated to historical authenticity and spatiotemporal accuracy;
[0015] b) It has no ability to perceive user intent or focus of interest, and only engages in one-way interaction.
[0016] Existing basic augmented display technologies suffer from a disconnect between perception, cognition, and presentation. The system cannot understand the true appearance and events of the current physical location at a specific historical moment, nor can it understand what the user is currently paying attention to or what information they need. This results in a superficial user experience and an inability to achieve deep cultural cognition.
[0017] Therefore, there is an urgent need in this field for an innovative technical solution that integrates high-precision spatiotemporal perception, contextual intelligent understanding, and personalized interaction to solve the core problems of static and fragmented presentation of historical information, rough integration of virtual and real elements, and shallow and passive user interaction in existing AR cultural tourism applications. This solution should enhance the system's ability to understand and dynamically respond to the physical environment, historical context, and user cognitive state in a unified manner, and support new smart tour guide and educational applications that are deeply immersive and integrate knowledge and action for cultural heritage. Summary of the Invention
[0018] Therefore, the purpose of this invention is to provide a context-aware historical spatiotemporal perspective system and a multi-dimensional context-aware computing method. It adopts an innovative embodied intelligent agent architecture driven by spatiotemporal anchors, multi-dimensional context-aware computing, and cloud-edge-device collaboration to achieve deep integration of the physical world, historical information, and user intent, and enhance the system's ability to provide personalized and immersive interactive guidance. It solves the three core problems of spatiotemporal misalignment, cognitive barriers, and static interaction in existing AR cultural tourism applications.
[0019] To achieve the above objectives, the present invention provides a context-aware historical spatiotemporal perspective system, comprising: a cloud-based intelligent central layer, an edge computing node layer, and a user embodied terminal layer deployed using a cloud-edge-device architecture;
[0020] The user embodied terminal layer is used to collect user multimodal perception data and generate user context summaries in real time, as well as receive and present historical content that blends the virtual and real worlds.
[0021] The edge computing node layer connects the user embodied terminal layer and the cloud intelligent hub layer. It is used to receive the user context summary, perform high-precision positioning and spatiotemporal anchor point matching, and forward the matching result and the user context summary to the cloud intelligent hub layer. It also receives cloud instructions, calculates virtual and real registration parameters, and sends them to the user embodied terminal layer.
[0022] The cloud-based intelligent central layer is used to retrieve historical knowledge and perform multi-dimensional context-aware calculations based on the received matching results and context summaries, generate personalized historical narrative instructions and related content, and send them down to the edge computing node layer.
[0023] Furthermore, the user-embodied terminal layer includes:
[0024] A multimodal sensing array is used to acquire multimodal sensing data, including environmental images, pose data, eye movements, physiological data, and speech data.
[0025] A lightweight real-time context understanding module is used to integrate eye-tracking and pose data to calculate the user's three-dimensional interest focus in the real world, and to calculate the real-time interest intensity and user cognitive load on the interest focus based on a preset dynamic attention model, generating a structured user context summary.
[0026] An immersive rendering and interaction engine is used to receive virtual and real registration parameters and virtual content from the edge computing node layer, perform augmented reality rendering, and capture user interaction commands.
[0027] A role-based application interface is used to provide user role selection and task management functions, and to inject the selected role information into the user context summary.
[0028] Furthermore, the edge computing node layer includes:
[0029] High-precision positioning and dynamic mapping services are used to fuse multi-source positioning signals to provide users with centimeter-level real-time pose and dynamically update it in response to environmental changes.
[0030] The spatiotemporal anchor point matching and registration engine is used to query and match the optimal spatiotemporal anchor point in the locally cached historical spatiotemporal knowledge subgraph based on the user's real-time pose and the target's historical time, and calculate the rendering and registration parameters required for virtual-real fusion based on the spatiotemporal anchor point matching results and real-time environmental data.
[0031] The local context aggregation and forwarding module receives user context summaries from one or more terminals, performs regional-level situation aggregation, and packages and uploads the aggregated context summaries and the spatiotemporal anchor point matching results to the cloud intelligent hub layer; and distributes the narrative instructions and content issued by the cloud intelligent hub layer to the corresponding terminals.
[0032] High-bandwidth content edge caching is used to pre-store or cache high-precision historical 3D models and media assets from the cloud, which can be called at high speed by the spatiotemporal anchor matching registration engine and the terminal rendering engine.
[0033] Furthermore, the cloud-based intelligent hub layer includes:
[0034] A multimodal hierarchical spatiotemporal knowledge base is used to organize and store historical data in a multidimensional structure based on objects, events, spatiotemporal, and media, with spatiotemporal anchor points as the core index.
[0035] The historical spatiotemporal knowledge enhancement model is connected to the knowledge base and is used to receive service requests containing user context and spatiotemporal anchors, retrieve related knowledge from the knowledge base, and perform deep spatiotemporal reasoning and character-based content generation.
[0036] The narrative generation and global scheduling engine is connected to the historical spatiotemporal knowledge enhancement model. Based on the reasoning results of the historical spatiotemporal knowledge enhancement model and the user context, it generates personalized narrative clues, interactive tasks and scheduling instructions, and coordinates the presentation of virtual content among multiple users.
[0037] Preferably, the spatiotemporal anchor point STA i Defined as a quintuple of data as shown in the following formula:
[0038]
[0039] In the formula, G i The geographic coordinate bounding box is defined using latitude, longitude, and elevation coordinates; T i For the effective time interval [t] start ,t end ];C i For associated historical 3D scene content, including multi-LOD models and materials; M i Metadata, derived from historical events and document citations; L i For spatiotemporal association links with other anchor points;
[0040] All spacetime anchors STA i A spatiotemporal graph network is constructed in the multimodal hierarchical spatiotemporal knowledge base.
[0041] Preferably, when receiving the user context summary and performing high-precision positioning and spatiotemporal anchor point matching, the process includes:
[0042] Receive real-time pose data and target historical time from the user terminal;
[0043] Based on the pose data and the target historical time, query the historical spatiotemporal knowledge base to find matching candidate spatiotemporal anchor points;
[0044] Calculate the matching score between each candidate anchor point and the current pose and target time, and select the optimal spatiotemporal anchor point as the registration benchmark;
[0045] Based on the real-time spatial relationship between the user and the optimal spatiotemporal anchor point, historical scene content at different levels of detail is dynamically scheduled.
[0046] Real-time ambient lighting information is collected and used to drive the historical scene content to perform consistent lighting rendering, while simultaneously calculating the virtual and real occlusion relationship.
[0047] Furthermore, the matching score S based on spatial distance and temporal proximity is calculated according to the following formula. match Select the anchor point STA with the highest score. k As the current registration benchmark:
[0048]
[0049] In the formula, d represents the pose. u .location to STA k Geometric distance from the center; T center For STA k The time center point; α is the spatial weighting factor.
[0050] This invention also provides a multi-dimensional context-aware computing method, executed in the cloud-based intelligent central layer of the aforementioned context-aware historical spatiotemporal perspective system, comprising the following steps:
[0051] The input multimodal perception data from the terminal within the received time slice Δt includes eye-tracking data, pose data, audio data, and environmental data.
[0052] By integrating eye-tracking data and pose data, the user's three-dimensional gaze point in the real-world coordinate system is calculated through ray casting.
[0053] The three-dimensional gaze point is associated with objects in the historical spatiotemporal knowledge base and identified as the current object of interest.
[0054] Based on the continuous fixation duration and pupil diameter changes of the object of interest, the real-time interest intensity S is calculated using a dynamic attention model. interest (t):
[0055] S interest (t)=β*S interest (t-1)+(1-β)*f(T fixation D pupil (6)
[0056] In the formula, T fixation For historical object O j Continuous fixation duration; D pupil β is the pupil diameter change; β is the attenuation factor used to simulate the natural decay of interest; f(·) is the normalization function that maps the physiological signal to an intensity increment.
[0057] Based on the user's voice features, gaze point distribution entropy, and head micro-motion variance, estimate the user's real-time cognitive load;
[0058] The output consists of a contextual summary vector composed of the object of interest, real-time interest intensity, and real-time cognitive load, which is used to drive personalized narrative decisions.
[0059] Preferably, the user's real-time cognitive load is estimated using the following formula:
[0060] L cog (t)=γ1*Speechrate (t)+γ2*Gaze entropy (t)+γ3*Motion var (Pose t (7)
[0061] In the formula, Speech rate (t) represents the number of voice questions asked per unit time; Gaze entropy (t) represents the degree of disorder in the gaze point switching between multiple objects; Motion var (Pose t γ1, γ2, and γ3 are the variances of minute head tremors, used to indicate fatigue or discomfort; γ1, γ2, and γ3 are the weighting coefficients of each factor.
[0062] The present invention also provides an augmented reality cultural heritage tour device, based on the aforementioned context-aware historical spatiotemporal perspective system; matched to the user end, it adopts a head-mounted augmented reality display device to carry the functional modules of the user embodied terminal layer, providing the user with an immersive historical spatiotemporal perspective experience.
[0063] The context-aware historical spatiotemporal perspective system and multi-dimensional context-aware computation method disclosed in this application have at least the following advantages:
[0064] 1) System architecture innovation based on embodied intelligent agents: A cloud-edge-device collaborative embodied intelligent agent architecture for cultural heritage perspective is proposed. The intelligent agent is not a chatbot that exists in virtual space, but is embodied in real physical cultural heritage scenarios. It provides services to users through continuous interaction with the real world, and reasonably allocates historical reasoning, real-time computing and lightweight perception tasks, and realizes a complex historical spatiotemporal interactive experience under limited terminal computing power.
[0065] 2) Innovation of dedicated historical and cultural knowledge base: The multimodal hierarchical spatiotemporal knowledge base (MHST-KB) with spatiotemporal anchor points (STA) as the core is proposed. It organizes discrete historical data in a structured and related way in the spatiotemporal dimension, laying the foundation for high-precision spatiotemporal query and content scheduling.
[0066] 3) Innovative Spatiotemporal Intelligent Registration Method: A dynamic 3D registration method driven by spatiotemporal anchors is proposed, which upgrades the traditional registration based on visual features to semantic registration based on the dual dimensions of "geographic location + historical time", realizing accurate and adaptive fusion of virtual content across time and space.
[0067] 4) Innovation of multi-dimensional context-aware computing model: Propose a dynamic attention-guided multi-dimensional context-aware computing model, which for the first time integrates multi-modal signals such as user eye movement, voice, and pose, and quantifies the intensity of user interest in specific historical objects and overall cognitive load in real time, realizing machine understanding of user cognitive context. Attached Figure Description
[0068] Figure 1 This is a schematic diagram of the framework of the context-aware historical spatiotemporal perspective system of the present invention.
[0069] Figure 2 Workflow for a context-aware historical spatiotemporal perspective system.
[0070] Figure 3 Flowchart of the spatiotemporal anchor point matching method.
[0071] Figure 4 This is the multi-dimensional context-aware computing method used in this invention. Detailed Implementation
[0072] The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0073] like Figure 1 As shown, one embodiment of the present invention provides a context-aware historical spatiotemporal perspective system, comprising: a cloud-based intelligent central layer, an edge computing node layer, and a user embodied terminal layer deployed using a cloud-edge-device architecture;
[0074] The user embodied terminal layer is used to collect user multimodal perception data and generate user context summaries in real time, as well as receive and present historical content that blends the virtual and real worlds.
[0075] The user-embodied terminal layer includes:
[0076] A multimodal sensing array is used to acquire multimodal sensing data, including environmental images, pose data, eye movements, physiological data, and speech data. These include visual perception: using an RGBD camera for environmental scanning, object recognition, and gesture capture; environmental perception: using fiber optics and temperature and humidity sensors to sense changes in the physical environment; inertial and pose perception: using an inertial navigation sensor (IMU) to provide high-speed self-motion data; and physiological and intention perception: using an eye tracker to capture gaze points, a directional microphone to capture voice commands, and a heart rate sensor to monitor excitement / fatigue states.
[0077] A lightweight real-time context understanding module is used to fuse eye-tracking and pose data to calculate the user's 3D focus of interest in the real world. Based on a preset dynamic attention model, it calculates the real-time interest intensity and user cognitive load on the focus of interest, generating a structured user context summary. This application employs an extremely lightweight object detection model (such as a pruned and quantized MobileNet SSD) to perform real-time analysis of RGB images, identifying rough categories such as "rock wall" and "ancient tree." Combining the detection bounding box with the gaze point improves the accuracy of the association.
[0078] The immersive rendering and interaction engine receives virtual and real registration parameters and virtual content from the edge computing node layer, performs augmented reality rendering, and captures user interaction commands, including natural human-computer interaction: supporting the recognition and response of various interaction methods such as voice dialogue, gesture interaction, and gaze selection.
[0079] The role-based application interface provides users with role selection and task management functions, and injects the selected role information into the user context summary. The role-based application interface acts as the user's digital travel companion, providing functions such as role selection, task logs, and management of digital collectibles (such as collected travelogue fragments), and serves as the software carrier for the entire experience.
[0080] The edge computing node layer connects the user embodied terminal layer and the cloud intelligent hub layer. It is used to receive the user context summary, perform high-precision positioning and spatiotemporal anchor point matching, and forward the matching result and the user context summary to the cloud intelligent hub layer. It also receives cloud instructions, calculates virtual and real registration parameters, and sends them to the user embodied terminal layer. This includes:
[0081] 1) High-precision positioning and dynamic mapping services are used to fuse multi-source positioning signals, providing users with centimeter-level real-time pose and dynamically updating in response to environmental changes. High-precision positioning / dynamic mapping services are the foundation of centimeter-level spatial perception, fusing visual SLAM, LiDAR point cloud prior maps, UWB and GNSS signals to achieve the following functions:
[0082] Seamless indoor and outdoor positioning: Provides stable and continuous centimeter-level user pose data;
[0083] Dynamic environment update: Real-time detection and update of changes such as temporary obstacles and crowd gathering areas in high-precision maps. The multi-positioning signal fusion in this application is to fuse different granularities according to the positioning accuracy of different positioning signals. For example, signals with lower positioning accuracy are used for global layout, and then high-precision signals are used for local layout. This hierarchical setting realizes the fusion and matching of positioning accuracy at different granularities. This process can be achieved using existing technologies.
[0084] 2) Spatiotemporal anchor point matching and registration engine, which is used to query and match the optimal spatiotemporal anchor point in the local cached historical spatiotemporal knowledge subgraph based on the user's real-time pose and the target's historical time, and calculate the rendering and registration parameters required for virtual-real fusion based on the spatiotemporal anchor point matching results and real-time environmental data;
[0085] The spacetime anchor point STA i Defined as a quintuple of data as shown in the following formula:
[0086] STA i ={G i ,T i Ci M i ,L i} (1)
[0087] In the formula, G i The geographic coordinate bounding box is defined using latitude, longitude, and elevation coordinates; T i For the effective time interval [t] start ,t end ];C i For associated historical 3D scene content, including multi-LOD models and materials; M i Metadata, derived from historical events and document citations; L i For spatiotemporal association links with other anchor points;
[0088] All spacetime anchors STA i A spatiotemporal graph network is constructed in the multimodal hierarchical spatiotemporal knowledge base.
[0089] Preferably, when receiving the user context summary and performing high-precision positioning and spatiotemporal anchor point matching, the process includes: receiving real-time pose data and target historical time from the user terminal;
[0090] Input the user's current 6-DOF pose (Pose) u and the target historical event T selected / inferred target The system executes a query in MHSTKB:
[0091] STA candidate =Query_MHST-KB(G≈Pose) u .location,T target ∈T i (2)
[0092] Returns a set of candidate anchor points.
[0093] (2) Optimal anchor point matching
[0094] Based on the pose data and the target historical time, query the historical spatiotemporal knowledge base to find matching candidate spatiotemporal anchor points;
[0095] Calculate the matching score between each candidate anchor point and the current pose and target time, and select the optimal spatiotemporal anchor point as the registration benchmark; calculate the matching score S based on spatial distance and temporal proximity according to the following formula. match Select the anchor point STA with the highest score. k As the current registration benchmark:
[0096]
[0097] In the formula, d represents the pose. u.location to STA k Geometric distance from the center; T center For STA k The time center point; α is the spatial weighting factor.
[0098] Based on the real-time spatial relationship between the user and the optimal spatiotemporal anchor point, the historical scene content at different levels of detail is dynamically scheduled; real-time ambient lighting information is collected, and the historical scene content is driven to perform lighting-consistent rendering, while simultaneously calculating the virtual-real occlusion relationship.
[0099] (3) Content scheduling and adaptive rendering
[0100] According to users and STA k Real-time distance and viewpoint, dynamically scheduling historical 3D scene content at different levels of detail (LOD) C i Simultaneously, environmental HDR lighting information is collected. env And thereby drive the virtual historical scene C i Global illumination shading is used to achieve consistency between virtual and real lighting. The occlusion relationship between virtual and real objects is calculated in real time using a depth buffer.
[0101] 3) Local context aggregation and forwarding module: receives user context summaries from one or more terminals, performs regional-level situation aggregation, and packages and uploads the aggregated context summaries and the spatiotemporal anchor matching results to the cloud intelligent hub layer; and distributes the narrative instructions and content issued by the cloud intelligent hub layer to the corresponding terminals.
[0102] 4) High-bandwidth content edge caching is used to pre-store or cache high-precision historical 3D models and media assets from the cloud, which can be called at high speed by the spatiotemporal anchor matching registration engine and the terminal rendering engine.
[0103] The cloud-based intelligent central layer is used to retrieve historical knowledge and perform multi-dimensional context-aware calculations based on the received matching results and context summaries, generating personalized historical narrative instructions and related content, and then distributing them to the edge computing node layer. This includes:
[0104] 1) A multimodal hierarchical spatiotemporal knowledge base, used to organize and store historical data in a multidimensional structure of objects, events, spatiotemporal, and media, with spatiotemporal anchors as the core index;
[0105] The multimodal hierarchical spatiotemporal knowledge base is a structured historical digital twin, using "spatiotemporal anchors" as the core index and organizing data according to a four-dimensional model of "object-event spatiotemporal medium".
[0106] Object layer: 3D models, attributes, and relationship maps of entities such as buildings, landforms, and cultural relics from different historical periods.
[0107] Event layer: related historical events, figures' activities, and social change chains.
[0108] Spatiotemporal Layer: Manages spatiotemporal anchor points (STAs). Each STA is a unique binding between geographic coordinates and a timestamp, such as: {Coordinates: Welcoming Pine of Huangshan, Time: Spring of 1616, Associated Content: Model of Welcoming Pine of the Ming Dynasty, Records of Xu Xiake}
[0109] Media layer: Related original materials such as documents, ancient paintings, old photographs, and audio recordings.
[0110] 2) Narrative Generation and Global Scheduling Engine
[0111] The narrative generation and global scheduling engine acts as the chief writer and director for personalized user experiences. After receiving context summaries from the edge layer, it generates dynamic narratives.
[0112] Dynamic narrative generation: Based on user roles, interests, and cognitive load, retrieve, reorganize, and generate coherent narrative clues and mission plots from MHSTKB.
[0113] Cross-user coordination and scheduling: In multi-user scenarios, coordinate the presentation of virtual content by different users to avoid conflicts, and publish collaborative tasks for groups.
[0114] Long-term user model updates: Aggregate user interaction data from multiple rounds to continuously optimize their interest profiles and preference models.
[0115] 3) A historical spatiotemporal knowledge enhancement model, connected to the knowledge base, is used to receive service requests containing user context and spatiotemporal anchors, retrieve related knowledge from the knowledge base, and perform deep spatiotemporal reasoning and character-based content generation. This historical spatiotemporal knowledge enhancement model is a domain knowledge understanding and reasoning engine. Based on a general large language model, it incorporates massive amounts of historical documents, local chronicles, and professional papers for incremental pre-training and instruction fine-tuning, enabling it to possess the following functions:
[0116] Deep spatiotemporal reasoning: Understanding and reasoning about complex spatiotemporal propositions, such as "What artificial repairs did the mountain paths undergo before and after Xu Xiake's ascent of Huangshan?"
[0117] Character-driven language generation: Capable of generating dialogue and content in the style, knowledge level, and writing style of historical figures such as Xu Xiake.
[0118] Multimodal understanding and summarization: Understanding user-uploaded images and audio, and associating them with relevant content in the knowledge base. In practical implementation, one can choose...
[0119] This application selects an open-source general-purpose large language model with 7 billion parameters (such as HistLLM or a similar architecture) as its base.
[0120] We collected a large amount of unstructured texts related to the target cultural heritage field (such as ancient Chinese architecture and historical geography), including local chronicles, academic papers, biographies of historical figures, and travelogues. After deduplication and cleaning, we formed a plain text corpus of about 50GB.
[0121] Using standard causal language modeling objectives, the base model is further pre-trained on the corpus to learn domain-relevant vocabulary, facts, and expression styles.
[0122] Approximately 10,000 instruction-response pairs were manually written. Instruction templates included phrases like, "Suppose you are Xu Xiake, introduce [a certain aspect] of [the object] before you to tourists." Responses were generated mimicking Xu Xiake's writing style and knowledge level. Simultaneously, instructions incorporating spatiotemporal reasoning were constructed, such as, "Compare the appearance differences of [location A] at [time 1] and [time 2]."
[0123] The low-rank matrix injected into the model's attention module during training significantly reduces training costs and avoids catastrophic forgetting.
[0124] The fine-tuned HistLLM can be deployed in the cloud and encapsulated as an API service that can receive structured requests (including roles, contexts, and STA_ID). This part is a relatively mature technology and will not be elaborated on here.
[0125] Taking the example of a user impersonating "Xu Xiake" touring Huangshan and gazing at a rock face, this illustrates how the system works collaboratively to complete a full closed loop of perception and response in real time. Figure 2 As shown, the system workflow is as follows:
[0126] (1) User terminal collects multimodal raw data
[0127] a) User action: The user (playing the role of Xu Xiake) wears AR glasses, stops near Qingluan Bridge in Huangshan, and stares at the vermilion rock wall for a long time;
[0128] b) System Action: The multimodal sensing array at the user terminal layer is activated, and synchronous data acquisition begins.
[0129] Visual data: RGBD cameras capture color and depth images of the scene in front.
[0130] Pose data: The IMU (Inertial Measurement Unit) provides real-time rotation and acceleration data of the device.
[0131] Physiological and Intent Data: Eye trackers accurately record the coordinates of the pupil center to calculate the gaze point in the screen coordinate system; microphones monitor ambient audio.
[0132] Environmental data: The light sensor detects the current ambient light intensity and color temperature.
[0133] (2) Lightweight real-time context understanding on the terminal
[0134] The system initiates a process where the terminal's built-in lightweight real-time context understanding module starts up and performs rapid local processing on the raw data:
[0135] 3D gaze point calculation: By fusing IMU pose data and eye-tracking gaze point coordinates, and using raycasting, the 2D screen gaze point is converted into real-world 3D spatial coordinates P_gaze.
[0136] Interest focus association: Based on visual data, it is preliminarily determined that P_gaze falls on a landscape object named "Vermilion Rock".
[0137] Interest Intensity Calculation: The module's built-in dynamic attention model calculates the user's real-time interest intensity value S_interest for the cinnabar rock wall based on the duration of continuous fixation and changes in pupil diameter. This value significantly increases.
[0138] Generate a context summary: Package the above information into a structured "context summary" data package, the content of which is roughly as follows: {User ID:User01, Role: Xu Xiake, Timestamp: T, Location: Qingluan Bridge, Object of Interest: Vermilion Rock Wall, Interest Intensity: 0.92, Cognitive Load Estimation: Low}
[0139] (3) Contextual data is uploaded to edge nodes
[0140] The system operates by having the terminal upload a lightweight context summary data packet to the nearest edge computing node via a wireless network (such as 5G / WiFi 6).
[0141] (4) Edge node localization and spatiotemporal anchor point matching
[0142] System action: High-precision fusion positioning and dynamic mapping service at the edge node layer is initiated.
[0143] Precise positioning: Combining the approximate location uploaded by the terminal, local visual SLAM feature points, and a pre-set high-precision point cloud map, the user's current centimeter-level 6DoF precise pose (Pose_u) is calculated.
[0144] Spatiotemporal anchor matching: The real-time spatiotemporal anchor matching engine quickly queries and matches the most relevant spatiotemporal anchors (STAs) in the locally cached spatiotemporal knowledge subgraph based on Pose_u and the Ming Dynasty time context selected by the user. For example, STA_Qingluan Bridge_1616 (representing the Qingluan Bridge scene when Xu Xiake traveled there in 1616).
[0145] (5) Edge node forwarding and cloud request
[0146] The system action involves the edge node encapsulating the context summary along with the matched spatiotemporal anchor ID, and uploading it as a complete service request to the cloud-based intelligent hub layer via a high-speed network.
[0147] (6) Cloud-based intelligent reasoning and content generation
[0148] System actions involve the coordinated work of various modules in the cloud-based central layer:
[0149] Knowledge Retrieval: The Multimodal Hierarchical Spatiotemporal Knowledge Base (MHSTKB) retrieves all knowledge related to the received STA_Qingluanqiao_1616 and the Zhushayan Wall object ID, including descriptions of geological features from the Ming Dynasty, excerpts from Xu Xiake's "Travels in Huangshan Diary," related historical paintings, and high-precision 3D models of the rock wall.
[0150] Deep reasoning and narrative generation: The HistLLM (Hist LLM) model for enhancing historical spatiotemporal knowledge is activated. Based on the "Xu Xiake" character setting, the user's high level of interest, and the retrieved knowledge, it reasones: "The user is focusing on observing the typical features of the Danxia landform of Huangshan in the Ming Dynasty from the perspective of Xu Xiake, and needs to be provided with an in-depth empirical interpretation that is consistent with his scientific expedition identity."
[0151] Subsequently, the narrative generation engine generates a personalized narrative instruction based on this, such as: "In the voice of Xu Xiake, combined with the original text of 'Travels in Huangshan Diary,' explain the rock characteristics of the cinnabar wall and its significance in the geological structure of Huangshan."
[0152] (7) Cloud-based decisions are distributed to the edge.
[0153] The system takes action by packaging the narrative instructions, along with the associated 3D model data (such as highlightable rock strata) retrieved from MHSTKB and the text / audio content (generated narration), and sending them back to the edge computing nodes.
[0154] (8) Real-time computing and content scheduling at edge nodes
[0155] The system operates as follows: upon receiving instructions, the edge nodes process two key tasks in parallel:
[0156] Real-time 3D registration calculation: The spatiotemporal anchor point registration engine calculates the precise transformation matrix and lighting rendering parameters for overlaying the virtual rock structure model onto the real rock wall based on the user's precise pose (Pose_u) and ambient lighting data, ensuring a seamless blend of virtual and real elements.
[0157] Content scheduling: Quickly retrieve resources such as audio explanations and 3D models from edge cache or receive them in real time from cloud data streams, and get them ready.
[0158] (9) Rendering commands are sent to the terminal
[0159] The system operates by sending the calculated registration parameters, the virtual content data stream to be rendered, and the playback instructions to the user's AR glasses terminal with minimal latency from the edge node.
[0160] (10) Immersive presentation and interactive loop of the terminal
[0161] a) System actions: The immersive rendering and interaction engine at the user terminal layer performs the final operation.
[0162] AR rendering: Based on the received registration parameters, a virtual wireframe diagram of the rock strata structure is accurately overlaid and blended with historical annotations onto the real cinnabar rock wall in the user's field of vision.
[0163] Multi-channel feedback: Through bone conduction headphones, a narration generated by HistLLM and synthesized with the voice of "Xu Xiake" is played: "Do you see this Red Cliff? I wrote a record here, 'The rocks are majestic and expansive, and the springs are high and scattered'... This is the beginning of the Danxia landform of Huangshan."
[0164] b) User experience: Within a delay of less than one second, users see digital annotations “growing” on the rock face in front of them and hear timely explanations from their digital travel companion “Xu Xiake”, completing a perfect intelligent interactive loop of “perception, understanding, decision-making, and response”.
[0165] The above system workflow description fully demonstrates the advantages of cloud-edge-device collaboration: complex AI reasoning and large-scale knowledge retrieval are completed in the cloud, high-precision positioning and real-time rendering calculations are completed at the edge, and sensitive perception and final presentation are completed on the terminal. Each of the three performs its own function, working together to support a smooth and immersive historical and spatial perspective experience.
[0166] This invention also provides a multi-dimensional context-aware computing method, executed in the cloud-based intelligent central layer of the aforementioned context-aware historical spatiotemporal perspective system, comprising the following steps:
[0167] The input multimodal sensing data from the terminal within the received time slice Δt includes eye-tracking data, pose data, audio data, and environmental data; the input vector includes:
[0168] I t ={Gaze t Pose t Audio t ,Env t} (4)
[0169] In the formula, Gaze t For eye-tracking data, including the screen coordinates of the fixation point and pupil diameter; Pose t Head 6DoF pose; Audio tFor audio features, such as whether someone is speaking, keyword extraction; Env t To simplify environmental characteristics, including ambient light and noise levels.
[0170] By fusing eye-tracking and pose data, the user's 3D gaze point in the real-world coordinate system is calculated using ray casting; P gaze =RayCast[CamerMatrix(Pose)] t Gaze t [.screen_point] (5)
[0171] The three-dimensional gaze point is associated with objects in the historical spatiotemporal knowledge base to determine the current object of interest; that is, the spatially closest object to P is found in MHSTKB. gaze Historical object O j Calculate the current time t for O j Intensity of interest S interest (t):
[0172] Based on the continuous fixation duration and pupil diameter changes of the object of interest, the real-time interest intensity S is calculated using a dynamic attention model. interest (t):
[0173] S interest (t)=β*S interest (t-1)+(1-β)*f(T fixation D pupil (6)
[0174] In the formula, T fixation For historical object O j Continuous fixation duration; D pupil β is the pupil diameter change; β is the attenuation factor used to simulate the natural decay of interest; f(·) is the normalization function that maps the physiological signal to an intensity increment.
[0175] Based on the user's voice features, gaze point distribution entropy, and head micro-motion variance, estimate the user's real-time cognitive load;
[0176] Preferably, the following formula is used to estimate the user's real-time cognitive load, which serves as the basis for adjusting the amount of information delivered:
[0177] L cog (t)=γ1*Speech rate (t)+γ2*Gaze entropy (t)+γ3*Motion var (Pose t (7)
[0178] In the formula, Speechrate (t) represents the number of voice questions asked per unit time; Gaze entropy (t) represents the degree of disorder in the gaze point switching between multiple objects; Motion var (Pose t γ1, γ2, and γ3 are the variances of minute head tremors, used to indicate fatigue or discomfort; γ1, γ2, and γ3 are the weighting coefficients of each factor.
[0179] Output a context summary vector consisting of the object of interest, real-time interest intensity, and real-time cognitive load, used to drive personalized narrative decisions. The context summary vector consists of the historical object, interest intensity, and cognitive load value [O j ,S interest (t),L cog (t)].
[0180] The narrative generation engine makes decisions based on this scenario summary vector, when S interest (t) is high and L cog When (t) is low, push in-depth professional interpretation; when L is low, push in-depth professional interpretation. cog When (t) is high, it is simplified to key tips or suggestions to rest.
[0181] The present invention also provides an augmented reality cultural heritage tour device, based on the aforementioned context-aware historical spatiotemporal perspective system; matched to the user end, it adopts a head-mounted augmented reality display device to carry the functional modules of the user embodied terminal layer, providing the user with an immersive historical spatiotemporal perspective experience.
[0182] Taking "a user selecting the character Xu Xiake and using augmented reality devices to explore Huangshan Mountain in Anhui Province in depth" as a typical application scenario, a specific embodiment of the "context-aware historical spatiotemporal perspective system and multi-dimensional context-aware computing method" described in this invention is illustrated below:
[0183] Scenario Overview: A user wearing AR glasses, at a service station at the foot of Huangshan Scenic Area, selects to play the role of "Xu Xiake," a Ming Dynasty geographer and traveler, through a terminal application. Based on the character's knowledge graph, behavioral patterns, and historical background, the system overlays Ming Dynasty exploration narratives and scientific research tasks onto the real Huangshan landscape, providing a deeply immersive cultural tour.
[0184] Detailed workflow description:
[0185] (1) Character binding and scene initialization
[0186] a) User Operations
[0187] The user selects "Xu Xiake" in the App role library. The system loads the exclusive profile of this role: Role_XuXiake = {Knowledge Graph (KG): including knowledge points from his travel notes "A Record of a Journey to Huangshan", exploration preferences (landforms, water systems, strange pine trees), and the geographical cognitive level in the Ming Dynasty; Language Model (LM): a dialogue style with a mixture of classical and vernacular Chinese, emphasizing evidence; Task Template (T): a scientific exploration task with the main line of "observation, record, survey".}
[0188] b) System response
[0189] The agent is activated in the tone of Xu Xiake: "I am Xu Hongzu from Jiangyin. I have always been fond of seeking secluded and beautiful places. Today, I am traveling with you in Huangshan. Let's carefully observe the rock bones and cloud veins of this mountain. Would you like to follow me to explore the realm of 'several red cliff cliffs, with wonders at every step'?" At the same time, the edge node starts to load the high-precision three-dimensional real-scene map of the Huangshan area and the associated Ming Dynasty time-space anchor point data.
[0190] (2) Time-space anchor point trigger and dynamic registration
[0191] a) User action
[0192] The user walks along the mountain climbing path towards the direction of "Guest-Greeting Pine".
[0193] b) System perception and registration
[0194] Terminal real-time positioning: Determine that the user is near "Qingluan Bridge"
[0195] Edge node query: Query MHSTKB and find that there is a key time-space anchor point STA_QingluanBridge_1616 at this location, associated with the record of Xu Xiake's visit to this place on the third day of the second lunar month in 1616
[0196] Dynamic registration algorithm: Immediately work after registration. Based on the current user's perspective, call the Ming Dynasty scene data associated with STA_QingluanBridge_1616, such as a more primitive mountain path and different vegetation landscapes, and use the consistent lighting rendering technology to accurately overlay the virtual scene with the real cliffs and streams. Through the AR glasses, the user can see that the stone bridge style is more primitive and the mountain path is faintly visible
[0197] (3) Multi-dimensional context perception and personalized narrative generation
[0198] a) Context capture
[0199] The user stops at the side of Qingluan Bridge, and the eye movement tracking module of the AR glasses detects that he has been staring at the "Vermilion Rock Wall" on the right for a long time.
[0200] b) Real-time calculation
[0201] Interest intensity: The dynamic attention model calculates that the interest value S_interest for "cinnabar rock wall" increases rapidly based on the fixation duration T_fixation = 8s.
[0202] Cognitive load: Based on the user's steady breathing rhythm (analyzed via microphone) and static posture, it is determined that their L_cognitive is low and they are in a state of focused observation.
[0203] c) Intelligent Response
[0204] The narrative engine generates a deep response based on the context summary (character: Xu Xiake, object of interest: Zhushayan Wall, S_interest: high, L_cognitive: low).
[0205] The virtual image of Xu Xiake (or only his voice) appears and says: "Do you see this red cliff? I wrote a record here, 'The rocks are majestic and expansive, and the springs are high and scattered, in thousands of streams...' This is the beginning of the Danxia landform of Huangshan. Its stone texture and stratification are very different from those of the back mountain."
[0206] Meanwhile, in the user's AR view, geological stratification lines and key sentences from Ming Dynasty travelogues are highlighted and overlaid on the rock face.
[0207] (4) Role-based task-driven and exploration-guided
[0208] a) Task Issuance
[0209] Based on the quest chain of the Xu Xiake character and the user's current location, the intelligent agent issues a character-specific side quest:
[0210] Years ago, while searching for a water source, I followed the sound of this stream upstream and found an unnamed waterfall. Would you be willing to follow suit and head to the "Human-shaped Waterfall" ahead, noting its width and the depth of the pool beneath? Upon completion, you will receive a fragment of a travelogue.
[0211] b) Guidance and Assistance
[0212] The system projects virtual "footprint" light trails on the ground as path guides and provides a virtual "ruler" tool that floats beside the user to complete the measurement task.
[0213] c) Dynamic adjustment
[0214] During the journey, the weather suddenly changed, and the sensors detected a decrease in light and an increase in humidity (change in Env_t). The context-aware model predicted that the user's L_cognitive might be affected by increased environmental stress.
[0215] The system dynamically simplifies navigation information, and the Xu Xiake intelligent agent suggests: "A mountain rain is coming, so hurry up. First and foremost, see the waterfall; secondly, take measurements."
[0216] (5) Results recording and experience closed loop
[0217] a) Task completed
[0218] Users arrive at the waterfall and use virtual tools to complete measurements. The system automatically records the data and unlocks mission rewards:
[0219] An exclusive AR video of Xu Xiake's memories of the waterfall, visualized using HistLLM-generated content, along with digital fragments of Xu Xiake's Travels, are added to the user's personal "travelogue".
[0220] b) Experience Generation
[0221] After the tour, the system automatically generates a "Xu Xiake Scientific Expedition Log," which integrates the user's walking route for the day, the "geological features" discovered, the measurement tasks completed, and the collected travelogue fragments. It also generates a short travelogue in Xu Xiake's writing style, which the user can share and review.
[0222] This embodiment fully demonstrates how the present invention combines cutting-edge AR and AI technologies with profound historical and cultural connotations to create a next-generation immersive cultural experience that transcends existing guided tour models.
[0223] Obviously, the above embodiments are merely illustrative examples for clear explanation and are not intended to limit the implementation. Those skilled in the art will recognize that other variations or modifications can be made based on the above description. It is neither necessary nor possible to exhaustively list all possible implementations here. However, obvious variations or modifications derived therefrom are still within the scope of protection of this invention.
Claims
1. A context-aware historical spatiotemporal perspective system, characterized in that, include: It adopts a cloud-edge-device architecture, consisting of a cloud-based intelligent hub layer, an edge computing node layer, and a user-embodied terminal layer. The user embodied terminal layer is used to collect user multimodal perception data and generate user context summaries in real time, as well as receive and present historical content that blends the virtual and real worlds. The edge computing node layer connects the user embodied terminal layer and the cloud intelligent hub layer. It is used to receive the user context summary, perform high-precision positioning and spatiotemporal anchor point matching, and forward the matching result and the user context summary to the cloud intelligent hub layer. It also receives cloud instructions, calculates virtual and real registration parameters, and sends them to the user embodied terminal layer. The cloud-based intelligent central layer is used to retrieve historical knowledge and perform multi-dimensional context-aware calculations based on the received matching results and context summaries, generate personalized historical narrative instructions and related content, and send them down to the edge computing node layer.
2. The context-aware historical spatiotemporal perspective system according to claim 1, characterized in that, The user-embodied terminal layer includes: A multimodal sensing array is used to acquire multimodal sensing data, including environmental images, pose data, eye movements, physiological data, and speech data. A lightweight real-time context understanding module is used to integrate eye-tracking and pose data to calculate the user's three-dimensional interest focus in the real world, and to calculate the real-time interest intensity and user cognitive load on the interest focus based on a preset dynamic attention model, generating a structured user context summary. An immersive rendering and interaction engine is used to receive virtual and real registration parameters and virtual content from the edge computing node layer, perform augmented reality rendering, and capture user interaction commands. A role-based application interface is used to provide user role selection and task management functions, and to inject the selected role information into the user context summary.
3. The context-aware historical spatiotemporal perspective system according to claim 2, characterized in that, The edge computing node layer includes: High-precision positioning and dynamic mapping services are used to fuse multi-source positioning signals to provide users with centimeter-level real-time pose and dynamically update it in response to environmental changes. The spatiotemporal anchor point matching and registration engine is used to query and match the optimal spatiotemporal anchor point in the locally cached historical spatiotemporal knowledge subgraph based on the user's real-time pose and the target's historical time, and calculate the rendering and registration parameters required for virtual-real fusion based on the spatiotemporal anchor point matching results and real-time environmental data. The local context aggregation and forwarding module receives user context summaries from one or more terminals, performs regional-level situation aggregation, and packages and uploads the aggregated context summaries and the spatiotemporal anchor point matching results to the cloud intelligent hub layer; and distributes the narrative instructions and content issued by the cloud intelligent hub layer to the corresponding terminals. High-bandwidth content edge caching is used to pre-store or cache high-precision historical 3D models and media assets from the cloud, which can be called at high speed by the spatiotemporal anchor matching registration engine and the terminal rendering engine.
4. The context-aware historical spatiotemporal perspective system according to claim 1, characterized in that, The cloud-based intelligent hub layer includes: A multimodal hierarchical spatiotemporal knowledge base is used to organize and store historical data in a multidimensional structure based on objects, events, spatiotemporal, and media, with spatiotemporal anchor points as the core index. The historical spatiotemporal knowledge enhancement model is connected to the knowledge base and is used to receive service requests containing user context and spatiotemporal anchors, retrieve related knowledge from the knowledge base, and perform deep spatiotemporal reasoning and character-based content generation. The narrative generation and global scheduling engine is connected to the historical spatiotemporal knowledge enhancement model. Based on the reasoning results of the historical spatiotemporal knowledge enhancement model and the user context, it generates personalized narrative clues, interactive tasks and scheduling instructions, and coordinates the presentation of virtual content among multiple users.
5. The context-aware historical spatiotemporal perspective system according to any one of claims 14, characterized in that, The spatiotemporal anchor point STA i Defined as a quintuple of data as shown in the following formula: STA i ={G i ,T i ,C i ,M i ,L i } (1) In the formula, G i The geographic coordinate bounding box is defined using latitude, longitude, and elevation coordinates; T i For the effective time interval [t] start ,t end ];C i For associated historical 3D scene content, including multi-LOD models and materials; M i Metadata, derived from historical events and document citations; L i For spatiotemporal association links with other anchor points; All spacetime anchors STA i A spatiotemporal graph network is constructed in the multimodal hierarchical spatiotemporal knowledge base.
6. The context-aware historical spatiotemporal perspective system according to claim 5, characterized in that, The edge computing node layer, when receiving the user context summary and performing high-precision positioning and spatiotemporal anchor point matching, includes: Receive real-time pose data and target historical time from the user terminal; Based on the pose data and the target historical time, query the historical spatiotemporal knowledge base to find matching candidate spatiotemporal anchor points; Calculate the matching score between each candidate anchor point and the current pose and target time, and select the optimal spatiotemporal anchor point as the registration benchmark; Based on the real-time spatial relationship between the user and the optimal spatiotemporal anchor point, historical scene content at different levels of detail is dynamically scheduled. Real-time ambient lighting information is collected and used to drive the historical scene content to perform consistent lighting rendering, while simultaneously calculating the virtual and real occlusion relationship.
7. The context-aware historical spatiotemporal perspective system according to claim 6, characterized in that, The matching score S based on spatial distance and temporal proximity is calculated using the following formula. match Select the anchor point STA with the highest score. k As the current registration benchmark: In the formula, d represents the pose. u .location to STA k Geometric distance from the center; T center For STA k The time center point; α is the spatial weighting factor.
8. A multi-dimensional context-aware computing method, executed in the cloud-based intelligent central layer of the context-aware historical spatiotemporal perspective system described in any one of claims 17, characterized in that, Includes the following steps: The input multimodal perception data from the terminal within the received time slice Δt includes eye-tracking data, pose data, audio data, and environmental data. By integrating eye-tracking data and pose data, the user's three-dimensional gaze point in the real-world coordinate system is calculated through ray casting. The three-dimensional gaze point is associated with objects in the historical spatiotemporal knowledge base and identified as the current object of interest. Based on the continuous fixation duration and pupil diameter changes of the object of interest, the real-time interest intensity S is calculated using a dynamic attention model. interest (t): S interest (t)=β*S interest (t-1)+(1-β)*f(T fixation ,D pupil ) (6) In the formula, T fixation For historical object O j Continuous fixation duration; D pupil β is the pupil diameter change; β is the attenuation factor used to simulate the natural decay of interest; f(·) is the normalization function that maps the physiological signal to an intensity increment. Based on the user's voice features, gaze point distribution entropy, and head micro-motion variance, estimate the user's real-time cognitive load; The output consists of a contextual summary vector composed of the object of interest, real-time interest intensity, and real-time cognitive load, which is used to drive personalized narrative decisions.
9. The multidimensional context-aware computing method according to claim 8, characterized in that: The user's real-time cognitive load is estimated using the following formula: L cog (t)=γ1*Speech rate (t)+γ2*Gaze entropy (t)+γ3*Motion var (Pose t ) (7) In the formula, Speech rate (t) represents the number of voice questions asked per unit time; Gaze entropy (t) represents the degree of disorder in the gaze point switching between multiple objects; Motion var (Pose t γ1, γ2, and γ3 are the variances of minute head tremors, used to indicate fatigue or discomfort; γ1, γ2, and γ3 are the weighting coefficients of each factor.
10. An augmented reality cultural heritage tour guide device, characterized in that, The context-aware historical spatiotemporal perspective system according to any one of claims 17; and a head-mounted augmented reality display device for carrying the functional modules of the user embodied terminal layer, providing the user with an immersive historical spatiotemporal perspective experience.