Large and small model and multi-agent collaborative text and travel interactive narrative guide method and system

By employing a cultural tourism interactive narrative guidance method that utilizes large and small models and multi-agent collaboration, personalized narrative content is generated using user data and knowledge graphs. This solves the problem of rigidity in existing guidance methods and enhances the immersive experience for tourists.

CN121880618APending Publication Date: 2026-04-17BEIJING WANTU SIRUI TECH CO LTD
0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
BEIJING WANTU SIRUI TECH CO LTD
Filing Date
2025-12-25
Publication Date
2026-04-17

AI Technical Summary

Technical Problem

Existing cultural tourism tour guide methods cannot generate personalized narrative content based on tourists' real-time location, behavior, and environment, resulting in fragmented immersive experiences and difficulty in creating deep emotional resonance.

Method used

By employing a large-scale model and a multi-agent collaborative approach, the system acquires users' spatial, visual, and interactive data, combines them with a pre-built cultural tourism knowledge graph, generates a set of narrative behaviors that match the user's state, and renders multimodal guided tour content.

Benefits of technology

It enables intelligent adaptive narrative interaction based on tourists' real-time location, behavior, and environment, thereby improving the quality of tourists' immersive cultural and tourism experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121880618A_ABST
    Figure CN121880618A_ABST
Patent Text Reader

Abstract

The invention provides a big and small model and multi-agent collaborative text and travel interaction narrative guide method and system. The method comprises the following steps: acquiring current spatial data, visual data and interaction data of a user and historical portrait data of the user; generating a first current state vector of the user on the basis of the spatial data, the visual data and the interactive data in combination with a pre-constructed text travel knowledge graph, and generating a second current state vector of the user on the basis of the first current state vector of the user, the historical portrait data and the text travel knowledge graph; generating a narrative speech set matched with the current state of the user by adopting a pre-constructed large model for macroscopic narrative planning and a small model for generating microcosmic speech for a specific role; and generating and rendering the multi-modal navigation content based on the narrative line set. Through the method disclosed by the invention, the limitation of a fixed narrative mode can be broken through, intelligent self-adaptive narrative interaction based on the real-time position, behavior and environment of the tourist is realized, and the quality of immersive travel experience of the tourist is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of artificial intelligence technology, and in particular to a method and system for interactive narrative guidance in cultural tourism that utilizes large and small models and multi-agent collaboration. Background Technology

[0002] Currently, cultural tourism experiences are transforming from traditional static viewing to dynamic, immersive, and interactive experiences that deeply integrate contextual perception and interactive feedback. The core of this trend lies in using technology to enhance visitor participation and emotional experience, achieving an organic combination of physical space and digital storytelling.

[0003] However, current common guided tour methods, such as fixed-point audio guides and QR code scanning for text, images, audio, and video introductions, are still mostly passively triggered services. Their typical drawback is that once tourists arrive at a preset point, the system can only broadcast or display pre-recorded or edited fixed content. This type of content presentation is rigid and the narrative is linear and simplistic, unable to dynamically generate coherent and personalized narrative content based on the tourist's real-time location, behavior, and environment. This results in a fragmented tour experience, making it difficult to create a deep sense of immersion and emotional resonance.

[0004] Therefore, how to break through the limitations of fixed narrative patterns and achieve intelligent adaptive narrative interaction based on tourists' real-time location, behavior, and environment, thereby improving the quality of tourists' immersive cultural and tourism experience, has become a technical problem that urgently needs to be solved by those skilled in the art. Summary of the Invention

[0005] In view of this, this disclosure proposes a cultural tourism interactive narrative guidance method and system based on large and small models and multi-agent collaboration. This method can break through the limitations of fixed narrative patterns and realize intelligent adaptive narrative interaction based on tourists' real-time location, behavior and environment, thereby improving the quality of tourists' immersive cultural tourism experience.

[0006] According to a first aspect of this disclosure, a method for interactive narrative guidance in cultural tourism using a large-scale model and multi-agent collaboration is provided, comprising: Acquire the user's current spatial data, visual data, interaction data, and the user's historical profile data; Based on the spatial data, the visual data, and the interaction data, combined with the pre-constructed cultural tourism knowledge graph, a first current state vector of the user is generated. The first current state vector includes the user ID, the user's current main associated location, the relative orientation to the main associated location, the narrative object of interest, the interaction intent, the current main narrative location of interest, and a set of narrative atom candidate sets associated with the current main narrative location and the narrative object. Based on the user's first current state vector, the historical profile data, and the cultural tourism knowledge graph, a set of narrative speech and behavior matching the user's current state is generated through a pre-constructed large model for macro narrative planning, a small model for generating micro speech and behavior for specific roles, and the collaboration of multiple intelligent agents. Based on the aforementioned set of narrative words and actions, multimodal guided tour content is generated and rendered.

[0007] In one possible implementation, constructing the cultural tourism knowledge graph includes: Obtain the narrative script of the cultural tourism industry, extract entities and relations from the narrative script to obtain the entity attribute set and the relation triple set, and construct an initial knowledge graph based on the entity attribute set and the relation triple set; For each location entity in the initial knowledge graph, the corresponding geographical coordinates and geofences are bound to obtain the spatiotemporal knowledge graph; The narrative script is decomposed to obtain multiple independently triggerable and interconnected narrative atoms, and each narrative atom is integrated into the spatiotemporal knowledge graph to obtain the cultural tourism knowledge graph.

[0008] In one possible implementation, when generating the user's first current state vector based on the spatial data, the visual data, the interaction data, and a pre-constructed cultural tourism knowledge graph, the process includes: Feature extraction is performed on the spatial data, the visual data, and the interaction data to obtain the user's current main associated location, relative orientation to the main associated location, narrative object of interest, and interaction intent; The user's ID, the user's current primary associated location, the relative orientation to the primary associated location, the narrative object of interest, and the interaction intent are combined in sequence to obtain the user's initial current state vector; Based on the user's current main associated locations, narrative objects of interest, and the cultural tourism knowledge graph, determine the user's current main narrative locations of interest and the candidate set of narrative atoms associated with the current main narrative locations and the narrative objects; The user's first current state vector is obtained by integrating the current main narrative location of interest to the user and the narrative atom candidate set into the initial current state vector.

[0009] In one possible implementation, when generating a set of narrative behaviors matching the user's current state based on the user's first current state vector, the historical profile data, and the cultural tourism knowledge graph, using a pre-constructed large model for macro-narrative planning and a small model for generating micro-level behaviors for specific roles, the process includes: Based on the user's first current state vector, the historical profile data, and the cultural tourism knowledge graph, a second current state vector of the user is generated; Based on the user's second current state vector, the large model is used to perform macro-narrative planning analysis, so as to select the optimal narrative atom from the narrative atom candidate set and generate a set of agents to be activated associated with the optimal narrative atom; Using the small model, corresponding dialogue content is generated for each agent in the set of agents to be activated, resulting in a set of narrative speech and behavior that matches the user's current state.

[0010] In one possible implementation, generating the user's second current state vector based on the user's first current state vector, the historical profile data, and the cultural tourism knowledge graph includes: The current main narrative location is extracted from the user's first current state vector, and the subgraph structure associated with the current main narrative location is extracted from the cultural and tourism knowledge graph; The user's second current state vector is obtained by integrating the subgraph structure and the historical profile data into the user's first current state vector.

[0011] In one possible implementation, when generating and rendering multimodal guided tour content based on the narrative speech and behavior set, the process includes: Iterate through the dialogue content in the aforementioned collection of narrative words and actions; For the current dialogue content encountered during the traversal, visual content, audio content, and AR rendered content are generated based on the current dialogue content, and the visual content, audio content, and AR rendered content are rendered synchronously. After the traversal is complete, the generation and rendering of each dialogue content in the narrative speech and behavior set is realized.

[0012] In one possible implementation, after generating and rendering multimodal guided tour content based on the narrative speech and behavior set, the method further includes: Collect multi-dimensional interaction data during user usage, and iteratively update the narrative planning strategy of the cultural tourism knowledge graph and the large model based on the multi-dimensional interaction data.

[0013] According to a second aspect of this disclosure, a cultural tourism interactive narrative guide system based on a large-scale model and multi-agent collaboration is provided, comprising: The data acquisition module is used to acquire the user's current spatial data, visual data, interaction data, and the user's historical profile data; The state vector generation module is used to generate the user's first current state vector based on the spatial data, the visual data, the interaction data, and a pre-constructed cultural tourism knowledge graph. The first current state vector includes the user ID, the user's current main associated location, the relative orientation to the main associated location, the narrative object of interest, the interaction intent, the current main narrative location of interest, and a set of narrative atom candidate sets associated with the current main narrative location and the narrative object. The narrative planning module is used to generate a set of narrative behaviors that match the user's current state, based on the user's first current state vector, the historical profile data, and the cultural tourism knowledge graph, using a pre-built large model for macro narrative planning and a small model for generating micro-level behaviors for specific roles. The content rendering module is used to generate and render multimodal guided tour content based on the narrative speech and behavior set.

[0014] According to a third aspect of this disclosure, a cultural tourism interactive narrative guide device with a large-scale model and multi-agent collaboration is provided, comprising: a processor; a memory for storing processor-executable instructions; wherein the processor is configured to execute the method described in the first aspect of this disclosure.

[0015] According to a fourth aspect of this disclosure, a non-volatile computer-readable storage medium is provided that stores computer program instructions thereon, wherein the computer program instructions, when executed by a processor, implement the method described in the first aspect of this disclosure.

[0016] This disclosure provides a method and system for interactive narrative guidance in cultural tourism using a large-scale model and multi-agent collaboration. The method includes: a data acquisition module for acquiring the user's current spatial data, visual data, interaction data, and historical profile data; a state vector generation module for generating the user's first current state vector based on the spatial data, visual data, and interaction data, combined with a pre-constructed cultural tourism knowledge graph, wherein the first current state vector includes the user ID, the user's current main associated locations, the relative orientation to the main associated locations, the narrative objects of interest, the interaction intent, the current main narrative locations of interest, and a candidate set of narrative atoms associated with the current main narrative locations and narrative objects; a narrative planning module for generating a set of narrative behaviors matching the user's current state based on the user's first current state vector, historical profile data, and cultural tourism knowledge graph, through a pre-constructed large model for macro-narrative planning, a small model for generating micro-level behaviors for specific roles, and collaboration among multiple agents; and a content rendering module for generating and rendering multimodal guided tour content based on the set of narrative behaviors. The method disclosed herein can break through the limitations of fixed narrative patterns and achieve intelligent adaptive narrative interaction based on tourists' real-time location, behavior, and environment, thereby improving the quality of tourists' immersive cultural and tourism experience.

[0017] Other features and aspects of this disclosure will become clear from the following detailed description of exemplary embodiments with reference to the accompanying drawings. Attached Figure Description

[0018] The accompanying drawings, which are included in and form part of this specification, illustrate exemplary embodiments, features, and aspects of this disclosure together with the specification and serve to explain the principles of this disclosure.

[0019] Figure 1 A flowchart illustrating an immersive cultural tourism interactive narrative guide method according to an embodiment of the present disclosure; Figure 2 A schematic block diagram of an immersive cultural tourism interactive narrative guide device according to an embodiment of the present disclosure is shown; Figure 3 A schematic block diagram of an immersive cultural tourism interactive narrative guide device according to an embodiment of the present disclosure is shown. Detailed Implementation

[0020] Various exemplary embodiments, features, and aspects of this disclosure will now be described in detail with reference to the accompanying drawings. The same reference numerals in the drawings denote elements that have the same or similar functions. Although various aspects of the embodiments are shown in the drawings, they are not necessarily drawn to scale unless specifically indicated otherwise.

[0021] The term “exemplary” as used herein means “serving as an example, embodiment, or illustration.” Any embodiment illustrated herein as “exemplary” is not necessarily to be construed as superior to or better than other embodiments.

[0022] Furthermore, to better illustrate this disclosure, numerous specific details are set forth in the following detailed description. Those skilled in the art will understand that this disclosure can be practiced without certain specific details. In some instances, methods, means, components, and circuits well known to those skilled in the art have not been described in detail in order to highlight the main points of this disclosure.

[0023] <Method Implementation> Figure 1 A flowchart illustrating a cultural tourism interactive narrative guide method based on a size model and multi-agent collaboration according to an embodiment of this disclosure is shown. Figure 1 As shown, the method includes steps S1100-S1400.

[0024] S1100 acquires the user's current spatial data, visual data, interaction data, and the user's historical profile data.

[0025] Specifically, the user's current spatial data, visual data, and interaction data are generated in real time within the user's app. The specific generation process for each type of data is as follows: (1) Spatial data: The user's app obtains the user's current coarse latitude and longitude position via the Beidou module; the user's current inertial sensor data is obtained via the three-axis accelerometer, gyroscope and magnetometer in the inertial measurement unit (IMU); the user's current position coordinates are obtained by fusing the user's current coarse latitude and longitude position and the inertial sensor data through the Kalman filter algorithm. And the current orientation vector of the phone in space ; Set the user's current location coordinates And the current orientation vector of the phone in space The vector formed { , This serves as the user's current spatial data.

[0026] (2) Visual data: This part of the data is captured by the user's app at a pre-set sampling rate (e.g., 1 frame per second) from the phone's camera in real time. Keyframe images at the current moment are then captured from the video stream and used as the visual data for that moment. By analyzing the visual data, the narrative object that the user is currently interested in can be determined.

[0027] (3) Interaction Data: This type of data mainly originates from information generated by various interactive events initiated by the user in the user-side APP. For example, when a user uses the voice input function of the user-side APP to input the current interaction intent, the APP will convert the user's voice content into corresponding text, and this text transcription result will serve as the user's current interaction data. Similarly, when a user performs a click operation on the user-side APP, the event information triggered by each click will serve as the user's current interaction data. Subsequent analysis of the interaction data will determine the user's true current interaction intent.

[0028] After the user-side app generates the user's current spatial data, visual data, and interaction data, it packages the acquired data and uploads it to the system executing the method disclosed herein. In this way, the system can obtain the user's current spatial data, visual data, and interaction data in a timely manner.

[0029] Furthermore, the system includes a user database that records each user's long-term interests and historical interaction records. When executing the method disclosed herein, the long-term interests and historical interaction records of the current user can be directly retrieved from the user database, and historical profile data of the current user can be generated and obtained based on these records. .

[0030] S1200 generates the user's first current state vector based on the user's current spatial data, visual data, and interaction data, combined with a pre-constructed cultural tourism knowledge graph. The first current state vector includes the user ID, the user's current primary associated location, the relative orientation to the primary associated location, the narrative object of interest, the interaction intent, the current primary narrative location of interest, and a candidate set of narrative atoms associated with the current primary narrative location and narrative object.

[0031] First, it's important to note that a cultural tourism knowledge graph needs to be constructed before performing this step. In one possible implementation, the construction process of this cultural tourism knowledge graph is as follows: First, we acquire narrative scripts from the cultural and tourism industry, extract entities and relationships from these scripts to obtain sets of entity attributes and relation triples. Based on these extracted sets, we construct an initial knowledge graph. The narrative scripts for the cultural and tourism industry primarily originate from unstructured text data from various channels. This data includes, but is not limited to, historical documents, local chronicles, academic papers, and professional content provided by domain experts. This data contains profound historical and cultural background and local characteristics, providing solid content support for the cultural and tourism industry.

[0032] In one feasible implementation, when extracting entities and relations from the narrative script to obtain the set of entity attributes and the set of relation triples, a verification mechanism for the extraction results and an iterative prompt word optimization mechanism are used. This aims to extract structured knowledge with high precision from complex, archaic narrative scripts. The specific extraction steps are as follows: First, based on the core concepts of the cultural tourism industry (such as people, places, events, cultural relics, works, etc.) and the relationships between these core concepts (such as occurring in, participating in, creating, inscribing, etc.), an ontology schema for the cultural tourism industry is defined, and a dedicated initial prompt template is designed for each type of entity and relationship included in the ontology schema. For example, the initial prompt template for extracting entities such as "historical events" is: "Please find all historical events from the following text and output them in JSON format, including fields such as 'event name', 'occurrence time', 'main participants', and 'occurrence location'." Secondly, based on the initial prompt word templates involved, a large model is used as an efficient information extraction engine to perform batch knowledge extraction on each input narrative script. When the large model outputs each entity and relation triple, it simultaneously generates a confidence score reflecting the certainty of its judgment. The system then automatically extracts a structured initial set of entity attributes (with confidence scores) and a set of relation triples (subject, relation, object, confidence score).

[0033] Then, the initial entity attribute set and especially the relation triple (subject, relation, object) set are quality checked. If the quality check fails, the corresponding initial prompt word template is iteratively optimized based on the quality issues detected. Based on the iteratively optimized prompt word template, the corresponding entities or relations are extracted again and quality checked again until the extracted entity attribute set and relation triple set pass the quality check. The entity attribute set and relation triple set that pass the quality check are output as the final result.

[0034] In one possible implementation, the quality verification of the initial entity attribute set and relation triplet set includes: checking whether the confidence level of each entity and relation in the initial entity attribute set and relation triplet set is lower than the low confidence threshold; checking whether each entity and relation in the initial entity attribute set and relation triplet set is ambiguous; checking whether each entity and relation in the initial entity attribute set and relation triplet set conflicts with entities and relations obtained from other data sources (such as authoritative databases or structured event timelines) (e.g., contradictory descriptions of the occurrence time of the same event); if any of the above quality problems exist, the quality verification is deemed to have failed.

[0035] If the quality verification fails, the system dynamically generates and appends corresponding optimization instructions to the prompt word template used in this extraction based on the identified quality problem type, thereby achieving iterative optimization of the prompt word template. For the three typical quality problems—low confidence, ambiguity, and data conflict—the generation of optimization instructions follows preset rules and a template mapping mechanism: To address the issue of low confidence: When the confidence score of the model output attached to an entity or relation triple is lower than a set threshold, the generated optimization instructions focus on requiring the model to review and mine evidence for the low-confidence content. For example: "For entities or relations with low confidence, please re-examine the context and look for more direct descriptions or evidence to support the judgment. If the evidence is insufficient, consider discarding or labeling it as uncertain." To address the issue of ambiguity: When an entity is identified as having vague references (such as "this building" or "this company") or a general relationship is identified, the generated optimization instructions aim to guide the model to perform contextual disambiguation and specification. For example: "If there are vague references or general references in the text, please combine the semantics of the context to perform disambiguation and clarify the specific entity name or the precise type of relationship it refers to." Regarding the issue of data conflicts: When the extraction results contradict the data from an authoritative external data source, the generated optimization instructions will specify the calibration basis and standardization requirements. For example: "Please note the differences between the expressions of dates, place names, etc. in the text and the standard knowledge base. Please use [the specific data source name, such as 'Chinese Historical Chronology'] as the benchmark for information calibration and output it in a standard format (such as the Gregorian calendar)." The optimization instruction generation mechanism is automatically triggered by the system according to the preset quality problem classification rules. The generation process of each instruction is as follows: first, the problem type is matched to the corresponding instruction template library, then the key information in the current problem context (such as entity type, conflict fields, etc.) is extracted and the variables in the template are filled in, and finally a targeted and executable natural language instruction is formed.

[0036] Through the aforementioned quality checks and multi-round iterative optimization of the prompt word template, the accuracy and robustness of the prompt word template can be significantly improved, thereby further enhancing the output quality of the entity attribute set and relation triple set. Based on this, a high-quality initial knowledge graph can be constructed.

[0037] Second, for each location entity in the initial knowledge graph, its corresponding geographical coordinates and geofences are bound to obtain a spatiotemporal knowledge graph. Specifically, for each location entity in the initial knowledge graph, its name is matched with a high-precision GIS map database to obtain its corresponding precise latitude and longitude coordinates, which are then bound to the location entity. Based on this, dynamic geofences are set and bound to location entities with narrative value to achieve location-based automated narrative triggering. These geofences are typically based on the latitude and longitude coordinates of the location entity. Centered on, with radius The circular region is mathematically defined as: In this formula, Represents the i-th location entity in the knowledge graph. These are the geographic coordinates of any point on the Earth's surface. It is a location entity The central latitude and longitude coordinates, It is a location entity The set trigger radius (e.g., R=10 meters for a cultural relic pavilion; R=50 meters for a square). It is a function that calculates the geographical distance between two points. This definition provides the core basis for subsequently determining whether a user has entered a specific narrative scene.

[0038] Thus, a spatiotemporal knowledge graph has been constructed, with entities as nodes, relationships as edges, and including spatial coordinates and geofence attributes for location entities. An example fragment of this spatiotemporal knowledge graph is shown below: Node (V): {Entity ID: "E001", Entity Type: "Person", Attribute: {"Name": "Li Bai", "Dynasty": "Tang"}}, {Entity ID: "L001", Entity Type: "Location", Attribute: {"Name": "Yellow Crane Tower", "Coordinates": [114.305, 30.554],"Trigger Radius": 50}}, {Entity ID: "Evt001", Entity Type: "Event", Attribute: {"Name": "Poem on Yellow Crane Tower"}}. (This likely refers to a specific node or event.) This represents the set of all entity vertices of type "location" in a knowledge graph.

[0039] Edge (E): {Relation ID: "R001", Type: "Creation", Start Node: "E001", Target Node: "P001"}, {Relation ID: "R002", Type: "Occurred at", Start Node: "Evt001", Target Node: "L001"}. Here, P001 represents the entity node of the poem "Farewell to Meng Haoran at Yellow Crane Tower".

[0040] Third, the narrative script is decomposed to obtain multiple independently triggerable but interconnected narrative atoms. These narrative atoms are then integrated into a spatiotemporal knowledge graph to obtain a cultural tourism knowledge graph. This cultural tourism knowledge graph is stored in the system for retrieval and use when executing the method disclosed herein. The specific steps are as follows: First, the original chronological narrative script needs to be meticulously broken down into several appropriately sized narrative atoms that can be triggered independently but are interconnected. These narrative atoms are the basic building blocks of the entire narrative system.

[0041] Specifically, each narrative atom is a self-contained structured data unit, including key fields such as atom ID, trigger condition, associated entity set, preset multimedia resource ID, completion condition, and possible links to subsequent narrative atom IDs. The atom ID identifies and distinguishes different narrative atoms. The trigger condition (e.g., "User enters Geofence(L_Yellow Crane Tower)") determines the activation condition of the narrative atom. The associated entity set (e.g., poet Li Bai, Yellow Crane Tower, the poem "Farewell to Meng Haoran at Yellow Crane Tower") determines the main content of the narrative atom. The preset multimedia resource ID points to images, audio, video, and other materials associated with the narrative atom, used to present a rich sensory experience to the user upon triggering. The completion condition determines whether the narrative content of the narrative atom has been fully presented to the user. Subsequent narrative atom IDs provide possible path guidance for the advancement of the narrative flow, enabling various narrative atoms to be linked in a certain logical order, forming a coherent narrative chain. In general, each narrative atom can be viewed as a special dynamic node that encapsulates various information such as triggering logic, related entities, media resources, and narrative flow, aiming to provide basic support for the subsequent generation of a complete and smooth narrative flow.

[0042] Secondly, semantic analysis is carried out on each narrative atom obtained from the decomposition. Based on the results of the semantic analysis, multi-dimensional semantic links between each narrative atom and entity nodes in the spatiotemporal knowledge graph are constructed. Then, each narrative atom is integrated into the spatiotemporal knowledge graph, and finally, a cultural tourism knowledge graph is constructed.

[0043] It's important to note that the multi-dimensional semantic links established between narrative atoms and entity nodes essentially constitute a relational network that combines logical structure with semantic depth. This not only systematically connects the key elements of the narrative but also defines the logical driving path of the narrative, thereby ensuring that the narrative content unfolds and progresses in an orderly manner within a clear structure.

[0044] In one possible implementation, when performing semantic analysis on each narrative atom obtained from the decomposition, constructing multi-dimensional semantic links between each narrative atom and entity nodes in the spatiotemporal knowledge graph based on the semantic analysis results, and then integrating each narrative atom into the spatiotemporal knowledge graph, the following steps may be included: First, by performing semantic analysis on the narrative atom, the entities, entity types, and potential subsequent narrative atoms included in the narrative atom are identified; then, relational edges matching the entity type are selected to link the narrative atom to the corresponding entity nodes in the spatiotemporal knowledge graph; at the same time, link relationships between the narrative atom and subsequent narrative atoms are created.

[0045] For example, a narrative atom node (NA_001) about "Li Bai's poem" describes that "the poet Li Bai composed 'Farewell to Meng Haoran at Yellow Crane Tower' at Yellow Crane Tower." Semantic analysis of this narrative atom identifies the following entities and entity types: the location entity "Yellow Crane Tower," the person entity "Li Bai," and the work entity "Farewell to Meng Haoran at Yellow Crane Tower." The identified potential subsequent narrative atom is NA_002 "Meng Haoran's response." Therefore, when integrating the narrative atom node NA_001 into the spatiotemporal knowledge graph, the following linking operations will be included: (1) Establish spatial triggering link: By creating a "triggered by" type relationship edge that matches the location entity, NA_001 is linked to the location entity node "Yellow Crane Tower" in the spatiotemporal knowledge graph. This relationship edge is the spatial basis for subsequent judgment on whether the user enters the narrative scene.

[0046] (2) Binding the core role link: By creating a relationship edge of type "behavioral subject" that matches the character entity, NA_001 is linked to the character entity node "Li Bai" in the spatiotemporal knowledge graph. This relationship edge defines the core execution role of this narrative atom.

[0047] (3) Linking the narrative object: By creating a relation edge of the type "involving works" that matches the work entity, NA_001 is linked to the work entity node "poem 'Farewell to Meng Haoran at Yellow Crane Tower'" in the spatiotemporal knowledge graph. This relation edge clarifies the core content or object around which the narrative unfolds.

[0048] (4) Preset narrative flow link: By creating a "potential follow-up" type relationship edge, NA_001 is linked to another subsequent narrative atom node such as NA_002 "Meng Haoran's response". This variable relationship constitutes a non-linear narrative branch that can be explored by the user.

[0049] The core innovation of this step lies in dynamically integrating these narrative atoms into the aforementioned spatiotemporal knowledge graph. This refined method of integrating composite links fundamentally transforms static historical knowledge (entities and relationships) into dynamic narrative capabilities that can be reasoned and invoked by the system. It enables the knowledge graph to evolve from a passively queried fact base into a narrative engine that can proactively respond to the environment and possess state transition logic. When a user's state meets the triggering condition of a narrative atom, the system can immediately locate that narrative atom and all its contextual associations by traversing these composite links, thereby activating it and placing it into the candidate narrative set.

[0050] By constructing the initial knowledge graph, integrating the location coordinates of various entities and geofences, and integrating the narrative atoms, the final cultural tourism knowledge graph can be constructed. This cultural tourism knowledge graph G=(V,E,A), where V is the set of entity vertices, E is the set of edges, and A is the set of attributes.

[0051] After constructing the cultural tourism knowledge graph, it is stored in a graph database. To support subsequent location-aware queries, all location entities in the cultural tourism knowledge graph are... latitude and longitude coordinates A spatial index is established based on its trigger radius. This way, when the user's location coordinates... When inputting data, first use the current user's location coordinates. Determine the user's current primary associated location Then, the latitude and longitude coordinates of various locations can be quickly retrieved using the index. and its trigger radius ,calculate and The distance between them will be used to determine the nearest trigger radius. The location serves as the primary narrative location. ; Obtain the current main narrative location All related narrative atoms are linked to achieve precise and immediate narrative connections. Among these, the current main narrative location... The calculation formula is as follows: In the formula, Representative for location entities Geofencing is set.

[0052] After constructing the aforementioned cultural tourism knowledge graph, the operation of generating the user's first current state vector based on spatial data, visual data, and interaction data, combined with the pre-constructed cultural tourism knowledge graph, can be performed. This operation can specifically include the following steps: S1210 extracts features from spatial data, visual data, and interactive data to obtain the user's current main associated location, relative orientation to the main associated location, narrative object of interest, and interactive intent.

[0053] First, it's important to note that to effectively improve the accuracy and precision of feature extraction, specialized small models were configured for different data types, including spatial feature extraction models, visual feature extraction models, and interactive feature extraction models. These small models are specifically designed and optimized to more accurately extract feature data that influences narrative reasoning from their respective data types, thus ensuring the accuracy and reliability of subsequent analysis. This targeted design not only improves the efficiency of feature extraction but also further enhances the credibility and practicality of the results.

[0054] When using a spatial feature extraction model to extract features from spatial data to obtain the user's current primary associated location and its relative orientation to the primary associated location, the following steps may be included: Obtaining the user's current spatial data { , The input is fed into a spatial feature extraction model, which will extract the user's current location coordinates. Matching the user with the coordinates of preset points of interest in the internal map, that is, determining the preset point of interest closest to the user through Euclidean distance calculation as the user's current primary associated location. (e.g., location_id: "Yellow Crane Tower"). Meanwhile, based on... Calculate the user and the main associated locations relative orientation (e.g., facing: "main entrance of the pavilion"). Finally, { , } as the structured extraction result of spatial features.

[0055] When using a visual feature extraction model to extract features from visual data to identify narrative objects of interest to the user, the following steps can be included: Input the acquired visual data into the visual feature extraction model. This model will focus on extracting specific objects with narrative value (such as unique architectural components, cultural relics, or QR code identifiers) from the visual data within the cultural tourism knowledge graph, and input a list of identified features. (e.g., [feature_id:"flying eaves feature",confidence:0.95]), using this as the narrative object that the user is currently interested in. The visual feature extraction model is built upon a finely tuned lightweight image recognition model, MobileNetV3, and is used to analyze and infer keyframe images in visual data. This model does not perform general object recognition, but rather focuses on identifying specific objects with narrative value within the cultural and tourism knowledge graph.

[0056] When using an interaction feature extraction model to extract features from interaction data and obtain the user's current interaction intent, the acquired interaction data is first input into the interaction feature extraction model. This model will then perform the following operations to extract the user's current interaction intent. The specific steps are as follows: First, based on a pre-defined intent classification dictionary and semantic matching rules, the intent category corresponding to the interaction data is determined. Here, interaction data refers to text data input by the user that reflects their interaction intent. This text data includes text data directly input by the user and text transcription results converted from the user's voice input.

[0057] First, it's important to note the pre-defined cultural and tourism intent classification dictionary within the system. This dictionary includes intent categories such as "inquiring about history," "biographical information," "navigation guidance," and "requesting recommendations," and provides a rich set of keywords, synonyms, and typical question templates for each intent category. The typical question templates specify the keywords that each intent category might include. For example, for "inquiring about history," keywords could include "history," "origin," "when it was built," and "who built it." Furthermore, the system pre-configures semantic matching rules for each intent category (e.g., "sentences containing both 'history' and 'building' are highly likely to belong to the 'inquiring about history' intent" intent). This allows for the subsequent determination of the intent category corresponding to the interactive data by combining the pre-defined cultural and tourism intent classification dictionary and semantic matching rules.

[0058] When determining the intent category corresponding to interaction data based on a preset intent classification dictionary and semantic matching rules: First, the user-input interaction data is segmented into words; then, the segmented results are matched with keywords in the graph classification dictionary. Based on the keyword matching results and preset semantic matching rules, the probability that the user's current interaction data belongs to each intent category is calculated, and the intent category with the highest probability is determined as the current user's intent category. .

[0059] Secondly, the entities mentioned in the interaction data are identified, and entity links are established between the identified entities and entity nodes in the cultural tourism knowledge graph to obtain the entity linking results. Specifically, the entities mentioned in the interaction data are first identified; then, the identified entity names (such as "Li Bai" and "Yellow Crane Tower") are precisely linked or fuzzily matched with entity nodes in the cultural tourism knowledge graph (for example, linking the user's mention of "Li Taibai" to the standard entity "Poet Li Bai" through an alias table). The set of all entity IDs linked in the cultural tourism knowledge graph is taken as the entity linking results. Among them, entity link results The value is a list containing the IDs of all successfully linked entities.

[0060] Finally, the intent category and entity link result node are merged, and the merged result is taken as the user's current interaction intent. For example, for the user input "What is the story between Li Bai and the Yellow Crane Tower?", the current interaction intent is output. This can be represented as: { "Inquiring about the relationship between people and places", :["E001","L001"]}. This interaction intent It not only includes what the user "wants to do" but also "to whom," providing a clear interactive context for subsequent narrative decisions.

[0061] S1220, the user's ID, current primary associated location, relative orientation to the primary associated location, narrative object of interest, and interaction intent are combined sequentially to obtain the user's initial current state vector. The expression for the initial current state vector is as follows: In the formula, This is the user's initial current state vector. For the user's current primary associated location, relative to the main associated locations, For the narrative object that the user is currently interested in, This represents the user's current interaction intent, with the subscript 't' representing the current timestamp.

[0062] At this point, the system has completed the process of converting raw sensor data (including spatial data, visual data, and interaction data) into a standardized context vector (i.e., the user's initial current state vector). The transformation of the current state vector, and the initial current state vector. It encapsulates the low-latency perception results processed by a lightweight model and serves as the initial signal for subsequent modules to make decisions.

[0063] Furthermore, in order to assign an initial current state vector For deeper semantic understanding, the cultural and tourism knowledge graph G will be further semantically enhanced. For details of the semantic enhancement process, please refer to steps S1230 and S1240.

[0064] S1230, based on the user's current primary associated locations, narrative objects of interest, and cultural tourism knowledge graph, determine the user's current primary narrative locations of interest, as well as a candidate set of narrative atoms associated with the current primary narrative locations and narrative objects. Specifically, this may include the following steps: First, based on the user's current primary associated location Together with the cultural tourism knowledge graph G, determine the current main narrative locations that users are interested in. The specific process is described above and will not be repeated here.

[0065] Then, the current main narrative locations that the user is interested in are retrieved from the cultural tourism knowledge graph G. and narrative object Each associated narrative atom is considered, and the set of associated narrative atoms is taken as a candidate set of currently available narrative atoms. .

[0066] S1240 integrates the user's current main narrative locations and narrative atom candidate sets into the initial current state vector, obtaining the user's first current state vector. Specifically, the first current state vector reflecting the user's current state is obtained through semantic enhancement of the cultural tourism knowledge graph G. It can be as follows: The first current state vector after semantic enhancement This module encompasses the complete contextual content from the raw signal to semantic information, serving as the direct and sole input for the multi-agent narrative decision-making engine in step S1300, which coordinates large and small models, to make intelligent decisions. Through the processing flow of this module, a precise, real-time, and highly semantic mapping from the physical world to the digital world is achieved.

[0067] S1300, based on the user's first current state vector, historical profile data, and cultural tourism knowledge graph, generates a set of narrative behaviors that match the user's current state through a pre-built large model for macro narrative planning, a small model for generating micro-level behaviors for specific roles, and collaboration among multiple intelligent agents.

[0068] In this system, "intelligent agent" refers to the procedural abstraction of entity nodes with narrative behavior capabilities in the cultural tourism knowledge graph. During the initialization phase, the system creates corresponding intelligent agent agents for specific types of entity nodes in the knowledge graph (such as historical figures, legendary creatures, or even anthropomorphic cultural relics or buildings). Each intelligent agent agent encapsulates at least the following elements: Identity and knowledge: Bind to a specific entity node in the knowledge graph, inheriting all the attributes and relationships of that entity, constituting its "memory" and "cognition".

[0069] Behavioral Model: Associated with a small language model fine-tuned for the specific role, which learns the language style, knowledge scope and behavioral patterns of the corresponding historical figure or role, and is the core engine for generating "speech and behavior".

[0070] Interaction Interface: Receives context-rich instructions from the macro narrative planner (large model) and calls its behavior model to generate specific dialogue content that conforms to the current narrative scene and its own role setting.

[0071] The specific steps can be as follows: S1310, based on the user's first current state vector, historical profile data, and cultural tourism knowledge graph, generate the user's second current state vector. Specifically, this may include the following steps: First, based on the user's first current state vector Extract the current main narrative location And extract the locations related to the current main narrative from the cultural tourism knowledge graph G. Related subgraph structure The subgraph structure Includes locations related to the current main narrative. Related figures, historical events, cultural relics, and other entities, as well as the rich relationships between these entities.

[0072] Second, based on subgraph structure And the historical portrait data obtained above The first current state vector integrated into the user In this process, the user's second current state vector is obtained. Among them, the second current state vector The expression is as follows: In the formula, That is, the location of the current main narrative. Related subgraph structures.

[0073] From the above explanation, we can clearly understand that the second current state vector after further enhancement processing... It has fully integrated all the key elements needed in the decision-making process. These elements not only cover the first current state vector It contains all environmental information related to the user's current state, as well as background knowledge that significantly impacts decision-making. And historical profile data that can reflect the current personalized characteristics of users. It can be said that the second current state vector The design and construction of the system fully considered the needs for multi-dimensional information integration, thus laying a solid foundation for the rationality and accuracy of subsequent grand narrative planning.

[0074] S1320, based on the user's second current state vector, employs a large model for macro-narrative planning analysis to select the optimal narrative atom from the candidate narrative atom set and generate a set of agents to be activated associated with the optimal narrative atom. Here, the large model refers to a pre-trained LLM model with macro-narrative planning analysis capabilities, which, upon inputting the second current state vector... The optimal narrative action will then be automatically output. This narrative action includes the optimal narrative atoms. and the optimal narrative atom The associated set of agents to be activated .

[0075] Specifically, the macro-narrative planning analysis of a large model may include the following steps: First, from the second current state vector Extracting the candidate set of narrative atoms .

[0076] Then, based on the second current state vector , as a candidate set of narrative atoms Each narrative atom in Calculate an adaptability score Among them, the rating The calculation formula is as follows: in, The cosine similarity function is used. It is an embedding model that maps text to a vector space. The item measures the narrative atoms With the current situation (i.e., the second current state vector) The semantic matching degree of ). Represents historical user profile data Calculated narrative atoms The probability of matching user preferences. Measuring Narrative Atoms Compared to user interaction history The novelty of the experience avoids repetitive experiences. It is a weighting hyperparameter used to balance the importance of various indicators.

[0077] Next, the narrative atom with the highest score is selected as the optimal narrative atom. That is, the optimal narrative atom The calculation formula is as follows: Finally, multi-agent scheduling is performed to integrate the cultural tourism knowledge graph G with the optimal narrative atoms. The associated entity nodes are identified as agents to be activated, and the set of all agents to be activated is used as the optimal narrative atom. The associated set of agents to be activated It should be noted here that the optimal narrative atom It is associated with the entity nodes in the knowledge graph that need to "appear". Each entity node is associated with a pre-created agent proxy instance that encapsulates a special small model, so that it enters a ready state and is ready to generate narrative content. This agent that is ready to generate narrative content is the agent to be activated.

[0078] Among them, the set of agents to be activated The expression is as follows: .

[0079] In the formula, To be with the optimal narrative atom Associate the agent to be activated.

[0080] Furthermore, in generating a set of agents to be activated... Afterwards, the large model will also provide... Each agent to be activated in Generate a context-rich, customized character cue. This lays the groundwork for generating narrative speech and behavior in small models that highly match the character setting and the current context. For example, for the "poet" agent to be activated, the generated character prompt might be: "You (character: Li Bai) are at the Yellow Crane Tower, facing the beautiful scenery of the Yangtze River, a modern tourist (who can be...)" (Implying his potential interest in poetry) Approach him. Greet him proactively with a poetic phrase, naturally mentioning the river view before you, in a tone that is both bold and slightly wistful. S1330 uses a small model, which is a set of agents to be activated. Each intelligent agent to be activated Generate corresponding dialogue content This allows us to obtain a set of narrative behaviors that match the user's current state.

[0081] First, it should be noted that each agent to be activated... It was not done directly by a large model, but by a lightweight small language model that was fine-tuned on the corpus for that specific role. To drive it. Each It is specifically designed to mimic the language style and behavior patterns of specific historical figures or characters, generating corresponding dialogue content. .

[0082] Specifically, for the set of agents to be activated Each intelligent agent to be activated : Obtain customized character prompts generated for the large model Character prompts The system identifies the proposed dialogue roles and obtains small models that match those roles. Character hints Input to the smallest model In the middle, through small models Generate conversational content that best matches the intended dialogue roles. The dialogue content The calculation expression is as follows: in, It is the conditional probability distribution defined by the fine-tuned small model. When generating dialogue content for the agent to be activated, the small model not only focuses on the relevance of the generated dialogue content, but also considers the influence of the proposed role's personality.

[0083] Each agent to be activated Dialogue content By combining them in order, we obtain the narrative speech and behavior set D that matches the user's current state. The recording format of the narrative speech and behavior set D is as follows: From the above, we can understand that for each intelligent agent to be activated... The roles proposed are different, and the small models used to generate the dialogue content for these agents to be activated are also different. Specifically, when the agents to be activated involve multiple proposed roles, multiple small models are invoked simultaneously to perform the work. Through the division of labor and cooperation among these small models, the entire system is able to handle concurrent interactions between multiple roles, and its response speed is much faster than frequently calling large language models (LLMs). This mechanism of utilizing multiple small models working together effectively improves the system's efficiency in handling multi-role interaction tasks, avoids the slow speed problems that may be caused by frequent calls to large language models, and ensures that the system can respond quickly and accurately when facing complex multi-role scenarios.

[0084] In another possible implementation, to ensure the interactivity, logical consistency, and narrative fluency of the dialogue content in the narrative speech and action set D, the following operations can be included when generating the narrative speech and action set D: First, simple message passing is allowed between the agents to be activated. Specifically, the agents to be activated... Through small models The generated dialogue content As the next intelligent agent to be activated Small model used This serves as one of the input contexts, thereby ensuring interactivity between the various intelligent dialogue contents to be activated.

[0085] Secondly, consistency checks are performed on the dialogue content within the narrative speech and action set D. This consistency check process utilizes a consistency scoring function driven by a large model. Complete, used to quantify the content of these conversations in a given context The degree of logical consistency and narrative fluency within the text. Among these, the consistency scoring function... As shown below: In the formula, is the Sigmoid function, which maps the scores to the (0,1) interval; W and b are learnable parameters within the LLM model; It combines the entire narrative set of words and actions (D) with the current situation. The vector representation after joint encoding.

[0086] During consistency verification, the consistency scoring function is used first. The consistency score corresponding to the narrative speech and action set D is calculated. If the consistency score is lower than a preset threshold, the large model will analyze the degree of deviation between the dialogue content $Dialogue_{A_{i}}$ of each agent and the overall narrative context $Context_{LLM}$, and locate the main problematic role causing the inconsistency. Subsequently, the large model will adjust the role prompts for the problematic role. Alternatively, it may instruct its associated smaller models to regenerate the dialogue content based on the revised context. This process iterates until the consistency score of the narrative speech set D is greater than or equal to a preset threshold, thus satisfying the consistency score rule.

[0087] Through the aforementioned dialogue content interaction mechanism and overall consistency distribution mechanism, it is ensured that even if different dialogue content is generated by multiple dispersed small models, the final narrative presented is a coherent whole.

[0088] Another possible implementation allows for the management and state updating of narrative branches. Specifically, this applies to narrative nodes with multiple possible developments, where one narrative node is associated with multiple narrative atoms. The large model will generate multiple narrative atoms based on the context. The branching options are available for the user to choose from. The system records and generates a candidate set of narrative atoms based on the user's choices. In this way, subsequent narrative states will transition to the corresponding narrative branches based on the user's choices. Finally, based on the candidate set of narrative atoms generated from the user's choices... Determine the optimal narrative atom for execution and optimal narrative atoms The associated set of agents to be activated And see above for the set of agents to be activated. Output the set of narrative statements and actions, D, that satisfy the consistency scoring rule. Simultaneously update the user state record to prepare for the next narrative decision loop.

[0089] The innovation of step S1300 lies in constructing a highly efficient collaborative architecture that uses a large model as the "director" for overall narrative planning and multiple small models as "actors" for micro-level speech and behavior generation. This collaborative architecture can utilize the large model's intent understanding and creative narrative capabilities to ensure the richness and cultural depth of the content, while leveraging the lightweight and fast response characteristics of the small models to achieve low-latency responses to user location, behavior, and interaction. This provides users with a highly personalized and smooth immersive narrative experience on lightweight devices.

[0090] S1400 generates and renders multimodal guided tour content based on the narrative speech and action set. Specifically, this may include the following steps: traversing each dialogue content in the narrative speech and action set; for the currently traversed dialogue content, generating visual content, audio content, and AR rendered content based on the current dialogue content, and simultaneously rendering the visual content, audio content, and AR rendered content; after traversal, the generation and rendering of each dialogue content in the narrative speech and action set is completed. In fact, the visual content here mainly refers to image content.

[0091] To address the fragmented nature of simple image retrieval, this step employs a conditional latent diffusion algorithm based on a large model to analyze the dialogue content within the narrative speech-action set D. This generates visual content that is consistent in style and highly relevant to the dialogue. This visual content can include historical scene reconstructions, character designs, etc.

[0092] The conditional latent diffusion algorithm learns latent features from the context using a latent space and calculates simplified representations of these features to discover patterns. Specifically, in the latent space, a precondition c is set as a guide to progressively denoise and generate the latent encoding z of the target visual content. Here, the precondition c consists of two parts: one is the dialogue content co-generated by the size model. Secondly, there is the unified visual style encoding, such as the style of Tang Dynasty ink painting. The denoising process is handled by a UNet network. The objective function for predicting noise is defined as follows: in, It is an encoder used to process the target visual content (i.e., the real image that needs to be displayed). Encoding as a latent representation , yes At time step A noisy version. This objective function Used for training the UNet network During the training phase, the training data includes real images x and their corresponding text descriptions and style labels. Image x is encoded into a latent representation z by an encoder, and then fed into the network along with random time steps t, text conditions (corresponding to dialogue content), and style codes s for optimization.

[0093] During actual system reasoning, the visual content generation process is as follows: First, the current dialogue content is... The text is encoded into a conditional vector and combined with a predefined visual style code s to form a complete precondition c. Then, it is processed in the latent space with random Gaussian noise. Starting with the precondition c, and guided by the pre-trained UNet network, Multi-step iterative denoising is performed to obtain a latent encoding that meets condition c. Finally, it is decoded back into pixel space by a decoder to obtain the desired visual content.

[0094] Using this algorithm, the system can dynamically generate visual content that fits the context based on the text description of "Li Bai drinking alone at Yueyang Tower" and the specified traditional Chinese style.

[0095] In one possible implementation, the process of generating speech content based on the current dialogue content may include the following steps: First, determine the content of the current conversation. For each dialogue character, obtain the character embedding vector corresponding to that dialogue character. The content of the conversation and role embedding vector The input is fed into a pre-trained neural speech synthesis model to generate a Mel spectrogram that matches the voice features (including timbre, intonation, and rhythm features) of the dialogue character. The calculation formula for the Mel spectrogram generated by the neural speech synthesis model is shown below: in, The output is the Mel spectrum. It's a neural speech synthesis model; Text is the input of the current dialogue content. , This is the current conversation content. The corresponding role embedding vector includes information about the current dialogue content. The vocal characteristics of the proposed dialogue characters.

[0096] In this step, by The neural speech synthesis model introduces a learnable role embedding vector into this end-to-end TTS architecture. This allows for the injection of voice features that match the dialogue character during the generation of the Mel spectrogram, thus providing a foundation for the subsequent generation of natural speech that matches the voice features of the dialogue character.

[0097] It should be noted that the system pre-defines an initial character embedding vector for each character. The training samples (i.e., multiple dialogue contents defined for each role) are used. During the training of the neural speech synthesis model, for each role, the training data defined for that role and the initial role embedding vector defined for that role are input into the neural speech synthesis model to train it to generate Mel spectrograms for that specific role. During training, not only are the model parameters of the neural speech synthesis model updated, but the role embedding vectors for each role are also updated simultaneously. The process involves iterative updates. After training, the final neural speech synthesis model is able to accurately embed the character embedding vectors of different roles. It synthesizes voices with different timbres and temperaments that match the character settings.

[0098] Secondly, a lightweight vocoder converts the Mel spectrogram output from the previous step into waveform data, thereby obtaining natural speech that matches the voice characteristics of the dialogue characters. Thus, through the first two steps, a bold and unrestrained, powerful voice is synthesized for the "poet Li Bai," while a higher-pitched, faster-paced voice is synthesized for the "book boy" character, greatly enhancing the realism of the narrative and the distinctiveness of the characters.

[0099] Finally, based on the current dialogue content The system selects appropriate background music from the sound effects library to determine the emotional tone of the speech, dynamically adjusts the volume ratio between the natural speech and the background music, and then merges and outputs the final speech content.

[0100] It's important to note that dynamically adjusting the volume ratio between natural speech and background music aims to ensure a balance between speech clarity and ambiance. This balance problem can be formulated as an optimization problem: in, Indicates volume is Natural speech, Indicates volume is Background music, It is a sharpness evaluation function. It is an emotional impact assessment function. As weight, This represents the maximum acceptable total volume for the device. By solving this problem, the system can adaptively configure the optimal volume ratio between natural speech and background music for different scenarios.

[0101] Furthermore, the purpose of generating AR rendered content based on the current dialogue content is to stably "place" virtual objects at specific locations within the real-world image captured by the phone, achieving the overlay of virtual objects and the real-world image. This can specifically include the following steps: First, obtain the agent to be activated corresponding to the current dialogue content, and then obtain the virtual object defined in the narrative atom bound to the agent to be activated from the cultural tourism knowledge graph.

[0102] Secondly, based on the rotation matrix and translation vector of the camera coordinate system relative to the world coordinate system, the acquired virtual objects are transformed into 2D pixel coordinates on the mobile phone screen and superimposed on the real image currently captured on the mobile phone to obtain AR rendered content.

[0103] It should be noted that this disclosure employs a lightweight visual-inertial odometry method that integrates visual feature points and IMU data to continuously track the 6-DoF pose of the mobile phone phase, i.e., the rotation matrix of the camera coordinate system relative to the world coordinate system. Translation vector The rotation matrix of the camera coordinate system relative to the world coordinate system is tracked. After translating the vector, the virtual object can be converted into 2D pixel coordinates on the mobile phone screen.

[0104] Specifically, let's assume the virtual object's default coordinates in the world coordinate system are... Then its coordinates in the mobile phone camera coordinate system This can be obtained through rigid body transformation: Subsequently, using the camera intrinsic parameter matrix K (which can be obtained through calibration), 2D pixel coordinates projected onto the phone screen symbol This indicates equality in homogeneous coordinates. This formula implements a mapping from the 3D world to the 2D screen. It innovatively utilizes the binding relationship between predefined virtual objects in narrative atoms and world coordinates, such as the coordinates of where a wine glass should be placed on the stone table. Above, combined with real-time pose provided by visual-inertial odometry This enables stable registration and accurate perspective overlay of virtual objects within the mobile phone camera view.

[0105] In one possible implementation, the simultaneous rendering of visual content, audio content, and AR rendered content may include the following steps: First, it's important to clarify that to ensure precise time synchronization of the generated municipal bureau content, audio content, and AR rendered content, the system defines a global rendering clock and timestamps each frame of data. The rendering scheduler schedules different rendering channels based on the clock. The synchronization problem can be modeled as minimizing the rendering time of all output channels i. With ideal presentation time The sum of absolute errors: By using algorithms such as predicting rendering time, preloading resources, and setting synchronization buffers, the system ensures that users see the correct images, virtual objects, and hear the corresponding voice at the correct time, thereby guaranteeing the coherence of the generated narrative content.

[0106] In one possible implementation, after generating and rendering multimodal guided tour content based on the narrative speech and behavior set, it also includes: collecting multidimensional interaction data during user use, and iteratively updating the narrative planning strategy of the cultural tourism knowledge graph and the big model based on the multidimensional interaction data.

[0107] The core of this step is to leverage real-world user interaction data and continuously optimize the performance of a large model as a narrative decision engine using a reinforcement learning framework. This step establishes a closed loop from user feedback to model parameters, enabling the system to learn from practice and continuously improve the accuracy of user intent recognition and the quality of narrative generation. Specifically, it may include the following steps: First, multi-dimensional interactive data collection and reward signal construction.

[0108] This step records multi-dimensional data throughout the entire user-system interaction cycle. This data primarily includes: a) the duration of user dwell time on a single narrative atom. b) Explicit user ratings of the narrative experience For example, 1-5 stars; c) The user's completion rate of the narrative path recommended by the system. ; and d) system response delays recorded from the large model narrative planning decision-making process. The raw data itself has different dimensions and needs to be standardized. Then, the system integrates these standardized indicators into a scalarized reward signal. This signal is used to evaluate the quality of a narrative decision in subsequent steps. The specific form of the reward function is as follows: In this formula, It is the normalized user rating. The function is used to smoothly compress the dwell time to the [0,1] interval. It is a reference duration used for normalization. It's a weighting factor used to balance the importance of different goals such as user experience, completion, and system performance. This composite reward... It can more comprehensively and stably measure the overall effect of a single narrative decision.

[0109] Second, policy gradient optimization and proximal policy optimization algorithms.

[0110] The core of this step is to utilize the reward signal obtained in the previous step to optimize the narrative decision-making strategy of the large model. The narrative decision-making process of the large model is modeled as a Markov decision process. The large model, based on the state... That is, the first current state vector mentioned above. Make an action That is, selecting narrative atoms And schedule agents to be activated, then the environment transitions to a new state. And generate rewards .

[0111] The optimization goal of the system is to find a strategy This refers to the parameterization strategy of large models to maximize expected cumulative reward. ,in, It is an interactive trajectory. This is the discount factor. Since the action space is discrete and can choose different narrative atoms, and the policy itself is a complex neural network, this step uses a proximal policy optimization algorithm to perform stable and efficient policy updates. By limiting the step size of each policy update, drastic oscillations during training are avoided. Its core optimization objective function is: in, It is the probability ratio of the new strategy to the old strategy. It is the advantage function estimate at time step t, which measures the performance of the state. Take action below The degree of superiority or inferiority relative to the average level. This is a hyperparameter used to control the range of pruning. By continuously collecting interaction data, calculating the advantage function, and minimizing the loss function after this pruning, the system can iteratively update the policy parameters of the large model. This makes them inclined to make narrative decisions that yield higher cumulative rewards.

[0112] Third, the incremental evolution mechanism of knowledge graphs and case bases.

[0113] Reinforcement learning primarily optimizes decision-making strategies, but the system's knowledge base also needs to evolve accordingly. This step embeds successful interaction experiences into the system's knowledge framework. When the cumulative reward of a narrative sequence exceeds a success threshold, the sequence is marked as a successful case and stored in the historical case database. The system uses keyword extraction and vectorization techniques to generate a semantic index for each case, facilitating rapid similar case retrieval during subsequent narrative decision-making in large models.

[0114] Meanwhile, the cultural tourism knowledge graph G is also dynamically evolving. The system will statistically analyze each narrative atom. success rate That is, the number of successful attempts divided by the total number of triggers. When When a narrative atom consistently exceeds a dynamic threshold, the system increases the weight or priority of its corresponding entity node in the cultural and tourism knowledge graph. Conversely, for narrative atoms that consistently perform poorly, their weight is reduced, and in extreme cases, they may even be temporarily disabled. Through this mechanism, the cultural and tourism knowledge graph G evolves from a static knowledge base into a dynamic, self-optimizing intelligent knowledge structure that reflects the effectiveness of narrative strategies.

[0115] This disclosure provides a method for interactive narrative guidance in cultural tourism using a combination of large and small models and multi-agent collaboration. The method includes: acquiring the user's current spatial data, visual data, interaction data, and historical profile data; generating the user's first current state vector based on the spatial data, visual data, and interaction data, combined with a pre-constructed cultural tourism knowledge graph. The first current state vector includes the user ID, the user's current main associated locations, the relative orientation to the main associated locations, the narrative objects of interest, the interaction intent, the current main narrative locations of interest, and a candidate set of narrative atoms associated with the current main narrative locations and narrative objects; generating a set of narrative behaviors matching the user's current state based on the user's first current state vector, historical profile data, and cultural tourism knowledge graph, through a pre-constructed large model for macro-narrative planning, a small model for generating micro-level behaviors for specific roles, and collaboration among multiple agents; and generating and rendering multimodal guided tour content based on the narrative behavior set.

[0116] In this disclosure, the generation of narrative speech and behavior sets comprehensively considers multiple factors, including spatial data reflecting the user's current location, visual data reflecting the various scenes the user can currently see, interactive data reflecting the user's current intentions, and historical profile data reflecting the user's past preferences and habits. Through the full integration and analysis of this data, the generated narrative speech and behavior sets possess a powerful function of adaptive adjustment based on the user's real-time state data. This means that regardless of changes in the user's real-time state, such as changes in location, gaze shift, or behavioral patterns, the narrative speech and behavior sets can quickly react and adjust accordingly. This results in the generation of highly coherent and personalized narrative content. The generation of this personalized narrative content significantly improves the immersive experience for tourists during cultural tourism, allowing them to more deeply integrate into the cultural tourism scene and fully appreciate the charm of cultural tourism projects.

[0117] <Device Embodiment> Figure 2 This diagram illustrates a schematic block diagram of a cultural tourism interactive narrative guide system based on a size model and multi-agent collaboration according to an embodiment of the present disclosure. Figure 2 As shown, the device 100 includes: Data acquisition module 110 is used to acquire the user's current spatial data, visual data, interaction data, and the user's historical profile data; The state vector generation module 120 is used to generate a first current state vector of the user based on the spatial data, the visual data, the interaction data, and a pre-constructed cultural and tourism knowledge graph. The first current state vector includes the user ID, the user's current main associated location, the relative orientation to the main associated location, the narrative object of interest, the interaction intent, the current main narrative location of interest, and a set of narrative atom candidate sets associated with the current main narrative location and the narrative object. The narrative planning module 130 is used to generate a set of narrative speech and behavior that matches the user's current state, based on the user's first current state vector, the historical profile data, and the cultural tourism knowledge graph, using a pre-built large model for macro narrative planning and a small model for generating micro speech and behavior for specific roles. The content rendering module 140 is used to generate and render multimodal guided tour content based on the narrative speech and behavior set.

[0118] <Equipment Example> Figure 3 This diagram illustrates a schematic block diagram of a cultural tourism interactive narrative guide device based on a size model and multi-agent collaboration according to an embodiment of the present disclosure. Figure 3 As shown, the cultural tourism interactive narrative guide device 200, which utilizes a large-scale model and multi-agent collaboration, includes a processor 210 and a memory 220 for storing executable instructions of the processor 210. The processor 210 is configured to implement any of the aforementioned large-scale model and multi-agent collaboration cultural tourism interactive narrative guide methods when executing the executable instructions.

[0119] It should be noted here that the number of processors 210 can be one or more. Furthermore, the immersive cultural tourism interactive narrative guide device 200 of this embodiment may also include an input device 230 and an output device 240. The processors 210, memory 220, input device 230, and output device 240 can be connected via a bus or other means, without specific limitations here.

[0120] The memory 220, as a computer-readable storage medium, can be used to store software programs, computer-executable programs, and various modules, such as the programs or modules corresponding to the size model and the multi-agent collaborative cultural tourism interactive narrative guide method in this disclosure embodiment. The processor 210 executes various functional applications and data processing of the immersive cultural tourism interactive narrative guide device 200 by running the software programs or modules stored in the memory 220.

[0121] Input device 230 can be used to receive input digital numbers or signals. These signals may include key signals related to user settings and function control of the device / terminal / server. Output device 240 may include a display device such as a screen.

[0122] <Storage Medium Examples> According to a fourth aspect of this disclosure, a non-volatile computer-readable storage medium is also provided, on which computer program instructions are stored, which, when executed by processor 210, implement the cultural tourism interactive narrative guide method of any of the aforementioned size models and multi-agent collaboration.

[0123] The various embodiments of this disclosure have been described above. These descriptions are exemplary and not exhaustive, and are not limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein is chosen to best explain the principles, practical applications, or technical improvements to the technology in the market, or to enable others skilled in the art to understand the embodiments disclosed herein.

Claims

1. A method for interactive narrative guidance in cultural tourism using a large-scale model and multi-agent collaboration, characterized in that, include: Acquire the user's current spatial data, visual data, interaction data, and the user's historical profile data; Based on the spatial data, the visual data, and the interaction data, combined with the pre-constructed cultural tourism knowledge graph, a first current state vector of the user is generated. The first current state vector includes the user ID, the user's current main associated location, the relative orientation to the main associated location, the narrative object of interest, the interaction intent, the current main narrative location of interest, and a set of narrative atom candidate sets associated with the current main narrative location and the narrative object. Based on the user's first current state vector, the historical profile data, and the cultural tourism knowledge graph, a set of narrative speech and behavior matching the user's current state is generated through a pre-constructed large model for macro narrative planning, a small model for generating micro speech and behavior for specific roles, and the collaboration of multiple intelligent agents. Based on the aforementioned set of narrative words and actions, multimodal guided tour content is generated and rendered.

2. The method according to claim 1, characterized in that, The construction of the cultural tourism knowledge graph includes: Obtain the narrative script of the cultural tourism industry, extract entities and relations from the narrative script to obtain the entity attribute set and the relation triple set, and construct an initial knowledge graph based on the entity attribute set and the relation triple set; For each location entity in the initial knowledge graph, the corresponding geographical coordinates and geofences are bound to obtain the spatiotemporal knowledge graph; The narrative script is decomposed to obtain multiple independently triggerable and interconnected narrative atoms, and each narrative atom is integrated into the spatiotemporal knowledge graph to obtain the cultural tourism knowledge graph.

3. The method according to claim 2, characterized in that, When generating the user's first current state vector based on the spatial data, the visual data, the interaction data, and a pre-constructed cultural tourism knowledge graph, the process includes: Feature extraction is performed on the spatial data, the visual data, and the interaction data to obtain the user's current main associated location, relative orientation to the main associated location, narrative object of interest, and interaction intent; The user's ID, the user's current primary associated location, the relative orientation to the primary associated location, the narrative object of interest, and the interaction intent are combined in sequence to obtain the user's initial current state vector; Based on the user's current main associated locations, narrative objects of interest, and the cultural tourism knowledge graph, determine the user's current main narrative locations of interest and the candidate set of narrative atoms associated with the current main narrative locations and the narrative objects; The user's first current state vector is obtained by integrating the current main narrative location of interest to the user and the narrative atom candidate set into the initial current state vector.

4. The method according to claim 1, characterized in that, When generating a set of narrative behaviors matching the user's current state based on the user's first current state vector, the historical profile data, and the cultural tourism knowledge graph, using a pre-constructed large model for macro-narrative planning and a small model for generating micro-level behaviors for specific roles, the process includes: Based on the user's first current state vector, the historical profile data, and the cultural tourism knowledge graph, a second current state vector of the user is generated; Based on the user's second current state vector, the large model is used to perform macro-narrative planning analysis, so as to select the optimal narrative atom from the narrative atom candidate set and generate a set of agents to be activated associated with the optimal narrative atom; Using the small model, corresponding dialogue content is generated for each agent in the set of agents to be activated, resulting in a set of narrative speech and behavior that matches the user's current state.

5. The method according to claim 4, characterized in that, When generating the user's second current state vector based on the user's first current state vector, the historical profile data, and the cultural tourism knowledge graph, the process includes: The current main narrative location is extracted from the user's first current state vector, and the subgraph structure associated with the current main narrative location is extracted from the cultural and tourism knowledge graph; The user's second current state vector is obtained by integrating the subgraph structure and the historical profile data into the user's first current state vector.

6. The method according to claim 1, characterized in that, When generating and rendering multimodal guided tour content based on the aforementioned set of narrative speech and actions, the process includes: Iterate through the dialogue content in the aforementioned collection of narrative words and actions; For the current dialogue content encountered during the traversal, visual content, audio content, and AR rendered content are generated based on the current dialogue content, and the visual content, audio content, and AR rendered content are rendered synchronously. After the traversal is complete, the generation and rendering of each dialogue content in the narrative speech and behavior set is realized.

7. The method according to claim 1, characterized in that, After generating and rendering multimodal guided tour content based on the aforementioned set of narrative words and actions, the process further includes: Collect multi-dimensional interaction data during user usage, and iteratively update the narrative planning strategy of the cultural tourism knowledge graph and the large model based on the multi-dimensional interaction data.

8. A cultural tourism interactive narrative guide system based on large-scale models and multi-agent collaboration, characterized in that, include: The data acquisition module is used to acquire the user's current spatial data, visual data, interaction data, and the user's historical profile data; The state vector generation module is used to generate the user's first current state vector based on the spatial data, the visual data, the interaction data, and a pre-constructed cultural tourism knowledge graph. The first current state vector includes the user ID, the user's current main associated location, the relative orientation to the main associated location, the narrative object of interest, the interaction intent, the current main narrative location of interest, and a set of narrative atom candidate sets associated with the current main narrative location and the narrative object. The narrative planning module is used to generate a set of narrative behaviors that match the user's current state based on the user's first current state vector, the historical profile data, and the cultural tourism knowledge graph, through a pre-constructed large model for macro narrative planning, a small model for generating micro-level behaviors for specific roles, and the collaboration of multiple intelligent agents. The content rendering module is used to generate and render multimodal guided tour content based on the narrative speech and behavior set.

9. A cultural tourism interactive narrative guide device featuring a large-scale model and multi-agent collaboration, characterized in that, include: processor; Memory used to store processor-executable instructions; The processor is configured to implement the method of any one of claims 1 to 7 when executing the executable instructions.

10. A non-volatile computer-readable storage medium storing computer program instructions thereon, characterized in that, When the computer program instructions are executed by the processor, they implement the method described in any one of claims 1 to 7.