VR-based navigation system interaction method and server
By constructing a pair of virtual intelligent agents that engage in interactive games, analyzing scene data, and optimizing wayfinding paths based on user behavior, the problem of single wayfinding paths in existing technologies is solved, achieving a personalized and intelligent wayfinding experience.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- 贵州轻工职业大学
- Filing Date
- 2026-05-25
- Publication Date
- 2026-06-26
AI Technical Summary
Existing virtual reality wayfinding technology cannot adaptively adjust to the actual spatial semantics of the scene and the dynamic behavior of different users, resulting in rigid and monotonous wayfinding paths that fail to create a highly attractive and personalized tour experience.
By constructing a virtual agent pair that engages in a game of adversarial interaction, including an environment-building agent and a narrative-eroding agent, spatial structure and semantic label information are parsed from scene scanning data to generate dynamic wayfinding paths. The paths are then iteratively corrected and the anchor point set is updated through adversarial game interaction, and the wayfinding paths are optimized by combining user behavior data.
It achieves deep coupling between wayfinding paths and virtual information, enhances the intelligence and personalization of wayfinding interaction in virtual reality scenes, and provides immersive, narrative-coherent, and clearly expressive fusion rendering.
Smart Images

Figure CN122289619A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of computer intelligent technology, specifically to a VR-based wayfinding system interaction method and server. Background Technology
[0002] Virtual reality (VR) technology is now widely used in scenarios such as virtual exhibitions, digital museums, and immersive training. In virtual reality scenarios, wayfinding systems guide users to explore and acquire knowledge in a structured way within a three-dimensional space by planning browsing routes and displaying directional information or explanatory content at key locations.
[0003] Existing virtual reality wayfinding technologies typically rely on predefined fixed paths and manually preset virtual information labels. System developers need to model the scene beforehand, manually set all nodes on the wayfinding path, and statically bind virtual information such as text, images, or models to be displayed to each node. When a user enters the virtual reality scene, the system strictly guides the user's movement according to a predetermined order, and simply triggers and displays the bound content when the user reaches a preset node. However, the above method cannot adaptively adjust to the actual spatial semantics of the scene and the dynamic behavior of different users, resulting in rigid and monotonous wayfinding paths. The presented virtual information is disconnected from the user's spatial context and real-time wayfinding process, making it difficult to create a highly attractive and personalized tour experience. Summary of the Invention
[0004] The purpose of this invention is to provide a VR-based wayfinding system interaction method and server to solve the problems mentioned in the background art.
[0005] This invention provides an interactive method for a VR-based wayfinding system, comprising:
[0006] Obtain a set of scene scan data of the target virtual reality scene, and parse spatial structure information, semantic tag distribution information and guide trigger point location information from the set of scene scan data;
[0007] Create a virtual agent pair that engages in a game of interaction, comprising an environment-building agent and a narrative-eroding agent. Generate an environment memory model based on the spatial structure information and configure it to the environment-building agent. Determine a fictional anchor point generation strategy based on the semantic tag distribution information and configure it to the narrative-eroding agent.
[0008] Combining the scene scanning data set, the current user's navigation task flow, and historical user behavior data, the initial navigation path and a set of fictional anchor points are generated through the interactive game virtual intelligent agent pair. The interactive game virtual intelligent agent pair is driven to engage in adversarial game, causing the environment building agent to correct the initial navigation path based on the set of fictional anchor points and the narrative erosion agent to update the set of fictional anchor points based on the corrected navigation path. The process is iterated alternately until a preset equilibrium condition is met, at which point the game navigation path and the target set of fictional anchor points are output.
[0009] Assign a matching fictional anchor from the target fictional anchor set to each guidance node in the game guidance path and generate a fictional anchor activation instruction, and combine the guidance node sequence and its associated fictional anchor activation instructions into a narrative guidance path;
[0010] In the target virtual reality scene, the navigation interaction interface is rendered based on the narrative navigation path, and the interface moves sequentially to each navigation node. When a navigation node is reached, the corresponding virtual information content is retrieved according to the activation command of the virtual anchor point associated with the navigation node. The virtual information content is superimposed on the corresponding spatial position of the scene, and the rendered screen with virtual information is output to the user's virtual reality display device.
[0011] This invention provides an interactive server for a wayfinding system, comprising:
[0012] A processor; a storage device on which a computer program is stored; a network interface for providing network communication functions; when the computer program is executed by the processor, the processor implements the above-described VR-based navigation system interaction method.
[0013] The present invention provides a readable storage medium on which a program or instruction is stored, and when the program or instruction is executed by a processor, it implements the above-described VR-based guide system interaction method.
[0014] Compared with existing technologies, the beneficial effects of this invention are as follows: By constructing a game-theoretic virtual agent pair comprising an environment-building agent and a narrative-eroding agent, spatial structure information, semantic tag distribution information, and wayfinding trigger point location information are analyzed from the scene scanning data set as the basis for the game. This allows the environment memory model generated by the environment-building agent to dynamically antagonize the fictional anchor point generation strategy determined by the narrative-eroding agent. After combining the current user's wayfinding task flow and historical user behavior data, this game-theoretic virtual agent pair iteratively corrects the initial wayfinding path and updates the set of fictional anchor points until equilibrium is reached, breaking through the limitations of static path planning. The output game-theoretic wayfinding path and the target set of fictional anchor points achieve deep coupling between physical space wayfinding logic and enhanced narrative content. Furthermore, by allocating and activating the target set of fictional anchor points to wayfinding nodes to form a narrative wayfinding path, and accurately retrieving superimposed virtual information content when the user arrives, a highly immersive, narratively coherent, and clearly defined wayfinding intent-driven fused rendering screen is presented to the user, improving the intelligence and personalization level of wayfinding interaction in virtual reality scenes. Attached Figure Description
[0015] To more clearly illustrate the technical solutions in the embodiments or related technologies of this application, the drawings used in the description of the embodiments or related technologies will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0016] Figure 1 This is a flowchart illustrating a VR-based wayfinding system interaction method provided in an embodiment of this application.
[0017] Figure 2 This is a schematic diagram of the basic structure of a wayfinding system interactive server provided in an embodiment of this application.
[0018] Figure 3 This is a functional block diagram of a wayfinding system interactive device provided in an embodiment of this application. Detailed Implementation
[0019] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of this application.
[0020] Please see Figure 1 , Figure 1This is a flowchart of a VR-based wayfinding system interaction method provided in an embodiment of this application. The method can be executed by the wayfinding system interaction server, or by the wayfinding system interaction server and the server together. The method includes steps 110-150.
[0021] This invention can be applied to spatial navigation and information presentation scenarios in virtual reality environments, and is illustrated using VR advertising navigation as an exemplary embodiment. In a VR advertising navigation scenario, a user wearing a virtual reality display device enters a virtual shopping mall scene constructed using 3D digital technology. The target virtual reality scene refers to the 3D digital scene of the virtual shopping mall. The navigation system's interactive server generates a navigation path that balances spatial navigation efficiency and advertising narrative immersion by executing the following method steps, and outputs an enhanced rendering screen that integrates virtual advertising information to the user's virtual reality display device.
[0022] Step 110: Obtain the scene scan data set of the target virtual reality scene, and parse the spatial structure information, semantic tag distribution information and guide trigger point location information from the scene scan data set.
[0023] The scene scan data set is a standard data package generated after offline scanning and structured annotation of the 3D digital scene of the virtual shopping mall. The data package is organized in a hierarchical data block manner. The top layer is divided into three independent storage areas: point cloud geometry, semantic annotation, and trigger point configuration. Each area header contains a type identifier and a check code.
[0024] During parsing, the point cloud geometric blocks are first located, and the total number of vertices and faces is read. The vertex coordinate sequence and face index sequence, represented in single-precision floating-point, are then sequentially obtained. Next, the maximum and minimum coordinate values are counted along the X, Y, and Z axes to construct a scene axial bounding box. The diagonal length is calculated as a normalization scaling factor, and all vertices are scaled uniformly to achieve spatial scale normalization. Then, triangle vertices are extracted according to face indices, and the face normals are calculated using the cross product of vectors. Adjacent faces with a normal angle less than a threshold are aggregated into candidate geometric primitives. Finally, based on conditions such as normal direction, height, and area, categories such as floor, wall, column, and ceiling are determined to generate spatial structure information.
[0025] Simultaneously, a record array is extracted from the semantic annotation block. Each record contains a tag content string (including the main category, subcategory, and description text) and anchor position coordinates. Based on this, a three-dimensional spatial index tree is constructed to quickly retrieve neighboring semantic tags, thus forming semantic tag distribution information. Finally, the trigger point identifier, normalized coordinates, and trigger radius are parsed from the trigger point configuration block to provide spatial references for key interaction locations and generate wayfinding trigger point location information.
[0026] Step 120: Create a virtual agent pair that engages in a game between an environment-building agent and a narrative-eroding agent; generate an environment memory model based on the spatial structure information and configure it to the environment-building agent; determine a fictional anchor point generation strategy based on the semantic tag distribution information and configure it to the narrative-eroding agent.
[0027] The interactive game virtual agent pair consists of two AI entities with different behavioral strategies. The environment-building agent is responsible for memorizing, understanding, and planning paths to the virtual scene's spatial structure, while the narrative-eroding agent is responsible for analyzing user interests and needs, generating fictional narrative content, and formulating delivery strategies. In a cyclical game, the two agents both antagonize and adapt to each other: the environment-building agent tends to maintain the simplicity and efficiency of the path, while the narrative-eroding agent tends to inject more narrative-guiding nodes into the path. Their decisions are continuously coordinated through a pre-set equilibrium mechanism, ultimately achieving a balance between path efficiency and narrative richness.
[0028] Step 121: Extract spatial structure information from the scene scan data set, perform spatial connectivity domain segmentation on the spatial structure information, and generate a spatial connectivity domain partitioning result composed of passable spatial units.
[0029] This step extracts the boundaries of continuous areas that users can freely walk on from the floor-like plane through 3D voxelization and connectivity analysis. First, a voxelized grid map is constructed, with the elevation range of the floor-like plane in normalized space as the central layer, extending vertically upwards and downwards by a predetermined human height to form voxelized intervals. For each grid cell, ray casting is used to detect whether it intersects with or is occluded by walls or pillars; if it is not occluded and the vertical distance between the bottom of the grid and the floor plane is less than a preset step height threshold, it is marked as a passable grid; otherwise, it is marked as an obstacle. Then, a seed-filling algorithm is used to recursively visit adjacent passable grids in six directions, starting from unvisited passable grids, aggregating them into a connected component and assigning a unique spatial unit identifier.
[0030] After traversing all accessible grids, each connected region constitutes a accessible spatial unit. Boundary seam detection is performed on adjacent spatial units. If a corresponding accessible grid pair exists on the boundary and there are no obstacles between them, a spatial passage is determined to exist between the two units, and the center coordinates of the entrances at both ends of the passage and the minimum width of the cross-section are recorded. Finally, the spatial connected region partitioning result is obtained, which includes a list of each accessible spatial unit and its grid index.
[0031] Step 122: Assign a spatial unit identifier to each passable spatial unit, and generate a spatial connectivity topology diagram between each passable spatial unit based on the spatial adjacency relationship in the spatial structure information.
[0032] Based on the spatial connectivity partitioning results output in step 121, an undirected weighted graph structure is constructed using each traversable spatial unit as a node in the graph theory model. Each node in the graph is uniquely named using a spatial unit identifier, and its geometric center coordinates and volume are appended as node attributes. The edge set of the graph is derived from spatial adjacency relationships: for each pair of traversable spatial units with a spatial passage detected in step 121, an undirected edge is established between the two nodes. The edge attributes store the center coordinates of the entrances at both ends of the spatial passage, the normalized length of the central curve of the spatial passage, and the minimum width of the cross-section. The edge weight is calculated using the ratio of the length of the central curve of the spatial passage to the minimum width of the cross-section, and this ratio is used as a measure of the passage cost.
[0033] In constructing the graph structure, redundant edge pruning is performed: when three or more nodes form a fully connected subgraph, the edges with the lowest weights are retained to ensure graph connectivity, while redundant edges with significantly higher weights are deleted to ensure the graph structure's acyclic tree-like tendency and reduce the search space for subsequent path searches. The final graph structure data is serialized and stored in adjacency list format. Each node's adjacency list entry records the spatial unit identifier, edge weight, and channel entry coordinate triplet for all adjacent nodes. This adjacency list is the spatially connected topology graph. Simultaneously, the adjacency list's index metadata, including the total number of nodes, the total number of edges, and the overall connectivity component check results of the graph, is also persistently written.
[0034] Step 123: Project each semantic tag in the semantic tag distribution information to the corresponding accessible spatial unit according to its spatial location coordinates, establish the mapping relationship between the accessible spatial unit and the semantic tag, and generate a spatial unit semantic binding relationship table.
[0035] The semantic tag projection operation uses the graph nodes established in step 122 as the spatial unit reference system. A mapping structure is created with spatial unit identifiers as keys and semantic tag information lists as values. Each semantic tag in the semantic tag distribution information is traversed, and its three-dimensional coordinates of the tag anchor position are taken as the query point. The query point is compared with the Euclidean distance of the geometric center coordinates of all accessible spatial units. At the same time, a fast spatial inclusion detection is performed by combining the spatial volume and geometric bounding box of each accessible spatial unit: it is determined whether the query point falls inside the bounding box of an accessible spatial unit. If it does, the accessible spatial unit is taken as the belonging unit; if the query point does not fall inside any bounding box, the geometric center of the nearest unit is taken as the belonging unit. After determining the belonging unit, the tag content string and the tag anchor position coordinates are extracted from the semantic tags. The main category identifier and subcategory identifier of the tag content string are stored as independent fields, and they are added to the semantic tag information list corresponding to the belonging unit along with the complete descriptive text and anchor coordinates.
[0036] After traversal, the semantic tag information list of each passable spatial unit is grouped according to the main category identifier, and the number of tags in each group is counted to form an ordered semantic tag index structure. This mapping structure is the spatial unit semantic binding relationship table, which is stored in the form of serialized key-value pairs. The key-value part contains group statistics to accelerate semantic queries.
[0037] Step 124: Extract the coordinates of each wayfinding trigger point from the wayfinding trigger point location information, map each wayfinding trigger point coordinate to the passable space unit where it is located, and mark the passable space unit that aggregates wayfinding trigger points as the wayfinding trigger space unit set.
[0038] Each trigger point record in the wayfinding trigger point location information includes a unique trigger point identifier, 3D position parameters, and a trigger radius parameter. The 3D position parameters of each trigger point are extracted as query coordinates, and a spatial projection operation similar to step 123 is performed: First, the unit to which each trigger point coordinate belongs is determined among all accessible spatial units in the spatial connectivity partitioning result. The belonging determination uses a comparison logic between the distance from the trigger point coordinates to the geometric center of the accessible spatial unit and the radius of the unit's bounding sphere; if the distance is less than the bounding sphere radius, the unit is the belonging unit. After determining the belonging unit, a trigger identifier field is appended to the attributes of that belonging unit, and the unique trigger point identifier and trigger radius parameter are written to the trigger attribute list. The trigger attribute list is initialized when the first trigger point is added, and multiple trigger points may be added subsequently.
[0039] For trigger points whose trigger radius parameters cover the boundary of adjacent passable spatial units, an indirect trigger identifier and a reference to the main trigger unit must also be added to that adjacent unit. After completing the spatial mapping of all trigger points, all passable spatial units with non-empty trigger attribute lists are selected, and the spatial unit identifiers of these units are aggregated into a set, which is the wayfinding trigger spatial unit set. This set is stored in array form, and each element of the array contains a spatial unit identifier and a reference to the list of all trigger point identifiers within that unit.
[0040] Step 125: Organize the spatial connectivity domain partitioning result, the spatial connectivity topology graph, the spatial unit semantic binding relationship table, and the guidance trigger spatial unit set into an environment memory model data structure, and serialize the environment memory model data structure into an environment memory model.
[0041] The environmental memory model data structure adopts a top-level container design, containing four member fields, which respectively carry the four types of spatial environment description data generated in the aforementioned steps.
[0042] Field 1 stores the spatial connectivity partitioning results, encapsulating the bounding box data of traversable spatial units, the raster index list, and the original data of spatial adjacency channels into binary blocks, which are then written after encoding using a general serialization protocol. Field 2 stores the adjacency list representation of the spatial connectivity topology graph. The node array and edge array are sequentially serialized into structured text. Each node entry contains an identifier, center coordinates, and a list of adjacent edges; each edge entry contains an endpoint identifier, edge weight, and channel entry coordinates. Field 3 stores the spatial unit semantic binding relationship table, expanding the key-value pair set into a record array. Each record contains a spatial unit identifier, a semantic label category group identifier, a label content string array, and an anchor coordinate array. Field 4 stores the identifier array of the wayfinding trigger spatial unit set and the corresponding trigger attribute list references.
[0043] After the container structure is completed, it is encoded into a byte sequence using a binary serialization protocol. A cyclic redundancy check (CRC) code is calculated and appended to the header as an integrity verification field. Finally, it is compressed using a lossless compression algorithm to generate a binary format environment memory model file, with a version number and generation timestamp appended to the end of the file.
[0044] Step 126: Load the preset environment construction agent basic template, and inject the environment memory model into the perception memory component of the environment construction agent basic template to obtain the environment construction agent.
[0045] The environment-building intelligent agent basic template adopts a multi-component collaborative architecture, including three core components: perception and memory, path planning, and behavior evaluation.
[0046] The Perception Memory component reserves a data interface for the environment memory storage area, used to load external structured environment representation data and support spatial location queries. The loading process first reads the configuration file and component definition script, instantiates three components, and establishes a message passing channel; then, it loads the environment memory model file, decompresses and deserializes it to restore the top-level container structure. From field one, it extracts the spatial connectivity partitioning results, restores the correspondence between the raster index and the region boundary, and loads it into the spatial raster cache of the Perception Memory component. From field two, it extracts the topology graph node and edge data, reconstructs the adjacency list, and loads it into the topology cache of the Perception Memory component and the graph search acceleration cache of the path planning component. From field three, it extracts the semantic binding relationship table and loads it into the semantic label cache in the form of a hash index. From field four, it extracts the wayfinding trigger spatial unit set and loads it into the trigger point marking cache.
[0047] After data injection is complete, the perception memory component performs verification: checking the consistency of cell identifiers between the grid buffer and the topology buffer, verifying the existence of spatial cell identifiers in the semantic label buffer, checking the validity of the normalized range of the trigger point marker buffer coordinates, and marking the ready state of the bit component after the verification passes.
[0048] Step 127: Extract the distribution features of user dwell time in different spatial locations and the interest association features of triggered semantic tags from the historical user behavior data.
[0049] Historical user behavior data is extracted from the user behavior log database by user identifier partition. Each log entry contains a timestamp, user identifier, normalized 3D spatial coordinates, and trigger operation identifier fields, the latter recording the operation type code and associated semantic tag identifier. During processing, logs within a preset backtracking window are filtered by the current user identifier and sorted in ascending order by timestamp. For extracting dwell time distribution features, the passable area of the scene is first divided into voxel grids consistent with previous data. The log sequence is traversed to determine the grid to which each record belongs. Dwell time is calculated based on the difference in timestamps between adjacent logs. For multiple visits to the same grid, the dwell time and frequency are accumulated, generating a raw statistical table with the grid index as the key.
[0050] Spatiotemporal smoothing is then performed: spatially, the cumulative duration of each grid cell is weighted and averaged with respect to its 3D neighborhood using an inverse distance weighting method; temporally, earlier records are assigned exponential decay weights, with the decay coefficient positively correlated with the time interval from the current date, resulting in a smoothed distribution of dwell time. Interest-related feature extraction involves extracting semantic tag identifiers from log operation identifiers and matching them with scene semantic tags. Upon successful matching, the interaction frequency counter for the corresponding category is incremented. For each semantic tag category, the ratio of the current user's trigger count to the total number of trigger counts for that category in the system is calculated. A Bayesian smoothing method is used to incorporate prior pseudo-counts for confidence correction, eliminating the influence of sample differences to form interest-related features.
[0051] Step 128: Construct a user interest space mapping function based on the dwell time distribution features and the interest association features. The user interest space mapping function is used to map user interest categories to spatial regions in the scene.
[0052] The construction process uses each interest category in the interest association features as the input parameter of the mapping function, and the output function value is the spatial response intensity distribution matrix corresponding to that interest category. The dimension of the spatial response intensity distribution matrix is exactly the same as the dimension of the voxel grid matrix used to divide the scene space, and each element in the matrix corresponds to the spatial coverage of a voxel grid. For the input interest category Ck, the construction process traverses all voxel grids in space. For the voxel grid at index (u, v, w), the smooth dwell time value Tstay(u, v, w) of that grid is read from the dwell time distribution features. This value is then normalized to its global maximum value by dividing it by the maximum smooth dwell time value of all voxel grids in the dwell time distribution features.
[0053] The association probability value Passoc(u, v, w) between the interest category Ck and the semantic labels triggered in the voxel grid is read from the interest association features. This association probability value is obtained by weighting the trigger frequencies of all labels belonging to the interest category Ck and whose semantic labels are projected onto the voxel grid. The comprehensive interest response intensity R(u, v, w) of the voxel grid (u, v, w) is obtained by weighted fusion: R(u, v, w) = αspace × Tstay norm (u, v, w) + αinterest × Passoc(u, v, w), where Tstay norm The normalized dwell time value is represented by αspace and αinterest, which are the spatial behavior weight coefficient and interest semantic weight coefficient, respectively. The sum of the two is a constant of 1, and αinterest is greater than αspace to emphasize the higher information value of the user's active interaction intention compared to passive dwell time.
[0054] After performing the above calculations on all voxel grids, all R values for the corresponding interest category Ck are organized into a matrix according to the spatial arrangement of the voxel grids, thus obtaining the spatial response intensity distribution matrix corresponding to interest category Ck. The user interest space mapping function internally maintains a function mapping table, using the interest category identifier as the key and the corresponding spatial response intensity distribution matrix as the value, supporting fast lookup by interest category.
[0055] Step 129: Divide the scene space covered by the semantic tag distribution information into multiple candidate anchor point region grids, use the user interest space mapping function to filter out spatial grids with interest response intensity exceeding a preset filtering threshold from the candidate anchor point region grids to obtain a fictional anchor point candidate location library, and pair the fictional anchor point candidate location library with the corresponding narrative tag material to encapsulate it into a fictional anchor point generation strategy.
[0056] In this embodiment of the invention, the scene space is first divided into regular grids with a resolution that is the same as or coarser than that of the voxelization in step 121 to reduce computational complexity. Each grid serves as a candidate anchor point region grid, with its spatial location represented by the normalized 3D coordinates of its center point. The coordinates of all grid center points and their spatial indices are extracted to form a set of candidate anchor point region grids. For each grid, the spatial response intensity distribution matrix of each interest category in the user interest space mapping function is called based on the center point coordinates. The comprehensive interest response intensity value of the point under each interest category is calculated by trilinear interpolation. If the intensity value of any category exceeds a preset screening threshold, the grid is marked as selected, and the interest category label and intensity value that made it selected are recorded. These are summarized into a set of virtual anchor point candidate entries consisting of center point coordinates, interest labels, and intensity values. All selected entries form a virtual anchor point candidate location library.
[0057] Next, the pairing of fictional anchors with narrative tag materials is performed: each record in the narrative tag material library contains a material identifier, narrative type, content data package reference path, and a list of applicable interest tags. For each entry in the candidate library, its interest category tag is extracted, and all records in the material library whose applicable tag list contains that tag are retrieved; if there are multiple matches, they are sorted in descending order of narrative quality score, and the one ranked first is selected as the paired material; if the scores are the same, the smaller size is selected first, and the material identifier and reference path are written back to the candidate entry. After all pairings are completed, the candidate location library and pairing relationships are combined into a fictional anchor generation strategy data body, and the array is clustered and sorted by interest category tag.
[0058] Step 1210: Load the preset narrative erosion agent basic template, write the fictional anchor point generation strategy into the anchor point decision component of the narrative erosion agent basic template, and obtain the narrative erosion agent.
[0059] The internal component architecture of the Narrative Erosion Agent's basic template comprises three functional components: the Anchor Point Decision Component, the Interest Analysis Component, and the Narrative Adversarial Component. The Anchor Point Decision Component is the core policy decision-making unit of the Narrative Erosion Agent. Internally, it contains a decision policy table data structure, which stores candidate anchor point information and decision rule parameters in key-value pairs. The loading process reads the Narrative Erosion Agent's basic template configuration from the template storage, instantiates the three components, and establishes their mutual references. After instantiation, the Anchor Point Decision Component first performs a self-check process to confirm that the decision policy table is writable.
[0060] Subsequently, the virtual anchor point generation strategy data body generated in step 129 is read, and the virtual anchor point candidate location library array and its pairing relationships are extracted by deserialization. For each candidate entry in the array, a record is created in the decision strategy table, using the candidate anchor point spatial grid index as the key, and storing the candidate location coordinates, interest category label, intensity value, and unique identifier of the paired material as the key-value pair.
[0061] After the data writing is complete, the decision strategy table is re-indexed, with an inverted index structure added based on interest category tags to accelerate subsequent filtering and retrieval by interest category. A read-only query connection is established between the interest analysis component and the decision strategy table, allowing the interest analysis component to quickly query the interest category distribution information in the strategy table when it receives user behavior data. The narrative adversarial component obtains read and write permissions to the decision strategy table, enabling it to dynamically add, delete, and modify fictitious anchor points during the game.
[0062] Step 130: Combining the scene scan data set, the current user's navigation task flow, and historical user behavior data, the initial navigation path and a set of fictional anchor points are generated through the interactive game virtual intelligent agent pair. The adversarial game between the interactive game virtual intelligent agent pair is driven so that the environment construction intelligent agent corrects the initial navigation path based on the set of fictional anchor points and the narrative erosion intelligent agent updates the set of fictional anchor points based on the corrected navigation path. The process is iterated alternately until a preset equilibrium condition is met, and then the game navigation path and the target set of fictional anchor points are output.
[0063] This step generates a wayfinding path that balances efficiency and immersion through a game-like competition between an environment-building agent and a narrative-eroding agent. First, the environment-building agent generates an initial wayfinding path based on the sequence and priority of target points in the wayfinding task flow, using a spatial connectivity topology graph. It then selects the path with the highest semantic richness as the backbone based on semantic binding relationships. Next, the narrative-eroding agent analyzes historical behavior data using an interest space mapping function to generate interest heatmap regions and selects candidate locations for fictional anchors. These fictional anchors are then matched with narrative materials to form a set of fictional anchors. The two agents then engage in a cyclical competition: the environment-building agent forcibly inserts fictional anchors into the path and replans it, while the narrative-eroding agent dynamically adds, deletes, or modifies anchor content based on the semantic matching degree of the new path segments. When the change in path nodes and the increase in anchor information remain below the tolerance threshold for several consecutive rounds, equilibrium is reached, and the game-like wayfinding path and the target set of fictional anchors are output.
[0064] Step 131: The intelligent agent constructs the environment to parse the sequence of target points to be visited and the access priority order of each target point from the navigation task flow. The sequence of target points consists of spatial location identifiers that the user needs to arrive at in sequence in the virtual scene.
[0065] The current user's navigation task flow is passed to the path planning component of the environment's intelligent agent in the form of structured data messages. The message body contains an array of target points, each of which is described by two fields: a target point spatial location identifier and an access priority value. The target point spatial location identifier corresponds to the identifier of a traversable spatial unit in the spatial connectivity topology graph, and the access priority value is a non-negative integer. If two target points have different priority values, the one with the smaller value has higher access urgency.
[0066] The path planning component performs a stable sorting operation on the target point array, with the sort key being the access priority value. The target points in the sorted array are arranged in descending order of priority. If multiple target points have the same priority value, their relative order in the original array remains unchanged. The sorted array is the sequence of target points to be visited. The path planning component also records the total number of target points in the sequence and whether any are marked as invalid because their spatial location identifiers cannot find a matching node in the spatial connectivity topology graph.
[0067] Step 132: The intelligent agent constructs the environment and calls the spatial connectivity topology graph in the environment memory model. Starting from the user's current spatial location, and taking the first unvisited target point in the target point sequence as the temporary endpoint, the agent performs spatial path search processing to generate the original path node sequence connecting the starting point to the temporary endpoint.
[0068] The path planning component first queries the user's current spatial coordinates from the perceptual memory component. Through coordinate matching, it determines the identifier of the currently accessible spatial unit where the user is located, serving as the starting node for the path search. It then sequentially extracts the first unvisited target point from the target point sequence and uses its spatial location identifier to locate the corresponding endpoint node in the spatial connectivity topology graph. Using the starting and endpoint nodes as input for the path search, it calls the A* search algorithm to calculate the shortest path in the graph structure.
[0069] The cost function of A* search consists of two parts: the actual path cost from the starting node to the current expanded node, which is the cumulative sum of the edge weights along the path; and the estimated cost from the current expanded node to the ending node, which is the normalized Euclidean distance between the geometric centers of the two nodes multiplied by the reciprocal of the scene's average traffic speed coefficient. During node expansion, the algorithm sequentially retrieves the identifiers and corresponding edge weights of all adjacent nodes from the current node's adjacency list. It checks if the adjacent node is in the closed list. If not, and the actual path cost from the current node to the adjacent node is lower than the recorded cost, the algorithm updates the parent node pointer and actual path cost of the adjacent node and adds it to the open list. The algorithm iteratively expands the node with the lowest overall cost from the open list until the ending node is added to the closed list or the open list is empty, indicating a search failure. After a successful search, the algorithm backtracks from the ending node along the parent node pointer to the starting node, recording the spatial unit identifier sequence of all nodes traversed along the backtracking path, thus generating the original path node sequence connecting the starting point to the temporary ending point. The path planning component performs pairwise checks on the node sequence to ensure that adjacent nodes in the sequence are indeed connected by corresponding edges in the spatial connectivity topology graph.
[0070] Step 133: The intelligent agent constructs an environment to calculate the semantic richness of each passable spatial unit along the original path node sequence according to the semantic binding relationship table of the spatial unit. The path with the highest semantic richness is selected as the initial wayfinding path backbone, and all target points are sequentially connected according to the access priority order to generate the initial wayfinding path.
[0071] A* search typically returns several alternatives with similar costs when multiple candidate paths exist. The environment-building agent invokes the path planning component to evaluate the semantic richness of all candidate paths. For a given candidate path, the list of accessible spatial unit identifiers traversed by the path is used as an index to query the spatial unit semantic binding relationship table and extract the semantic label records bound to each accessible spatial unit. The quantification of semantic richness comprehensively considers two dimensions: the total number of semantic labels and the diversity of semantic label categories. The specific calculation method is as follows: Let the number of unique semantic labels of all accessible spatial units traversed by a candidate path be Ndistinct, and the number of unique semantic label categories be Ccategory. The semantic richness score Score is calculated as the weighted product of the natural logarithm of Ndistinct and Ccategory, i.e., Score = ln(Ndistinct + 1) × (Ccategory) γ ), where γ is the category diversity amplification factor. After calculating the score for each candidate path, the path with the highest score is selected as the chosen path for that path segment.
[0072] Using the selected path segment as the current backbone, the currently processed target point is marked as visited, and this target point becomes the new starting point for the next path search. Steps 132 and 133 are repeated to connect the remaining unvisited target points in the target point sequence one by one, generating several path segments. After all path segments are generated, they are connected end-to-end in the order they were generated. At the connection points, duplicate traversable spatial unit identifiers between adjacent path segments are removed to obtain the initial wayfinding path. The initial wayfinding path is organized into an ordered list of traversable spatial unit identifiers, with the first element being the starting point, the last element being the final target point, and the intermediate elements arranged in the order they were traversed.
[0073] Step 134: Call the narrative erosion agent to receive the historical user behavior data and the scene scan data set, and use the user interest space mapping function in the fictional anchor point generation strategy to analyze the interest distribution in the historical user behavior data to obtain the interest heat distribution area.
[0074] The interest analysis component of the narrative erosion agent receives historical user behavior data and scene scan data sets from the user behavior log database and initiates the interest heatmap analysis process. First, the interest analysis component calls the user interest space mapping function associated with the decision strategy table through its connection interface with the anchor decision component. The interest analysis component parses the trigger operation identifier field in the historical user behavior data into a sequence of user interactions with semantic tags of various categories, extracting a set of interest categories that intersect with the current scene's semantic tag distribution information. For each interest category in the set, the user interest space mapping function is called to obtain the corresponding spatial response intensity distribution matrix. The spatial response intensity distribution matrices of all interest categories are then merged element-wise, taking the maximum value at each voxel grid position, to form a comprehensive response intensity matrix.
[0075] Spatial smoothing filtering is applied to the composite response intensity matrix. This smoothing filter uses a 3D Gaussian kernel for convolution. The standard deviation parameter of the 3D Gaussian kernel is adaptively adjusted based on the voxel mesh side length and the overall spatial scale of the scene. The standard deviation is the sum of a multiple of the voxel mesh side length and a small-scale scene constant. Each element in the convolved composite response intensity matrix represents the heat value of interest at the corresponding voxel mesh location.
[0076] Based on this, the interest analysis component performs boundary extraction of the thermal distribution region: it traverses the interest thermal value matrix, marking all voxel meshes with interest thermal values exceeding the adaptive thermal cutoff threshold as hotspot meshes. The thermal cutoff threshold is calculated as a linear combination of the mean and standard deviation of the interest thermal values of all elements in the matrix, i.e., threshold = mean + cutoff coefficient × standard deviation. Spatially adjacent hotspot meshes are aggregated into hotspot clusters through 3D connected component analysis. The outer boundary of each hotspot cluster is used to extract isosurfaces, and then these isosurfaces are vertically projected towards the ground to obtain a 2D boundary polygon. The vertex sequence of this polygon is arranged clockwise. The sequence of boundary polygons corresponding to all hotspot clusters constitutes the interest thermal distribution region.
[0077] Step 135: Invoke the narrative erosion agent to select fictional anchor locations located within the interest heat distribution area from the fictional anchor candidate location library, and extract fictional narrative content from the narrative tag material for each selected fictional anchor location to generate a fictional anchor set.
[0078] The anchor point decision component of the narrative erosion agent receives interest heatmap distribution data from the interest analysis component and begins the fictitious anchor point selection operation. The anchor point decision component traverses all fictitious anchor point candidate entries in its decision policy table and performs an inclusion test within the polygon for the candidate position coordinates of each entry.
[0079] Since both the candidate location coordinates and the boundary polygon of the thermal region are located in normalized space, the test only needs to process the two-dimensional coordinates projected onto the horizontal plane. The test algorithm uses the ray casting method: a semi-infinite ray is emitted horizontally from the projection point of the candidate location coordinates on the horizontal plane, and the number of intersections between the ray and each edge of the thermal region polygon is counted. If the number of intersections is odd, the candidate location is determined to be inside the thermal distribution area of interest. Candidate entries that meet the inclusion condition are marked as initially selected hit entries. The set of initially selected hit entries may contain a large number of candidate locations, requiring further simplification and screening. During simplification, the distance between each initially selected hit entry is calculated and a greedy sorting algorithm is used: the entry with the highest comprehensive interest response intensity value is selected as the first selected entry, and redundant entries whose distance from the selected entry is less than a preset minimum distance threshold are removed from the remaining entries; from the remaining entries, the entry with the highest intensity is selected again, and the removal operation is repeated until all entries are processed or removed.
[0080] The streamlined set of entries serves as the officially selected fictional anchor point locations. For each officially selected fictional anchor point location, based on the unique identifier of the paired material recorded in the entry, the corresponding fictional narrative content data package is extracted from the narrative tag material library. The fictional narrative content data package contains reference paths to 3D models or animation resources used for virtual scene rendering, text description data for display, and reference paths to synchronized audio resources. Depending on the content volume parameters, the data package also contains a multi-level material reference chain that adapts to the resolution. The spatial coordinates of the fictional anchor point location, the unique identifier of the paired material, and the corresponding fictional narrative content data package are combined into a fictional anchor point object. All fictional anchor points are then aggregated into a fictional anchor point set.
[0081] Step 136: Start an adversarial game round based on the initial guidance path and the set of fictional anchor points, and determine whether the preset equilibrium condition has been reached based on the statistical information after the end of each round of adversarial game. If so, output the game guidance path and the target set of fictional anchor points.
[0082] Step 1361: Start the adversarial game round. The intelligent agent constructs the environment to obtain the set of fictional anchor points for the current round. All the fictional anchor point positions in the set of fictional anchor points are used as mandatory nodes and inserted into the initial wayfinding path. The spatial path search process is called again to generate the adjusted wayfinding path.
[0083] This step involves the actions performed by the environment-building agent in each round of the adversarial game. The environment-building agent receives the serialized data of the set of fictional anchor points transmitted by the narrative erosion agent in the current round, deserializes it, and uses it to overwrite the mandatory node list maintained internally by the environment-building agent. Each mandatory node entry in the mandatory node list contains the coordinates of the fictional anchor point and the spatial unit identifier mapped from those coordinates. The mapping method involves performing an inclusion test between the coordinates of the fictional anchor point and the bounding boxes of each traversable spatial unit in the spatial connectivity partitioning result to determine the traversable spatial unit to which each fictional anchor point belongs.
[0084] The environment-building agent merges the list of mandatory nodes and the sequence of target points into a global waypoint sequence. The waypoints are ordered according to the user's current location, the mandatory nodes are sorted by distance, and then the sequence is added to the target point sequence. A multi-stop path search algorithm is then invoked, which extends the A* search to support multiple specified mandatory points. The multi-stop search uses two adjacent waypoints in the waypoint sequence as the start and end points of a sub-search, sequentially executing sub-path searches and concatenating the results. If adjacent mandatory nodes are spatially close, the algorithm automatically detects and eliminates path redundancy. After path generation, path smoothing is performed, replacing sharp turns at path corners with smooth transitions using Bézier curves or circular arcs. The smoothed, ordered sequence of passable spatial unit identifiers constitutes the adjusted wayfinding path for the current round.
[0085] Step 1362: Receive the adjusted navigation path through the narrative erosion agent, analyze all passable spatial units traversed by the adjusted navigation path, and for sections where the semantic tags of the passable spatial units match the narrative content of the fictional anchor positions below a preset attraction threshold, add new fictional anchor positions or modify the narrative content of existing fictional anchor positions to generate an updated set of fictional anchors.
[0086] The narrative adversarial component of the narrative erosion agent receives the adjusted wayfinding path data from the environment-constructed agent. The narrative adversarial component expands the sequence of traversable spatial unit identifiers in the adjusted wayfinding path, queries the semantic binding relationship table of each spatial unit, and obtains the set of semantic tag content strings bound to each traversable spatial unit. It then performs word segmentation on all semantic tag content strings and extracts keywords, using a pre-trained word embedding model to map the keywords into fixed-dimensional semantic vectors.
[0087] For each fictional anchor in the current set of fictional anchors, keywords from the descriptive text in its fictional narrative content data package are extracted and mapped to narrative content vectors in the same semantic space. Cosine similarity is used as the matching metric function to calculate the cosine similarity between the narrative content vector and the corresponding semantic tag keyword vector of the accessible spatial unit. Anchor-spatial unit matching pairs with cosine similarity values below a preset attraction threshold are marked as low-attractivity segments. For low-attractivity segments, the narrative adversarial component executes the following strategy: it attempts to retrieve other candidate anchors with higher matching degrees to the semantic tags of the spatial unit from the candidate entries in the decision strategy table of that accessible spatial unit. If a candidate with a higher matching degree is retrieved and this candidate is not occupied in the adjacent accessible spatial units of the currently adjusted wayfinding path, the original fictional anchor is replaced with the new candidate anchor, or a new fictional anchor is added to the adjacent spatial grid of that accessible spatial unit and assigned a corresponding high-matching narrative tag material. If no replacement candidate is available, the original anchor remains unchanged. After completing all segment judgments and processing, all modified or added virtual anchor point objects are packaged into an updated virtual anchor point set.
[0088] Step 1363: The environment-constructing agent continues to adjust the navigation path according to the updated set of fictional anchor points, and the narrative erosion agent continues to update the set of fictional anchor points according to the newly adjusted navigation path, repeatedly alternating between the adversarial adjustment process.
[0089] A closed-loop iterative process is formed: the updated set of fictitious anchor points output from step 1362 of the previous round is sent back to the environment-building agent via the inter-process communication channel, serving as the input for the new round of step 1361, driving the path to be adjusted again; while the newly adjusted navigation path output from step 1361 is fed back to the narrative erosion agent, triggering the update of the anchor point set in the new round of step 1362. Message interaction between the two agents adopts a preset request-response protocol, with the protocol message header containing the round number, the sender agent type identifier, and the message body checksum. During alternating execution, the round number is monitored in real time. When the round number exceeds the preset maximum number of game rounds, the iteration is forcibly terminated and the current result is output to prevent infinite loops.
[0090] Step 1364: After each round of confrontation, record the change in the node sequence of the adjusted wayfinding path and the information increment of the updated set of fictitious anchor points. When the change in the node sequence is lower than the preset change tolerance for several consecutive rounds and the information increment is lower than the preset increment tolerance for several consecutive rounds, it is determined that the preset equilibrium condition has been reached.
[0091] After each round of the game, the environment-building agent and the narrative-erosion agent calculate the change in their respective decision outputs, which are then aggregated to the game flow controller. The method for calculating the change in node sequences is as follows: Take the node identifier sequence of the adjusted wayfinding path in the current round and the node identifier sequence of the adjusted wayfinding path in the previous round. Use a dynamic programming algorithm to find the length of the longest common subsequence between the two sequences. Then calculate the node sequence change as 1 - (length of the longest common subsequence × 2) / (current round sequence length + previous round sequence length). The method for calculating information increment is as follows: Count the number of anchor points in the current round's fictional anchor point set that have changed compared to the previous round. Changes include three types of operations: addition, deletion, and narrative content modification. Take the quotient of the number of changes divided by the total number of fictional anchor points in the current round.
[0092] The game-playing process controller internally uses a counter that increments only when the change in the node sequence is below a preset change tolerance and the information increment is simultaneously below a preset increment tolerance; if either criterion is not met, the counter is reset to zero. When the counter value reaches a preset number of consecutive stable rounds, the adversarial game is considered to have reached a preset equilibrium condition. The preset change tolerance, preset increment tolerance, and preset number of consecutive stable rounds are all adjustable parameters read from the configuration file during system initialization.
[0093] Step 1365: The adjusted navigation path when the preset equilibrium condition is reached is taken as the game navigation path, and the updated set of virtual anchor points is output simultaneously as the target set of virtual anchor points.
[0094] When the equilibrium condition is determined, the game flow controller retrieves a complete copy of the latest adjusted guidance path's node sequence from the current memory space of the environment-constructing agent and persists it as the game guidance path. Simultaneously, it retrieves a complete copy of the latest updated set of fictional anchor points from the memory space of the narrative erosion agent and persists it as the target set of fictional anchor points. The game guidance path and the target set of fictional anchor points are packaged into the same game result data packet, along with procedural statistics such as the total number of game rounds, the round number where equilibrium was reached, and the final node sequence change and information increment value, and output to subsequent processing steps.
[0095] Step 140: Assign a matching fictional anchor from the target fictional anchor set to each guidance node in the game guidance path and generate a fictional anchor activation instruction, and combine the guidance node sequence and its associated fictional anchor activation instructions into a narrative guidance path.
[0096] This step integrates the game-generated navigation path and the set of fictional anchor points into an executable narrative navigation path. First, nodes are extracted sequentially from the game-generated navigation path to generate a sequence of identifiable navigation nodes; simultaneously, the spatial coordinates and narrative content of the target fictional anchor points are extracted. Then, based on the principle of minimizing spatial distance, a unique fictional anchor point is matched to each navigation node, and the corresponding presentation style template is determined according to the content type tags of the material (e.g., panoramic video, 3D model). Next, combining the material volume parameters and the estimated movement time between nodes, the playback duration is calculated, the activation start and end times are set, and this is encapsulated into a fictional anchor point activation instruction containing invocation, presentation, and timing control commands. Finally, navigation nodes are bound to activation instructions to construct a narrative navigation path data structure containing spatial location, directional guidance, and activation instructions. Timing compatibility checks are used to eliminate overlapping activation times of adjacent nodes, ensuring the orderly presentation of superimposed virtual information.
[0097] Step 141: Extract all path nodes in the game guidance path, generate guidance node identifiers for each path node in traversal order, and generate a guidance node sequence.
[0098] The game-themed wayfinding paths are stored as an ordered list, where each element is a path node object. Each path node object contains a passable spatial unit identifier field and the 3D coordinates of its geometric center in normalized space. The wayfinding node sequence generation process involves creating an empty wayfinding node sequence container and then traversing it element by element in the game-themed wayfinding path list. For the Kth path node, a wayfinding node description object is constructed. This description object contains an automatically generated wayfinding node identifier, the original passable spatial unit identifier and center 3D coordinates of the path node, and a sequence position number determined by the traversal index K. The wayfinding node identifier is generated using an encoding rule that concatenates a prefix string with a fixed-length zero-padding numeric sequence number. Each wayfinding node description object is sequentially appended to the wayfinding node sequence container. After complete traversal, the container contains the complete wayfinding node sequence.
[0099] Step 142: Extract the three-dimensional spatial coordinates of each fictional anchor point in the target fictional anchor point set and the corresponding fictional narrative content material.
[0100] Each fictional anchor object in the target fictional anchor set contains three core fields: a candidate location coordinate field, stored as a normalized 3D coordinate triplet; a unique material identifier field, stored as a string; and a fictional narrative content material reference field, which can be used to retrieve the complete material data package from the material management database. This step iterates through the target fictional anchor set, extracting all candidate location coordinates into a location coordinate array, extracting the unique material identifiers into a material identifier array, and storing the fictional narrative content material data packages returned from the material management database in a batch query based on the material identifiers into a material data package array, ensuring that the three arrays are linked by the same index.
[0101] Step 143: For each wayfinding node in the wayfinding node sequence, determine the spatial distance between the spatial coordinates of the wayfinding node and the spatial coordinates of each virtual anchor point, and take the virtual anchor point with the smallest spatial distance as the initial matching virtual anchor point of the wayfinding node.
[0102] For the m-th wayfinding node in the wayfinding node sequence, obtain the coordinate position (Xnd) of the wayfinding node in the normalized space from its description object. m Ynd m Znd m Traverse the array of position coordinates. For the nth virtual anchor point in the array, its coordinates are (Xanc...). n Yanc n Zanc n Calculate the square of the spatial distance between the two. Find the value of DistSq for the current wayfinding node m. mn A fictitious anchor point n that obtains the minimum value min Record this fictitious anchor point as the initial matching fictitious anchor point of the wayfinding node m, and simultaneously record DistSq. mn_min As a matching cost, the above pairing operation is performed on all wayfinding nodes in the wayfinding node sequence to form a mapping table from wayfinding node identifiers to the initial matched fictitious anchor points.
[0103] Step 144: For the initial matched fictional anchor point, obtain the content type tag and content volume parameter of the fictional narrative content material of the initial matched fictional anchor point, and determine the presentation style template of the fictional narrative content material when it is superimposed in the virtual scene according to the content type tag.
[0104] Extract the fictional anchor object for each matching pair from the initial matching fictional anchor mapping table obtained in step 143, and read the metadata information of the fictional narrative content material from the corresponding index position in the material data package array. The metadata information includes a content type tag field and a content volume parameter field. The content type tag field is a predefined enumerated string, which can take several types, such as panoramic video, 3D model animation, or 2D graphic pop-up. The content volume parameter field contains two values: the estimated total playback duration of the material and the material data volume. The presentation style template library is a configuration library that stores presentation templates that correspond one-to-one with the content type tags.
[0105] The panoramic video type corresponds to the dome overlay template, which specifies that virtual information is rendered as a dome texture map onto the inner surface of a virtual sphere centered at a fictional anchor point. The 3D model animation type corresponds to the spatial anchoring instantiation template, which specifies that the 3D model is instantiated at the fictional anchor point with the scene's local coordinate system as the reference origin and the skeletal animation plays in a loop. The 2D graphic pop-up type corresponds to the screen space following template, which specifies that graphic information is rendered as a transparent panel in the user's screen space and the panel's orientation adjusts with the viewpoint. The corresponding presentation style template is indexed based on the content type tag, and the template identifier and template parameter structure are written into the context of the fictional anchor point activation instruction construction process.
[0106] Step 145: Based on the content volume parameters and the estimated path time between navigation nodes, calculate the playback duration of the fictional narrative content material, and combine the playback duration to set the activation start time and activation end time to obtain the time sequence activation parameter group.
[0107] The estimated total playback duration of the material in the content volume parameter represents the time required for complete playback. For the current node and its successor nodes in the guide node sequence, the Euclidean distance between the center coordinates of the two nodes is calculated as the path length, which is then divided by the user's preset average browsing speed to obtain the estimated path movement time. This time is compared with the estimated total playback duration of the material, and the smaller value is taken as the playback duration of the fictional narrative content material for that node; if the current node is the last node in the sequence, the estimated total playback duration of the material is directly used as the playback duration. The activation start time is set to the moment when the user arrives at the boundary of the accessible space unit where the node is located and their line of sight is towards the guide node, and the activation end time is the start time plus the playback duration. These two are encapsulated into a timing activation parameter group to guide the timing scheduling of the rendering pipeline.
[0108] Step 146: Encapsulate the presentation style template and the timing activation parameter group into a fictional anchor activation instruction for the navigation node. The fictional anchor activation instruction includes instructions for calling fictional narrative content materials, instructions for presentation in a visual form, and instructions for timing control in a visual form.
[0109] The fictional anchor activation command is represented by a structured command data body, which is divided into three field groups: the invocation command field group stores the unique identifier of the fictional narrative content material and the storage path of the material file, which is used by the rendering engine to read and load the actual content resources; the presentation command field group stores the presentation style template identifier and template parameter structure serialization data obtained from step 144, and the template parameters include the overlay position offset, scaling factor, and opacity parameter; the timing control command field group stores the timing activation parameter group obtained from step 145, specifically including the activation start time, activation end time, and playback loop marker. Each guide node identifier is bound to its corresponding fictional anchor activation command data body.
[0110] Step 147: Based on the navigation node sequence and the fictional anchor activation instruction, determine the narrative navigation path data structure, and generate the narrative navigation path after completing the temporal compatibility review of the narrative navigation path data structure.
[0111] Step 1471: Bind each waypoint in the waypoint node sequence to the virtual anchor point activation instruction corresponding to that waypoint node to generate an activation instruction association mapping table.
[0112] Using the guide node identifier as the mapping key and the serialized data block of the virtual anchor activation instruction data body generated in step 146 as the mapping value, a key-value pair mapping structure is constructed. This mapping structure is also appended with a list of keys arranged in the sequence order of the guide nodes as a sequential index to retain the node traversal order information.
[0113] Step 1472: According to the arrangement order of the guide node sequence, the data bound in the activation instruction association mapping table is expanded sequentially to generate a narrative guide path data structure. The narrative guide path data structure records the position of each guide node in space, the direction of the next guide node, and the virtual anchor point activation instruction to be executed at that guide node.
[0114] Create a top-level data structure for the narrative wayfinding path, containing an array of wayfinding node entries, the length of which equals the number of nodes in the wayfinding node sequence. Iterate through the wayfinding node sequence sequentially, filling each wayfinding node entry according to its position in the sequence. Each entry contains: a current position field, filled with the 3D coordinates of the wayfinding node; a direction indicator field, calculated and stored based on the difference between the coordinates of the current wayfinding node and the coordinate vectors of its successor; and an activation command field, retrieved and filled using an activation command association mapping table. Fields are left blank if the first node has no predecessor and the last node has no successor.
[0115] Step 1473: Perform a time-series compatibility review on the narrative navigation path data structure. When it is detected that the activation instructions of the virtual anchor points bound to adjacent navigation nodes overlap and conflict in the activation time window, adjust the activation start time of the latter navigation node to after the activation end time of the former navigation node to eliminate the time conflict of virtual information superposition.
[0116] Iterate through the array of wayfinding node entries in the narrative wayfinding path data structure. For the p-th wayfinding node entry and the p+1-th wayfinding node entry, extract the timing activation parameter group from their respective activation instruction fields. Compare the activation end time of the previous wayfinding node with the activation start time of the next wayfinding node on the timeline. If the activation end time is later than the activation start time, a time window overlap conflict is determined. The handling of the overlap conflict is: reassign the activation start time in the activation instruction of the next wayfinding node to the sum of the activation end time of the previous wayfinding node and a preset minimum time interval constant, and correspondingly postpone the activation end time of the next wayfinding node by the same time offset. Perform the review and adjustment operation on all adjacent node pairs in sequence until the time windows of all adjacent nodes have no overlap. Persist in storing the narrative wayfinding path data structure after the timing compatibility review, thus generating the narrative wayfinding path.
[0117] Step 150: In the target virtual reality scene, the navigation interaction interface is rendered based on the narrative navigation path, and the system moves to each navigation node in sequence. When a navigation node is reached, the corresponding virtual information content is retrieved according to the activation command of the virtual anchor point associated with the navigation node. The virtual information content is superimposed on the corresponding spatial position of the scene, and the rendered screen with virtual information is output to the user's virtual reality display device.
[0118] Step 151: Start the main rendering loop of the target virtual reality scene and create a wayfinding interaction interface container in the user interface layer of the rendering pipeline. The wayfinding interaction interface container is used to carry wayfinding indicator lines and virtual information overlay layers.
[0119] The rendering pipeline is built on the graphics processor hardware interface, and its rendering process is divided into scene geometry layer, lighting and shadow layer, and user interface layer according to the rendering order. In the user interface layer initialization phase of the rendering pipeline, a wayfinding interactive interface container object is instantiated. This container object consists of two rendering surfaces: the lower layer is the wayfinding line rendering surface, configured to use a self-illuminating shader without lighting, supporting variable line width and dynamic color gradient characteristics; the upper layer is the virtual information overlay layer, configured to support alpha blending and a transparent rendering surface with depth testing turned off. Its rendering order is arranged after the scene geometry layer rendering is completed to ensure that the virtual information is always visible.
[0120] Step 152: Read the sequence of navigation nodes in the narrative navigation path, and generate a three-dimensional navigation indicator line in the virtual scene that points from the user's current location to the first navigation node in the navigation node sequence, based on the user's current location and the first navigation node in the navigation node sequence.
[0121] Load the array of wayfinding node entries from the narrative wayfinding path data structure, and extract the first element as the target wayfinding node. Obtain the current viewpoint coordinates of the user's virtual avatar as the starting point of the wayfinding line, and the current position coordinates of the target wayfinding node entry as the ending point of the wayfinding line. Generate a Bézier curve between the starting and ending points as the geometry of the wayfinding wayfinding line. The control point setting strategy for the Bézier curve is as follows: at the starting point, extend a preset starting distance in the positive direction along the user's current line of sight vector as the first control point; at the ending point, extend a preset ending distance in the negative direction along the unit direction vector pointing from the target wayfinding node to its subsequent wayfinding nodes as the second control point. Discretize the Bézier curve into a sequence of polyline segments and upload it to the vertex buffer of the graphics processor.
[0122] Step 153: Control the virtual camera to translate the viewpoint along the direction indicated by the three-dimensional guide line according to the preset smooth movement strategy, and move the virtual camera from the current viewpoint position to the spatial position of the first guide node.
[0123] The virtual camera's movement is controlled by an easing function, which employs a smooth start to suppress sudden acceleration. The total translation distance is divided into time slices, and the virtual camera's translational displacement and orientation within each time slice are interpolated based on the easing function's output parameters and the tangent direction of the guide lines. Simultaneously, the field of view is dynamically adjusted: during translational acceleration phases, the field of view is slightly reduced to minimize dizziness, and during deceleration phases, the field of view is restored.
[0124] Step 154: When the virtual camera arrives at the first guide node, trigger the virtual anchor point activation command bound to the guide node, and extract the virtual information content identifier to be called, the presentation style template, and the timing activation parameter group from the virtual anchor point activation command.
[0125] The rendering pipeline detects the distance between the virtual camera's position coordinates and the current target guide node's coordinates in each frame. When the distance is less than a preset trigger distance threshold and the virtual camera's motion is in a deceleration and stop phase, it determines that the guide node has been reached. The corresponding virtual anchor activation command data body is queried in the activation command association mapping table using the guide node identifier. The virtual information content identifier in the call command field group, the presentation style template identifier in the presentation command field group, and the timing activation parameter group in the timing control command field group are then parsed and extracted.
[0126] Step 155: Search for the corresponding virtual information content data package in the virtual information material repository based on the virtual information content identifier, and determine the display format of the virtual information content data package based on the presentation style template.
[0127] The virtual information material repository contains pre-stored data packages of various advertising content. Using the virtual information content identifier extracted in step 154 as the query key, the corresponding virtual information content data package is retrieved from the repository and asynchronously loaded into memory. Based on the presentation style template identifier, the corresponding rendering instruction set is read: if it is a dome overlay template, the virtual information content is rendered as the inner surface texture of a sphere centered at the anchor point; if it is a spatial anchoring instantiation template, a 3D model is loaded at the anchor point and skeletal animation is played; if it is a screen space following template, the position of the virtual information panel in the screen coordinate system is calculated, causing it to continuously track and follow the user's gaze direction.
[0128] Step 156: Based on the time-series activation parameter group, the virtual information content data packet is rendered at the corresponding spatial location in the virtual scene. The virtual information content data packet is then rendered and mixed with the original geometric data and lighting information of the scene to generate a virtual information overlay frame.
[0129] Upon activation, for dome and spatial anchoring templates, a rendering command for the 3D model is submitted to the rendering pipeline. During rendering, template testing is enabled to handle occlusion between the model and scene geometry. For screen space templates, textures are mapped to quadrilateral primitives on transparent rendering surfaces, and blending factors are calculated to make the virtual information edges semi-transparent. In the rendering blending phase, the rendered virtual information is blended pixel-by-pixel with the already rendered scene color buffer, with the blending weights correlated to the opacity parameter.
[0130] Step 157: During the virtual information presentation process, continuously receive user eye tracking data, parse the user eye direction vector, and calculate the deviation angle between the user eye direction vector and the three-dimensional guide line.
[0131] The eye-tracking component continuously sends the user's current gaze direction vector as a data stream, with its data frame rate independent of the rendering frame rate. In each rendering frame, the most recently received gaze direction vector is acquired and its dot product is performed with the direction vector from the user's viewpoint to the nearest point on the wayfinding line. The inverse cosine of the dot product result is then used to obtain the deviation angle.
[0132] Step 158: When the deviation angle is detected to exceed the preset visual attraction threshold, adjust the display attributes of the three-dimensional guide line to enhance visual stimulation.
[0133] If the deviation angle value exceeds the preset visual attraction threshold for multiple drawing frames, the indicator line visual enhancement mechanism is triggered. Adjustments to display properties include: increasing the brightness multiplier of the indicator line material's luminous color channel; increasing the scrolling rate of the dynamic texture in the indicator line vertex shader; and increasing the particle jet density of the halo particle emitters around the indicator line. Property adjustments are achieved by updating shader uniform variables to the graphics processor.
[0134] Step 159: When the virtual camera leaves the current navigation node and moves to the next navigation node, the triggering and overlay operations are repeatedly executed according to the activation instruction of the virtual anchor point bound to the next navigation node, until all navigation nodes in the narrative navigation path are traversed, and the rendered screen is continuously output to the user's virtual reality display device.
[0135] When it is detected that the next target node has switched to the next navigation node in the narrative navigation path, the current virtual information content resources are unloaded, the occupied memory is released, and the resource loading and rendering process corresponding to the virtual anchor activation instruction of the next navigation node is triggered in sequence. This process is repeated until all navigation nodes have been traversed and the corresponding virtual information has been presented.
[0136] In an optional embodiment, after step 150, the method further includes:
[0137] Step 210: Obtain the eye-tracking data stream fed back by the user's virtual reality display device, parse the user's gaze dwell coordinate sequence and pupil diameter fluctuation sequence from the eye-tracking data stream, and extract the superimposed virtual information content in the rendered image corresponding to the user's gaze dwell coordinate sequence.
[0138] The raw data of user eye movement includes the three-dimensional coordinates of the origin of the gaze, the unit vector of the gaze direction, and the pupil diameter in millimeters. The raw data is denoised, and a moving median filter is used to eliminate transient anomalies caused by blink artifacts. By calculating the intersection of the gaze direction and the scene depth map, the gaze direction is converted into the coordinates of the user's gaze dwell point in the three-dimensional scene, arranged along the time axis to form a sequence of user gaze dwell coordinates. The pupil diameter data is interpolated and aligned with timestamps to form a pupil diameter fluctuation sequence. The correspondence between each coordinate point in the user gaze dwell coordinate sequence and the world space location of the virtual information overlay rendered in step 150 is queried. Overlayed virtual information content whose gaze dwell point falls within the presentation space of a certain virtual information content and whose dwell time exceeds the shortest duration for gaze discrimination is extracted as the content to be focused on.
[0139] Step 220: Determine the narrative immersion response characteristics of the current user to the superimposed virtual information content based on the user's gaze dwell coordinate sequence and the pupil diameter fluctuation sequence, and filter out inactive virtual anchors in the target virtual anchor set whose immersion response intensity does not reach the preset excitation threshold according to the narrative immersion response characteristics.
[0140] Narrative immersion response characteristics are measured from two dimensions: Spatially, the ratio of the total time a user's gaze lingers within the overlaid virtual information content's presentation area to the total duration of the entire presentation interval is used as attention concentration; physiologically, the deviation of the average pupil diameter fluctuation within the presentation interval from the baseline pupil diameter is used as cognitive load intensity. These two dimensions are then weighted and fused, with a weighting bias towards the physiological dimension, to form the immersion response intensity value. Virtual anchors that have not yet been assigned or have not actually triggered activation commands are extracted from the target set of virtual anchors. The theoretical immersion response intensity estimate for these anchors if they were located in the current user's spatial position is calculated. Anchors with estimated values below a preset activation threshold are marked as inactive virtual anchors.
[0141] Step 230: Based on the spatial position of the inactive virtual anchor point in the target virtual reality scene, retrieve the spatial connectivity topology graph in the environment memory model and generate a supplementary navigation path segment connecting the current user's spatial position with the inactive virtual anchor point.
[0142] Obtain the identifier of the currently accessible spatial unit where the user is located and the identifier of the accessible spatial unit to which the coordinates of the inactive virtual anchor point belong. Perform path search in the spatial connectivity topology graph using these two identifiers as start and end nodes to generate supplementary navigation path fragments. The path fragments contain an ordered list of accessible spatial unit identifiers and the length of each segment.
[0143] Step 240: Inject the supplementary wayfinding path fragment into the narrative wayfinding path to obtain an enhanced narrative wayfinding path, and update the wayfinding node sequence in the wayfinding interaction interface container according to the enhanced narrative wayfinding path to guide the current user to move along the supplementary wayfinding path fragment and trigger the virtual information content corresponding to the inactive fictional anchor point.
[0144] The supplementary wayfinding path fragments are appended to the head of the remaining untraversed node sequence of the narrative wayfinding path according to their insertion positions, forming a new node sequence, i.e., the enhanced narrative wayfinding path. The wayfinding node sequence currently in use by the wayfinding interaction interface container is updated with the enhanced narrative wayfinding path, so that the wayfinding indicator lines and virtual camera movement paths in the next stage automatically guide to the supplementary wayfinding path fragments.
[0145] As an example, after step 150, the method further includes:
[0146] Step 310: Collect hand posture sensing data associated with the user's virtual reality display device and perform hand skeletal joint tracking and gesture action segmentation to obtain gesture interaction segments and their corresponding hand movement trajectory features and hand shape change features during the activation period of each superimposed virtual information content in the rendered screen.
[0147] The hand's joints are captured in 3D rotation and translation data at each moment using a controller or hand-tracking camera to construct a hand skeleton tree pose sequence. An action segmentation algorithm based on velocity thresholding and pause detection is used to divide the continuous hand skeleton tree pose sequence into independent gesture action segments according to the joint velocity zero-crossing points and local minima. The start and end timestamps are cross-referenced with the activation periods of each virtual information content to obtain a set of gesture interaction segments within the activation period of each superimposed virtual information content. Motion trajectory features and hand shape change features are extracted for each gesture interaction segment: motion trajectory features include the curvature maxima and total length of the gesture translation trajectory; hand shape change features include the amplitude of finger bending changes and the temporal coding sequence of the five-finger opening and closing states.
[0148] Step 320: Match the hand movement trajectory features and the hand shape change features with a preset gesture semantic mapping table to obtain narrative feedback semantic tags that characterize the user's interest type and interest response level to the corresponding superimposed virtual information content.
[0149] The pre-defined gesture semantic mapping table is a lookup table that maps typical hand kinematic features to several narrative feedback semantic tags. During matching, the weighted Euclidean distance between the motion trajectory features and hand shape change features of the current gesture interaction segment and the feature vectors of each entry in the mapping table is calculated. The entry with the smallest distance is selected as the matching result, and the corresponding interest tendency type and interest response level quantification value are output.
[0150] Step 330: Collect all narrative feedback semantic tags of the superimposed virtual information content, and generate a user interest offset vector with the interest tendency type as the dimension and the statistical distribution of the interest response degree as the value.
[0151] Using all possible interest types as vector dimension coordinate axes, calculate the weighted average and standard deviation of the interest response degree corresponding to all narrative feedback semantic tags belonging to that dimension for each dimension. Use the weighted average as the offset of that dimension and the reciprocal of the standard deviation as the confidence weighting coefficient of that dimension to form the user interest offset vector.
[0152] Step 340: Extract the user interest space mapping function from the fictional anchor point generation strategy of the narrative erosion agent, and use the user interest offset vector to adjust the mapping weights of interest categories and spatial grids in the user interest space mapping function to generate an adaptive interest mapping function.
[0153] The normalized offsets of each dimension of the user interest offset vector are used as adjustment factors. For the spatial response intensity distribution matrix corresponding to each interest category in the original user interest space mapping function, the linear stretching coefficient of the interest category adjustment factor is multiplied by the matrix elements corresponding to each interest category. The stretching coefficient is determined by the sign and absolute value of the offset. Positive offsets amplify the intensity, while negative offsets compress the intensity. The mapping function is then reconstructed after adjustment.
[0154] Step 350: Traverse each candidate anchor spatial grid in the virtual anchor candidate location library, calculate the expected activation intensity of each candidate anchor spatial grid through the adaptive interest mapping function, and select candidate anchor spatial grids with expected activation intensity higher than the preset narrative trigger threshold to form an updated candidate anchor set.
[0155] Re-traverse the virtual anchor candidate location library constructed in step 129, use the adaptive interest mapping function to calculate the new expected activation intensity value of each candidate anchor for the current user interest state, filter out candidate anchors with intensity values higher than the system's preset narrative trigger threshold, and form an updated candidate anchor set.
[0156] Step 360: Send a re-game trigger command to the environment building agent, pass the updated candidate anchor point set and the user interest offset vector to the environment building agent, and jointly generate a reconstructed narrative navigation path with the narrative erosion agent to replace the narrative navigation path in the current navigation interaction interface container for continued rendering.
[0157] Send a re-game instruction to trigger the environment building agent and the narrative erosion agent to re-execute a round of mutual game iteration with the current user's latest spatial location, the updated candidate anchor point set, and the user's interest offset vector as input. The final output is to reconstruct the narrative navigation path and replace the current path container.
[0158] Optionally, after step 150, the method further includes:
[0159] Step 410: Assign a timestamp to each rendering frame of the rendering sequence output to the user's virtual reality display device, parse the virtual information content identifier contained in the virtual information overlay in each rendering frame, determine the activation interval of each virtual information content by comparing the appearance and disappearance boundaries of the virtual information content identifier across frames, calculate the actual experience duration based on the timestamp difference of the activation interval to obtain the actual experience duration distribution, and sort according to the start time of the activation interval to obtain the actual experience sequence.
[0160] The rendering pipeline iterates through the sequence of rendered frames, reading all virtual information content identifiers contained in each frame from its rendering metadata. Using each unique virtual information content identifier as an index, the timestamps of its first and last appearances in the frame sequence are recorded as activation intervals, and the difference between these two timestamps is taken as the actual experience duration. The actual experience sequence is obtained by sorting all virtual information content activation interval start timestamps in ascending order, and then grouping each actual experience duration by virtual information content category to obtain the actual experience duration distribution.
[0161] Step 420: Retrieve historical narrative experience records that match the target virtual reality scene identifier from the historical user behavior data and perform pattern mining to obtain the historical experience sequence and historical experience duration distribution.
[0162] The retrieval of historical narrative experience records is filtered using target virtual reality scene identifiers and time windows. For the filtered records, multiple sets of historical experience sequence sequences and historical experience duration distributions under different time windows are obtained using the same method as in step 410. The multiple sets of data are merged, and a consensus historical experience sequence sequence is constructed according to the probability of occurrence of each virtual information content identifier. The historical experience duration distribution is statistically analyzed according to the average duration of each virtual information content identifier.
[0163] Step 430: The actual experience sequence is compared with the historical experience sequence using sequence editing cost to output the experience sequence offset. The actual experience duration distribution is compared with the historical experience duration distribution using distribution difference metric to output the experience duration offset vector. The two are then fused to obtain the user experience mode offset feature.
[0164] The experience sequence offset is obtained by normalizing the sequence edit distance by the length of the historical experience sequence. The value of each dimension of the experience duration offset vector is the difference between the actual experience duration and the historical average duration of the corresponding virtual information content identifier, divided by the historical duration standard deviation. The fusion operation adds the normalized offset as an additional dimension to the vector and concatenates it horizontally with the experience duration offset vector to obtain the user experience pattern offset feature.
[0165] Step 440: Retrieve the pairing relationship between the candidate location library of fiction anchor points and the narrative tag material in the fiction anchor point generation strategy of the narrative erosion agent. According to the offset semantic direction indicated by the user experience mode offset feature, replace the original narrative tag material in the pairing relationship that is positively correlated with the offset semantic direction with the corrected narrative tag material with opposite narrative attributes, generate the corrected pairing relationship and write it back to the fiction anchor point generation strategy to update the strategy parameters.
[0166] The semantic offset direction is analyzed, specifically the semantic feature vectors of the virtual information content categories with the largest positive offset in experience duration and the semantic feature vectors of the categories with the largest negative offset in experience duration. The average difference between these two semantic vectors is used as the offset semantic vector. For each material in a pairing relationship, the cosine similarity between its semantic vector and the offset semantic vector is calculated. If the similarity is higher than the positive association threshold, a new material with a semantic vector opposite to that of the paired material is selected from the material library for replacement. After replacing all relevant pairs, the results are written back to the decision strategy table.
[0167] Step 450: The narrative erosion agent and the environment construction agent are triggered to re-execute the mutual game-playing adversarial iteration based on the updated fictional anchor point generation strategy, with the current user space location and the navigation task flow as input. After generating the regression narrative navigation path, it is loaded into the navigation interaction interface container to execute the narrative navigation rendering of the next cycle.
[0168] The narrative erosion agent is notified that the fictional anchor point generation strategy has been changed. After reloading the decision strategy table, the narrative erosion agent, together with the environment building agent, re-executes the complete process of steps 130 to 150 to generate a regression narrative navigation path that replaces the current path and puts it into rendering.
[0169] In this embodiment of the invention, the game-playing virtual agent pair is the core inventive point, and the relevant algorithms of the environment-building agent and the narrative erosion agent it contains are described below.
[0170] The core algorithm architecture for the environment-building intelligent agent adopts a combination of a hierarchical spatial encoder based on graph attention networks and a sequential path planner based on policy gradients. This hierarchical spatial encoder consists of three stacked graph attention operations. Each layer of graph attention operations performs weighted message passing on each node in the spatially connected topology graph based on the features of its neighboring nodes. The input to the first layer of graph attention operations is the initial geometric feature vector of each traversable spatial unit, including the geometric center coordinates, spatial volume, and diagonal length of the bounding box of the spatial unit. After linear transformation, this vector is mapped to a hidden layer vector of dimension d_e and then enters the attention aggregation. The attention weight is obtained by concatenating the source node vector and the target node vector and calculating the normalized attention score through a single-layer feedforward network. After aggregating the neighboring node information, each node obtains an updated node representation through residual connections and layer normalization. The second layer of graph attention operations introduces spatial edge features as attention bias terms. Specifically, the edge features are normalized values of edge weights and channel widths. When calculating the attention score, the edge features are linearly transformed and added to the node attention score, enabling the graph attention operation to perceive the strength of the traversal cost of edges. The third layer of graph attention operation adopts a multi-head attention mechanism, which executes H independent attention operations in parallel. The node representations output by each group are concatenated along the feature dimension and then restored to the final d_e-dimensional node encoding vector through the output linear layer. The node encoding vectors output by these three layers of graph attention operation contain not only the geometric attributes of each spatial unit itself, but also the topological relationship and semantic context with the surrounding area, which together constitute the spatial structure encoding parameters after training.
[0171] The sequence path planner employs a decoder network based on a long short-term memory (LSTM) architecture. At each decoding time step, the decoder network takes into account the node encoding vector of the current spatial unit and the sequence of generated partial path nodes. It maintains the historical state latent vector of path planning through the LSM unit, and the attention pointer module calculates the selection probability distribution on the node encoding vectors of all currently unvisited spatial units. The attention pointer module uses the current latent vector output by the LSM unit as the query vector and performs a contracted dot product attention calculation with the node encoding vectors of all candidate spatial units. The probability value of each candidate spatial unit being selected as the next path node is output through a flexible maximum function.
[0172] During training, a strategy combining self-supervised learning and reinforcement learning is employed: In the pre-training phase, the optimal path of A* search for randomly sampled start-end point pairs in the scene is used as the supervision label. The cross-entropy loss function is used to clone the behavior of the sequence path planner, making the path probability distribution output by the decoder network approximate the one-hot encoding distribution of the nodes on the optimal path. The learning rate in this phase decays from the initial value using cosine annealing scheduling, training continues until the path accuracy on the validation set of random start-end point pairs covered by the spatial connectivity topology graph converges. In the reinforcement learning fine-tuning phase, a weighted combination of semantic richness score and path cost is used as the reward function. Specifically, the reward value equals the semantic richness score of the spatial units traversed by the path minus the total path cost multiplied by the penalty coefficient. The policy gradient algorithm is used to update the parameters of the sequence path planner. A baseline value is introduced during gradient estimation to reduce variance. The baseline value is output by a value network with the same structure as the policy network. The value network takes the node encoding vector of the current spatial unit and the path state as input and outputs an estimated value of the expected cumulative reward. The value network and the policy network are trained and updated synchronously using the mean squared error loss function.
[0173] The parameter configurations after training include setting the number of graph attention layers of the hierarchical spatial encoder to 3, the hidden layer dimension d_e to a preset value, the number of multi-head attention heads H to a preset value, the hidden state dimension of the long short-term memory unit to a preset value, and independent configuration values of the same order of magnitude as d_e. The temperature coefficient of the attention pointer module of the sequence path planner is fixed to a low value during the inference phase to ensure the determinism and reproducibility of path selection.
[0174] II. The core algorithm architecture of the narrative erosion agent adopts a combined design of a user interest encoder based on a deep interest network and an anchor decision-maker based on a dual deep Q-network. The user interest encoder consists of three sequentially connected components: an input feature layer, an interest embedding layer, and an attention pooling layer. The input feature layer receives user interaction sequences divided into time windows from historical user behavior data. Each sequence element is in the form of a triple, containing the spatial grid index of the triggering operation, the triggering semantic label category identifier, and the normalized value of the dwell time. The input feature layer performs embedding operations on the spatial grid index and the semantic label category identifier respectively, mapping the discrete index to a dense embedding vector of dimension d_i. The normalized dwell time is used as a scaling factor and multiplied element-wise with the embedding vector to output a weighted interaction vector sequence.
[0175] The interest embedding layer contains multiple interest capsule units. Each interest capsule unit maintains a learnable interest center vector. For each vector in the input weighted interaction vector sequence, the interest embedding layer calculates the cosine similarity between the vector and the interest center vectors of each interest capsule, and dynamically routes the interaction vectors to the corresponding interest capsules based on the similarity. The number of routing iterations is fixed at a preset number of rounds. In each round of routing, the interest center vectors of the interest capsules are updated by weighting the routing weights. Finally, each interest capsule outputs an interest representation vector. The interest representation vectors of all interest capsules are concatenated along the feature dimension to obtain a multi-interest representation matrix.
[0176] The attention pooling layer uses the user's navigation task flow features in the current scene as the query vector. These features are obtained by transforming the spatial distribution statistical vector of the target point sequence in the navigation task flow through a fully connected layer. In the attention pooling layer, the query vector interacts with the multi-interest representation matrix, calculating the attention correlation score between the interest capsule's interest representation vector and the current navigation task. The multi-interest representation matrix is then weighted and summed to output a user interest context vector. This user interest context vector integrates the entangled information of the user's historical behavioral preferences and the current navigation task objective. After training, the number of interest capsules in the interest embedding layer is set to a preset value, the embedding dimension d_i is set to a preset value, and the number of routing iteration rounds is set to a preset value. The anchor point decision-maker is built based on a dual-deep Q-network architecture, specifically comprising two sets of feedforward neural networks with identical structures and asynchronously updated parameters: an online Q-network and a target Q-network. The online Q-network receives a state vector, which is a concatenation of the user interest context vector, the semantic label statistics vector of the traversable spatial units along the current navigation path, and the coverage index of the currently activated set of virtual anchors. After processing through two fully connected hidden layers, it outputs the action value of each candidate anchor spatial grid. The activation functions of the fully connected hidden layers are all non-linear activation functions, and the number of hidden layer neurons decreases layer by layer. The target Q-network has the same structure as the online Q-network but its parameter updates are delayed. During training, the online Q-network is updated using gradient descent with a mean squared error loss function. The target value in the loss function uses the target Q-network's estimate of the action value of the next state. The specific calculation of the temporal difference target is the immediate reward plus a decay factor multiplied by the maximum action value output by the target Q-network. The immediate reward is defined comprehensively based on the user interaction strength and path coverage improvement after the virtual anchor is activated. The interaction strength is quantified by the user's gaze dwell time and spatial proximity after the virtual anchor is triggered, and the path coverage is measured by the increase in the number of semantic label categories traversed by the path after inserting the virtual anchor.
[0177] Training employs an experience replay mechanism, randomly sampling mini-batch state transition quadruples from the experience buffer for parameter updates. At the end of each training step, the parameters of the online Q-network are weighted into the target Q-network parameters using a soft update method with a preset scaling factor. After training convergence, the anchor decision-maker parameters are configured as follows: the number of fully connected hidden layers in both the online Q-network and the target Q-network is 2; the number of neurons in each layer decreases from the input state dimension to the action space dimension; the experience buffer capacity is a preset value; the decay factor is a preset value; and the soft update scaling factor is a preset value. During the inference phase, the anchor decision-maker uses only the online Q-network to determine the selection and placement of fictitious anchors using a greedy action selection strategy.
[0178] This invention constructs a game-theoretic virtual agent pair comprising an environment-building agent and a narrative-eroding agent. It uses spatial structure information, semantic tag distribution information, and wayfinding trigger point location information from scene scan data sets as the basis for the game, creating a dynamic adversarial relationship between the environment memory model generated by the environment-building agent and the fictional anchor point generation strategy determined by the narrative-eroding agent. By combining the current user's wayfinding task flow and historical user behavior data, this game-theoretic virtual agent pair iteratively modifies the initial wayfinding path and updates the set of fictional anchor points until equilibrium is reached, overcoming the limitations of static path planning. The output game-theoretic wayfinding path and the target set of fictional anchor points achieve deep coupling between physical space wayfinding logic and enhanced narrative content. Furthermore, by allocating and activating the target set of fictional anchor points to wayfinding nodes to form a narrative wayfinding path, and accurately retrieving overlaid virtual information content when the user arrives, it presents the user with a highly immersive, narratively coherent, and clearly defined wayfinding intent-driven fused rendering, enhancing the intelligence and personalization of wayfinding interaction in virtual reality scenes.
[0179] Please see Figure 2 The figure is a schematic diagram of the basic structure of a wayfinding system interaction server 200 provided in an embodiment of this application. The wayfinding system interaction server 200 includes: a processor 201; a storage device 202 on which a computer program 2020 is stored; and a network interface 203 for providing network communication functions. When the computer program 2020 is executed by the processor 201, the processor 201 implements any of the VR-based wayfinding system interaction methods described above.
[0180] Please see Figure 3 This application provides a functional block diagram of a wayfinding system interactive device, which includes:
[0181] The scanning data acquisition module is used to acquire a set of scene scanning data of the target virtual reality scene, and to parse spatial structure information, semantic tag distribution information and guide trigger point location information from the set of scene scanning data;
[0182] The agent pair creation module is used to create a virtual agent pair that engages in a game between an environment-building agent and a narrative-eroding agent. It generates an environment memory model based on the spatial structure information and configures it to the environment-building agent. It also determines a fictional anchor point generation strategy based on the semantic tag distribution information and configures it to the narrative-eroding agent.
[0183] The adversarial game-driving module is used to combine the scene scanning data set, the current user's guidance task flow, and historical user behavior data to generate an initial guidance path and a set of fictional anchor points through the adversarial game-playing virtual intelligent agent pairs. It drives the adversarial game-playing of the adversarial game-playing virtual intelligent agent pairs so that the environment building intelligent agent corrects the initial guidance path based on the set of fictional anchor points and the narrative erosion intelligent agent updates the set of fictional anchor points based on the corrected guidance path. The module iterates alternately until a preset equilibrium condition is met, and then outputs the game-playing guidance path and the target set of fictional anchor points.
[0184] The wayfinding path combination module is used to assign a matching fictional anchor from the target fictional anchor set to each wayfinding node in the game wayfinding path and generate a fictional anchor activation instruction, and combine the wayfinding node sequence and its associated fictional anchor activation instructions into a narrative wayfinding path.
[0185] The VR navigation rendering module is used to perform navigation interaction interface rendering processing based on the narrative navigation path in the target virtual reality scene, and move to each navigation node in sequence. When a navigation node is reached, the corresponding virtual information content is retrieved according to the virtual anchor point activation instruction associated with the navigation node, the virtual information content is superimposed on the corresponding spatial position of the scene, and the rendered screen with virtual information is output to the user's virtual reality display device.
[0186] Based on the above, a readable storage medium is provided, on which a program or instructions are stored, and when the program or instructions are executed by a processor, the steps of the above method are implemented.
[0187] Furthermore, it should be noted that this application also provides a computer program product, which may include a computer program that can be stored in a computer-readable storage medium. The processor of the wayfinding system interaction server reads the computer program from the computer-readable storage medium, and the processor can execute the computer program, causing the wayfinding system interaction server to perform the aforementioned... Figure 1 The methods described in the corresponding embodiments are already known, and therefore will not be repeated here. Furthermore, the beneficial effects of using the same method will also not be repeated. For technical details not disclosed in the computer program product embodiments related to this application, please refer to the description of the method embodiments of this application.
[0188] It should be noted that the various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For the systems or apparatus disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the descriptions are relatively simple, and relevant parts can be referred to the method section.
Claims
1. A VR-based wayfinding system interaction method, characterized in that, include: Obtain a set of scene scan data of the target virtual reality scene, and parse spatial structure information, semantic tag distribution information and guide trigger point location information from the set of scene scan data; Create a virtual agent pair that engages in a game of interaction, comprising an environment-building agent and a narrative-eroding agent. Generate an environment memory model based on the spatial structure information and configure it to the environment-building agent. Determine a fictional anchor point generation strategy based on the semantic tag distribution information and configure it to the narrative-eroding agent. Combining the scene scanning data set, the current user's navigation task flow, and historical user behavior data, the initial navigation path and a set of fictional anchor points are generated through the interactive game virtual intelligent agent pair. The interactive game virtual intelligent agent pair is driven to engage in adversarial game, causing the environment building agent to correct the initial navigation path based on the set of fictional anchor points and the narrative erosion agent to update the set of fictional anchor points based on the corrected navigation path. The process is iterated alternately until a preset equilibrium condition is met, at which point the game navigation path and the target set of fictional anchor points are output. Assign a matching fictional anchor from the target fictional anchor set to each guidance node in the game guidance path and generate a fictional anchor activation instruction, and combine the guidance node sequence and its associated fictional anchor activation instructions into a narrative guidance path; In the target virtual reality scene, the navigation interaction interface is rendered based on the narrative navigation path, and the interface moves sequentially to each navigation node. When a navigation node is reached, the corresponding virtual information content is retrieved according to the activation command of the virtual anchor point associated with the navigation node. The virtual information content is superimposed on the corresponding spatial position of the scene, and the rendered screen with virtual information is output to the user's virtual reality display device.
2. The method according to claim 1, characterized in that, The process of creating a virtual agent pair that engages in a game of interaction, comprising an environment-building agent and a narrative-eroding agent, generating an environment memory model based on the spatial structure information and configuring it to the environment-building agent, and determining a fictional anchor point generation strategy based on the semantic tag distribution information and configuring it to the narrative-eroding agent, includes: Spatial structure information is extracted from the scene scan data set, and spatial connectivity domain segmentation is performed on the spatial structure information to generate a spatial connectivity domain partitioning result composed of passable spatial units; Assign a spatial unit identifier to each passable spatial unit, and generate a spatial connectivity topology diagram between each passable spatial unit based on the spatial adjacency relationship in the spatial structure information. Each semantic tag in the semantic tag distribution information is projected to the corresponding passable spatial unit according to its spatial location coordinates, and a mapping relationship between the passable spatial unit and the semantic tag is established to generate a spatial unit semantic binding relationship table. Extract the coordinates of each wayfinding trigger point from the wayfinding trigger point location information, map the coordinates of each wayfinding trigger point to the accessible space unit where it is located, and mark the accessible space unit that aggregates wayfinding trigger points as the wayfinding trigger space unit set; The spatial connectivity domain partitioning result, the spatial connectivity topology graph, the spatial unit semantic binding relationship table, and the guidance triggering spatial unit set are organized into an environmental memory model data structure, and the environmental memory model data structure is serialized into an environmental memory model. Load the preset environment construction agent basic template, and inject the environment memory model into the perception memory component of the environment construction agent basic template to obtain the environment construction agent; Extract the distribution characteristics of user dwell time in different spatial locations and the interest association characteristics of triggered semantic tags from the historical user behavior data; A user interest space mapping function is constructed based on the dwell time distribution characteristics and the interest association characteristics. The user interest space mapping function is used to map user interest categories to spatial regions in the scene. The scene space covered by the semantic tag distribution information is divided into multiple candidate anchor point region grids. The user interest space mapping function is used to filter out the space grids whose interest response intensity exceeds the preset filtering threshold to obtain a virtual anchor point candidate location library. The virtual anchor point candidate location library is then paired with the corresponding narrative tag material and encapsulated as a virtual anchor point generation strategy. Load the preset narrative erosion agent basic template, write the fictional anchor point generation strategy into the anchor point decision component of the narrative erosion agent basic template, and obtain the narrative erosion agent.
3. The method according to claim 1, characterized in that, The process involves combining the scene scan data set, the current user's navigation task flow, and historical user behavior data. An initial navigation path and a set of fictional anchor points are generated through the interactive game-theoretic agent pair. This drives the adversarial game between the interactive game-theoretic agent pair, causing the environment-building agent to modify the initial navigation path based on the set of fictional anchor points, and the narrative erosion agent to update the set of fictional anchor points based on the modified navigation path. This iterative process continues until a preset equilibrium condition is met, at which point the game-theoretic navigation path and the target set of fictional anchor points are output. This includes: The intelligent agent constructed by the environment parses the sequence of target points to be visited and the access priority order of each target point from the navigation task flow. The sequence of target points consists of spatial location identifiers that the user needs to arrive at in sequence in the virtual scene. The intelligent agent constructs the environment and calls the spatial connectivity topology graph in the environment memory model. Starting from the user's current spatial location, and taking the first unvisited target point in the target point sequence as the temporary endpoint, it performs spatial path search processing to generate the original path node sequence connecting the starting point to the temporary endpoint. The intelligent agent constructed by the environment calculates the semantic richness of each passable spatial unit along the original path node sequence according to the semantic binding relationship table of the spatial unit, selects the path with the highest semantic richness as the initial guidance path backbone, and sequentially connects all target points according to the access priority order to generate the initial guidance path. The narrative erosion agent is invoked to receive the historical user behavior data and the scene scan data set. The user interest space mapping function in the fictional anchor point generation strategy is used to analyze the interest distribution in the historical user behavior data to obtain the interest heat distribution region. The narrative erosion agent is invoked to select fictional anchor locations located within the interest heat distribution area from the fictional anchor candidate location library, and to extract fictional narrative content from the narrative tag material for each selected fictional anchor location to generate a set of fictional anchors. The adversarial game round is initiated based on the initial guidance path and the set of fictional anchor points. The game guide path and the target set of fictional anchor points are output based on the statistical information after each round of adversarial game.
4. The method according to claim 3, characterized in that, The adversarial game round is initiated based on the initial guidance path and the set of fictitious anchor points. The system then determines whether the preset equilibrium condition has been reached based on statistical information after each round. If so, it outputs the game guidance path and the target set of fictitious anchor points, including: Initiate an adversarial game round, construct an agent through the environment to obtain the set of fictional anchor points for the current round, take all the positions of the fictional anchor points in the set as mandatory nodes, insert them into the initial wayfinding path, and re-call the spatial path search process to generate the adjusted wayfinding path. The narrative erosion agent receives the adjusted navigation path, analyzes all passable space units traversed by the adjusted navigation path, and for sections where the semantic tags of passable space units match the narrative content of the fictional anchor positions below a preset attraction threshold, adds new fictional anchor positions or modifies the narrative content of existing fictional anchor positions to generate an updated set of fictional anchors. The environment-constructed agent continues to adjust the navigation path according to the updated set of fictional anchor points, and the narrative erosion agent continues to update the set of fictional anchor points according to the newly adjusted navigation path, repeatedly alternating between the adversarial adjustment process; After each round of confrontation, the node sequence change of the adjusted wayfinding path and the information increment of the updated set of fictitious anchors are recorded. When the node sequence change is lower than the preset change tolerance for several consecutive rounds and the information increment is lower than the preset increment tolerance for several consecutive rounds, the preset equilibrium condition is determined to be reached. The adjusted navigation path when the preset equilibrium condition is reached is taken as the game navigation path, and the updated set of virtual anchor points is output simultaneously as the target set of virtual anchor points.
5. The method according to claim 1, characterized in that, The step of assigning a matching fictional anchor from the target fictional anchor set to each guidance node in the game-theoretic guidance path and generating a fictional anchor activation instruction, and combining the guidance node sequence and its associated fictional anchor activation instructions into a narrative guidance path, includes: Extract all path nodes in the game guidance path, generate guidance node identifiers for each path node in traversal order, and generate a guidance node sequence. Extract the three-dimensional spatial coordinates of each fictional anchor point in the target set of fictional anchor points and the corresponding fictional narrative content material; For each wayfinding node in the wayfinding node sequence, determine the spatial distance between the spatial coordinates of the wayfinding node and the spatial coordinates of each virtual anchor point, and take the virtual anchor point with the smallest spatial distance as the initial matching virtual anchor point of the wayfinding node. For the initial matched fictional anchor point, obtain the content type tag and content volume parameter of the fictional narrative content material of the initial matched fictional anchor point, and determine the presentation style template of the fictional narrative content material when it is superimposed in the virtual scene according to the content type tag. Based on the content volume parameters and the estimated path time between navigation nodes, the playback duration of the fictional narrative content material is calculated, and the activation start time and activation end time are set in combination with the playback duration to obtain the time sequence activation parameter group. The presentation style template and the timing activation parameter group are encapsulated into a fictional anchor activation instruction for the navigation node. The fictional anchor activation instruction includes instructions for calling fictional narrative content materials, instructions for presentation in a visual form, and instructions for timing control in a visual form. Based on the navigation node sequence and the fictional anchor activation instruction, the narrative navigation path data structure is determined, and the narrative navigation path is generated after completing the temporal compatibility review of the narrative navigation path data structure.
6. The method according to claim 5, characterized in that, The process of determining the narrative navigation path data structure based on the navigation node sequence and the fictitious anchor point activation instruction, and generating the narrative navigation path after completing a temporal compatibility review of the narrative navigation path data structure, includes: Each wayfinding node in the wayfinding node sequence is bound to the virtual anchor point activation instruction corresponding to that wayfinding node, and an activation instruction association mapping table is generated. According to the arrangement order of the guide node sequence, the data bound in the activation instruction association mapping table is unfolded in sequence to generate a narrative guide path data structure. The narrative guide path data structure records the position of each guide node in space, the direction of the next guide node, and the virtual anchor point activation instruction to be executed at that guide node. The narrative navigation path data structure is subjected to a time sequence compatibility review. When it is detected that the activation instructions of the virtual anchor points bound to adjacent navigation nodes overlap and conflict in the activation time window, the activation start time of the latter navigation node is adjusted to be after the activation end time of the former navigation node in order to eliminate the time conflict of virtual information superposition. The narrative navigation path data structure, after undergoing time-series compatibility review, is persistently stored to obtain the narrative navigation path.
7. The method according to claim 1, characterized in that, The process of rendering a navigational interface based on the narrative navigation path in the target virtual reality scene, and sequentially moving to each navigation node, involves retrieving corresponding virtual information content according to the activation command of the virtual anchor point associated with that navigation node, overlaying the virtual information content onto the corresponding spatial position of the scene, and outputting a rendered image infused with virtual information to the user's virtual reality display device, including: Start the main rendering loop of the target virtual reality scene, and create a wayfinding interactive interface container in the user interface layer of the rendering pipeline. The wayfinding interactive interface container is used to carry wayfinding indicator lines and virtual information overlay layers. Read the sequence of navigation nodes in the narrative navigation path, and generate a three-dimensional navigation indicator line in the virtual scene that points from the user's current location to the first navigation node in the navigation node sequence, based on the user's current location and the first navigation node in the navigation node sequence. Control the virtual camera to translate the viewpoint along the direction indicated by the three-dimensional guide line according to the preset smooth movement strategy, and move the virtual camera from the current viewpoint position to the spatial position of the first guide node; When the virtual camera arrives at the first wayfinding node, the virtual anchor point activation command bound to the wayfinding node is triggered, and the virtual information content identifier to be called, the presentation style template, and the activation start time are extracted from the virtual anchor point activation command. The corresponding virtual information content data package is retrieved from the virtual information material repository based on the virtual information content identifier, and the display format of the virtual information content data package is determined based on the presentation style template. Based on the activation start time, the virtual information content data packet is presented at the corresponding spatial location in the virtual scene. The virtual information content data packet is then rendered and mixed with the original geometric data and lighting information of the scene to generate a virtual information overlay frame. During the virtual information presentation process, user eye tracking data is continuously received, the user eye direction vector is parsed, and the deviation angle between the user eye direction vector and the three-dimensional guide line is calculated. When the deviation angle is detected to exceed the preset visual attraction threshold, the display attributes of the three-dimensional wayfinding line are adjusted to optimize the color vibrancy and dynamic flashing frequency of the three-dimensional wayfinding line. When the virtual camera leaves the current navigation node and moves to the next navigation node, the triggering and overlay operations are repeatedly executed according to the activation instruction of the virtual anchor point bound to the next navigation node, until all navigation nodes in the narrative navigation path are traversed, and the rendered screen is continuously output to the user's virtual reality display device.
8. The method according to any one of claims 1-7, characterized in that, After outputting the rendered image to the user's virtual reality display device, it also includes: The eye-tracking data stream fed back by the user's virtual reality display device is obtained, and the user's gaze dwelling coordinate sequence and pupil diameter fluctuation sequence are parsed from the eye-tracking data stream. The superimposed virtual information content corresponding to the user's gaze dwelling coordinate sequence in the rendered image is extracted. Based on the user's gaze dwell coordinate sequence and the pupil diameter fluctuation sequence, the narrative immersion response characteristics of the current user to the superimposed virtual information content are determined, and according to the narrative immersion response characteristics, inactive virtual anchors whose immersion response intensity does not reach the preset activation threshold are selected from the target virtual anchor set. Based on the spatial location of the inactive virtual anchor point in the target virtual reality scene, the spatial connectivity topology graph in the environment memory model is retrieved to generate a supplementary navigation path segment that connects the current user's spatial location with the inactive virtual anchor point; The supplementary wayfinding path fragment is injected into the narrative wayfinding path to obtain an enhanced narrative wayfinding path, and the wayfinding node sequence in the wayfinding interaction interface container is updated according to the enhanced narrative wayfinding path to guide the current user to move along the supplementary wayfinding path fragment and trigger the virtual information content corresponding to the inactive fictional anchor point.
9. An interactive server for a wayfinding system, characterized in that, include: A processor; a storage device having a computer program stored thereon; a network interface for providing network communication functions; when the computer program is executed by the processor, the processor enables the processor to implement the VR-based wayfinding system interaction method as described in any one of claims 1-8.
10. A readable storage medium, characterized in that, The readable storage medium stores a program or instructions that, when executed by a processor, implement the VR-based guide system interaction method as described in any one of claims 1-8.