AR navigation system and method based on visual language model

The AR navigation system, which combines the visual language model with SLAM and AR visual language memory model, solves the shortcomings of existing AR navigation technology in positioning accuracy, semantic understanding and long-term memory, realizes high-precision positioning, multimodal interaction and real-time navigation, and improves the intelligence and stability of the AR navigation system in complex environments.

CN120744017APending Publication Date: 2025-10-03HANGZHOU DIANZI UNIV +1
View PDF 0 Cites 13 Cited by

Patent Information

Application Number
CN202510831107.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-20
Publication Date
2025-10-03

AI Technical Summary

Technical Problem

Existing AR navigation technology has shortcomings in positioning accuracy, semantic understanding, interactive experience and long-term memory, and cannot meet the needs of augmented reality devices for intelligent, low-latency AR navigation services with spatiotemporal semantic understanding capabilities.

Method used

An AR navigation system based on a visual language model is adopted, combining pure visual SLAM with a visual language model. Video data is collected through AR devices to build a point cloud map of a multi-floor environment. The AR visual language memory model is used to generate a long-term memory database. Combined with a lightweight AR visual language navigation auxiliary matching model for query and path planning, high-precision positioning, multimodal semantic interaction, and efficient path planning are achieved.

Benefits of technology

It achieves high-precision positioning and dynamic mapping in complex indoor and outdoor environments, supports multi-floor navigation, has multimodal semantic interaction capabilities, can respond to complex semantic queries in real time, meet real-time navigation needs, and improve navigation intelligence and stability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120744017A_ABST
    Figure CN120744017A_ABST
Patent Text Reader

Abstract

The invention discloses an AR navigation system and method based on a visual language model, and the method comprises the steps: firstly carrying out video data collection, and constructing a memory database; secondly, querying a navigation target most related to a natural language request of a user in the constructed memory database to obtain a current AR equipment pose and a target pose; and then solving the shortest path from the current pose to the target pose by using the current AR equipment pose and the target pose according to the point cloud map, and optimizing the path direction to finally obtain an optimized path. And finally, guiding a user to move along the planned path through view superposition path indication and voice prompt in AR equipment by utilizing the optimized path, and updating the point cloud map and the memory database. According to the invention, high-precision real-time positioning and sparse point cloud map construction can be realized only by camera input in indoor and outdoor complex environments with low GPS precision, accurate navigation is carried out, and the flexibility and intelligent level of navigation are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the interdisciplinary technical field of augmented reality (AR) navigation and multimodal artificial intelligence, and specifically relates to an AR intelligent navigation system and method that integrates pure visual simultaneous localization and mapping (Visual SLAM), a vision-language model (VLM), and a retrieval-augmented generation with large language model (RAG-LLM). The system is suitable for realizing long-term temporal memory construction, multi-round natural language interaction, and efficient path planning in scenarios that rely solely on cameras. Background Art

[0002] With the rapid development of AR glasses, mobile phones, and service robots in the field of spatial intelligence, user demands for navigation systems are shifting from traditional spatial positioning to a deeper understanding and interaction with spatiotemporal context. In practical applications, users expect to initiate cross-temporal, multimodal, and complex queries through natural language, such as "Find the nearest coffee break area on my current floor," "Where is my parking location," or "Trace the red handbag display I passed two hours ago," requiring long-term memory retrieval and semantic reasoning. These advanced AR navigation functions require systems with powerful spatiotemporal modeling capabilities, multimodal data fusion technology, and real-time augmented reality interaction capabilities. Ultimately, they aim to provide users with precise spatial intelligent navigation services through AR devices such as AR glasses and mobile phones.

[0003] However, the existing technical system has obvious deficiencies in the following three aspects:

[0004] (1) Limited positioning accuracy and memory capacity: Traditional AR navigation systems generally rely on GPS and static preset maps. Especially in indoor environments, GPS positioning accuracy decreases significantly (the error often exceeds 5 meters), making it impossible to determine the floor on which the user is located, and unable to respond to changes in the dynamic environment in real time. At the same time, such systems usually only cache perception information for a short period of time (10-30 minutes), lack the long-term memory capacity of user behavior and environmental status, and are difficult to support retrospective queries.

[0005] (2) Visual language navigation has high computational complexity and weak generalization ability: Although navigation algorithms based on visual language understanding (such as Vision-Language Navigation, VLN) can realize the perception of the immediate environment and target parsing, their training is highly dependent on the target environment, resulting in unstable performance when transferred to real scenes. In addition, when processing large-scale scenes containing N spatial nodes, the computational complexity of cross-modal alignment is O(N 2 ), it is difficult to meet the real-time navigation requirements with an interaction delay of less than 1 second.

[0006] (3) Large language models have poor adaptability in long-term memory: Although mainstream large language models can maintain semantic coherence in short conversations, their fixed-length context windows limit the long-term retention of information about multiple objects and locations in complex scenes. On the one hand, a large amount of key information in the environment (such as object locations, path nodes, etc.) cannot be accurately memorized; on the other hand, the rolling memory mechanism causes key historical information to be overwritten over time, significantly affecting the recall accuracy of memory during navigation.

[0007] In summary, current AR navigation technology has systematic shortcomings in dynamic environment modeling of multi-layer indoor and outdoor scenes, long-term spatiotemporal memory retention, and real-time response efficiency. It cannot meet the needs of augmented reality devices for intelligent, low-latency AR navigation services with spatiotemporal semantic understanding capabilities, and has become a core technical bottleneck that urgently needs to be broken through. Summary of the Invention

[0008] The present invention aims to solve the above problems and proposes an augmented reality navigation system and method based on a visual language model. The core idea of ​​the system is: first, the system uses an AR device equipped with a camera, such as a robot, AR glasses or a mobile phone, to move in the target environment to collect video data, and combines pure visual SLAM with a visual language model to extract the semantics of objects and scene spatial information in indoor and outdoor multi-story environments. The AR visual language memory model (AR-VLM) proposed in this invention is used to 2 ) to build a long-term memory database that can be efficiently retrieved. Subsequently, users can initiate navigation requests through natural language. The system uses the lightweight AR visual language navigation auxiliary matching model (ARN-Pilot) designed by this invention to retrieve the most relevant target location in the memory library, and generates the optimal path through visual positioning and path planning algorithms, and finally presents the navigation results in real time in the AR interface. The system consists of the following five modules:

[0009] (1) Pure visual map building module: Use the camera of AR devices such as AR glasses or mobile phones to capture real-time video stream V s , using monocular vision SLAM algorithms such as ORB-SLAM3 combined with a monocular depth estimation model to estimate the device pose in real time (for example, calculate the pose of the device at the kth second of the video as p k ) and build an AR navigation point cloud map M. If there are multiple floors, build a point cloud map for each floor;

[0010] (2) Memory building module: Memory building module: Memory building module: Memory building module: Memory building module: Memory building module: Memory building module s The sampling point at the kth second) is then used to calculate the Δt-second continuous video frame sequence nearby. Using the AR visual language memory model AR-VLM proposed by the present invention 2 Generate natural language descriptions of the environment, stores, road conditions, and other scene contents k , and is mapped to a vector v through a semantic encoding model such as mxbai k , with timestamp t k , corresponding to the pose p k 、Floor k Jointly construct real-time location semantic information memory entries for AR navigation (v k ,p k ,l k ,t k ) and stored in the AR navigation memory database D;

[0011] (3) Query processing module: In the memory database D, according to the user's natural language navigation request Q, the lightweight AR visual language navigation auxiliary matching model ARN-Pilot designed by the present invention is used to parse its semantic intent and construct the query embedding vector e Q , for all semantic vectors v in D k Perform vector similarity search, select the top n most relevant memory entries and feed them back to LLM for multiple rounds of filtering, and finally lock the optimal target pose p based on the current AR device position. j and its corresponding memory content;

[0012] (4) Path planning module: based on the optimal target pose p j The point cloud map M obtained by the pure visual map construction module is used to estimate the current AR device pose p i , and combined with the visual positioning algorithm based on image feature matching to enhance accuracy, an occupancy grid map M′ on the possible path is constructed to represent the walkable area. If there are multiple floors, the floors involved are merged into a grid map through stairs, elevators, etc. The A* algorithm is used to solve p in the walkable area. i To the target pose p j The shortest path P is traced along the negative gradient and smoothed with B-spline to generate a smooth shortest path P * ;

[0013] (5) Navigation interaction module: smooth the shortest path P * It is superimposed on the user's AR field of view of the AR device to generate a real-time navigation guidance interface, using visual arrows, path highlights, etc. to intuitively display the next moving direction, ensuring that the user can smoothly reach the target posture according to the planned path. j At the same time, real-time video is recorded during navigation to update the point cloud map M and memory database D, including the latest real-time road conditions, to facilitate the response to subsequent new user needs.

[0014] The navigation system operation process of the present invention mainly includes four stages: data collection and memory construction, AR navigation target query and analysis, position matching and path planning, navigation execution and human-computer interaction, covering the following steps:

[0015] Step S1: Data collection and memory construction. In this step, the smart terminal (such as AR glasses or mobile phone) moves in the target environment and continuously collects video data. The AR device pose p is estimated in real time based on the image sequence in the video data. k , used to build or update the AR navigation point cloud map M, record the spatial position of each key frame, and use OCR text recognition technology to recognize the road sign text and determine the floor l where the AR device is located k Then, the system will continue to play video segments for Δt seconds. Input the AR visual language memory model AR-VLM proposed by this invention 2 , the output corresponds to the natural language description L k If there are multiple floors, a description of stairs or elevators must be included. Call a semantic encoding model (such as mxbai) to convert the natural language description L k Encoded as a semantic vector v k , and finally construct the real-time location semantic information memory entry for AR navigation (v k ,p k ,l k ,t k ) and stored in the AR navigation memory database D, where t k is the timestamp.

[0016] Step S2: AR navigation target query parsing. In this step, the memory database D constructed in step S1 is used to query the navigation target most relevant to the user's natural language request. Specifically, the user submits a navigation request Q in text or voice. The lightweight AR visual language navigation auxiliary matching model ARN-Pilot designed by the present invention is used to perform multiple rounds of reasoning on the request Q. The current AR device pose p is obtained by combining the visual positioning algorithm based on image matching. i , select the optimal target pose p from the memory database D j As the navigation target. If multiple floors are involved, the location of the staircase or elevator must also be retrieved.

[0017] Step S3: Position matching and path planning. This step uses the current AR device pose p obtained in step S2 i and the optimal target pose p jBased on the point cloud map M constructed in step S1, a two-dimensional occupancy grid map M′ on the possible path is generated, marking the walkable and obstacle areas. If there are multiple floors, the floors involved are merged into a grid map through stairs, elevators, etc. The A* algorithm is used to solve the problem of the current pose p in M′. i To the target pose p j The shortest path P is obtained, and the path direction is optimized by the negative gradient descent method. The path P is further smoothed using the B-spline curve, and the optimized path P is finally obtained. * .

[0018] Step S4: Navigation execution and human-computer interaction. This step uses the optimized path P obtained in step S3 * In the AR device, the user is guided along the planned path, climbing stairs, or taking an elevator by superimposing path instructions and voice prompts in the field of view. Simultaneously, the camera continues to capture the current video image to obtain the latest environmental visual information for real-time positioning in step S3, as well as updating the point cloud map M and memory database D in step S1, thereby responding to new navigation requests in subsequent user interactions.

[0019] The present invention proposes an AR navigation system and method based on a visual language model, which addresses the shortcomings of existing technologies in positioning accuracy, semantic understanding, interactive experience, and long-term memory, and has the following significant beneficial effects:

[0020] High-precision positioning and dynamic mapping capabilities: Based on pure visual SLAM (such as ORB-SLAM3) and a monocular depth estimation algorithm, this system achieves high-precision real-time positioning and sparse point cloud mapping using only camera input in complex environments, both indoors and outdoors, where GPS accuracy is low. It can effectively adapt to dynamic scene changes without requiring pre-installed static maps or external positioning devices. Furthermore, it supports AR navigation across multiple floors in novel environments, such as finding a car or store in a newly visited shopping mall.

[0021] Multimodal semantic interaction and instant navigation feedback: This system integrates a visual language model with a retrieval-enhanced large language model to parse complex semantic queries posed by users in natural language and provide real-time responses. It also supports a "walk-and-record" mode, where the AR device automatically collects environmental data as the user moves and generates new semantic memory entries that are dynamically written to the database. Users can perform semantic queries and route navigation at any time based on the latest memory database, forming a natural language-driven closed-loop human-computer interaction, greatly enhancing the flexibility and intelligence of navigation.

[0022] Long-term memory retention and efficient path planning: The system builds a multimodal memory database containing semantic vectors, spatial locations, and timestamps to store and manage users' long-term behavior trajectories, supporting cross-temporal queries of historical scenarios. Combining image matching with the A* path search algorithm, the system rapidly generates smooth paths on the constructed occupancy grid map, with computational complexity significantly lower than traditional cross-modal matching methods. This ensures sub-second navigation feedback, meeting the needs of real-time interactive applications.

[0023] In summary, the present invention breaks through the core bottlenecks of traditional AR navigation systems in terms of perception accuracy, semantic understanding, spatiotemporal memory, and human-computer interaction, and significantly improves the navigation intelligence, stability, and practicality of augmented reality terminals in dynamic and complex environments, especially in multi-story indoor and outdoor environments. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] Figure 1 Schematic diagram of the overall system architecture;

[0025] Figure 2 System technical steps flow chart;

[0026] Figure 3 AR Visual Language Memory Model AR-VLM 2 Structural diagram;

[0027] Figure 4 Schematic diagram of the refined path decision module in the lightweight AR visual language navigation-assisted matching model ARN-Pilot. DETAILED DESCRIPTION

[0028] In order to more clearly illustrate the specific implementation method of the present invention and show the features and advantages of the present invention, non-limiting embodiments of the present invention are further described in detail below with reference to the accompanying drawings.

[0029] It should be noted that the following embodiments use a mobile phone with a camera as an AR device to schematically illustrate the basic ideas, processes and technical principles of the present invention. In actual implementation, a variety of different types of devices with visual perception and navigation display functions can be selected, including AR glasses, robots, etc.

[0030] Example 1: Model Detailed Design

[0031] Attachment Figure 1 and attached Figure 2 The main steps of the AR navigation system and method based on the visual language model in the embodiment of the present invention are illustrated by way of example. Figure 2 As shown, the AR navigation system and method based on the visual language model in this embodiment may include step S1, step S2, step S3 and step S4.

[0032] Step S1, data collection and memory construction, may include S1-1 and S1-2.

[0033] Step S1-1, visual feature extraction and SLAM positioning, using pure visual SLAM algorithms such as ORB-SLAM3 to capture the real-time video stream V of the AR device s Process it, extract the feature points, and calculate the real-time pose p of the AR device k , merge the feature point sets of each sampling point to build the AR navigation point cloud map M.

[0034] Furthermore, in step S1-1, ORB feature point detection can be used, such as for real-time video stream V s The kth frame image I k Extract ORB feature point set F k ={f1,f2,...,f N}, where N = 2000 is the number of feature points.

[0035] Furthermore, in step S1-1, the AR device pose p can be calculated by the PnP algorithm. k , the optimization goal is to minimize the reprojection error:

[0036]

[0037] where X i is a 3D map point, x i is the corresponding 2D observation, and π is the projection function.

[0038] Furthermore, in step S1-1, based on the feature points and point cloud map obtained above, the feature points of the key frame are triangulated into a 3D point cloud M={X j |j=1,2,...,C}. The point cloud size C is about 10 in a medium-sized scene (such as a single-story shopping mall). 4 ~10 5 Experiments show that C = 10 5 The point cloud can achieve a positioning error of <0.5m in a 10,000㎡ scene.

[0039] Furthermore, in step S1-1, if the scene contains multiple floors, a point cloud map is constructed for each floor separately.

[0040] Step S1-2, video semantic encoding and memory, real-time video stream V captured by AR device s Segment, and then use the AR visual language memory model AR-VLM proposed in this invention to analyze each segment 2 Construct an environment description text as the semantic information memory of the corresponding position of the segment, encode it into a semantic vector, and store it in the AR navigation memory database.

[0041] Furthermore, in step S1-2, the real-time video stream V can be s Segmentation, for example, from the t k Continuous Δt=5 seconds video segment starting from second Enter AR-VLM 2 Model, generate natural language description L k (e.g. "There is a red sofa on the left and the elevator entrance is 10 meters ahead"), corresponding to the t k Real-time location semantic information memory of the second's location.

[0042] Furthermore, in step S1-2, the mxbai encoder can be used to generate semantic vectors for the memory, and L k Mapped to a 768-dimensional vector v k =E(L k )∈R 768 , where E is the encoding function.

[0043] Furthermore, in step S1-2, the real-time location semantic information memory entry (v k ,p k ,l k ,t k ) is stored in database D, where p k The device pose output by SLAM.

[0044] Step S2, AR navigation target query and resolution, may include S2-1 and S2-2.

[0045] In step S2-1, AR navigation user intention analysis is performed, using the lightweight AR visual language navigation auxiliary matching model ARN-Pilot designed by the present invention to analyze the text or voice query input by the user to generate a structured query that is more suitable for navigation target retrieval.

[0046] Furthermore, in step S2-1, the query Q = "find the red handbag counter I passed by two hours ago" can be input, and ARN-Pilot can be used to generate a structured query Q′ = {action: "locate", target: "red handbag counter", time constraint: "t k ≤t now -2×60×60"}, where t now The current time in seconds.

[0047] Step S2-2, AR navigation memory retrieval matching. The lightweight AR visual language navigation auxiliary matching model ARN-Pilot designed by the present invention generates answers and evaluates the degree of matching based on user queries based on retrieval content and its own knowledge. At the same time, it performs reasoning and analysis on the navigation conditions of multiple floors and finds the optimal navigation target based on the user's personalized preference needs.

[0048] Furthermore, in step S2-2, a simplified way of ARN-Pilot is to embed the query into e Q =E(Q)∈R 768 and all v in D k Calculate cosine similarity As the comprehensive matching degree, select s k The top n=5 candidate memory entries of >η are used as navigation targets, where η is the minimum similarity threshold.

[0049] Furthermore, in step S2-2, if the target obtained by the query is not on the same floor as the starting position, it is necessary to enter step S2-1 again to generate a query for steps and stairs as waypoints, and perform the memory retrieval matching of step S2-2 again.

[0050] Furthermore, in step S2-2, the large model can be used to analyze the user's route preference selection and set road sections that need to be avoided (such as sections with excessive pedestrian flow).

[0051] Step S3, location matching and path planning, may include S3-1 and S3-2.

[0052] Step S3-1, occupancy grid map construction, determine the real-time position and floor of the AR device, project the point cloud M of the floor where the current AR device is located onto a two-dimensional plane, and generate a grid map M′ with a resolution of r = 0.1m.

[0053] Furthermore, in step S3-1, a visual positioning algorithm based on image matching (such as SuperPoint, etc.) can be used to determine the real-time position and floor of the AR device according to the degree of feature similarity between the image captured by the camera in real time during navigation and the memory image (with location information).

[0054] Furthermore, in step S3-1, a semantic segmentation model such as SegFormer-B2 can be used to segment the ground area, determine the plane where the ground is located, and cut out the part from the ground plane to the plane at a fixed height from the point cloud M. The part is projected onto a two-dimensional plane to generate a grid map M′ as the walkable area:

[0055]

[0056] Where w(X) is the confidence weight of point X, τocc is the occupation threshold, which can be set to 0.3.

[0057] Step S3-2, AR navigation optimal path planning, calculate a shortest path P in the walkable area of ​​the grid map M′, and generate an AR navigation path P that is more in line with user comfort * .

[0058] Furthermore, in step S3-2, the A* algorithm can be used to search for a shortest path P. The specific method is as follows: ′ = 0 area defines the graph node, and the cost function f(n) = g(n) + h(n) is optimized to solve the path, where f(n) is the total expected cost, g(n) is the total moving cost from the current point to the starting point, and h(n) is the expected cost from the current point to the target point.

[0059] Furthermore, in step S3-2, the solved path can be further smoothed using B-spline, that is, the path P is parameterized into a B-spline curve. Among them C i is the control point, and p=3 is the order. Using this method, we can get an AR navigation path P that is more in line with user comfort. * .

[0060] Step S4, navigation execution and human-computer interaction, may include S4-1 and S4-2.

[0061] Step S4-1, visual perception and AR feedback, real-time rendering of the navigation path P* through the AR interface. At the same time, the camera of the AR device continues to capture the current video screen, analyzes and updates the map and memory in real time.

[0062] Furthermore, in step S4-1, the navigation path P* can be displayed using dynamic visual elements (such as gradient color path bands and distance-sensing arrows) and spatialized voice prompts (3D sound positioning) to guide the user to move, climb stairs, or take an elevator. If the system detects that the user has deviated from the path for more than 5 seconds, it automatically increases the intensity of the path display and vibrates to remind the user.

[0063] Furthermore, in step S4-1, while the AR navigation is being visualized, the camera of the AR device continues to capture the current video image and enters step S1 again to update the map M and the memory database D, thereby responding to new navigation requests in subsequent user interactions.

[0064] Step S4-2, AR multimodal human-computer interaction, integrates real-time voice, text, button, gesture and other interactive interfaces on the AR device to obtain user personalized needs or the latest navigation goals at any time and respond in real time.

[0065] Furthermore, in step S4-2, users can initiate new requests at any time using natural language (e.g., "Take me to the nearest restroom," "Take the elevator instead of the stairs," or "Take a less crowded route"). The AR device accurately captures these requests through its microphone and real-time speech recognition algorithm. The lightweight LLM, combined with the current environmental context (e.g., a temporary sign indicating "Restroom under repair"), instantly interprets the request and generates a structured navigation task.

[0066] Furthermore, in step S4-2, for complex requests (such as "find a restaurant that does not require queuing"), the system can provide a preview of multiple options through an AR floating window (such as "Restaurant A is 50 meters away and there are many people queuing at the door"), and the user can confirm the selection through voice or other interactive methods.

[0067] Furthermore, in step S4-2, if the user provides a completely new navigation target, the process re-enters step S1.

[0068] Example 2: AR-VLM 2 Model detailed design

[0069] AR visual language memory model AR-VLM proposed by this invention 2 Mainly used for real-time video stream V captured by AR devices s The existing Multimodal Large Language Model (MLLM) is capable of understanding the semantics of open-world videos, but it is trained using online datasets and lacks real-world spatial perception. Directly using or fine-tuning the MLLM is insufficient to meet the requirements for generating real-time location semantic memory for AR navigation.

[0070] like Figure 3 As shown in the figure, to address the problem that the existing multimodal large model (MLLM) lacks real space perception capabilities in AR navigation, the present invention proposes a spatiotemporal-semantic joint coding network, which uses a dual-stream heterogeneous fusion architecture of spatial and semantic streams, and performs semantic reorganization under geometric constraints through a neural symbol alignment layer to generate navigation memory with both physical accuracy and scene adaptability. Specifically:

[0071] (1) Spatial flow: The attention weight matrix of the Transformer in the MLLM is voxelized to realize the spatiotemporal attention mechanism of AR navigation, and the AR navigation spatiotemporal understanding Transformer is composed to parse the video stream captured by the AR device in real time, build a multi-level attention 3D semantic field, and perform multi-scale spatial voxel-level semantic segmentation on the key attention locations of AR navigation (such as indoor stairs, shopping mall revolving doors, road signs, etc.).

[0072] (2) Semantic flow: Using decoupled language modeling, we design a large multimodal language model for AR navigation, and use AR navigation memory to generate prompts, breaking down open vocabulary into spatial atomic concepts (such as "escalator", "fire hydrant", "revolving door") and dynamic relational predicates (such as "bypass", "warning") to avoid language bias in network datasets. At the same time, we integrate AR navigation semantic analysis tools, such as using OCR to recognize text on road signs, designing a real-time path structure diagram, introducing learnable attention parameters between the diagram and spatial atomic concepts, and training AR-VLM. 2 The model determines the floor where the current AR device is located. k Other navigation related information.

[0073] In the spatial stream, this paper proposes to naturally extend the Transformer attention mechanism to the 3D voxel spatiotemporal domain, fundamentally improving the spatial perception capabilities of AR navigation devices. Unlike traditional methods that process video frames only on a plane, the model first maps RGB-D video frames to a sparse voxel space using CLIP features and 3D position encoding. Then, a sparse-dilated multi-level architecture is adopted. This not only fuses local sparse attention with dilation attention, but also embeds scale switching from the top-level overall to the bottom-level details in the voxel transformer: the upper layer captures the overall geometric outline of the shopping mall atrium, while the lower layer identifies the edge structure of each shelf layer. In addition, by constructing a cross-frame spatiotemporal attention mechanism, the system distinguishes between dynamic occlusions, such as "a cleaning cart is stationary" and "a luggage cart is being pushed," achieving a parallel understanding of dynamic and static spatial structures. The output of this module is a high-precision three-dimensional semantic field with structural semantic annotations, which not only supports physical location positioning, but also provides semantic-level structural understanding, fundamentally making up for the lack of spatial perception in traditional MLLM.

[0074] In the semantic stream, the present invention designs a network architecture based on a symbolic-geometric alignment layer. This architecture deconstructs linguistic cues into a three-dimensional feature vector consisting of "function description + method description + dynamic attributes" and injects them into the spatial model through this alignment layer, achieving a deep fusion of linguistic and spatial information. This design enables linguistic cues to not only complete semantics but also become active participants in driving entity understanding and navigation operations, thereby improving the model's navigation capabilities in complex indoor environments. Specifically, the symbolic-geometric alignment layer is responsible for establishing a direct mapping between linguistic symbols and the spatial voxel grid, aligning the functional description in the cues with their spatial locations. Symbolic nodes (such as "revolving door") are embedded into target voxels in the voxel grid through structural graph attention, enabling the localization of linguistic signals in three-dimensional space. Subsequently, a multi-scale feature cross-fuser receives 3D semantic field features from the spatial stream, parses and encodes the low-level fine geometry, mid-level local structure, and high-level overall layout, and fuses them with the symbolic semantics derived from the cues. The network architecture based on the symbolic-geometric alignment layer constructs a graph attention mechanism based on slot attention to capture the interaction between candidate paths and spatial structure. It also performs self-supervised optimization through a cross-modal contrastive learning strategy. Furthermore, in order to avoid the language bias existing in the network corpus, the system adopts decoupled language modeling, designs an AR navigation multimodal large language model, and uses AR navigation memory to generate prompts to decompose open vocabulary into spatial atomic concepts (such as "escalator", "fire hydrant", "revolving door") and dynamic relationship predicates (such as "bypass", "warning"), and integrates AR navigation semantic analysis tools, such as combining OCR to perform text recognition on road signs along the way, and updating the real-time semantic information of the 3D semantic field; designing a real-time path structure diagram, which establishes an alignment relationship with the spatial atomic concept through learnable attention parameters, thereby assisting the network architecture based on the symbol-geometry alignment layer to judge, such as the current floor l k In this way, the model is able to perform semantic reasoning at different granularity levels, thereby generating reliable memories that support multi-level, function-oriented navigation requirements in complex indoor environments.

[0075] Example 3: Detailed Design of the ARN-Pilot Model

[0076] Because traditional large language models are designed for general tasks, directly using them for the specific task of AR navigation memory matching in the present invention wastes most of the weight and makes it difficult to understand actual navigation relationships (e.g., points that are actually close may be separated by a wall). This embodiment proposes a lightweight AR visual language navigation auxiliary matching model, ARN-Pilot, for real-time memory matching in AR navigation. It is used to query the AR navigation memory database D for the optimal navigation target based on real-time environmental conditions, user navigation target descriptions, and personalized needs. The present invention designs a multi-layer topology perception architecture for the ARN-Pilot model to perform on-demand reasoning and matching of the optimal navigation target, and designs a distillation-projection joint compression method to train the ARN-Pilot model based on the large language model.

[0077] Specifically, the ARN-Pilot model designs a three-tiered topology-aware architecture: The bottom layer pre-constructs a multi-layered building graph network, whose nodes are encoded physical coordinates and semantic features (e.g., "fire door on the east side of the 3rd floor"). The weights of the edges in the graph network include actual path distance, navigability, and dynamic traffic flow coefficients. The middle layer deploys a lightweight graph topology encoder based on the constructed building graph network. This encoder generates reachability vectors for location nodes through neighborhood aggregation, thus constructing a multi-tiered topology-aware architecture. The top layer, based on this multi-tiered topology-aware architecture, designs a dynamic decision tree as a refined path decision module and a topology constraint retrieval algorithm. This algorithm jointly infers user intent, the real-time environment, and the graph topology to achieve personalized navigation. Training adopts a three-stage progressive strategy: first, a navigation target matching student model is pre-trained using a distillation-projection joint compression method. Then, contrastive learning is used to optimize the semantic-topological joint embedding space. Finally, reinforcement learning is used to fine-tune the decision tree weights.

[0078] Furthermore, the specific process of using the distillation-projection joint compression method to train the ARN-Pilot model based on the large language model proposed in the present invention is as follows: first, the large language model LLM (AR navigation target matching teacher model) is used to generate the decision trajectory of the AR navigation target matching thinking tree as a soft label, and through knowledge distillation, the ARN-Pilot model based on the Graph Neural Network (GNN) is guided to learn the reasoning path; then the topological projection loss function L designed by the present invention is used to calculate the loss path of the ARN-Pilot model. project Enforce the lightweight graph topology encoder to align with graph topology semantics:

[0079] L project =‖φ(GNN(p))-φ(LLM(p))‖2

[0080] Where φ is the mapping function of the semantic-topological joint embedding space. Its core is to align the graph structure information of the physical space with the semantic information of the language space into a unified vector space through learning. p is the predicted position. The ARN-Pilot model uses a layered activation mechanism during deployment. When the user's movement speed is detected to be greater than 1.2m / s (the movement threshold), it automatically switches to a minimalist mode, retaining only the single-layer GNN to calculate the key path nodes.

[0081] Furthermore, the path decision module in the ARN-Pilot model implements refined navigation based on a multi-layer topology perception architecture on the basis of a multi-layer building graph network: the building topology layer predefines cross-floor traffic rules (such as elevator locations) and generates a global accessibility vector through graph convolution; the floor topology layer analyzes the physical connection relationship of the actual walkable path to obtain candidate paths, and uses a gated graph attention network to dynamically weight edge weights (such as the impact of pedestrian flow on traffic speed); the real-time perception layer integrates temporary obstacle information input by AR devices to construct an incremental graph update mechanism. The topological information of these three layers (building topology layer, floor topology layer, and real-time perception layer) is jointly input into the Mind Tree reasoning engine. Figure 4 In the three-dimensional evaluation of distance, time, and personalized constraints shown in the figure, the topology perception module provides structured features for each dimension (such as the building topology layer constrains cross-layer path selection, the floor topology layer corrects dynamic passage errors, and the real-time perception layer optimizes demand-oriented decision output). These features are respectively implemented through the learnable path scoring module, the dynamic traffic prediction module, and the user preference analysis module, and finally through the learnable weighted decision aggregator Ψ(λ dis ·λ topo ,λ time ,λ custom ) outputs the optimal path, where λ dis is the distance score, λ topo is the topological confidence, and time Path estimation time score, λ custom is the degree of satisfaction of user preference weights. The parameters of the path scoring module, dynamic traffic prediction module, user preference analysis module, and weighted decision aggregator ψ can be learned using a linear fully connected layer. During training, topological adversarial sample enhancement (such as randomly shielding key nodes) is used to improve the robustness of the ARN-Pilot model to graph structure damage. The path scoring module injects the actual edge weights of the graph network, the traffic predictor learns cyclical laws (such as lunchtime cafeteria congestion) based on the time patterns of historical memory entries, and the user preference analysis module extracts walking habits through latent space clustering. The ARN-Pilot model uses the curriculum learning strategy Curriculum Learning during training, gradually transitioning from simple path decisions to complex scenarios with multiple constraints.

[0082] Furthermore, during the contrastive learning phase to optimize the semantic-topological joint embedding space for AR navigation, this paper proposes a hierarchical contrastive learning framework to achieve fine alignment of semantic and topological features. This framework samples positive and negative pairs and hierarchical negative samples, constructing a joint contrast task of "semantic pairs + topological pairs" from a multimodal perspective. This ensures that the generated embeddings maintain the semantic consistency of the linguistic description while also strengthening the relationships expressed by physical adjacency and path constraints in the graph structure. Specifically, during the same / different sample contrast optimization process, the model focuses on capturing the commonalities between concepts such as "elevator" and "stairs," which have similar functions but different usage. Contrastive learning brings the semantic embeddings of these concepts closer together, enhancing their relevance in the navigation semantic space and improving the ability to identify functional alternatives or complementary paths during navigation route recommendations. Furthermore, a hierarchical contrastive learning loss reflects sensitivity to differences between these concepts, such as the different spatial connectivity, usability, and path accessibility of "elevator" and "stairs." This enables the model to distinguish these differences in topological embeddings and optimize the selection of detailed path planning. This framework employs layer-by-layer temperature parameter adjustment and a hierarchical negative sampling mechanism. Specifically, in contrastive learning at different levels, different temperature parameters are set to control the smoothness of the similarity distribution. This allows the model to capture fine-grained path structure relationships at the bottom layer, enhance the ability to identify semantic connections between nodes at the middle layer, and maintain consistency in overall navigation planning and global semantic coherence at the top layer. Through this multi-level, multi-perspective contrastive learning, the model is able to form representations in the embedding space that both reflect semantic commonalities and distinguish differences in methods, significantly improving its ability to understand and judge diverse path options in actual navigation tasks.

[0083] Furthermore, a topology-constrained retrieval algorithm is designed to solve the personalized navigation target matching problem. When a user queries a target, a k-hop reachable subgraph of the current node is first generated. A large model for AR navigation matching evaluation is designed, and the navigation path is constructed into a graph structure. After encoding it using GNN, it is injected into the above-mentioned ARN-Pilot model in the form of a multi-tuple. The node semantic matching degree and path topology matching degree are simultaneously calculated within the subgraph range, and the comprehensive matching degree is calculated:

[0084]

[0085] in is the semantic matching degree, is the path topology matching degree, α and β are the weights of the two, v q and v k are the graph vertices corresponding to user queries and memory items in the memory bank, respectively, e c and e k For the corresponding navigation edge. Introducing the adaptive threshold generator Only keep the overall matching degree candidate targets, among which is the environmental complexity, is a learnable weight, and σ is a sigmoid function. During training, adversarial examples (such as positions that are close but unreachable) are constructed to force the model to learn topological similarity features.

[0086] Example 4: Indoor multi-floor scenario application

[0087] In the indoor shopping mall navigation scenario, after the user enters the mall wearing AR glasses, steps S1 and S2 are started. The system continuously collects environmental video through the front camera for 60 minutes, constructs a sparse map M containing 12,000 3D feature points, and generates 35 semantic memory entries and stores them in the database D. When the user issues a voice query "I want to drink coffee", the system uses the thought chain to perform reasoning preprocessing on the query and obtains the query text "Where is the nearest coffee break area on the current floor?" First, the query text is converted into a 768-dimensional vector e using the mxbai encoder. Q , calculated by cosine similarity with all v in database D k After setting the matching threshold η = 0.7, the system selects the three memory entries with the highest similarity, which are:

[0088] v 103 (t=8:30): "There is a Starbucks 20 meters south of Atrium 3, with 70% vacant seats." Matching degree s k =0.82;

[0089] v 88 (t=8:15): "There is a coffee vending machine at the end of the east corridor on the first floor", matching degree s k =0.63;

[0090] v 121 (t=8:45): "Free coffee is available in the lounge area next to elevator No. 5 on the 2nd floor." Matching degree: s k =0.58;

[0091] where v 103 The match score is 0.82, and its corresponding description is "There is a Starbucks 20 meters south of Atrium 3, with a seat vacancy rate of 70%." Combined with the current AR device pose p i =(120.5,45.2,0), the system excludes options that are too far away and finally locks the target posture

[0092] In step S3, position matching and path planning, the system projects the point cloud M generated by ORB-SLAM3 onto the ground plane and constructs a 200×150 grid map M′ with a 0.1-meter accuracy. The occupancy probability of each grid is calculated using the occupancy grid algorithm, where the threshold τ of the feature point confidence weight w(X) is occ Set to 0.3. Then the A* algorithm was used for path search, using Euclidean distance as the heuristic function, and the path from p was completed in 18 milliseconds. i to p j The system further uses a 32.7-meter-long path planning to smooth the path. To improve the user experience, the system uses a 3rd-order B-spline curve to smooth the path, with the control point spacing set to 0.5 meters and ensuring that the maximum curvature does not exceed 0.15 meters. -1 ,Finally, an optimized path of 33.1 meters in length is obtained, and all turning radii are larger than 1.2 meters.

[0093] During navigation execution and human-computer interaction in step S4, the AR navigation interface displays a green turn arrow and translucent blue path guidance in real time, generating a turn prompt with a direction of π / 4 at 3 meters from the user. The system operates stably at a frame rate of 60fps. When a pedestrian is detected 2 meters ahead, local path replanning is triggered, with response latency kept within 200 milliseconds. A voice prompt, "Destination 28 meters ahead, path clear," enhances the interactive experience. Simultaneously, the camera continuously captures images, which are analyzed by the memory building module, updating real-time road conditions to address subsequent user needs.

[0094] Furthermore, if the starting and ending locations are not on the same floor, the path planning module first retrieves the semantic memory entry for the elevator / staircase from the memory database D (e.g., "There is an upward escalator at the end of the east corridor on the 1st floor") and inserts it into the path sequence as a required waypoint. During the occupancy grid map construction phase, the system associates and maps the multi-layer point cloud map using floor identifiers. When the user approaches a waypoint, visual positioning calibration is initiated, and the postural offset after crossing floors is confirmed through ORB feature matching. After completing the floor switch, the path planning module automatically loads the grid map M′ of the target floor and, using B-spline curves, generates a three-dimensional composite path after crossing floors. It then overlays 3D floor markers and cross-floor guidance animations (e.g., an upward arrow and a virtual escalator model) through the AR interface, while simultaneously announcing a directional voice prompt, "You are about to reach the 2nd floor. Please take the escalator up," to ensure continuous cross-floor navigation.

[0095] Furthermore, if the user raises a new demand of "I want to save money" through the navigation interaction module during the navigation process, the system will input the query into the query processing module again and obtain v through the large language model reasoning. 121That is, "Free coffee is available in the rest area next to Elevator No. 5 on the 2nd floor" is the latest and most relevant option. The navigation target is reset, the path planning module re-determines the navigation path, and the navigation interaction module displays the updated route.

[0096] Furthermore, if the user raises a new requirement of "I want to avoid crowds" through the navigation interaction module during the navigation process, the system will input the query into the query processing module again, and avoid congested sections based on the historical traffic data of the road conditions obtained by the visual language model analysis. The path planning module will re-determine the navigation path, and then the navigation interaction module will display the updated route.

[0097] Example 5: Outdoor Parking Lot Navigation

[0098] In the parking lot car search scenario, the system needs to handle long-term memory retrieval across time periods and path planning in a dynamic environment. When the user parked the car 3 hours ago, the onboard camera had recorded the parking location information (black SUV, license plate number A6B). When the user returns to the parking lot to query "Where is my parking location?", the system first filters the memory entries with timestamps within the last 3 hours in database D, and retrieves a total of 12 relevant records. Through the multimodal verification mechanism, the system found that v 2056 The text description of the entry "black SUV, license plate number A6B" has a semantic similarity of 0.91 with the query. The YOLOv8-nano license plate detection model is then used to perform visual verification on the historical keyframes, obtaining a confidence score of 0.92. Finally, the pose of the target in the SLAM local coordinate system is confirmed to be p j =(215.7,180.3).

[0099] Furthermore, the system adopts a hybrid positioning solution that integrates GPS and vision, takes points with high GPS positioning reliability as reference points, and uses the visual positioning algorithm to calculate the relative positions between other points and these reference points, and converts the local coordinates p j The coordinates were converted to WGS84 (39.9042°n, 116.4074°E), and visual relocalization was performed by matching ORB feature points between the current frame and historical keyframes. A minimum of 15 matching pairs of points was required, and after optimization using the PnP algorithm, the positioning error was kept within 1.2 meters. During path planning, the system detected two new rows of roadblocks in Area B (with a point cloud cluster radius greater than 0.8 meters) in real time and immediately updated the occupancy grid map, marking the affected areas (180:185, 90:95) as untraversable. This necessitated replanning of the original 52-meter path. The system generated a new 58-meter path within 22 milliseconds, guiding the user through the detour with a voice prompt, "Construction ahead, please turn right," combined with an AR turn arrow.

[0100] Furthermore, the system continuously updates the environment status and stores the dynamic event "2025-05-05 14:30 Construction fence in channel 2, area B" as a new memory entry. new It is stored in the database, and the probability of walkability in the semantic map is corrected to 0.2. Compared with traditional solutions, the present invention shows significant advantages in parking lot scenarios: it supports long-term memory recall of more than 3 hours, while traditional solutions can usually only cache 30 minutes of data; it has real-time detection and dynamic replanning capabilities without relying on preset map updates; through the visual-text joint verification mechanism, the cross-modal retrieval accuracy is increased from less than 60% of the traditional solution to more than 90%. In a parking lot environment with weak GPS signals, the system ensures the reliability and safety of navigation through the weighted fusion of visual positioning and GPS (weight ratio of 0.7:0.3), combined with a dynamic obstacle detection range of 5 meters in radius.

Claims

1. An AR navigation method based on a visual language model, characterized in that: The following steps are involved: Step S1: collect video data and build a memory database; Step S2: Query the navigation target most relevant to the user's natural language request in the constructed memory database D, and obtain the current AR device pose and the target pose; Step S3: Using the current AR device pose and the target pose, according to the point cloud map M, the shortest path from the current pose to the target pose is solved, and the path direction is optimized to finally obtain the optimized path; Step S4: Using the optimized path, in the AR device, the user is guided to move along the planned path by overlaying path instructions and voice prompts in the field of view, and the point cloud map M and memory database D are updated.

2. The AR navigation method based on the visual language model according to claim 1, characterized in that: The step S1 is specifically implemented as follows: the smart terminal moves in the target environment and continuously collects video data; based on the image sequence in the video data, it estimates the AR device posture p in real time. k , used to build or update the AR navigation point cloud map M, record the spatial position of each key frame, and use OCR text recognition technology to recognize the road sign text and determine the floor l where the AR device is located k ; Then, every consecutive Δt seconds of video segment Enter the AR Visual Language Memory Model (AR-VLM). 2 , the output corresponds to the natural language description L k ; If there are multiple floors, include descriptions of stairs or elevators; call the semantic encoding model to convert the natural language description L k Encoded as a semantic vector v k , and finally construct the real-time location semantic information memory entry for AR navigation (v k ,p k ,l k ,t k ) and stored in the AR navigation memory database D, where t k is the timestamp.

3. The AR navigation method based on the visual language model according to claim 2, characterized in that: The AR visual language memory model AR-VLM 2 , using a dual-stream heterogeneous fusion architecture of spatial and semantic streams, and performing semantic reorganization under geometric constraints through a neural symbol alignment layer, to generate navigation memory with both physical accuracy and scene adaptability. The specific implementation is as follows: Spatial Stream: Voxelize the attention weight matrix of the Transformer in the multimodal large language model (MLLM) to implement the spatiotemporal attention mechanism for AR navigation. This constitutes the AR navigation spatiotemporal understanding Transformer to parse the video stream captured by the AR device in real time, construct a multi-level attention 3D semantic field, and perform multi-scale spatial voxel-level semantic segmentation on the location of AR navigation attention. Semantic flow: Using decoupled language modeling, we design a large multimodal language model for AR navigation. We use AR navigation memory to generate prompts and decompose open vocabulary into spatial atomic concepts and dynamic relational predicates. At the same time, we integrate AR navigation semantic analysis tools, design a real-time path structure diagram, introduce learnable attention parameters between the path structure diagram and spatial atomic concepts, and train the AR-VLM. 2 The model determines the floor where the current AR device is located. k Navigation related information.

4. The AR navigation method based on the visual language model according to claim 3, characterized in that: In the spatial stream, RGB-D video frames are first mapped to a sparse voxel space through CLIP and 3D position encoding. A sparse-expanding multi-level architecture is then used to fuse local sparse attention with expanded attention. The sparse-expanding multi-level architecture proposes extending the Transformer attention mechanism naturally to the 3D voxel spatiotemporal domain, resulting in a voxel Transformer. Scale switching from the top-level overall to the bottom-level details is also embedded in the voxel Transformer. Furthermore, a cross-frame spatiotemporal attention mechanism is constructed within the voxel Transformer to distinguish dynamic occlusions and achieve a parallel understanding of dynamic and static spatial structures. The final output is a 3D semantic field feature with structural semantic annotations. In the semantic stream, a network architecture based on the symbol-geometry alignment layer is designed to deconstruct the language prompts into three-dimensional feature vectors of function descriptions, method descriptions and dynamic attributes, and inject them into the spatial model through the alignment layer to achieve deep fusion of language and spatial information. Specifically, the symbol-geometry alignment layer is responsible for establishing a direct mapping between language symbols and spatial voxel grids, and aligning the functional descriptions in the language prompts with the spatial positions one by one: the symbol nodes are embedded into the target voxels of the voxel grid through the structural graph attention to achieve the positioning of the language signal in the three-dimensional space; then, the multi-scale feature cross-fuser receives the 3D semantic field features of the spatial stream, parses and encodes them respectively, and fuses them with the symbolic semantics derived from the prompt encoding; the network architecture based on the symbol-geometry alignment layer constructs a slot-based attention network. Attention's graph attention mechanism captures the interaction between candidate paths and spatial structures, and on the other hand, performs self-supervised optimization through a cross-modal comparative learning strategy; using decoupled language modeling, a multimodal large language model for AR navigation is designed. With the help of AR navigation memory generation prompts, open vocabulary is decomposed into spatial atomic concepts and dynamic relational predicates, and AR navigation semantic analysis tools are integrated to perform text recognition on road signs along the way and update the real-time semantic information of the 3D semantic field; a real-time path structure diagram is designed, and the path structure diagram and spatial atomic concepts are aligned through learnable attention parameters to assist the network architecture based on the symbol-geometry alignment layer in judging navigation-related status information.

5. The AR navigation method based on the visual language model according to claim 4, characterized in that: The specific implementation of step S2 is as follows: querying the navigation target most relevant to the user's natural language request in the constructed memory database D, proposing a navigation request Q for the user, performing multiple rounds of reasoning on the request Q using the lightweight AR visual language navigation auxiliary matching model ARN-Pilot, and combining the current AR device pose p obtained by the visual positioning algorithm with image matching. i , select the optimal target pose p from the memory database D j as a navigation target.

6. The AR navigation method based on the visual language model according to claim 5, characterized in that: The lightweight AR visual language navigation-assisted matching model, ARN-Pilot, queries the AR navigation memory database D for the optimal navigation target based on real-time environmental conditions, user navigation target descriptions, and personalized needs. ARN-Pilot uses a multi-layer topology perception architecture to perform on-demand reasoning and matching of the optimal navigation target, and employs a distillation-projection joint compression method for training based on a large language model. The specific implementation is as follows: The ARN-Pilot model designs a three-tiered topology-aware architecture: The bottom layer pre-builds a multi-layered building graph network, where nodes are encoded physical coordinates and semantic features, and edge weights include actual path distance, travel difficulty, and dynamic pedestrian flow coefficients. The middle layer deploys a lightweight graph topology encoder based on the constructed building graph network. This encoder generates reachability vectors for location nodes through neighborhood aggregation, creating a multi-tiered topology-aware architecture. The top layer designs a dynamic decision tree as the path decision module based on a multi-layer topology-aware architecture, and designs a topology constraint retrieval algorithm to jointly reason about user intentions, real-time environment and graph topology to achieve personalized navigation; the training adopts a three-stage progressive strategy: first, the navigation target matching student model is pre-trained using the distillation-projection joint compression method, then contrastive learning is used to optimize the semantic-topological joint embedding space, and finally reinforcement learning is used to fine-tune the decision tree weights.

7. The AR navigation method based on the visual language model according to claim 6, characterized in that: The path decision module implements navigation based on a multi-layer building graph network and a multi-layer topology perception architecture: the building topology layer predefines cross-floor traffic rules and generates a global reachability vector through graph convolution; the floor topology layer analyzes the physical connection relationship of the actual walkable path to obtain candidate paths, and uses a gated graph attention network to dynamically weight the edge weights; the real-time perception layer integrates temporary obstacle information input by the AR device to construct an incremental graph update mechanism; the topology information of the building topology layer, floor topology layer and real-time perception layer are jointly input into the thinking tree reasoning engine. In the three-dimensional evaluation of distance, time and personalized constraints, the topology perception module provides structured features for each dimension, and scores the paths through learnable paths. The optimal path is output by a learnable weighted decision aggregator using a module, a dynamic traffic prediction module, and a user preference analysis module. The parameters of the path scoring module, the dynamic traffic prediction module, the user preference analysis module, and the weighted decision aggregator Ψ are learned using a linear fully connected layer. Topological adversarial sample enhancement is used during training to improve the robustness of the ARN-Pilot model. The path scoring module injects the actual edge weights of the graph network, the traffic predictor learns cyclical laws based on the temporal patterns of historical memory entries, and the user preference analysis module extracts walking habits through latent space clustering. The ARN-Pilot model uses a curriculum learning strategy during training, gradually transitioning from simple path decisions to complex scenarios with multiple constraints. When a user queries a target, the topology constraint retrieval algorithm first generates a k-hop reachable subgraph of the current node. A large model for AR navigation matching evaluation is designed, and the navigation path is constructed into a graph structure. After encoding using GNN, the structure is injected into the ARN-Pilot model in the form of a multi-tuple. The node semantic matching degree and path topology matching degree are simultaneously calculated within the subgraph, and the comprehensive matching degree is calculated: in is the semantic matching degree, is the path topology matching degree, α and β are the weights of the two, v q and v k are the graph vertices corresponding to user queries and memory items in the memory bank, respectively, e c and e k For the corresponding navigation edge; introduce the adaptive threshold generator Only keep the overall matching degree candidate targets, among which is the environmental complexity, is the learnable weight, and σ is the sigmoid function.

8. The AR navigation method based on the visual language model according to claim 7, characterized in that: The specific process of the distillation-projection joint compression method is as follows: first, the large language model LLM is used to generate the decision trajectory of the AR navigation target matching thinking tree as a soft label, and through knowledge distillation, the ARN-Pilot model based on the graph neural network GNN is guided to learn the reasoning path; then the topological projection loss function L is used to calculate the decision trajectory of the AR navigation target matching thinking tree as a soft label. project The lightweight graph topology encoder is forced to align with the graph topology semantics. The ARN-Pilot model uses a hierarchical activation mechanism during deployment. When the user's movement speed is detected to be greater than the movement threshold, it automatically switches to a minimalist mode, retaining only the single-layer GNN calculation key path nodes. In the stage of optimizing the semantic-topological joint embedding space, a hierarchical contrastive learning framework is proposed to achieve fine alignment of semantic and topological features. The hierarchical contrastive learning framework constructs joint contrast tasks of semantic pairs and topological pairs from a multimodal perspective by sampling positive and negative pairs and hierarchical negative samples. Specifically, in the optimization process of same / different sample comparison, the commonalities between concepts with the same function but different usage are captured, and the semantic embeddings of such concepts are brought closer through contrastive learning to enhance the relevance in the navigation semantic space. The sensitivity to the differences between such concepts is reflected through hierarchical contrastive learning loss. The framework adopts layer-by-layer temperature parameter adjustment and hierarchical negative sample mechanism, that is, in contrastive learning at different levels, different temperature parameters are set to control the smoothness of the similarity distribution, capture fine-grained path structure relationships at the bottom layer, enhance the semantic interconnection recognition ability between nodes at the middle layer, and maintain the consistency of the overall navigation planning and global semantic coherence at the high layer.

9. The AR navigation method based on the visual language model according to claim 8, characterized in that: The step S3 is specifically implemented as follows: using the obtained current AR device posture p i and the optimal target pose p j ,Based on the point cloud map M, a two-dimensional occupancy grid map M′ on the path is generated, marking the walkable and obstacle areas; If there are multiple floors, the floors involved are merged into a grid map through stairs and elevators; the A* algorithm is used to solve the problem of the current posture p in M′. i To the target pose p j The shortest path P is obtained, and the path direction is optimized by the negative gradient descent method. The path P is smoothed using the B-spline curve, and the optimized path P is finally obtained. * .

10. An AR navigation system based on a visual language model, used to implement the AR navigation method described in any one of claims 1 to 9, characterized in that: Includes the following modules: Pure visual map construction module: Captures real-time video streams, uses a monocular visual SLAM algorithm combined with a monocular depth estimation model to estimate the device position in real time and build an AR navigation point cloud map; Memory building module: For each sampling point in the point cloud map, the AR visual language memory model AR-VLM is used to store the continuous Δt seconds of video frame sequence near it. 2 Generate a natural language description and map it into a vector through a semantic encoding model, along with the timestamp, corresponding pose, and floor, and store it in the AR navigation memory database D; Query processing module: Based on the navigation request, the lightweight AR visual language navigation-assisted matching model ARN-Pilot is used to construct a query embedding vector in the memory database D. Vector similarity retrieval is performed on the semantic vectors in D. The optimal target location and its corresponding memory content are then located based on the current AR device position. Path planning module: Estimate the current AR device pose p based on the optimal target position and point cloud map i , use A* algorithm to solve p in the walkable area i The shortest path P to the target position is traced along the negative gradient and smoothed with B-spline to generate the optimized path P * ; Navigation interaction module: superimposes the optimized path in the user's AR field of view of the AR device, generates a real-time navigation guidance interface, and records real-time video during the navigation process, updating the point cloud map and memory database D.

Citation Information

Cited By

  • AR intelligent labeling method and system based on display content

    CN121010978A

  • AR intelligent labeling method and system based on display content

    CN121010978B

  • Geological point cloud thinning method and system

    CN121053476A

  • Training and pushing integrated video analysis method and system based on multi-modal large model

    CN121074764A

  • Intelligent agent autonomous navigation method and apparatus, and electronic device

    CN121140806A