Complex instruction driven navigation method based on cross-modal ontology collaborative active perception
By building dynamic grid maps and semantic maps in unknown environments and using the Transformer model to process complex instructions, the problem of navigation to a specified location is solved, and the success rate and sequential decision navigation performance of navigation tasks are improved.
Patent Information
- Application Number
- CN202510309938.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-17
- Publication Date
- 2025-06-17
AI Technical Summary
The prior art is difficult to effectively navigate to a designated location in unknown environments, especially when facing multi-step, fuzzy natural language instructions, and lacks generalization capabilities in new environments.
A complex instruction-driven navigation method based on cross-modal ontology collaborative active perception is adopted. By constructing a global grid map and a local active-aware semantic map with real-time dynamic accumulation, complex instructions are disassembled into multiple subtasks, and the multimodal features and logical values of candidate points are calculated through the Transformer model to determine the best candidate points.
Effectively completing multi-step, fuzzy instructions navigation tasks in unknown environments enhances generalization capabilities in new environments and improves the understanding and execution performance of complex instructions.
Smart Images

Figure CN120160611A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of visual language navigation, and particularly relates to a complex instruction-driven navigation method based on cross-modal ontology collaborative active perception. Background Art
[0002] Visual Language Navigation (VLN for short) requires an agent to navigate to a specified location according to natural language instructions. When the agent is initialized to an area unrelated to the instructions, especially in an unknown environment, enabling it to explore the area related to the instructions and reach the target location remains a major challenge. To enhance the agent's ability to explore the area related to the instructions in the scene, previous studies usually rely on imitation learning for natural language instructions to train the agent [1, 2, 3, 4, 5], but this method limits the generalization ability of the agent in a new environment.
[0003] Recent studies have utilized open-vocabulary object detection models [6, 7], such as OwlViT [8], or pre-trained vision-language models [9, 10] to establish the association between vision and language. However, if the instructions used for matching are always complete sentence instructions, these methods may lead to inaccurate region recognition results. For example, for the instruction "Turn slightly to the left and bypass the round table and chair. Then wait there.", it is necessary to disassemble the instruction before exploring the scene related to this instruction.
[0004] For a long time, researchers have proposed methods such as soft common sense constraints
[11] , dynamic extraction of constraint words
[12] , and human-machine dialogue [13, 14] to address the unknown environment VLN tasks with unclear goals and ambiguous descriptions. However, the above methods cannot simultaneously solve the following challenges:
[0005] (i) The generalization problem of random node initialization: It is very difficult to find the area related to the instructions in an unknown environment;
[0006] (ii) The task is complex and the description is ambiguous: Although the model is good at locating specific objects, it is difficult to navigate to the position required by natural language instructions, especially instructions containing relative position information and ambiguous descriptions;
[0007] (iii) Sequential decision-making navigation for multi-step tasks: If the robot only matches the observed scene with the entire instruction while ignoring the temporal information of the completed part of the instruction, it may be attracted by interfering objects during task execution and deviate from the route.
[0008] Prior Art Documents:
[0009] [1] Michael Chang, Arjun Gupta, and Saurabh Gupta. Semantic visual navigation by watching youtube videos. In Advances in Neural Information Processing Systems, pages 4283–4294, 2020.
[0010] [2] Rohan Chitnis, Tom Silver, ByungHyun Kim, et al. Camps: Learning context-specific abstractions for efficient planning in factored mdps. In Conference on Robot Learning, pages 64–79. PMLR, 2021.
[0011] [3] Nikhil Gireesh, Dinesh A S Kiran, Subhajit Banerjee, et al. Object goal navigation using data regularized q-learning. In 2022 IEEE 18th International Conference on Automation Science and Engineering (CASE), pages 1092–1097. IEEE, 2022.
[0012] [4] Tom Silver, Rohan Chitnis, Abigail Curtis, et al. Planning with learned object importance in large problem instances using graph neural networks. In Proceedings of the AAAI Conference on Artificial Intelligence, pages 11962–11971, 2021.
[0013] [5] Hao Wang, Albert GH Chen, Xiangyu Li, et al. Find what you want: learning demand-conditioned object attribute space for demand-driven navigation. Advances in Neural Information Processing Systems, 36, 2024.
[0014] [6] V.S. Dorbala, G. Sigurdsson, R. Piramuthu, et al. Clip-nav: Using clip for zero-shot vision-and-language navigation. arXiv preprint arXiv:2211.16649, 2022.
[0015] [7] Shantanu Y Gadre, Mitchell Wortsman, Gabriel Ilharco, et al. Cows on pasture: Baselines and benchmarks for language-driven zero-shot object navigation. In Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition, pages 23171–23181, 2023.
[0016] [8] Matthias Minderer, Xiaohua Zhai, Ingmar Daunhawer, Karsten Ferrante, Cordelia Schmid, Neil Houlsby, Alexey Dosovitskiy, and Georg Heigold. Simple open-vocabulary object detection with vision transformers. arXiv preprint arXiv:2205.06230, 2022.
[0017] [9] Wei Cai, Shuran Huang, Gang Cheng, et al. Bridging zero-shot object navigation and foundation models through pixel guided navigation skill. In 2024 IEEE International Conference on Robotics and Automation (ICRA), pages 5228–5234. IEEE, 2024.
[0018]
[10] Dai Y, Peng R, Li S, et al. Think, act, and ask: Open-world interactive personalized robot navigation[C] / / 2024 IEEE International Conference on Robotics and Automation(ICRA). IEEE, 2024: 3296-3303.
[0019]
[11] Zhou K, Zheng K, Pryor C, et al. Esc: Exploration with soft common sense constraints for zero-shot object navigation[C] / / International Conference on Machine Learning. PMLR, 2023: 42829-42842.
[0020]
[12] Hu Z, Pan J, Fan T, et al. Safe navigation with human instructions in complex scenes[J]. IEEE Robotics and Automation Letters, 2019, 4(2): 753-760.
[0021]
[13] Gao X, Gao Q, Gong R, et al. Dialfred: Dialogue-enabled agents for embodied instruction following[J]. IEEE Robotics and Automation Letters, 2022, 7(4): 10049-10056.
[0022]
[14] Thomason J, Padmakumar A, Sinapov J, et al. Improving grounded natural language understanding through human-robot dialog[C] / / 2019 International Conference on Robotics and Automation(ICRA). IEEE, 2019: 6934-6941. Summary of the Invention
[0023] The purpose of the present invention is to overcome the defects of the above-mentioned existing technologies and provide a complex instruction-driven navigation method based on cross-modal ontology collaborative active perception, which can complete multi-step and fuzzy instruction navigation tasks in an unknown environment, and solve the problems of insufficient generalization ability of previous methods in a new environment and limitations in dealing with fuzzy instructions.
[0024] The purpose of the present invention can be achieved through the following technical solutions: A complex instruction-driven navigation method based on cross-modal ontology collaborative active perception, including the following steps:
[0025] S1. Set up a complex instruction visual language navigation task in a discrete simulation environment;
[0026] S2. Construct a globally accumulated real-time dynamic grid map during the navigation process, and construct a local active perception semantic map;
[0027] S3. Classify the three-dimensional points in the local perception semantic map, and screen out candidate points based on the classification results;
[0028] S4. Decompose the complex instruction into multiple subtasks and encode them, then perform spatio-temporal alignment processing with the features of the candidate points, calculate the correlation between different candidate points and each subtask, and determine the multi-modal features of the candidate points;
[0029] S5. Input the multi-modal features of the candidate points and the global grid map features into the Transformer model, calculate the logical value set of the candidate points through the attention mechanism, and determine the best candidate point.
[0030] Further, the specific process of step S1 is as follows:
[0031] Given a discrete simulation environment of MatterPort3D and a complex natural language form instruction specifying a target location, randomize the initial node of the agent and require the agent to navigate from the initial node to the target location.
[0032] Further, step S2 includes the following steps:
[0033] S21. In each time step of the navigation task, the agent captures panoramic RGB-D images from multiple perspectives. Each RGB-D image contains an RGB image and a Depth image (i.e., a depth image);
[0034] S22. Use the pre-trained CLIP-ViT-B / 32 model to process each RGB image to obtain image features;
[0035] Extract three-dimensional points in the camera coordinate system from each Depth image and project them onto the grid map, and continuously accumulate historical grid information;
[0036] S23. Use an open-vocabulary object detector to process the RGB and Depth images to generate a semantic map composed of voxel points with semantic features.
[0037] Further, the specific process of step S22 is as follows:
[0038] Let represent the set of RGB images at time step t, represent the corresponding set of depth maps. For each image I t,k , use the pre-trained CLIP-ViT-B / 32 model to process it and extract image features of size E dim from the initial size (H, W), where E dim represents the feature vector dimension of each image;
[0039] Each depth map d t,k is then downsampled from size (h, w) to the specified size (h', w'), and the agent selects key perspectives [v1,..., v m and depth features from key patch positions [p1,..., p n . The RGB-D image data at these positions is then converted from the camera coordinate system to the Cartesian coordinate system, with the current position of the agent as the origin of the new coordinate system and the current facing direction as the positive y-axis.
[0040] Further, the specific process of step S23 is as follows:
[0041] At each navigation time step t, the agent obtains the panoramic image I of the current position local,t =[I local,t (1),…,I local,t (s)] and the corresponding depth map D local,t =[d local,t (1),…,d local,t (s)]. Each image and the total instruction T are processed by the OwlViT model to obtain the correlation features between each image and the instruction, denoted as where A dim is the dimension of the correlation features;
[0042] The coordinates are transformed from the camera coordinate system to the environment coordinate system and then to the voxel coordinate system through a triple coordinate conversion system (TCCS). First, the depth map is converted into 3D points in the camera coordinate system and scaled according to the field of view (FOV) and resolution of the camera to obtain the coordinates in the camera space coordinate system:
[0043]
[0044] Then, the camera coordinate system is converted to the environment coordinate system through the rotation matrix, and the transformation is completed by combining the translation matrix. Finally, the points in the world coordinate system are obtained. Assuming that the yaw angle (θ h ) represents the angle of rotation around the z-axis, and the pitch angle (θ b ) represents the angle of rotation around the x-axis, the total rotation matrix is the product of the two angle rotation matrices R total :
[0045]
[0046] The translation matrix T translation of the current position of the agent in the environment is as follows, where (x, y, z) are the coordinates of the current position of the agent:
[0047]
[0048] The final homogeneous transformation matrix is represented as T transform =R total *T translation . Given the homogeneous coordinates P=(X, Y, Z, 1) of a 3D point in the camera coordinate system, its point in the world coordinate system is P′=T transform *P. Finally, all the points transformed into the world coordinate system are scaled into the voxel coordinate system, and the height is adjusted to the height of the agent. Each transformed point is stored as a node in the semantic map, including attributes such as its camera coordinate system coordinates, correlation with the instruction, and distance to the nearest candidate point.
[0049] Furthermore, step S3 specifically involves establishing edge relationships between voxel points based on the attribute information of nodes in the semantic map, classifying each voxel point into an occupied point, a boundary point, or a free point, and classifying the navigation phase into an exploration phase or a target localization phase.
[0050] Furthermore, the specific process of step S3 is as follows:
[0051] In the semantic map, the types of nodes are classified according to the agent's height and the depth value to the camera. Among them, points with a depth value within a preset threshold range from the agent are marked as occupied points of type.Occupied, representing obstacles; points with adjacent voxels but not adjacent on all four sides are classified as boundary points of type.Frontier; points without neighbors are discarded and marked as isolated points; the remaining points are marked as free points of type.Free;
[0052] When the agent detects a point of interest (PoI) in the RGB-D image, that is, when the relevance between the voxel point and the instruction is greater than the set threshold, it enters the target localization phase and selects the navigable point closest to each PoI as the candidate point set;
[0053] If no PoI is detected, the agent enters the exploration phase and preferentially selects the boundary point closest in distance to guide it to move towards the unexplored area;
[0054] Subsequently, a heuristic function and a priority queue are used to perform advanced filtering on the nodes, tracking the visited nodes, adding the historical nodes and the current node to the access queue, and preferentially processing the unvisited neighbors according to the heuristic value. During the filtering process, high-cost points are removed, and only the points with the lowest heuristic value are retained for expansion. The shortest distance from node u to node v is calculated as:
[0055]
[0056] where P uv represents the possible path from node u to node v, (i,j) are the points passed by this path, and w(i,j) is the heuristic cost from i to j.
[0057] Furthermore, the specific process of step S4 is as follows: Use the LlaVA (Large Language and Vision Assistant) model to decompose the natural language instruction into executable fragments {T1, T2, …, T q}, and then use the BERT (Bidirectional Encoder Representations from Transformers) model to encode each segment to obtain where E dim is the encoding dimension, and L i is the number of tokens in the i-th language segment;
[0058] At execution time step t, calculate the correlation between each candidate point and the instruction segment. For each segment T i , calculate the candidate point grid feature G t and the correlation matrix of the segment feature S i ; Assign higher priority to the unexecuted segments through the weight w i to obtain the weighted total correlation vector, where is the number of three-dimensional points recorded in the global grid map at the t-th step:
[0059]
[0060] Set the correlation feature corresponding to the i-th grid cell in where n i represents the number of features corresponding to the i-th grid cell, and weight the grid feature G t . If multiple candidate points are mapped to the same grid cell Cell i , then calculate the weighted sum of these features to obtain the final feature of this grid cell:
[0061]
[0062] Furthermore, the specific step S5 is to use the self-attention mechanism and cross-attention mechanism of the Transformer model to process the context relationship between the global grid feature and the candidate point feature, as well as the correlation between the instruction feature and the visual feature, and use the action generation module to calculate the best candidate point from the fused visual and text features.
[0063] Furthermore, the specific process of the step S5 is as follows:
[0064] Use the self-attention mechanism of the Transformer to calculate the dependency relationship between the global map feature G t and the candidate point feature , where r represents the number of candidate points;
[0065] Subsequently, the global map feature G t is merged with the candidate point feature C t and cross-attention is applied to calculate their correspondence with the instruction, generating the attention feature A of the first stage t,1 , capturing the alignment between the global map and the candidate points under the given instruction;
[0066] A t,1 is merged with the instruction feature into a new key-value pair (KV), and the candidate point feature and the current position information are merged into a new query (Q) to generate the attention feature A of the second stage t,2 , and finally the logical value set of the candidate points is output and the best candidate point is determined.
[0067] Compared with the prior art, the present invention has the following advantages:
[0068] The present invention constructs a real-time dynamically accumulated global grid map during the navigation process and constructs a local active perception semantic map; classifies the three-dimensional points in the local perception semantic map and filters the candidate points based on the classification results; disassembles the complex instruction into multiple subtasks and encodes them, and then performs spatio-temporal alignment processing with the features of the candidate points to calculate the correlation between different candidate points and each subtask, determining the multi-modal features of the candidate points; then inputs the multi-modal features of the candidate points and the global grid map features into the Transformer model, and calculates the logical value set of the candidate points and determines the best candidate point through the attention mechanism. Thus, a hybrid map combining the dynamically accumulated and growing global grid map and the local semantic map is used to efficiently store multi-modal scene information. On this basis, by disassembling complex instructions to align different subtask descriptions with the visual information of different scene regions, autonomous perception exploration and sequential decision-making navigation can be achieved, so as to complete the navigation task of multi-step and ambiguous instructions in an unknown environment, solving the problems of insufficient generalization ability of previous methods in new environments and limitations in dealing with ambiguous instructions.
[0069] The present invention enhances the generalization ability of the randomized initial node and more accurately locates the region related to the instruction by constructing a dynamically expanding global grid map and matching it with a semantic map that focuses on local information.
[0070] When constructing the global map, the present invention retains the point coordinates mapped in the global coordinate system before, and only updates the map with the three-dimensional points obtained from the new panoramic RGB-D image features, which helps to efficiently store historical trajectory information and enables the intelligent agent to reason based on previous observations.
[0071] The present invention uses a local semantic map to screen candidate points. The agent classifies the candidate points into boundary points, free points (points where walking is possible), and occupied points (points occupied by objects) according to the local semantic map. According to the candidate point type, it is divided into an exploration stage and a target localization stage (i.e., the harvesting stage). If there are points with a relevance to the instruction exceeding a certain threshold, target localization is performed and points are selected. If not, the nearest boundary point is selected, which can ensure the accuracy of the initial screening of candidate points.
[0072] The present invention enables the large language model LLaVA to understand the abstract semantics of complex natural language instructions, decomposes the instructions into multiple executable subtask segments, then encodes each segment of the instructions through the BERT model to generate a semantic vector representation consistent with the visual feature embedding dimension, calculates the correlation matrix between the screened candidate points and different subtask segments in the instructions, and obtains which subtask segment the candidate points are most relevant to. Then, the visual features of the corresponding candidate points are obtained from the global grid map, which can establish an accurate connection between the instruction segmentation and the environmental features, thereby improving the clarity and performance of executing instructions in an unknown environment.
[0073] The present invention uses the attention mechanism of the Transformer model to establish the relevance of multimodal information and calculates the best next candidate point. On the one hand, it uses the self-attention mechanism of the Transformer to calculate the global map feature G t and the candidate point feature C t the dependence relationship between them. On the other hand, it combines the global map feature G t and the candidate point feature C t and applies cross-attention to calculate their correspondence with the instruction to generate the attention feature A t,1 for capturing the alignment between the global map and the candidate points under the given instruction. Among them, the self-attention mechanism helps the model focus on the relationship between the candidate points and the trajectory history, while the cross-attention mechanism emphasizes the part of the candidate points more relevant to the instruction, which can ensure obtaining accurate best candidate points. BRIEF DESCRIPTION OF THE DRAWINGS
[0074] Figure 1 is a schematic diagram of the method flow of the present invention;
[0075] Figure 2 is a schematic diagram of the application framework of the embodiment;
[0076] Figure 3 is a schematic diagram of the process of constructing a local semantic map in the embodiment;
[0077] Figure 4 is a schematic diagram of the candidate point screening process in the embodiment;
[0078] Figure 5Schematic diagram of the navigation trajectory comparison between the method of the present invention and other existing methods in the embodiments;
[0079] Figure 6 Schematic diagram of spatially aligning different language segments with different regions in the embodiments. Detailed implementation manners
[0080] The present invention will be described in detail below with reference to the accompanying drawings and specific embodiments.
[0081] Embodiment
[0082] A complex instruction-driven navigation method based on cross-modal ontology collaborative active perception, comprising the following steps:
[0083] S1. Set a complex instruction visual language navigation task in a discrete simulation environment;
[0084] S2. Construct a globally accumulated grid map in real-time dynamics during the navigation process, and construct a local active perception semantic map;
[0085] S3. Classify the three-dimensional points in the local perception semantic map, and screen out candidate points based on the classification results;
[0086] S4. Decompose the complex instruction into multiple subtasks and encode them, then perform spatio-temporal alignment processing with the features of the candidate points, calculate the correlation between different candidate points and each subtask, and determine the multi-modal features of the candidate points;
[0087] S5. Input the multi-modal features of the candidate points and the global grid map features into the Transformer model, calculate the logical value set of the candidate points through the attention mechanism, and determine the best candidate point.
[0088] In this embodiment, the above solution is applied to build a zero-shot visual language navigation framework Ali-UI, as Figure 2 shown, mainly including: the local semantic map module uses the OwlViT model to associate natural language instructions with RGB images, extracts three-dimensional point clouds from depth images, and voxelizes these point clouds into a semantic map for task classification and candidate point selection;
[0089] The global grid map module aligns these three-dimensional points with the visual features encoded by CLIP-ViT-B / 32 to generate a top-down continuously accumulated grid map;
[0090] By performing a weighted max-pooling operation on the language segments encoded by BERT and the grid features, the correlation between the grid features and different language segments is further refined;
[0091] Finally, combining historical and current data, through two-layer Transformer processing with an attention mechanism, the final candidate points are generated.
[0092] The specific application process of this embodiment is as follows:
[0093] Step 1: This embodiment is evaluated on the Matterport simulator. 100 instructions from 10 unseen scenes are selected from each of the three indoor simulation environment visual language navigation datasets, namely R2R, REVERIE, and SOON, as the evaluation tasks. The evaluation metrics set in this embodiment are trajectory length (TL), navigation error (NE), success rate (SR), and standardized path length (SPL); object recognition success rate (RGS) and object recognition success rate penalized by path length (RGSPL). In addition, in order to evaluate the sequential decision-making navigation performance, this embodiment also introduces a new metric, sequence match degree (SMD), which counts the correctly ordered subtask pairs during the execution process and normalizes them by the total number of pairs. The higher the value, the better the sequential task execution.
[0094] When setting up the complex instruction visual language navigation task, given a discrete simulation environment of MatterPort3D and a complex natural language form instruction specifying the target location, randomize the initial node of the agent. Require the agent to navigate from the initial node to the target location. During the navigation process, obtain the RGB-D images of multiple panoramic views at each step of the agent and the corresponding camera parameters. Use the simulation function to obtain the candidate points existing in a certain image and their three-dimensional point positions in the environment. The purpose of this solution is to select the best candidate points around each step for the agent.
[0095] Step 2: When constructing the global dynamic grid map, the agent captures panoramic RGB-D images from 36 views at each time step t of the navigation task. Let represent the set of RGB images at time step t, represent the corresponding set of depth maps. Each image I t,k is processed using the pre-trained CLIP-ViT-B / 32 model to extract image features of size E dim from the initial size (H, W), where E dim represents the feature vector dimension of each image. Each depth map d t,k is downsampled from the size (h, w) to the specified size (h', w'). The agent selects from the key viewpoints [v1,…,v m and key patch positions [p1,…,p nExtract the features of each key depth map. The RGB-D image data at these key positions are extracted into 3D points in the camera coordinate system and then transformed from the camera coordinate system to the ego-centered Cartesian coordinate system, where the current position of the agent serves as the origin and the forward direction is the positive y-axis. This system enables the agent to construct and dynamically expand its map as it moves in the environment. The panoramic image features are processed by a deep learning model including a linear layer, a LayerNorm layer, and a BERT-based model. Different types of features, such as image features and position features, are transformed into a unified embedding space E dim , and concatenated. The average panoramic feature is obtained by calculating the average feature of all viewpoints as the panoramic feature of the current position. For each candidate point, the visual feature of the image corresponding to its viewpoint is selected as the visual feature of each candidate point. At each time step, when constructing the global map, this scheme retains the point coordinates mapped in the global coordinate system before and only updates the map with the 3D points obtained from the new panoramic RGB-D image features, which helps to efficiently store historical trajectory information and enables the agent to reason based on previous observations.
[0096] Step 3: To construct a local semantic map, this scheme installs an open-vocabulary object detector on the agent's camera to process the RGB and depth images and generate a semantic map composed of voxel points with semantic features. The agent distinguishes between the exploration phase and the harvesting phase based on the presence of interesting voxel points in the scene. At each navigation time step t, the agent obtains the panoramic image of the current position from the viewpoints containing candidate points, denoted as I local,t =[I local,t (1),…,I local,t (s)] and the corresponding depth map D local,t =[d local,t (1),…,d local,t (s)]. Each image and the total instruction T are processed by the OwlViT model to obtain the correlation features between each image and the instruction, denoted as where A dim is the dimension of the correlation features.
[0097] This scheme transforms the coordinates from the camera coordinate system to the environment coordinate system and then to the voxel coordinate system through a triple coordinate conversion system (TCCS). First, the depth map is converted into 3D points in the camera coordinate system and scaled according to the field of view (FOV) and resolution of the camera to obtain the camera space coordinates. The following is the coordinate conversion formula:
[0098]
[0099] Then, the camera coordinate system is transformed into the environmental coordinate system through a rotation matrix, and the transformation is completed by combining the translation matrix, and finally the points in the world coordinate system are obtained. Specifically, assuming that the yaw angle (θ h ) represents the angle of rotation around the z-axis, and the pitch angle (θ b ) represents the angle of rotation around the x-axis, then the total rotation matrix is the product of the two angle rotation matrices R total :
[0100]
[0101] The translation matrix T translation of the current position of the agent in the environment is as follows, where (x, y, z) are the coordinates of the current position:
[0102]
[0103] The final homogeneous transformation matrix can be expressed as T transform = R total * T translation . Given the homogeneous coordinates (X, Y, Z, 1) of a 3D point in the camera coordinate system, its point in the world coordinate system is P' = T transform * P. Finally, all the points transformed into the world coordinate system are scaled into the voxel coordinate system, and the height is adjusted to match the height of the agent. Each transformed point is stored as a node in the semantic map, and attributes such as its coordinates in the camera coordinate system, relevance to the instruction, and distance to the nearest candidate point are recorded. Based on this information, the edge relationship between voxel points is established, and each voxel point is classified. According to these classifications, it is judged whether each voxel is a boundary point or a point of interest (PoI, Point of Interest), and then the navigation stage is classified as the exploration or harvest stage. If there is a point of interest, it is judged as the harvest stage - that is, the target localization stage, otherwise it is the exploration (Explore) stage.
[0104] Figure 3 Shows how the agent constructs local semantic maps from different perspectives in this embodiment. First, the points extracted from the depth map are represented in the camera coordinate system, and then through calculating the camera rotation matrix, these points are transformed into the three-dimensional coordinates of the real world, and each three-dimensional point is matched with the nearest candidate point. Subsequently, the points in the camera coordinate system are voxelized into the semantic map, and the points are classified by type.
[0105] Step 4, as Figure 4As shown in the figure, in the semantic map, the types of nodes are classified according to the agent's height and the depth value to the camera. Points with a depth value within a certain threshold range from the agent are marked as occupied points of type.Occupied, representing obstacles; points with adjacent voxels but not surrounded by neighbors on all sides are classified as frontier points of type.Frontier; points without neighbors are discarded and marked as isolated points; the remaining points are marked as free points of type.Free. This solution uses the semantic map for the first candidate point screening. When the agent detects a point of interest (PoI) in the RGB-D image, that is, when the relevance between the voxel point and the instruction is greater than a certain threshold, it enters the harvest stage (Harvest), and selects the navigable point closest to each PoI as the candidate point set; if no PoI is detected, the agent enters the exploration stage (Explore) and preferentially selects the nearest frontier point to guide it to move towards the unexplored area. Subsequently, the present invention uses a heuristic function and a priority queue to further perform advanced screening on the nodes and track the visited nodes. Both the historical nodes and the current nodes are added to the visited queue, and the unvisited neighbors are preferentially processed according to the heuristic value. During the filtering process, points with high heuristic costs are removed, and only the points with the lowest heuristic value are retained for subsequent screening. The present invention takes the distance of the agent from the current position u, passing through different candidate points, to the target position v as the heuristic cost of different candidate points. Specifically, the shortest distance from node u to node v is calculated as follows, where P uv represents a possible path from node u to node v, (i, j) is the path from node i to node j on this path, and w(i, j) is the heuristic cost from i to j:
[0106]
[0107] The above navigation candidate point selection strategy based on the best-first search algorithm is shown in Table 1. Adopting this candidate point selection strategy enables the agent to dynamically adjust its actions according to the environment and the progress of the task.
[0108] Table 1
[0109]
[0110] After obtaining the points of interest (POI) and frontier points from the semantic map, this solution discards the following types of points: points that are farther from the target than the current point, points with a later order in the most matching segment but a farther distance to the target point, and points that are not visible in the perspective. Finally, the relevance between the filtered candidate points and the language segments is constructed and input into the Transformer layer to output the best candidate points.
[0111] Step 5: Use a large language model to disassemble the instructions and spatially align the disassembled language fragments with visual features. For example, the instruction "Leave the room through the left door, turn slightly left, pass by the round table, and then wait there" contains spatial descriptions (such as "left door"), order descriptions (such as "turn slightly left"), and abstract semantics (such as "wait there"). To address the challenges of temporal understanding and abstract semantic understanding in such instructions, this embodiment uses the LLaVA model to decompose the instructions into executable fragments {T1, T2, …, T q}, and then uses the BERT model to encode each fragment to obtain where E dim is the encoding dimension, and L i is the number of tokens in the i-th language fragment.
[0112] At execution time step t, calculate the correlation between each candidate point and the instruction fragment. For each fragment T i , calculate the correlation matrix between the candidate point grid feature G t and the fragment feature S i : Assign higher priority to unexecuted fragments through the weight w i to obtain the weighted total correlation vector, where is the number of three-dimensional points recorded in the global grid map at the t-th step:
[0113]
[0114] Set the correlation feature corresponding to the i-th grid cell in to i where n i represents the number of features corresponding to the i-th grid cell. Weight the grid feature G t . If multiple candidate points map to the same grid cell Cell i , then calculate the weighted sum of these features to obtain the final feature of this cell:
[0115]
[0116] That is to say, to ensure execution in the order of the instructions, this solution gives priority to the candidate points that best match the unexecuted segments, and then dynamically assigns weights to each segment feature, with the weights increasing as the segments appear in the order of the instructions.
[0117] Step 6: The action generator in this embodiment adopts a similar architecture to GridMM. The key idea is to use the self-attention mechanism of the Transformer to calculate the global map feature G t and the candidate point feature The dependencies among them, where r represents the number of candidate points. Subsequently, the global map feature G t is merged with the candidate point feature C t and cross-attention is applied to calculate their correspondence with the instruction, generating the attention feature A t,1 in the first stage, which is used to capture the alignment between the global map and the candidate points under the given instruction. Thus, the self-attention mechanism and the cross-attention mechanism of the Transformer model are used to process the context relationship between the global grid feature and the candidate point feature, as well as the correlation between the instruction feature and the visual feature. For further optimization, the model merges A t,1 with the instruction feature into a new key-value pair (KV), while the candidate point feature and the current position information are merged into a new query (Q), finally generating the attention feature A t,2 in the second stage. The self-attention mechanism helps the model focus on the relationship between the candidate points and the trajectory history, while the cross-attention mechanism emphasizes the part of the candidate points that is more relevant to the instruction. The final output of the network is a set of logical values of the candidate points, which serves as a guiding score for the action selection process, and the best candidate point is calculated from the fused visual and text features.
[0118] To verify the effectiveness of this solution, for the same complex instruction, this embodiment respectively uses this solution and the existing method to generate visual navigation trajectories, as Figure 5 shown, where the initial node for evaluation is close to the target position (blue), while the randomized initial node is far from the target position (green). The method proposed in this solution successfully generates a trajectory - first locates to the bedroom and then reaches the stairs, while the current state-of-the-art VLFM method fails (orange) - the agent keeps wandering in the storage room.
[0119] In addition, when the agent aligns different language segments with different regions, different from the existing method, as Figure 6 shown, the Ali-UI system of this solution matches the language segments with the regions, constructs a semantic map to determine which perspectives need to perform the "Harvest" task, and selects the candidate point that best matches the target segment.
[0120] In summary, this scheme constructs a dynamically accumulated and growing global grid map and local semantic map to store multimodal scene information, and uses a large model to disassemble complex instructions and align subtask descriptions with scene information, thereby improving the navigation success rate of complex instructions and sequential decision navigation performance of randomized initial nodes. By introducing a dynamically expanded global grid map and matching it with a semantic map that focuses on local information, the generalization ability of the randomized initial node is enhanced to more accurately locate the area related to the instruction. In order to improve the understanding of complex and ambiguous instructions, since it is observed that even large pre-trained models have difficulty decomposing verbal instructions into clear multi-step subtasks, this scheme establishes a matching mechanism between shorter but still abstract language fragments and different regions in the scene. To achieve sequential decision navigation, this scheme first filters candidate points according to the order of the most matching segment in the total instruction, and then calculates the correlation between the filtered candidate points and the complete instruction enhanced by local attention. This method improves the clarity and performance of executing instructions in unknown environments. This scheme has been extensively experimented on multiple authoritative complex instruction datasets R2R, REVERIE, and SOON, and achieved the best results in key navigation task indicators SR (success rate) and SRL (path length weighted success rate). Compared with the most cutting-edge Zero-Shot navigation technology VLFM, its performance on multiple datasets has improved by at least 19% in terms of success rate.
Claims
1. A complex instruction driven navigation method based on cross-modal ontology collaborative active perception, characterized in that: The following steps are involved: S1. Set up complex instruction visual language navigation tasks in a discrete simulation environment; S2, construct a real-time dynamic global grid map during navigation, and construct a local active perception semantic map; S3, classifying the three-dimensional points in the local perception semantic map, and screening candidate points based on the classification results; S4, decompose the complex instruction into multiple subtasks and encode them, then perform spatiotemporal alignment with the features of the candidate points, calculate the correlation between different candidate points and each subtask, and determine the multimodal features of the candidate points; S5. Input the multimodal features of the candidate points and the global grid map features into the Transformer model, calculate the logical value set of the candidate points through the attention mechanism, and determine the best candidate point.
2. According to claim 1, a complex instruction driven navigation method based on cross-modal ontology collaborative active perception is characterized in that: The specific process of step S1 is as follows: Given a MatterPort3D discrete simulation environment and a complex natural language instruction specifying the target location, the agent's initial node is randomized and the agent is required to navigate from the initial node to the target location.
3. According to claim 1, a complex instruction driven navigation method based on cross-modal ontology collaborative active perception is characterized in that: The step S2 comprises the following steps: S21. The agent captures panoramic RGB-D images from multiple perspectives at each time step of the navigation task, where each RGB-D image contains an RGB image and a Depth image, i.e., a depth image; S22, using the pre-trained CLIP-ViT-B / 32 model to process each RGB image to obtain image features; Extract the 3D points in the camera coordinate system from each Depth image and project them to the grid map, and continuously accumulate historical grid information; S23. Use an open vocabulary object detector to process RGB and Depth images to generate a semantic map consisting of voxel points with semantic features.
4. According to claim 3, a complex instruction driven navigation method based on cross-modal ontology collaborative active perception is characterized in that: The specific process of step S22 is as follows: set up represents the set of RGB images at time step t, Represents the corresponding depth map set, for each image I t,k , using the pre-trained CLIP-ViT-B / 32 model for processing, extracting a size of E from the initial size (H, W) dim The image features of E dim Represents the feature vector dimension of each image; Each depth map d t,k It will be downsampled from size (h, w) to the specified size (h′, w′), and the agent selects the key view [v1, ..., v m ] and from the key patch positions [p1, ...p n ], the RGB-D image data at these locations are then transformed from the camera coordinate system to a Cartesian coordinate system, where the agent’s current position is used as the origin of the new coordinate system and the current heading direction is the positive y-axis.
5. According to claim 4, a complex instruction driven navigation method based on cross-modal ontology collaborative active perception is characterized in that: The specific process of step S23 is as follows: At each navigation time step t, the agent obtains a panoramic image I of the current location local,t =[I local,t (1), ..., I local,t (s)] and the corresponding depth map D local,t =[d local,t (1), .., d local,t (s)], each image and the total instruction T are processed by the OwlViT model to obtain the correlation feature between each image and the instruction, denoted as Among them A dim It is the dimension of correlation characteristics; The coordinates are converted from the camera coordinate system to the environment coordinate system and then to the voxel coordinate system through the triple coordinate transformation system (TCCS). First, the depth map is converted to a 3D point in the camera coordinate system, and then scaled according to the camera's field of view (FOV) and resolution to obtain the camera space coordinate system coordinates: Then, the camera coordinate system is transformed into the environment coordinate system through the rotation matrix, and then the translation matrix is combined to complete the transformation, and finally the point in the world coordinate system is obtained. Assuming the yaw angle (θ h ) represents the angle of rotation around the z-axis, and the pitch angle (θ b ) represents the angle of rotation around the x-axis, then the total rotation matrix is the product of the two angle rotation matrices R total : The translation matrix T of the agent's current position in the environment translation As follows, where (x, y, z) are the coordinates of the agent's current position: The final homogeneous transformation matrix is represented as T transform =R total *T translation , given a 3D point with homogeneous coordinates P = (X, Y, Z, 1) in the camera coordinate system, its point in the world coordinate system is P′ = T transform *P,Finally, all points converted to the world coordinate system are scaled back to the voxel coordinate system and the height is adjusted to the height of the agent. Each converted point is stored as a node in the semantic map, containing attributes such as its camera coordinate system coordinates, its relevance to the instruction, and its distance to the nearest candidate point.
6. The complex instruction driven navigation method based on cross-modal ontology collaborative active perception according to claim 3 is characterized in that: The step S3 specifically establishes edge relationships between voxel points based on the attribute information of nodes in the semantic map, and classifies each voxel point into occupied points, boundary points or idle points, and classifies the navigation phase into an exploration phase or a target positioning phase.
7. The complex instruction driven navigation method based on cross-modal ontology collaborative active perception according to claim 6 is characterized in that: The specific process of step S3 is: In the semantic map, the type of node is classified according to the height of the agent and the depth value to the camera, where the point whose depth value with the agent is within the preset threshold is marked as occupied point type.Occupied, representing an obstacle; points with adjacent voxels but not surrounded by neighbors on all four sides are classified as boundary points type.Frontier; points without neighbors are discarded and marked as isolated points; the rest of the points are marked as free points type.Free; When the agent detects a point of interest (PoI) in the RGB-D image, that is, when the correlation between the voxel point and the instruction is greater than the set threshold, it enters the target positioning stage and selects the closest navigable point to each PoI as the candidate point set; If no PoI is detected, the agent enters the exploration phase and prioritizes the nearest boundary point, guiding it to move toward the unexplored area; Then, the heuristic function and priority queue are used to perform advanced screening of nodes, track visited nodes, add historical nodes and current nodes to the visit queue, and prioritize unvisited neighbors according to the heuristic value. During the filtering process, high-cost points are removed, and only the points with the lowest heuristic value are retained for expansion. The shortest distance from node u to node v is calculated as: Where P uv represents a possible path from node u to node v, (i, j) is the point passed by the path, and w(i, j) is the heuristic cost from i to j.
8. The complex instruction driven navigation method based on cross-modal ontology collaborative active perception according to claim 3 is characterized in that: The specific process of step S4 is: using the LlaVA model to decompose the natural language instruction into executable segments {T1, T2, ..., T q }, and then the BERT model is used to encode each fragment to obtain Where E dim is the encoding dimension, L i is the number of tokens in the i-th language segment; When executing time step t, the correlation between each candidate point and the instruction fragment is calculated. For each fragment T i , calculate the candidate point grid feature G t and the fragment feature S i The correlation matrix By weight w i Give higher priority to the fragments that have not been executed, so as to obtain a weighted total dependency vector, where is the number of 3D points recorded in the global grid map at step t: The correlation feature corresponding to the i-th grid unit in is set to Where n i Represents the number of features corresponding to the i-th grid unit, for the grid feature G t Weighted, if multiple candidate points Mapped to the same grid cell i , then the weighted sum of these features is calculated to obtain the final features of the grid unit:
9. The complex instruction driven navigation method based on cross-modal ontology collaborative active perception according to claim 8, characterized in that: The step S5 specifically uses the self-attention mechanism and cross-attention mechanism of the Transformer model to process the contextual relationship between the global grid features and the candidate point features, as well as the correlation between the instruction features and the visual features, and uses the action generation module to calculate the best candidate point from the fused visual and text features.
10. The complex instruction driven navigation method based on cross-modal ontology collaborative active perception according to claim 9, characterized in that: The specific process of step S5 is as follows: Use Transformer's self-attention mechanism to calculate the global map feature G t With candidate point features The dependency relationship between them, where r represents the number of candidate points; Then the global map feature G t With the candidate point feature C t Merge and apply cross attention to calculate their correspondence with the instruction to generate the first stage attention feature A t,1 ,captures the alignment of the global map and candidate points under a given instruction; A t,1 Merge with the instruction feature into a new key-value pair (KV), merge the candidate point feature and the current position information into a new query (Q), and generate the second stage attention feature A t,2 , and finally output the logical value set of candidate points and determine the best candidate point.
Citation Information
Cited By
Autonomous exploration method based on multi-modal data fusion and large language model
CN121500330A
An autonomous exploration method based on multi-modal data fusion and large language model
CN121500330B
Robot navigation method and device, electronic equipment, storage medium and program product
CN121632124A