Visual language navigation method and device based on large model and factor graph fusion
By employing a visual language navigation method that integrates large models and factor graphs, the robustness of visual language navigation in the real world is addressed. This enables robots to perform navigation tasks efficiently and accurately in unknown environments. Furthermore, by combining semantic and geometric optimization capabilities, the stability and adaptability of navigation are improved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- GUANGZHOU CHAOZHI MACHINERY CO LTD
- Filing Date
- 2026-02-28
- Publication Date
- 2026-05-29
AI Technical Summary
Existing technologies in visual language navigation suffer from a performance gap between simulation and reality. The model decision-making process lacks interpretability and is difficult to adapt to the real world. Furthermore, the symbol sequences output by large language models cannot be seamlessly integrated with the robot's underlying geometry-based continuous state estimation, resulting in insufficient robustness in complex physical environments.
A method based on large model and factor graph fusion is adopted. Natural language navigation instructions are parsed by a large language model to generate a language inference prior factor graph model, which is then fused with a factor graph model with localization and map building capabilities. The factor graph model is optimized using real observation data during robot navigation, and navigation target points are dynamically selected to enable the robot to navigate accurately in unfamiliar environments.
It enables robots to execute complex natural language commands in unknown environments with zero samples and high robustness. By combining semantic and geometric optimization capabilities, it improves the accuracy and stability of navigation and adapts to complex environmental changes.
Smart Images

Figure CN122108137A_ABST