Visual language navigation method and device based on large model and factor graph fusion

By employing a visual language navigation method that integrates large models and factor graphs, the robustness of visual language navigation in the real world is addressed. This enables robots to perform navigation tasks efficiently and accurately in unknown environments. Furthermore, by combining semantic and geometric optimization capabilities, the stability and adaptability of navigation are improved.

CN122108137APending Publication Date: 2026-05-29GUANGZHOU CHAOZHI MACHINERY CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
GUANGZHOU CHAOZHI MACHINERY CO LTD
Filing Date
2026-02-28
Publication Date
2026-05-29

AI Technical Summary

Technical Problem

Existing technologies in visual language navigation suffer from a performance gap between simulation and reality. The model decision-making process lacks interpretability and is difficult to adapt to the real world. Furthermore, the symbol sequences output by large language models cannot be seamlessly integrated with the robot's underlying geometry-based continuous state estimation, resulting in insufficient robustness in complex physical environments.

Method used

A method based on large model and factor graph fusion is adopted. Natural language navigation instructions are parsed by a large language model to generate a language inference prior factor graph model, which is then fused with a factor graph model with localization and map building capabilities. The factor graph model is optimized using real observation data during robot navigation, and navigation target points are dynamically selected to enable the robot to navigate accurately in unfamiliar environments.

Benefits of technology

It enables robots to execute complex natural language commands in unknown environments with zero samples and high robustness. By combining semantic and geometric optimization capabilities, it improves the accuracy and stability of navigation and adapts to complex environmental changes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122108137A_ABST
    Figure CN122108137A_ABST
Patent Text Reader

Abstract

The application discloses a visual language navigation method and device based on a large model and a factor graph fusion, wherein the method comprises the following steps: receiving an input natural language navigation instruction, analyzing the natural language navigation instruction by using a large language model to obtain key spatial relationship information, and generating a language inference prior factor graph model based on the key spatial relationship information; acquiring a first factor graph model with simultaneous localization and map building capability, and constructing a second factor graph model based on the first factor graph model and the language inference prior factor graph model; optimizing the second factor graph model by using real observation data in a navigation path detected by a robot in a navigation process to obtain an optimized second factor graph model; determining a navigation target point to which the robot is about to move based on the optimized second factor graph model, and controlling the robot to move to the navigation target point until the destination is reached. The method has the effect of performing a navigation task of a natural language instruction with zero samples and high robustness.
Need to check novelty before this filing date? Find Prior Art