Indoor Navigation Knowledge Graph for Abstract Instruction Reasoning
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current vision-language navigation technologies face challenges in understanding abstract instructions and making reasonable navigation decisions in indoor environments, due to limited cross-modal understanding and reasoning capabilities between visual and language modalities.
Innovation Solution
The proposed indoor navigation method enhances cross-modal understanding by extracting and fusing instruction features with visual features using a knowledge graph, and introduces entity knowledge reasoning to improve the decision-making process.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If vision-language navigation model uses basic visual and language features for navigation, then the system complexity is low, but the navigation decision accuracy deteriorates due to limited cross-modal understanding
Solution Approach 1:
The patent introduces a knowledge graph as an intermediary component that bridges visual and language modalities. The knowledge graph stores pre-defined spatial relationships and object associations, enabling the system to reason about cross-modal connections without directly complexifying the core navigation model. This mediator allows accurate navigation decisions by providing structured world knowledge that enhances both visual and language feature understanding.
Solution Approach 2:
The system performs preliminary action by pre-processing and storing spatial knowledge in the knowledge graph before navigation execution. Spatial relationships, object properties, and environmental structures are organized in advance, allowing the navigation model to query ready-made knowledge rather than computing complex relationships in real-time during navigation, thus improving accuracy without proportionally increasing system complexity.
2Adaptability or versatility
If the navigation system integrates knowledge graph reasoning, then the cross-modal understanding capability is improved, but the computational time and processing complexity increase
Solution Approach 1:
The knowledge graph is constructed and populated with spatial relationships and object knowledge in advance, before the navigation task begins. This pre-processing allows the system to quickly query and retrieve relevant spatial reasoning information during navigation without performing complex computations in real-time, thus enhancing cross-modal understanding while minimizing additional processing time.
Solution Approach 2:
The patent extracts and separates the spatial reasoning function into a dedicated knowledge graph module, independent from the main navigation decision-making process. By taking out the complex spatial relationship reasoning from the core navigation model and placing it in the knowledge graph, the system achieves versatile cross-modal understanding while keeping the navigation model itself computationally efficient and fast.
3Productivity
If the system uses only visual features from environment images, then the processing speed is fast, but the understanding of abstract instructions deteriorates
Solution Approach 1:
The patent merges visual features from environment images with language features from instructions and pre-defined knowledge from the knowledge graph. This combination allows the system to maintain fast processing by using efficient visual feature extraction while simultaneously improving instruction understanding through language feature integration and spatial reasoning from the knowledge graph, achieving both speed and accuracy.
Data Source
AI summary
An indoor navigation method is provided, including: receiving an instruction for navigation, and collecting an environment image; extracting an instruction room feature and an instruction object feature carried in the instruction, and determining a visual room feature, a visual object feature, and a view angle feature based on the environment image; fusing the instruction object feature and the visual object feature with a first knowledge graph representing an indoor object association relationship to obtain an object feature, and determining a room feature based on the visual room feature and the instruction room feature; and determining a navigation decision based on the view angle feature, the room feature, and the object feature.


