Indoor Navigation Knowledge Graph for Abstract Instruction Reasoning

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current vision-language navigation technologies face challenges in understanding abstract instructions and making reasonable navigation decisions in indoor environments, due to limited cross-modal understanding and reasoning capabilities between visual and language modalities.

Innovation Solution

The proposed indoor navigation method enhances cross-modal understanding by extracting and fusing instruction features with visual features using a knowledge graph, and introduces entity knowledge reasoning to improve the decision-making process.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If vision-language navigation model uses basic visual and language features for navigation, then the system complexity is low, but the navigation decision accuracy deteriorates due to limited cross-modal understanding

Engineering Contradiction:
Improvenavigation decision accuracyVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent introduces a knowledge graph as an intermediary component that bridges visual and language modalities. The knowledge graph stores pre-defined spatial relationships and object associations, enabling the system to reason about cross-modal connections without directly complexifying the core navigation model. This mediator allows accurate navigation decisions by providing structured world knowledge that enhances both visual and language feature understanding.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The system performs preliminary action by pre-processing and storing spatial knowledge in the knowledge graph before navigation execution. Spatial relationships, object properties, and environmental structures are organized in advance, allowing the navigation model to query ready-made knowledge rather than computing complex relationships in real-time during navigation, thus improving accuracy without proportionally increasing system complexity.

Inventive Principle:
Principle #10Preliminary action

2Adaptability or versatility

If the navigation system integrates knowledge graph reasoning, then the cross-modal understanding capability is improved, but the computational time and processing complexity increase

Engineering Contradiction:
Improvecross-modal understanding capabilityVSAvoidcomputational processing time
Core Design Contradiction:
Adaptability or versatilityVSLoss of time

Solution Approach 1:

The knowledge graph is constructed and populated with spatial relationships and object knowledge in advance, before the navigation task begins. This pre-processing allows the system to quickly query and retrieve relevant spatial reasoning information during navigation without performing complex computations in real-time, thus enhancing cross-modal understanding while minimizing additional processing time.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent extracts and separates the spatial reasoning function into a dedicated knowledge graph module, independent from the main navigation decision-making process. By taking out the complex spatial relationship reasoning from the core navigation model and placing it in the knowledge graph, the system achieves versatile cross-modal understanding while keeping the navigation model itself computationally efficient and fast.

Inventive Principle:
Principle #2Taking out (Extraction)

3Productivity

If the system uses only visual features from environment images, then the processing speed is fast, but the understanding of abstract instructions deteriorates

Engineering Contradiction:
Improveprocessing speedVSAvoidinstruction understanding accuracy
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The patent merges visual features from environment images with language features from instructions and pre-defined knowledge from the knowledge graph. This combination allows the system to maintain fast processing by using efficient visual feature extraction while simultaneously improving instruction understanding through language feature integration and spatial reasoning from the knowledge graph, achieving both speed and accuracy.

Inventive Principle:
Principle #5Merging (Combining)

Data Source

PatentUS12276507B2Indoor navigation method, indoor navigation equipment, and storage medium
Publication Date: 2025.04.15 BEIJING XIAOMI MOBILE SOFTWARE CO LTD
  • US12276507B2 patent drawing
  • US12276507B2 patent drawing
  • US12276507B2 patent drawing

AI summary

An indoor navigation method is provided, including: receiving an instruction for navigation, and collecting an environment image; extracting an instruction room feature and an instruction object feature carried in the instruction, and determining a visual room feature, a visual object feature, and a view angle feature based on the environment image; fusing the instruction object feature and the visual object feature with a first knowledge graph representing an indoor object association relationship to obtain an object feature, and determining a room feature based on the visual room feature and the instruction room feature; and determining a navigation decision based on the view angle feature, the room feature, and the object feature.