Indoor Navigation System for Visually Impaired Using Depth and Saliency Analysis

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current systems, such as the DenseCap model, fail to provide visually impaired individuals with effective indoor situational awareness and navigation due to a lack of distance and direction information, which are crucial for reconstructing scenes mentally and safely navigating indoor environments.

Innovation Solution

A wearable system comprising a motion sensor, image sensor, compass, and depth sensor that enhances image capture with angle and depth information, determines directional saliency, saliency at rest, and saliency in motion, and generates a virtual graph to provide audio or Braille instructions for safe navigation, incorporating distance and directional orientation.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Loss of information

If DenseCap model is used for scene description, then semantic understanding of visual scene is improved, but distance and direction information is lost

Engineering Contradiction:
Improvedistance and direction informationVSAvoidscene description accuracy
Core Design Contradiction:
Loss of informationVSMeasurement precision

Solution Approach 1:

The system segments the scene description into multiple components: semantic captions from DenseCap, depth information from depth sensor, and directional information from compass. These segmented components are processed separately and then integrated to provide complete spatial awareness.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system merges outputs from multiple independent systems: the DenseCap captioning model, depth sensor data, compass orientation, and motion sensor information are combined in the processor to create a unified situational awareness representation that includes both semantic and spatial information.

Inventive Principle:
Principle #5Merging (Combining)

2Ease of operation

If top-down view map is used for navigation, then path planning is improved, but indoor dynamic settings make top-down views non-viable

Engineering Contradiction:
Improvenavigation capabilityVSAvoidindoor environment adaptability
Core Design Contradiction:
Ease of operationVSAdaptability or versatility

Solution Approach 1:

Instead of providing a top-down map view that requires the user to mentally translate to first-person navigation, the system inverts the approach by providing first-person situational awareness directly aligned with the user's current orientation and movement, making navigation intuitive in dynamic indoor settings.

Inventive Principle:
Principle #13The other way round (Inversion)

Solution Approach 2:

The system transitions from the traditional 2D top-down map dimension to a first-person 3D spatial dimension that matches the user's perspective. Depth sensors and orientation data create a volumetric understanding of space around the user, enabling navigation without requiring map interpretation.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

3Loss of information

If audio-tone representations are used for obstacle detection, then simple obstacle warning is achieved, but natural language scene description is lost

Engineering Contradiction:
Improvenatural language descriptionVSAvoidlimited situational awareness
Core Design Contradiction:
Loss of informationVSObject-generated harmful factors

Solution Approach 1:

The system changes the parameter of information representation from simple audio tones to natural language captions enhanced with spatial parameters. The DenseCap model generates descriptive captions that are then enriched with depth and directional information, transforming basic obstacle warnings into comprehensive scene descriptions.

Inventive Principle:
Principle #35Parameter changes

4Reliability

If multiple sensors are integrated for comprehensive scene understanding, then situational awareness is improved, but device complexity increases

Engineering Contradiction:
Improvesituational awareness accuracyVSAvoidsensor integration complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The system employs a universal processor that handles multiple functions: processing DenseCap captions, integrating depth sensor data, combining compass orientation, and fusing motion sensor information. This multi-functional approach consolidates complexity into a single processing unit rather than requiring separate specialized systems.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Applied Scientific Principles

This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.

Function Achieved in This Case

The system effectively enhances situational awareness and navigation for visually impaired users by providing detailed audio or Braille instructions with distance and directional information, enabling them to safely navigate indoor spaces.

Implementation Method 1

a depth sensor. The depth sensor is configured to provide the depth/distance information

Methodology Applied
Scientific EffectTime of Flight: Time of Flight

Implementation Method 2

an image sensor configured to capture an image of the scene, in front of the user

Methodology Applied
Scientific EffectPhotoelectric Effect: Photoelectric Effect

Implementation Method 3

a motion sensor configured to detect motion of the user

Methodology Applied
Scientific EffectAccelerometer: Accelerometer

Data Source

PatentUS11955029B2System and method for indoor situational awareness and navigational aid for the visually impaired user
Publication Date: 2024.04.09 BHARATI VIVEK SATYA
  • US11955029B2 patent drawing
  • US11955029B2 patent drawing
  • US11955029B2 patent drawing

AI summary

A system and method for providing indoor situational awareness and navigational aid for the visually impaired user, is disclosed. The processor may receive input data. The processor may enhance the image based upon the angle and the depth information. The processor may determine “directional saliency”, “saliency at rest” and “saliency in motion” of the enhanced image of the scene to provide situational awareness and generate a virtual graph with a grid of nodes. The processor may probe each node in order to check whether or not the point corresponding to said node is on a floor and determine the shortest path to a destination in the virtual graph by only considering the points on the floor. The processor may convert the description of the shortest path and the scene into one or more of audio or Braille text instruction to the user.