Indoor Navigation System for Visually Impaired Using Depth and Saliency Analysis
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current systems, such as the DenseCap model, fail to provide visually impaired individuals with effective indoor situational awareness and navigation due to a lack of distance and direction information, which are crucial for reconstructing scenes mentally and safely navigating indoor environments.
Innovation Solution
A wearable system comprising a motion sensor, image sensor, compass, and depth sensor that enhances image capture with angle and depth information, determines directional saliency, saliency at rest, and saliency in motion, and generates a virtual graph to provide audio or Braille instructions for safe navigation, incorporating distance and directional orientation.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of information
If DenseCap model is used for scene description, then semantic understanding of visual scene is improved, but distance and direction information is lost
Solution Approach 1:
The system segments the scene description into multiple components: semantic captions from DenseCap, depth information from depth sensor, and directional information from compass. These segmented components are processed separately and then integrated to provide complete spatial awareness.
Solution Approach 2:
The system merges outputs from multiple independent systems: the DenseCap captioning model, depth sensor data, compass orientation, and motion sensor information are combined in the processor to create a unified situational awareness representation that includes both semantic and spatial information.
2Ease of operation
If top-down view map is used for navigation, then path planning is improved, but indoor dynamic settings make top-down views non-viable
Solution Approach 1:
Instead of providing a top-down map view that requires the user to mentally translate to first-person navigation, the system inverts the approach by providing first-person situational awareness directly aligned with the user's current orientation and movement, making navigation intuitive in dynamic indoor settings.
Solution Approach 2:
The system transitions from the traditional 2D top-down map dimension to a first-person 3D spatial dimension that matches the user's perspective. Depth sensors and orientation data create a volumetric understanding of space around the user, enabling navigation without requiring map interpretation.
3Loss of information
If audio-tone representations are used for obstacle detection, then simple obstacle warning is achieved, but natural language scene description is lost
Solution Approach 1:
The system changes the parameter of information representation from simple audio tones to natural language captions enhanced with spatial parameters. The DenseCap model generates descriptive captions that are then enriched with depth and directional information, transforming basic obstacle warnings into comprehensive scene descriptions.
4Reliability
If multiple sensors are integrated for comprehensive scene understanding, then situational awareness is improved, but device complexity increases
Solution Approach 1:
The system employs a universal processor that handles multiple functions: processing DenseCap captions, integrating depth sensor data, combining compass orientation, and fusing motion sensor information. This multi-functional approach consolidates complexity into a single processing unit rather than requiring separate specialized systems.
Applied Scientific Principles
This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.
Function Achieved in This Case
The system effectively enhances situational awareness and navigation for visually impaired users by providing detailed audio or Braille instructions with distance and directional information, enabling them to safely navigate indoor spaces.
Implementation Method 1
a depth sensor. The depth sensor is configured to provide the depth/distance information
Implementation Method 2
an image sensor configured to capture an image of the scene, in front of the user
Implementation Method 3
a motion sensor configured to detect motion of the user
Data Source
AI summary
A system and method for providing indoor situational awareness and navigational aid for the visually impaired user, is disclosed. The processor may receive input data. The processor may enhance the image based upon the angle and the depth information. The processor may determine “directional saliency”, “saliency at rest” and “saliency in motion” of the enhanced image of the scene to provide situational awareness and generate a virtual graph with a grid of nodes. The processor may probe each node in order to check whether or not the point corresponding to said node is on a floor and determine the shortest path to a destination in the virtual graph by only considering the points on the floor. The processor may convert the description of the shortest path and the scene into one or more of audio or Braille text instruction to the user.


