Aural Map Generation for Visually Impaired Navigation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Individuals who are visually impaired or cannot read visual information in their environment, such as street signs and advertisements, lack effective means to access and understand environmental visual information, as current solutions like Braille only provide information upon direct contact and do not convey broader environmental context.

Innovation Solution

A system that captures images of the environment using a video capture device, identifies environmental text, generates aural contextual indicators, and creates an aural map of the environment, allowing users to access and understand visual information through an aural output device.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If visual information is delivered through written communications (street signs, location indicators, alerts), then information delivery efficiency is improved, but accessibility deteriorates for visually impaired or illiterate individuals

Engineering Contradiction:
Improveinformation delivery efficiencyVSAvoidaccessibility to visually impaired individuals
Core Design Contradiction:
ProductivityVSAdaptability or versatility

Solution Approach 1:

The patent introduces an intermediary system consisting of image capture devices, processing systems, and audio output devices that mediate between visual environmental information and visually impaired users. The system captures images, identifies text and objects, and converts them into audio descriptions, serving as a bridge that makes visual information accessible without compromising the original visual communication methods

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent replaces the mechanical/visual system of direct visual perception with an electronic-acoustic system. Instead of relying on visual input mechanisms (eyes reading text), the system uses image capture devices, computer vision processing, and audio output to deliver information acoustically, substituting the visual-mechanical pathway with an electronic-acoustic one

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Adaptability or versatility

If Braille mechanisms are used for close-range information reading, then information accessibility is improved, but environmental awareness deteriorates due to lack of broader context

Engineering Contradiction:
Improveinformation accessibilityVSAvoidenvironmental context
Core Design Contradiction:
Adaptability or versatilityVSLoss of information

Solution Approach 1:

The patent transitions from two-dimensional close-range Braille contact to three-dimensional environmental mapping through audio spatialization. The system captures images of the broader environment, processes multiple objects and text elements, and presents them through spatial audio cues that convey directional and contextual information, adding dimensional context that Braille cannot provide

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Solution Approach 2:

The patent creates a multi-functional system that serves both close-range identification (like Braille) and broad environmental awareness simultaneously. The image capture device and processing system can identify specific text for detailed information while also providing overall environmental context through audio descriptions of multiple objects, locations, and text elements in the captured scene

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS11282259B2Non-visual environment mapping
Publication Date: 2022.03.22 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US11282259B2 patent drawing
  • US11282259B2 patent drawing
  • US11282259B2 patent drawing

AI summary

Aspects of the present invention provide an approach for non-visually mapping an environment. In an embodiment, a set of images that is within the field of view of the user is captured from a video capture device worn by the user. Environmental text that is within the set of images is identified. An aural contextual indicator that corresponds to the environmental text is then generated. This aural contextual indicator indicates the informational nature of the environmental text. An aural map of the environment is created using a sequence of the generated aural contextual indicators. This aural map is delivered to the user via an aural output device worn by the user in response to a user request.