Visual Image Representation via Depth-Aware Audio and Tactile Signals
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Blind and visually impaired individuals face challenges in mobility and orientation due to the inability of existing technologies to effectively convert visual information into usable sound or tactile signals for spatial awareness and object recognition.
Innovation Solution
A system and method that uses an imaging device to create a three-dimensional model of an environment, partitions it by surfaces, applies conversion functions to generate sound or tactile signals, and combines these signals to provide a user with a comprehensive auditory or tactile representation of object locations and characteristics, utilizing a data processing module and user interface for output.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If image-to-sound conversion is used to represent visual information for blind persons, then spatial awareness is improved, but training time for object-to-sound mapping is excessive
Solution Approach 1:
The patent segments the visual scene into multiple depth layers (foreground, midground, background) and assigns distinct audio characteristics to each layer. This segmentation allows blind users to quickly understand spatial relationships without extensive training, as each depth plane has its own recognizable audio signature rather than requiring learning of complex object mappings.
Solution Approach 2:
The patent transforms visual information by changing audio parameters (pitch, volume, spatial position) based on depth distance. Objects at different depths are represented by systematically varied audio parameters, creating an intuitive depth-perception system that reduces training time compared to direct object-to-sound mapping.
2Device complexity
If simple image-to-sound conversion is used, then device complexity is reduced, but sound resolution is insufficient for detailed object recognition
Solution Approach 1:
The patent adds a temporal dimension to the audio output by dynamically modulating sound characteristics based on depth information. This creates a multi-dimensional audio representation (spatial position + depth + time variation) that enhances object recognition resolution without requiring complex hardware modifications.
Solution Approach 2:
The patent introduces depth information as an intermediary layer between visual capture and audio output. This intermediary processing layer enriches the audio representation with spatial context, improving sound resolution for object recognition while keeping the overall system architecture relatively simple.
3Device complexity
If single-sense conversion (image-to-sound only) is used, then device complexity is minimized, but sensory stimulation is insufficient for comprehensive object recognition
Solution Approach 1:
The patent merges multiple sensory channels (auditory and tactile) into a unified depth-aware representation system. By combining vibration feedback with audio output, the system provides comprehensive object recognition information without significantly increasing device complexity, as both channels convey complementary spatial and textural information.
Solution Approach 2:
The patent creates a multi-functional system where the depth-based conversion framework serves multiple purposes: it provides spatial awareness through audio, texture information through vibration, and object identification through combined sensory input. This universal approach maximizes information delivery while maintaining system simplicity.
Data Source
Figure 1
Figure 2
Figure 3
AI summary
A method of representing visual images by alternative senses is provided herein. The method includes the following stages: obtaining an image of an environment having a background and physical objects distinguishable from the background; slicing the image into a plurality of slices; identifying, for each slice at a time, and slice by slice in a specified order, if at least a portion of the physical objects being contained within the slice; applying, in the specified order, a conversion function to each identified portion of the physical objects for generating a sound or tactile object-dependent signal; associating a sound or tactile location-dependent signal unique for each slice; superpositioning, in the specified order, each object- dependent signal with a respective location-dependent signal for creating a combined object-location signal; and outputting the combined object-location signal to a user via an interface in a form usable for a blind or visually impaired person.