AR Sound Source Identification via Audio-Visual Correlation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Visually impaired individuals face challenges in identifying the source of sounds in real-world environments, as existing technologies do not effectively correlate sounds with objects or actions outside controlled settings.
Innovation Solution
An augmented reality system that uses sound localization and machine learning to identify sound sources, generating both audio descriptions and magnified visual content for display, allowing users to correlate sounds with their visual counterparts.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of information
If visually impaired individuals rely on sound processing to understand real-world events, then they can detect sound sources, but they cannot correlate sounds with specific objects or actions
Solution Approach 1:
The patent combines audio processing and visual processing systems into a unified augmented reality framework. The audio system identifies sound sources while the visual system captures images, and these two streams are merged to provide correlated information about objects and actions, allowing visually impaired users to connect sounds with their visual counterparts.
Solution Approach 2:
The patent introduces an intermediary processing system that acts as a bridge between sound detection and object identification. This intermediary component analyzes both audio and visual data, correlates them together, and presents integrated information to the user, enabling the connection between sound sources and their corresponding objects or actions.
2Measurement precision
If existing technologies provide basic sound detection, then users can hear sounds, but they cannot effectively process and identify the source in complex real-world environments
Solution Approach 1:
The patent segments the complex task of sound source identification into distinct functional modules: audio processing module for detecting sounds, visual processing module for capturing and analyzing images, correlation module for matching sound with visual data, and output module for presenting information. This segmentation allows each module to specialize in one aspect while working together to solve the overall problem.
Solution Approach 2:
The patent creates a multi-functional augmented reality system that performs multiple functions simultaneously: detecting sounds, capturing images, identifying objects, analyzing actions, correlating audio-visual data, and providing multiple output formats. This universal system handles various real-world scenarios with a single integrated platform, reducing overall system complexity despite the advanced capabilities.
3Loss of information
If the system provides detailed information about sound sources, then users gain better understanding, but the information processing time increases
Solution Approach 1:
The patent implements preliminary action by pre-processing and pre-organizing both audio and visual data streams in real-time. The system continuously captures and segments sound data and image data, pre-identifies potential sound sources and objects, and prepares correlation frameworks in advance. When a sound event occurs, the pre-prepared data structures enable rapid matching and information retrieval without requiring extensive real-time computation.
Data Source
AI summary
Embodiments herein provide an augmented reality (AR) system that uses sound localization to identify sounds that may be of interest to a user and generates an audio description of the source of the sound as well as AR content that can be magnified and displayed to the user. In one embodiment, an AR device captures images that have the source of the sound within their field of view. Using machine learning (ML) techniques, the AR device can identify the object creating the sound (i.e., the sound source). A description of the sound source and its actions can outputted to the user. In parallel, the AR device can also generate AR content for the sound source. For example, the AR device can magnify the sound source to a size that is viewable to the user and create AR content that is then superimposed onto a display.


