AR Sound Source Identification via Audio-Visual Correlation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Visually impaired individuals face challenges in identifying the source of sounds in real-world environments, as existing technologies do not effectively correlate sounds with objects or actions outside controlled settings.

Innovation Solution

An augmented reality system that uses sound localization and machine learning to identify sound sources, generating both audio descriptions and magnified visual content for display, allowing users to correlate sounds with their visual counterparts.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Loss of information

If visually impaired individuals rely on sound processing to understand real-world events, then they can detect sound sources, but they cannot correlate sounds with specific objects or actions

Engineering Contradiction:
Improvecorrelation between sound and objectVSAvoidability to identify sound source
Core Design Contradiction:
Loss of informationVSEase of operation

Solution Approach 1:

The patent combines audio processing and visual processing systems into a unified augmented reality framework. The audio system identifies sound sources while the visual system captures images, and these two streams are merged to provide correlated information about objects and actions, allowing visually impaired users to connect sounds with their visual counterparts.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent introduces an intermediary processing system that acts as a bridge between sound detection and object identification. This intermediary component analyzes both audio and visual data, correlates them together, and presents integrated information to the user, enabling the connection between sound sources and their corresponding objects or actions.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Measurement precision

If existing technologies provide basic sound detection, then users can hear sounds, but they cannot effectively process and identify the source in complex real-world environments

Engineering Contradiction:
Improvesound source identification accuracyVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent segments the complex task of sound source identification into distinct functional modules: audio processing module for detecting sounds, visual processing module for capturing and analyzing images, correlation module for matching sound with visual data, and output module for presenting information. This segmentation allows each module to specialize in one aspect while working together to solve the overall problem.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent creates a multi-functional augmented reality system that performs multiple functions simultaneously: detecting sounds, capturing images, identifying objects, analyzing actions, correlating audio-visual data, and providing multiple output formats. This universal system handles various real-world scenarios with a single integrated platform, reducing overall system complexity despite the advanced capabilities.

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Loss of information

If the system provides detailed information about sound sources, then users gain better understanding, but the information processing time increases

Engineering Contradiction:
Improvecompleteness of sound source informationVSAvoidprocessing time
Core Design Contradiction:
Loss of informationVSLoss of time

Solution Approach 1:

The patent implements preliminary action by pre-processing and pre-organizing both audio and visual data streams in real-time. The system continuously captures and segments sound data and image data, pre-identifies potential sound sources and objects, and prepares correlation frameworks in advance. When a sound event occurs, the pre-prepared data structures enable rapid matching and information retrieval without requiring extensive real-time computation.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS20210209365A1Translating sound events to speech and ar content
Publication Date: 2021.07.08 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US20210209365A1 patent drawing
  • US20210209365A1 patent drawing
  • US20210209365A1 patent drawing

AI summary

Embodiments herein provide an augmented reality (AR) system that uses sound localization to identify sounds that may be of interest to a user and generates an audio description of the source of the sound as well as AR content that can be magnified and displayed to the user. In one embodiment, an AR device captures images that have the source of the sound within their field of view. Using machine learning (ML) techniques, the AR device can identify the object creating the sound (i.e., the sound source). A description of the sound source and its actions can outputted to the user. In parallel, the AR device can also generate AR content for the sound source. For example, the AR device can magnify the sound source to a size that is viewable to the user and create AR content that is then superimposed onto a display.