Video-Based Audio Source Direction Detection

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing video recording technologies face challenges in isolating and enhancing specific audio sources in casual settings or environments where careful microphone placement is not possible, such as sporting events or conference rooms, especially when multiple speakers are present.

Innovation Solution

The method involves determining the direction of an audio source in recorded video, filtering, and enhancing the audio to create a zooming effect by separating audio streams based on image regions and using phase correlation techniques to achieve beam-forming without prior knowledge of microphone placement, allowing users to select and amplify specific audio sources during video playback.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If beam forming techniques are used to isolate sounds from particular directions, then audio source isolation is improved, but the system requires prior knowledge of microphone placement and sophisticated setup

Engineering Contradiction:
Improveaudio source isolationVSAvoidmicrophone placement requirements
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The system automatically determines the direction of audio sources by analyzing the recorded video itself, without requiring external setup information or prior knowledge of microphone placement. The video content provides the reference data needed for audio enhancement, making the system self-sufficient and eliminating complex configuration requirements.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The video images serve as an intermediary between the audio signal and the enhancement process. By extracting visual information about speaker locations from the video frames, the system uses this intermediate data to guide audio source separation and directional filtering, bridging the gap between raw audio and targeted enhancement without requiring direct microphone placement knowledge.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Measurement precision

If directional microphones are used to capture sound from particular locations, then sound directionality is improved, but the system cannot adapt to multiple speakers in different locations without manual configuration

Engineering Contradiction:
Improvesound directionalityVSAvoidhandling multiple speakers
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The system dynamically adapts to different audio scenarios by continuously analyzing video content during playback. When multiple speakers are present, the system automatically identifies and tracks their positions across video frames, adjusting audio enhancement in real-time to follow the active speaker, making it versatile for various conference and meeting situations without manual reconfiguration.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The system performs preliminary analysis of the video content to identify potential audio sources and their locations before audio enhancement is applied. By pre-processing the video to extract spatial information about speakers and their movements, the system prepares the necessary data structures and parameters in advance, enabling rapid adaptation when speakers change positions or when multiple speakers become active.

Inventive Principle:
Principle #10Preliminary action

3Measurement precision

If multiple microphones are spaced apart for beam forming, then remote speaker capture is improved, but the user must identify the speaker during recording which reduces ease of use

Engineering Contradiction:
Improveremote speaker captureVSAvoidspeaker identification requirement
Core Design Contradiction:
Measurement precisionVSEase of operation

Solution Approach 1:

Instead of requiring speaker identification during recording as in traditional beam forming systems, the patent inverts the approach by performing speaker identification and audio source separation during video playback. The system analyzes the already-recorded video content to identify speakers and their directions, then applies audio enhancement retroactively, eliminating the need for complex real-time speaker identification during capture.

Inventive Principle:
Principle #13The other way round (Inversion)

Solution Approach 2:

The patent replaces the mechanical approach of physical microphone placement and real-time beam forming with a post-processing computational approach. By using digital signal processing on recorded audio combined with video analysis, the system substitutes complex mechanical speaker identification and positioning with automated computer vision and audio processing algorithms, greatly simplifying user operation.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

4Measurement precision

If microphones are placed near each speaker in conference rooms, then audio clarity for each speaker is improved, but the device complexity and setup requirements increase

Engineering Contradiction:
Improveaudio clarity per speakerVSAvoidmicrophone placement and synchronization
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent extracts the function of multiple strategically placed microphones from the physical setup and implements it through post-processing audio separation. Instead of requiring multiple microphones positioned near each speaker, the system takes a single or few omnidirectional microphones and uses computational methods to extract and separate individual speaker audio streams based on video-derived spatial information, eliminating complex physical setup while maintaining audio clarity.

Inventive Principle:
Principle #2Taking out (Extraction)

Applied Scientific Principles

This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.

Function Achieved in This Case

This approach enables users to intuitively select and enhance specific audio sources during video playback, improving audio clarity and reducing ambient noise, even in environments where standard beam-forming and tagging techniques are not applicable, providing an enhanced user experience.

Implementation Method 1

using phase correlation techniques to achieve beam-forming without prior knowledge of microphone placement

Methodology Applied
Scientific EffectPhase correlation:

Data Source

PatentUS10153002B2Selection of an audio stream of a video for enhancement using images of the video
Publication Date: 2018.12.11 INTEL CORP
  • US10153002B2 patent drawing
  • US10153002B2 patent drawing
  • US10153002B2 patent drawing

AI summary

An audio stream of a video is selected for enhancement using image of the video. In one example, audio streams in the video are identified and segregated. Points of interest and their locations are identified in the image of the video. The position of each audio stream is plotted to a location of a point of interest. A selection of a point of interest from the sequence of images is received. A plotted audio stream is selected based on the corresponding point of interest and the selected audio stream is enhanced.