Video Sound Source Isolation via Beamforming and Object Association

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

In video reproductions, listeners struggle to selectively control and manage the sound from multiple visual objects, as existing technologies do not allow for individual sound source manipulation, leading to difficulties in hearing specific sounds amidst background noise or personal preference.

Innovation Solution

A computer-implemented method that captures video and audio using multiple microphones, isolates sound sources using beamforming algorithms, and allows users to modify attributes like volume or pitch through inputs such as voice commands or gestures, associating sound sources with visual objects for precise control.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If multiple sound sources are captured simultaneously in a video, then the audio reproduces the real-world environment accurately, but listeners cannot selectively control or isolate individual sound sources

Engineering Contradiction:
Improveaudio accuracyVSAvoidsound source control
Core Design Contradiction:
ReliabilityVSEase of operation

Solution Approach 1:

The patent segments the mixed audio signal into individual sound source components using beamforming algorithms. Each sound source is isolated and processed independently, allowing listeners to selectively control individual sound sources while maintaining overall audio accuracy. The beamforming technique divides the audio field into directional beams, enabling precise separation of simultaneous sound sources.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces an intermediary processing layer between audio capture and playback that identifies, isolates, and tags individual sound sources. This intermediary system uses algorithms to separate mixed audio into distinct controllable streams, enabling selective manipulation of individual sound sources without affecting others, thus solving the control problem while preserving audio fidelity.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Ease of operation

If beamforming algorithms are used to isolate sound sources, then individual sound sources can be controlled separately, but the system complexity increases

Engineering Contradiction:
Improvesound source controlVSAvoidprocessing complexity
Core Design Contradiction:
Ease of operationVSDevice complexity

Solution Approach 1:

The patent implements a universal sound source isolation system that handles multiple sound sources simultaneously using the same beamforming algorithm framework. Rather than requiring separate processing for each sound source, the system uses a multi-functional approach where the beamforming algorithm automatically adapts to isolate any number of sound sources based on their spatial distribution, reducing overall system complexity.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The beamforming algorithm automatically identifies and isolates sound sources without requiring manual configuration or intervention. The system self-adjusts to the acoustic environment, automatically separating sound sources based on their directional characteristics. This self-service capability reduces the operational complexity despite the sophisticated processing required.

Inventive Principle:
Principle #25Self-service

3Ease of operation

If sound sources are isolated using multiple microphones and beamforming, then selective sound control is enabled, but the device configuration becomes more complex

Engineering Contradiction:
Improvesound source controlVSAvoidmicrophone array complexity
Core Design Contradiction:
Ease of operationVSDevice complexity

Solution Approach 1:

The patent combines multiple microphones into a unified beamforming array system that processes all inputs simultaneously. Rather than treating each microphone as a separate device requiring individual configuration, the system merges them into a single coordinated array that automatically performs sound source separation, reducing the perceived complexity for users while enabling sophisticated sound control.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent transitions from traditional single-channel audio processing to spatial audio processing by utilizing the dimensional information provided by the microphone array. The beamforming algorithm exploits the spatial dimension to separate sound sources based on their directional characteristics, enabling selective control without requiring additional physical devices beyond the microphone array.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Applied Scientific Principles

This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.

Function Achieved in This Case

Enables selective control of sound attributes from individual visual objects in videos, improving audio clarity and user experience by allowing users to manage sound sources based on their preferences, even when objects move or change positions.

Implementation Method 1

applying a beamforming algorithm to audio signals received by the two or more microphones to thereby form a beam pattern in which a sound source located in the beam pattern is isolated with respect to sound sources outside of the pattern

Methodology Applied
Scientific EffectBeamforming:

Implementation Method 2

changing pitch of the continuously isolated sound source as the visual object first moves towards and then away from a user, wherein the pitch changes from a higher pitch as the visual object approaches the user to a lower pitch as the visual object moves away from the user, thereby simulating a Doppler effect

Methodology Applied
Scientific EffectDoppler effect: Doppler Effect

Data Source

PatentUS11513762B2Controlling sounds of individual objects in a video
Publication Date: 2022.11.29 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US11513762B2 patent drawing
  • US11513762B2 patent drawing
  • US11513762B2 patent drawing

AI summary

A method for modifying a sound produced by a sound source in a video includes capturing video and audio of a scene is disclosed. Audio is captured using a microphone array. A sound source is isolated and a direction of arrival of the sound source with respect to a capture location is identified. One or more visual objects in the captured video are identified. One of the isolated sound sources is associated with one of the identified visual objects. An input identifying one of the isolated sound sources is received during playing of the captured video and audio. The input includes a command. Responsive to receiving the input, an attribute of the identified isolated sound source is modified. The input may identify a visual object associated with a sound source. A system and article of manufacture are also disclosed.