Acoustic Zooming Audio Enhancement via Beamforming

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

During video playback, audio related to an area of interest can be drowned out by environmental noise, as existing technologies do not effectively enhance audio while visually zooming into specific areas of a video.

Innovation Solution

The system employs a combination of microphones, a camera module, and an acoustic zooming controller that uses beamformers and neural networks to isolate and enhance the audio from a selected area of interest by transforming audio signals into the frequency domain, generating noise-suppressed signals, and combining beamformer signals based on the zoom area's tiles, thereby improving audio clarity.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If visual zooming into an area of interest is performed during video playback, then the visual clarity of the selected area is improved, but the audio related to that area remains drowned out by environmental noise

Engineering Contradiction:
Improvevisual clarityVSAvoidaudio noise
Core Design Contradiction:
Measurement precisionVSObject-affected harmful factors

Solution Approach 1:

The audio signal is segmented into multiple directional components using beamforming technology. The system divides the audio field into different spatial zones and extracts audio signals from specific directions corresponding to the visually zoomed area, separating target audio from environmental noise through spatial segmentation

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system applies different audio processing qualities to different spatial regions. The area of interest receives enhanced audio processing with higher gain and noise suppression, while other areas maintain normal audio levels, creating local quality differentiation that matches the visual zoom focus

Inventive Principle:
Principle #3Local quality

2Measurement precision

If beamformers and neural networks are used to isolate and enhance audio from a selected area, then audio clarity is improved, but device complexity increases

Engineering Contradiction:
Improveaudio clarityVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The system replaces traditional mechanical audio processing methods with computational approaches. Neural networks and digital signal processing algorithms substitute for physical acoustic filtering, enabling sophisticated audio isolation and enhancement through software-based beamforming and machine learning models

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The system introduces intermediate processing layers including beamformers, neural networks, and audio synthesis components that mediate between the raw microphone signals and the final enhanced audio output. These intermediaries break down the complex task of audio enhancement into manageable processing stages

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS11721354B2Acoustic zooming
Publication Date: 2023.08.08 SNAP INC
  • US11721354B2 patent drawing
  • US11721354B2 patent drawing
  • US11721354B2 patent drawing

AI summary

Method of performing acoustic zooming starts with microphones capturing acoustic signals associated with video content. Beamformers generate beamformer signals using the acoustic signals. Beamformer signals correspond respectively to tiles of video content. Each of the beamformers is respectively directed to a center of each of the tiles. Target enhanced signal is generated using beamformer signals. Target enhanced signal is associated with a zoom area of video content. Target enhanced signal is generated by identifying the tiles respectively having at least portions that are included in the zoom area, selecting beamformer signals corresponding to identified tiles, and combining selected beamformer signals to generate target enhanced signal. Combining selected beamformer signals may include determining proportions for each of the identified tiles in relation to the zoom area and combining selected beamformer signals based on the proportions to generate the target enhanced signal. Other embodiments are described herein.