Audio Source Enhancement via Spatio-Temporal Signature Analysis

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional audio processing techniques, such as audio beamforming, struggle to isolate and enhance the desired audio from a specific source-of-interest (SOI) in video recordings due to background noise and ambient sounds, requiring manual adjustments of directional microphones, which can be cumbersome and result in incomplete audio capture.

Innovation Solution

A method and system that processes audio data by identifying the SOI through user interaction or spatio-temporal signature analysis, decomposing audio signals, separating relevant components based on frequency, amplitude, and coherency, and selectively enhancing the SOI's audio while suppressing background noise during playback.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If audio beamforming technique is used to determine the direction of audio signals, then the directionality of audio capture is improved, but manual adjustment of directional microphones is required which increases operation complexity

Engineering Contradiction:
Improvedirection determination accuracyVSAvoidmicrophone adjustment convenience
Core Design Contradiction:
Measurement precisionVSEase of operation

Solution Approach 1:

The system automatically performs audio beamforming and source separation without requiring manual microphone adjustment. The electronic device autonomously processes audio signals from multiple microphones, identifies sources of interest, and enhances desired audio while suppressing background noise, eliminating the need for user intervention in directional control.

Inventive Principle:
Principle #25Self-service

2Measurement precision

If directional microphones are manually adjusted to capture audio from a specific direction, then the clarity of the desired audio source is improved, but audio signals from other directions cannot be recorded which reduces audio completeness

Engineering Contradiction:
Improveaudio source clarityVSAvoidaudio signal completeness
Core Design Contradiction:
Measurement precisionVSLoss of information

Solution Approach 1:

The system segments the mixed audio signal into multiple independent source components using blind source separation techniques. By decomposing the audio mixture into distinct sources and identifying sources of interest through spatio-temporal analysis, the system can selectively enhance specific audio sources while preserving and optionally playing back other audio components, thus maintaining audio completeness while improving clarity of selected sources.

Inventive Principle:
Principle #1Segmentation

3Loss of information

If conventional audio processing is used to capture all audio sources, then audio completeness is maintained, but the ability to isolate and enhance specific audio sources is reduced which decreases audio quality

Engineering Contradiction:
Improveaudio signal completenessVSAvoidaudio source isolation accuracy
Core Design Contradiction:
Loss of informationVSMeasurement precision

Solution Approach 1:

The system changes key parameters of the audio processing pipeline by implementing spatio-temporal signature analysis and blind source separation. These parameter changes enable the system to dynamically identify and isolate specific audio sources based on their spatial and temporal characteristics, significantly improving source separation accuracy while maintaining the ability to preserve or restore other audio components if needed.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS9318121B2Method and system for processing audio data of video content
Publication Date: 2016.04.19 SONY GROUP CORP
  • US9318121B2 patent drawing
  • US9318121B2 patent drawing
  • US9318121B2 patent drawing

AI summary

Various aspects of a method and system to process audio data are disclosed herein. In accordance with an embodiment, the method includes identification of a source-of-interest (SOI), via a user interface (UI), when video content is played back. The SOI is identified based on one or more parameters. An audio portion of the identified SOI is selectively enhanced when the video content is played back.