Audio Processing Device for Sound Scene Manipulation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing sound scene manipulation methods face challenges in efficiently altering the levels of individual sound sources in multi-microphone recordings, particularly with fewer microphones than sound sources, due to high computational demands and reliance on prior knowledge or intensive source separation techniques.

Innovation Solution

A method that generates auxiliary mixtures from fewer microphone recordings, allowing for independent level changes of sound sources without explicit separation, using a combination of beamforming and spectral modification techniques to create linearly independent signals for flexible sound scene manipulation.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If blind source separation (BSS) methods are used to extract individual sound sources from microphone recordings, then sound source extraction capability is improved, but computational complexity increases significantly making real-time applications difficult

Engineering Contradiction:
Improvesound source extraction capabilityVSAvoidcomputational complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent segments the sound scene manipulation task into two independent parts: (1) capturing the sound scene with multiple microphones to obtain mixture signals, and (2) manipulating individual sound source levels through post-processing of these mixtures without requiring full source separation. This avoids the computationally intensive BSS step while still enabling independent level control of sound sources.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

Instead of the traditional approach of separating sources first then manipulating levels, the patent inverts the process by directly manipulating sound source levels in the mixture domain using the relationship between microphone signals and sound source levels. This reverse approach eliminates the need for computationally heavy source separation while achieving the same level manipulation goal.

Inventive Principle:
Principle #13The other way round (Inversion)

2Measurement precision

If computational auditory scene analysis (CASA) is applied to analyze sound scenes mimicking human auditory system, then sound source identification capability is improved, but computational requirements become very high for real-life applications

Engineering Contradiction:
Improvesound source identification capabilityVSAvoidcomputational requirements
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent extracts only the essential information needed for level manipulation from the sound scene, specifically the relationship between microphone mixture signals and individual sound source levels. It discards the complex auditory feature extraction and source identification steps of CASA, keeping only the necessary computational elements for the specific task of level control.

Inventive Principle:
Principle #2Taking out (Extraction)

3Measurement precision

If beamforming with prior knowledge of sound source positions is used, then sound source separation capability is improved, but adaptability to dynamic scenarios decreases

Engineering Contradiction:
Improvesound source separation capabilityVSAvoidadaptability to dynamic scenarios
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The patent employs dynamic processing that adapts to changing sound scenes in real-time. The level manipulation parameters are continuously updated based on current mixture signals from multiple microphones, allowing the system to automatically adapt to moving sound sources and changing acoustic environments without requiring manual reconfiguration or prior knowledge of source positions.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The system dynamically changes the mixing parameters and level control parameters based on the current sound scene configuration. By adjusting the post-processing parameters in real-time according to the captured mixture signals, the system maintains effective sound source level control even as sources move or new sources appear, achieving both separation capability and adaptability.

Inventive Principle:
Principle #35Parameter changes

4Adaptability or versatility

If auxiliary mixtures are generated by combining multiple microphone recordings to enable independent level changes, then flexibility in sound scene manipulation is improved, but processing complexity increases

Engineering Contradiction:
Improveflexibility in sound scene manipulationVSAvoidprocessing complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent combines multiple microphone recordings into auxiliary mixtures with specific linear relationships, where each mixture preserves information about individual sound sources in different proportions. This merging creates the flexibility needed for independent level control while using simple linear combination operations that avoid the complexity of full source separation algorithms.

Inventive Principle:
Principle #5Merging (Combining)

Data Source

PatentEP2437517B1Sound scene manipulation
Publication Date: 2014.04.02 NXP BV
  • EP2437517B1 patent drawingFigure 1~2
  • EP2437517B1 patent drawingFigure 3
  • EP2437517B1 patent drawingFigure 4

AI summary

An audio-processing device. The device comprises: an audio input, for receiving one or more audio signals detected at respective microphones. Each of the audio signals comprises a mixture of a plurality of components, each component corresponding to a sound source. The device also comprises a control input, for receiving, for each sound source, a desired gain factor associated with the source, by which it is desired to amplify the corresponding component. The device further comprises an auxiliary signal generator, adapted to generate at least one auxiliary signal from the one or more audio signals, the at least one auxiliary signal comprising a different mixture of the components as compared with a reference one of the one or more audio signals; and a scaling coefficient calculator, adapted to calculate a set of scaling coefficients in dependence upon the desired gain factors and upon parameters of the different mixture, each scaling coefficient associated with one of the at least one auxiliary signal and optionally the reference audio signal. It also comprises an audio synthesis unit, adapted to synthesize an output audio signal by applying the scaling coefficients to the at least one auxiliary signal and optionally the reference audio signal and to combine the results. The scaling coefficients are calculated from the desired gain factors and the parameters of the different mixture such that the synthesized output signal provides the desired gain factor for each component.