Spatial Audio Stream Merging via Virtual Microphones

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current spatial audio technologies face limitations in modifying and merging sound scenes recorded with few microphones, as they struggle to efficiently represent and manipulate complex sound fields, especially when multiple sound sources are active simultaneously, and are inflexible regarding listening position and orientation.

Innovation Solution

An apparatus and method for generating a merged audio data stream using a demultiplexer to separate input audio streams into single-layer streams, a merging module that combines these streams based on cost values, and an information computation module to adjust audio signals for a virtual microphone, allowing for flexible sound scene manipulation and representation.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If object-based representations are used to represent sound scenes with N discrete audio objects, then flexibility at the reproduction side is improved, but difficulty in obtaining the representation from complex sound scenes recorded with few microphones increases

Engineering Contradiction:
Improveflexibility at reproduction sideVSAvoiddifficulty in obtaining representation
Core Design Contradiction:
Adaptability or versatilityVSDifficulty of detecting and measuring

Solution Approach 1:

The patent introduces virtual microphones as intermediary devices that are not physically present but computationally generated. These virtual microphones serve as mediators between the physical microphone array and the desired object-based representation, enabling the extraction of audio objects from complex sound scenes through virtual recording positions that can be strategically placed to optimize source separation and representation quality.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent creates virtual copies of microphone signals at multiple hypothetical positions in space. By generating these virtual microphone signals through signal processing, the system can represent the sound scene from multiple perspectives simultaneously, enabling object-based extraction without requiring physical microphones at those locations.

Inventive Principle:
Principle #26Copying

2Productivity

If spatial audio coding is applied to code information of different channels jointly, then coding efficiency is improved, but inability to modify the sound scene after loudspeaker signals are computed increases

Engineering Contradiction:
Improvecoding efficiencyVSAvoidability to modify sound scene
Core Design Contradiction:
ProductivityVSAdaptability or versatility

Solution Approach 1:

The patent segments the audio signal into discrete audio objects, each with its own spatial and temporal characteristics. This segmentation allows individual objects to be independently manipulated, modified, or repositioned after encoding, providing flexibility while maintaining the efficiency benefits of joint coding through the structured object-based representation.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces dynamic manipulation capabilities to the static encoded audio objects. The system allows real-time modification of audio objects including position changes, loudness adjustments, and spatial relocation, transforming the encoded representation from a fixed structure into a dynamically adjustable one that can adapt to different reproduction scenarios.

Inventive Principle:
Principle #15Dynamics

3Measurement precision

If traditional spatial sound recording techniques are used to capture sound fields, then reproduction of sound image at recording location is achieved, but inability to vary acoustic viewpoint and listening position increases

Engineering Contradiction:
Improveaccuracy of sound image reproductionVSAvoidability to vary acoustic viewpoint
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The patent adds the dimension of virtual recording positions to the traditional fixed microphone array setup. By computing virtual microphone signals at multiple hypothetical positions in 3D space, the system enables the listener to virtually move to different positions and viewpoints within the sound scene, adding spatial freedom to the traditional fixed perspective recording approach.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Data Source

PatentEP2786374B1Apparatus and method for merging geometry-based spatial audio coding streams
Publication Date: 2024.05.01 FRAUNHOFER GESELLSCHAFT ZUR FORDERUNG DER ANGEWANDTEN FORSCHUNG EV
  • EP2786374B1 patent drawingFigure 1
  • EP2786374B1 patent drawingFigure 2A
  • EP2786374B1 patent drawingFigure 2B

AI summary

An apparatus for generating a merged audio data stream is provided. The apparatus comprises a demultiplexer (180) for obtaining a plurality of single-layer audio data streams, wherein the demultiplexer (180) is adapted to receive one or more input audio data streams, wherein each input audio data stream comprises one or more layers, wherein the demultiplexer (180) is adapted to demultiplex each one of the input audio data streams having one or more layers into two or more demultiplexed audio data streams having exactly one layer, such that the two or more demultiplexed audio data streams together comprise the one or more layers of the input audio data stream. Furthermore, the apparatus comprises a merging module (190) for generating the merged audio data stream, having one or more layers, based on the plurality of single-layer audio data streams. Each layer of the input data audio streams, of the demultiplexed audio data streams, of the single-layer data streams and of the merged audio data stream comprises a pressure value of a pressure signal, a position value and a diffuseness value as audio data.