Spatial Audio Synthesis Using Microphone Array and Sound Profiler

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing audio reproduction techniques fail to accurately emulate the three-dimensional audio experience, particularly in virtual and augmented reality, due to oversimplified assumptions about sound propagation and the inability to account for multiple sound sources and listener positions, leading to inefficient and inaccurate audio processing.

Innovation Solution

The use of a microphone array and sound profiler with processing circuitry and memory to generate synthesized audio based on sound beam metadata, target listener location data, and spatial sound wave characteristics, applying Fast Fourier Transform and inverse FFT to produce timed sound coefficients that accurately emulate sound as heard by a listener at a specific position and orientation.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If multiple audio capturing devices are placed throughout the space to improve audio accuracy for multiple listeners, then the audio reproduction accuracy improves, but the device complexity and processing difficulty increase significantly

Engineering Contradiction:
Improveaudio reproduction accuracyVSAvoidnumber of capturing devices
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

A single microphone array is designed to serve multiple functions: it can capture audio for multiple listeners simultaneously, handle multiple sound sources, and adapt to different spatial configurations. The sound profiler processes the captured audio to generate personalized audio streams for each listener position, making the single array universal for serving the entire space rather than requiring dedicated devices for each listener.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The system changes parameters such as time delays, gain factors, and phase shifts dynamically based on listener position and sound source location. By adjusting these audio processing parameters in real-time, the system achieves accurate spatial audio reproduction for multiple listeners without needing multiple physical capturing devices, thus resolving the contradiction between accuracy and device complexity.

Inventive Principle:
Principle #35Parameter changes

2Measurement precision

If beamforming techniques are used to reproduce sound in predetermined directions, then the directional sound reproduction improves, but the ability to accurately represent sounds from multiple directions and positions deteriorates

Engineering Contradiction:
Improvedirectional sound reproductionVSAvoidmulti-directional sound representation
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The audio processing is segmented into multiple independent beamforming operations, each handling a specific sound source or listener position. The sound profiler divides the complex multi-directional audio scene into manageable segments that can be processed individually and then combined, allowing the system to maintain directional accuracy for each sound source while simultaneously representing sounds from multiple directions.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The beamforming directions and parameters are made dynamic rather than fixed. The system continuously adapts the beamforming parameters based on the positions of sound sources and listeners, allowing the directional reproduction to change in real-time. This dynamic adaptation enables accurate representation of sounds from multiple directions and positions without sacrificing directional precision.

Inventive Principle:
Principle #15Dynamics

3Productivity

If a simplified sphere-like propagation model is assumed for sound sources, then the processing complexity is reduced, but the audio accuracy for different listener positions and orientations deteriorates

Engineering Contradiction:
Improveprocessing efficiencyVSAvoidaudio accuracy for listener position
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The sound profiler performs preliminary processing of the captured audio signals, pre-calculating transfer functions and spatial characteristics for different listener positions. This preliminary action creates a library of pre-processed audio data that can be quickly retrieved and applied during real-time reproduction, maintaining high accuracy for different listener positions and orientations without the full processing complexity during actual playback.

Inventive Principle:
Principle #10Preliminary action

Applied Scientific Principles

This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.

Function Achieved in This Case

This solution provides a more accurate and immersive audio experience by accurately reconstructing sound from the perspective of a listener within a space, overcoming the limitations of existing solutions by accounting for the position and orientation of the listener relative to the sound source and space, enhancing the spatial audio reproduction.

Implementation Method 1

audio signals captured in a three-dimensional space

Methodology Applied
Scientific EffectAcoustic wave propagation: Sound

Implementation Method 2

applying Fast Fourier Transform and inverse FFT to produce timed sound coefficients

Methodology Applied
Scientific EffectFast Fourier Transform:

Implementation Method 3

the sounds produced by objects moving away from or towards each other may be heard differently due to the doppler effect

Methodology Applied
Scientific EffectDoppler effect: Doppler Effect

Data Source

PatentUS11881206B2System and method for generating audio featuring spatial representations of sound sources
Publication Date: 2024.01.23 INSOUNDZ
  • US11881206B2 patent drawing
  • US11881206B2 patent drawing
  • US11881206B2 patent drawing

AI summary

Systems and methods for spatially emulating a sound source. An apparatus includes a microphone array including microphones; and a sound profiler communicatively connected to the microphone array, the sound profiler including a processing circuitry and a memory which contains instructions that, when executed by the processing circuitry, configure the apparatus to: generate synthesized audio based on sound beam metadata, a sound profile, and target listener location data, wherein the sound beam metadata includes timed sound beams defining a directional dependence of a spatial sound wave, wherein the sound profile includes timed sound coefficients determined based on audio signals captured in a space wherein the target listener location data includes a position and an orientation, wherein the synthesized audio emulates sound that would be heard by a listener at the position and orientation of the target listener location data; and providing the synthesized audio for projection.