Spatial Audio Encoding via Machine Learning Compression
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Communication devices typically transmit audio as mono, lacking spatial cues, resulting in a dull listening experience for recipients, as they do not receive immersive spatial audio due to the absence of positional information.
Innovation Solution
A system utilizing a machine learning model to encode and decode audio signals from multiple microphones, transforming them into binaural or ambisonic formats that simulate the spatial audio experience, allowing for immersive listening by encoding spatial data and transmitting compressed audio between devices.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Device complexity
If audio is transmitted as mono from source device to target device, then data transmission is simplified and requires less bandwidth, but spatial audio experience is lost and listening experience becomes dull
Solution Approach 1:
The patent applies parameter changes by transforming audio from mono format to binaural/ambisonic formats through machine learning encoding and decoding processes. The system changes the audio parameter representation to include spatial information, converting flat mono audio into three-dimensional spatial audio that provides immersive listening experience while maintaining efficient transmission.
2Ease of manufacture
If spatial audio encoding is applied to capture positional information, then immersive listening experience is improved, but data transmission requirements and processing complexity increase
Solution Approach 1:
The patent uses a machine learning model as an intermediary between the source device and target device. The ML model encodes spatial audio information in a compressed representation that can be efficiently transmitted and then decoded at the target device. This intermediary approach allows spatial audio to be transmitted without requiring full high-fidelity audio channels, reducing overall system complexity while maintaining immersive experience.
3Measurement precision
If multiple microphones are used to capture spatial audio, then audio spatial information is improved, but device hardware complexity and cost increase
Solution Approach 1:
The patent creates a virtual copy of the spatial audio experience through machine learning encoding rather than physically replicating complex microphone arrays at every device. The source device with multiple microphones captures the spatial information once, and the ML model generates a compressed spatial representation that can be transmitted to any number of target devices, effectively copying the spatial audio experience without requiring each target device to have complex microphone hardware.
Data Source
AI summary
The technology generally relates to spatial audio communication between devices. For example, a first device and a second device may be connected via a communication link. The first device may capture audio signals in an environment through two or more microphones. The first device may encode the captured audio with spatial configuration data. The first device may transmit the encoded audio via the communication link to the second device. The second device may decode the encoded audio into binaural or ambisonic audio to be output by one or more speakers of the second device. The binaural or ambisonic audio may be converted into spatial audio to be output. The second device may output the binaural or spatial audio to create an immersive listening experience.


