Spatial Audio Encoding via Machine Learning Compression

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Communication devices typically transmit audio as mono, lacking spatial cues, resulting in a dull listening experience for recipients, as they do not receive immersive spatial audio due to the absence of positional information.

Innovation Solution

A system utilizing a machine learning model to encode and decode audio signals from multiple microphones, transforming them into binaural or ambisonic formats that simulate the spatial audio experience, allowing for immersive listening by encoding spatial data and transmitting compressed audio between devices.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Device complexity

If audio is transmitted as mono from source device to target device, then data transmission is simplified and requires less bandwidth, but spatial audio experience is lost and listening experience becomes dull

Engineering Contradiction:
Improveaudio transmission complexityVSAvoidlistening experience quality
Core Design Contradiction:
Device complexityVSEase of manufacture

Solution Approach 1:

The patent applies parameter changes by transforming audio from mono format to binaural/ambisonic formats through machine learning encoding and decoding processes. The system changes the audio parameter representation to include spatial information, converting flat mono audio into three-dimensional spatial audio that provides immersive listening experience while maintaining efficient transmission.

Inventive Principle:
Principle #35Parameter changes

2Ease of manufacture

If spatial audio encoding is applied to capture positional information, then immersive listening experience is improved, but data transmission requirements and processing complexity increase

Engineering Contradiction:
Improvelistening experience qualityVSAvoidencoding and transmission complexity
Core Design Contradiction:
Ease of manufactureVSDevice complexity

Solution Approach 1:

The patent uses a machine learning model as an intermediary between the source device and target device. The ML model encodes spatial audio information in a compressed representation that can be efficiently transmitted and then decoded at the target device. This intermediary approach allows spatial audio to be transmitted without requiring full high-fidelity audio channels, reducing overall system complexity while maintaining immersive experience.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Measurement precision

If multiple microphones are used to capture spatial audio, then audio spatial information is improved, but device hardware complexity and cost increase

Engineering Contradiction:
Improvespatial audio capture precisionVSAvoidmicrophone array complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent creates a virtual copy of the spatial audio experience through machine learning encoding rather than physically replicating complex microphone arrays at every device. The source device with multiple microphones captures the spatial information once, and the ML model generates a compressed spatial representation that can be transmitted to any number of target devices, effectively copying the spatial audio experience without requiring each target device to have complex microphone hardware.

Inventive Principle:
Principle #26Copying

Data Source

PatentUS12200465B2Spatial audio recording from home assistant devices
Publication Date: 2025.01.14 GOOGLE LLC
  • US12200465B2 patent drawing
  • US12200465B2 patent drawing
  • US12200465B2 patent drawing

AI summary

The technology generally relates to spatial audio communication between devices. For example, a first device and a second device may be connected via a communication link. The first device may capture audio signals in an environment through two or more microphones. The first device may encode the captured audio with spatial configuration data. The first device may transmit the encoded audio via the communication link to the second device. The second device may decode the encoded audio into binaural or ambisonic audio to be output by one or more speakers of the second device. The binaural or ambisonic audio may be converted into spatial audio to be output. The second device may output the binaural or spatial audio to create an immersive listening experience.