Spatial Procedural Audio with Machine Learning for Realistic XR Sound

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Traditional procedural audio techniques rely on physical characteristics and semi-empirical models, often failing to produce realistic audio with sufficient uniqueness and realism, especially in extended reality environments.

Innovation Solution

A method using a machine learning model, such as a Generative Adversarial Network (GAN), generates spatial procedural audio by processing noise signals to produce unique and realistic audio effects without relying on image data, incorporating spatial parameters like Direction of Arrival and diffuseness.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If traditional procedural audio techniques using physical characteristics and semi-empirical models are used, then the system can generate procedural audio, but the audio lacks sufficient realism and uniqueness

Engineering Contradiction:
Improverealism of audioVSAvoidcomplexity of audio generation system
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent replaces traditional mechanical/physical modeling approaches (semi-empirical models of objects) with a machine learning model (deep neural network). This substitution allows the system to learn realistic audio generation patterns from training data without relying on complex physical models, thereby improving realism while reducing system complexity.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The patent changes the fundamental parameters of the audio generation system by transitioning from fixed physical models to a data-driven neural network model with learnable parameters. The model takes noise signals as input and generates audio signals with spatial characteristics, allowing dynamic adjustment of audio properties without complex physical simulations.

Inventive Principle:
Principle #35Parameter changes

2Adaptability or versatility

If traditional procedural audio techniques relying on image data and empirical models are used, then spatial audio can be generated, but the system requires multiple input sources and complex processing

Engineering Contradiction:
Improvespatial audio generation capabilityVSAvoidnumber of input sources required
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent creates a universal audio generation model that can handle multiple audio scenarios and spatial configurations through a single deep neural network. The model is trained on diverse spatial audio data and can generate different types of sounds with appropriate spatial characteristics, eliminating the need for separate processing pipelines for different input sources.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent extracts and processes only the essential audio signal and spatial parameters through the neural network, removing the dependency on image data and other non-audio inputs. The model directly maps noise signals to spatial audio outputs, simplifying the input requirements while maintaining spatial audio generation capabilities.

Inventive Principle:
Principle #2Taking out (Extraction)

3Reliability

If procedural audio is generated using pre-recorded libraries, then consistent audio quality can be achieved, but the audio lacks uniqueness for different situations

Engineering Contradiction:
Improveaudio quality consistencyVSAvoiduniqueness of audio for given situation
Core Design Contradiction:
ReliabilityVSAdaptability or versatility

Solution Approach 1:

The patent introduces dynamic audio generation where the deep neural network processes different noise signals to generate unique audio outputs for each situation. The model maintains consistent quality through its trained parameters while adapting to different contexts by varying its input noise signals, enabling both consistency and uniqueness simultaneously.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

Instead of using fixed pre-recorded copies from libraries, the patent uses the neural network to synthesize new audio signals that replicate the characteristics of real-world sounds. The model learns from training data and generates novel audio instances that maintain quality consistency while being unique to each generation event.

Inventive Principle:
Principle #26Copying

Data Source

PatentUS12356166B1Method and system for generating spatial procedural audio
Publication Date: 2025.07.08 APPLE INC
  • US12356166B1 patent drawing
  • US12356166B1 patent drawing
  • US12356166B1 patent drawing

AI summary

A method performed by a programmed processor, the method including receiving an input noise signal, generating, using a machine learning model that has an input based on the input noise signal, a mono audio signal that includes a sound and a spatial parameter for the mono audio signal, and generating spatial audio data by spatially encoding the mono audio signal according to the spatial parameter.