Predictive Model Spatial Audio Generation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current systems for capturing and generating audio feedback for visual content, such as virtual reality environments, often rely on expensive and complex equipment, and typically produce one- or two-dimensional audio signals that fail to convey the location, depth, or position of sound sources, resulting in a limited immersive auditory experience for users.

Innovation Solution

A predictive model, such as a neural network, is used to generate spatial audio signals from non-spatial audio inputs, allowing for the conversion of mono or stereo audio into ambisonic audio that indicates the location of sound sources within a visual content environment, enabling a more immersive auditory experience by associating spatial audio with visual elements.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If conventional audio capture equipment is used, then the system is simple and inexpensive, but the audio output is limited to one-dimensional or two-dimensional signals that fail to convey spatial location

Engineering Contradiction:
Improvespatial audio localization accuracyVSAvoidaudio capture system complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent replaces complex mechanical audio capture systems with a computational approach using neural networks. Instead of using multiple microphones and complex spatial audio recording equipment, the system uses a predictive model that processes standard mono or stereo audio inputs and generates spatial audio representations through neural network transformations, substituting mechanical complexity with algorithmic processing

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The neural network serves as an intermediary between the simple audio input and the spatial audio output. It takes conventional audio signals as input, processes them through learned transformations, and produces spatial audio representations that convey location information without requiring complex capture hardware

Inventive Principle:
Principle #24Intermediary (Mediator)

2Adaptability or versatility

If expensive spatial audio capture equipment is used, then three-dimensional audio feedback is achieved, but the system becomes complex and unavailable for general use

Engineering Contradiction:
Improvespatial audio experience qualityVSAvoidcapture equipment complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent employs the principle of using simple, readily available audio capture devices (smartphones, standard microphones) instead of expensive specialized equipment. The system accepts standard audio inputs from ubiquitous devices and transforms them into spatial audio representations, making high-quality spatial audio experiences accessible without requiring costly or specialized hardware

Inventive Principle:
Principle #27Cheap short-living objects (Disposable)

Solution Approach 2:

The invention replaces complex mechanical spatial audio capture systems with a software-based neural network approach. Instead of using multiple microphones arranged in specific configurations or specialized spatial audio recorders, the system uses a predictive model that processes standard audio signals and generates spatial representations through computational transformations

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Ease of operation

If one-dimensional or two-dimensional audio feedback is output, then the system is simple and widely compatible, but the user cannot perceive the location or depth of sound sources

Engineering Contradiction:
Improveaudio system compatibilityVSAvoidspatial location information
Core Design Contradiction:
Ease of operationVSLoss of information

Solution Approach 1:

The patent applies dimensionality change by transforming one-dimensional or two-dimensional audio signals into three-dimensional spatial audio representations. The neural network processes conventional audio inputs and generates output that includes spatial location information, effectively adding the third dimension of spatial awareness while maintaining compatibility with standard audio playback devices

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

4Device complexity

If two-dimensional audio feedback is used, then the system remains simple and broadly compatible, but all audio content appears to originate from a single point in space

Engineering Contradiction:
Improveaudio processing complexityVSAvoidaudio spatial positioning accuracy
Core Design Contradiction:
Device complexityVSManufacturing precision

Solution Approach 1:

The system transitions from two-dimensional audio processing to three-dimensional spatial audio generation. The neural network analyzes the input audio and visual content to determine spatial positions of sound sources and generates corresponding spatial audio representations that accurately position sounds in three-dimensional space, eliminating the single-point origin limitation of two-dimensional audio

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Data Source

PatentUS10701303B2Generating spatial audio using a predictive model
Publication Date: 2020.06.30 ADOBE INC
  • US10701303B2 patent drawing
  • US10701303B2 patent drawing
  • US10701303B2 patent drawing

AI summary

Certain embodiments involve generating and providing spatial audio using a predictive model. For example, a generates, using a predictive model, a visual representation of visual content provideable to a user device by encoding the visual content into the visual representation that indicates a visual element in the visual content. The system generates, using the predictive model, an audio representation of audio associated with the visual content by encoding the audio into the audio representation that indicates an audio element in the audio. The system also generates, using the predictive model, spatial audio based at least in part on the audio element and associating the spatial audio with the visual element. The system can also augment the visual content using the spatial audio by at least associating the spatial audio with the visual content.