Neural Network Acoustic Analysis for Scene-Aware Audio

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional systems for rendering audio in virtual and augmented reality environments are inefficient, inflexible, and inaccurate due to the need for time-consuming impulse response capture and reliance on special equipment, leading to degraded immersion and sense of presence for users.

Innovation Solution

The use of neural networks to predict acoustic properties of a user environment based on simple audio recordings, eliminating the need for impulse response capture and enabling scene-aware audio rendering that accurately mimics the user's environment, using commodity devices like smartphones for efficient and flexible audio generation.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If conventional systems use impulse response capture to render audio in virtual/augmented reality environments, then audio rendering accuracy is improved, but system efficiency deteriorates due to time-consuming capture processes

Engineering Contradiction:
Improveacoustic property prediction accuracyVSAvoidaudio rendering efficiency
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The patent replaces the mechanical impulse response capture process with a neural network-based acoustic analysis system. The neural network processes ordinary audio recordings to predict acoustic properties, eliminating the need for specialized impulse response measurement equipment and procedures. This substitution maintains prediction accuracy while dramatically improving processing efficiency.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The system creates synthetic impulse responses by applying predicted acoustic properties to reference impulse responses. This copying approach allows the system to generate environment-specific audio characteristics without performing actual impulse response measurements, thereby improving efficiency while maintaining accuracy.

Inventive Principle:
Principle #26Copying

2Measurement precision

If conventional systems rely on special equipment for impulse response capture, then measurement precision is improved, but device complexity and ease of operation worsen

Engineering Contradiction:
Improveacoustic property measurement accuracyVSAvoidequipment requirements
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The neural network system is designed to process ordinary audio recordings from commodity devices, making the acoustic analysis capability universal and applicable to any device with a microphone. This eliminates the need for specialized measurement equipment while maintaining the ability to accurately predict acoustic properties of different environments.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The system enables acoustic analysis using audio recordings that users already capture with their existing devices. Instead of requiring separate measurement procedures and equipment, the system repurposes ordinary audio recordings for acoustic property prediction, simplifying the process and reducing device complexity requirements.

Inventive Principle:
Principle #25Self-service

3Manufacturing precision

If conventional systems use impulse response capture, then acoustic property accuracy is improved, but loss of time increases due to time-consuming capture and processing

Engineering Contradiction:
Improveacoustic property prediction accuracyVSAvoidimpulse response capture time
Core Design Contradiction:
Manufacturing precisionVSLoss of time

Solution Approach 1:

The neural network is pre-trained on a comprehensive dataset of audio recordings and corresponding acoustic properties. This preliminary training enables the system to make rapid predictions without requiring time-consuming impulse response capture and analysis during actual use. The heavy computational work is performed in advance during the training phase.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent replaces the time-consuming mechanical impulse response capture process with a computational neural network inference process. Once trained, the neural network can predict acoustic properties from ordinary audio recordings almost instantaneously, eliminating the time required for impulse response measurement and processing.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

4Measurement precision

If conventional systems require impulse response capture, then audio rendering accuracy is improved, but adaptability deteriorates due to inflexibility in deployment

Engineering Contradiction:
Improveenvironmental acoustic property accuracyVSAvoiddeployment flexibility
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The patent replaces the rigid impulse response capture methodology with a flexible neural network-based approach. The system can adapt to any environment by processing ordinary audio recordings captured in that environment, eliminating the need for standardized measurement procedures and specialized equipment. This substitution enables deployment in diverse settings without requiring controlled measurement conditions.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Data Source

PatentUS11812254B2Generating scene-aware audio using a neural network-based acoustic analysis
Publication Date: 2023.11.07 ADOBE INC
  • US11812254B2 patent drawing
  • US11812254B2 patent drawing
  • US11812254B2 patent drawing

AI summary

Methods, systems, and non-transitory computer readable storage media are disclosed for rendering scene-aware audio based on acoustic properties of a user environment. For example, the disclosed system can use neural networks to analyze an audio recording to predict environment equalizations and reverberation decay times of the user environment without using a captured impulse response of the user environment. Additionally, the disclosed system can use the predicted reverberation decay times with an audio simulation of the user environment to optimize material parameters for the user environment. The disclosed system can then generate an audio sample that includes scene-aware acoustic properties based on the predicted environment equalizations, material parameters, and an environment geometry of the user environment. Furthermore, the disclosed system can augment training data for training the neural networks using frequency-dependent equalization information associated with measured and synthetic impulse responses.