Deep Neural Network for Binaural Audio Virtual Rotation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Binaural audio recordings with fixed head orientations do not automatically adjust when the listener rotates their head, lacking a simple and cost-effective method to change sound sources' perceived locations in real-time.

Innovation Solution

A deep-learning based audio regression method using a deep neural network (DNN) processes 2-channel binaural audio signals and a selected rotation angle to generate a new binaural audio output signal, effectively simulating head rotation by extracting spherical location information and adjusting sound sources' perceived locations accordingly.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If binaural recordings are made with fixed head orientation, then the recording process is simple and cost-effective, but the sound source locations cannot be automatically adjusted when the listener rotates their head

Engineering Contradiction:
Improvesimplicity of recording processVSAvoidability to adjust sound source locations with head rotation
Core Design Contradiction:
Ease of operationVSAdaptability or versatility

Solution Approach 1:

The system performs preliminary action by pre-processing the binaural audio signal to extract spherical location information of sound sources before head rotation occurs. The deep neural network analyzes the fixed-head recording and identifies the spatial positions of sound sources, enabling subsequent virtual rotation without requiring re-recording or complex directional microphones.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent applies copying by creating a virtual copy of the binaural audio scene that can be independently manipulated. Instead of physically rotating the listener's head or re-recording the audio, the system generates a synthesized representation of the audio field and rotates this virtual copy to match the listener's head orientation, preserving the original recording's simplicity while adding adaptability.

Inventive Principle:
Principle #26Copying

2Measurement precision

If expensive re-recording or complex directional microphone arrays are used to change sound source locations, then the sound source locations can be accurately adjusted, but the cost and device complexity increase significantly

Engineering Contradiction:
Improveaccuracy of sound source locationVSAvoidcomplexity of recording system
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent replaces mechanical systems (physical rotation of microphones, complex directional arrays, or re-recording setups) with a computational approach. A deep neural network processes the fixed binaural recording and mathematically transforms the sound source locations to match head rotation, eliminating the need for expensive hardware changes while maintaining location accuracy.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The system achieves parameter changes by modifying the spatial parameters of the audio signal through computational transformation. The deep neural network adjusts the spherical location parameters of sound sources based on head rotation angle, effectively changing where sounds are perceived without altering the physical recording setup or requiring complex directional microphones.

Inventive Principle:
Principle #35Parameter changes

3Device complexity

If traditional signal processing methods are used to simulate head rotation, then the implementation is simpler, but the accuracy of sound source location transformation is insufficient

Engineering Contradiction:
Improvesimplicity of implementationVSAvoidaccuracy of sound source location transformation
Core Design Contradiction:
Device complexityVSMeasurement precision

Solution Approach 1:

The patent introduces an intermediary - a deep neural network - that bridges the simple fixed-head recording and the desired rotated audio output. This intermediary component analyzes the complex relationships between head orientation and sound source locations, processes the binaural signal accordingly, and generates accurate transformed audio that matches the listener's head position without requiring complex recording hardware.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS20240244389A1Deep learning solution for virtual rotation of binaural audio signals
Publication Date: 2024.07.18 INTEL CORP
  • US20240244389A1 patent drawing
  • US20240244389A1 patent drawing
  • US20240244389A1 patent drawing

AI summary

Techniques are provided herein for providing binaural sound signals that are virtually rotated to match head rotation, such that audio output to headphones is perceived to maintain its location relative to user when a user turns their head. In particular, techniques are presented to extract spherical location information already embedded in binaural signals to generate binaural sound signals that change to match head rotation. A deep-learning based audio regression method can use a 2-channel binaural audio signal and a rotation angle as input, and generate a new binaural audio output signal with the rotated environment corresponding to the rotation angle. The deep-learning based audio regression method can be implemented as a neural network, and can include deep learning operations, such as convolution, pooling, elementwise operation, linear operation, and nonlinear operation. A deep learning operation may be performed on internal parameters of the DNNs and one or more activations.