Spherical CNN Upsampling for HRTF Spatial Resolution

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing audio signal processing methods face challenges in simulating virtual spatial sounds with high precision due to variations in human anatomy, leading to inefficiencies in data volume and network bandwidth when transmitting high-resolution head-related transfer functions (HRTFs).

Innovation Solution

A method utilizing a spherical convolutional neural network to upsample a low-resolution HRTF into a high-resolution HRTF, increasing spatial resolution for more precise audio localization while maintaining efficient data transmission and storage.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If high-resolution HRTF is transmitted and stored, then audio localization precision is improved, but data volume and network bandwidth requirements increase

Engineering Contradiction:
Improveaudio localization precisionVSAvoiddata volume
Core Design Contradiction:
Measurement precisionVSQuantity of substance

Solution Approach 1:

The system performs up-sampling of low-resolution HRTF to high-resolution HRTF in advance using a pre-trained spherical CNN model, so that when audio processing is needed, the high-resolution data is already available locally without requiring large bandwidth transmission during actual use

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

Instead of transmitting the full high-resolution HRTF dataset, the system transmits a compact low-resolution HRTF and uses a trained spherical CNN model to generate the high-resolution version locally, creating a compressed representation that can be reconstructed when needed

Inventive Principle:
Principle #26Copying

2Reliability

If high-resolution HRTF is used, then virtual spatial sound simulation quality is improved, but storage requirements increase

Engineering Contradiction:
Improvevirtual spatial sound simulation qualityVSAvoidstorage requirements
Core Design Contradiction:
ReliabilityVSVolume of stationary object

Solution Approach 1:

The system changes the resolution parameter of HRTF dynamically - storing at low resolution and converting to high resolution only when needed for audio processing, thus optimizing the balance between quality and storage by adjusting the data representation based on operational requirements

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The low-resolution HRTF serves as a compact container that can be transformed into the full high-resolution HRTF when needed, similar to how a small doll contains a larger one, allowing the system to store minimal data while having access to comprehensive high-resolution data when required

Inventive Principle:
Principle #7Nested doll (Nesting)

3Loss of energy

If low-resolution HRTF is transmitted, then network bandwidth usage is reduced, but audio localization precision deteriorates

Engineering Contradiction:
Improvenetwork bandwidth usageVSAvoidaudio localization precision
Core Design Contradiction:
Loss of energyVSMeasurement precision

Solution Approach 1:

The spherical CNN model acts as an intermediary that receives the compact low-resolution HRTF and transforms it into high-resolution HRTF, bridging the gap between efficient transmission and high-precision processing without requiring direct transmission of large datasets

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS11477597B2Audio processing method and electronic apparatus
Publication Date: 2022.10.18 HTC CORP
  • US11477597B2 patent drawing
  • US11477597B2 patent drawing
  • US11477597B2 patent drawing

AI summary

An audio processing method is configured to upsample a first head related transfer function (HRTF) into a second HRTF. The first HRTF defines first audio feature values distributed at first intersection nodes in a spherical coordinate system over audio frequencies. The first intersection nodes are arranged with a first spatial resolution in the spherical coordinate system. The first HRTF is upsampled into the second HRTF by a spherical convolutional neural network model. The second HRTF defines a plurality of second audio feature values distributed at second intersection nodes in the spherical coordinate system over the audio frequencies. The second intersection nodes in the second HRTF are arranged with a second spatial resolution higher than the first spatial resolution of the first intersection nodes in the first HRTF.