Audio Signal Processing Apparatus for Immersive Spatial Sound

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current technologies face challenges in generating spatial sound in real-time for virtual reality devices with binaural rendering, particularly in reducing computational requirements while maintaining interactive user experiences.

Innovation Solution

An audio signal processing apparatus that receives and processes input audio signals by generating reflected sounds based on spatial information, using spectral modification filters and binaural parameter pairs to produce output audio signals that simulate virtual sound sources and environments with reduced computational load.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Manufacturing precision

If Room Impulse Response (RIR) filter with long length is used to reproduce spatial sound, then the realism and immersive quality of virtual reality audio is improved, but the computational load and memory requirements increase significantly

Engineering Contradiction:
Improvespatial sound reproduction qualityVSAvoidcomputational load
Core Design Contradiction:
Manufacturing precisionVSDevice complexity

Solution Approach 1:

The patent divides the long RIR filter into multiple shorter sub-filters arranged in a cascade structure. Each sub-filter processes a portion of the spatial sound reproduction task, allowing the system to achieve the effect of a long filter while using computationally manageable short filters. This segmentation reduces the computational complexity from O(N) for a single long filter to O(k×M) where k is the number of sub-filters and M is the length of each sub-filter, with M << N.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent transitions from using a single long temporal filter to using multiple short filters arranged in a spatial cascade configuration. This dimensional change allows the system to distribute the computational workload across multiple stages rather than requiring one massive computation, effectively trading temporal complexity for spatial arrangement of processing stages.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Manufacturing precision

If binaural rendering processes multiple target objects and channels, then the immersive quality and spatial accuracy are improved, but the power consumption and computational burden increase

Engineering Contradiction:
Improvespatial audio accuracyVSAvoidpower consumption
Core Design Contradiction:
Manufacturing precisionVSUse of energy by moving object

Solution Approach 1:

The patent segments the binaural rendering process into distinct stages: spatial positioning, RIR application, and binaural transformation. Each target object and audio channel is processed through this segmented pipeline independently, allowing for optimized resource allocation. The cascade of short RIR filters is applied once per object-channel pair rather than using a single long filter, reducing redundant computations and power consumption.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent applies partial action by using shorter RIR filters that capture the essential spatial characteristics without processing the entire long impulse response. This provides sufficient spatial audio accuracy for immersive VR experiences while avoiding the excessive computational burden of processing the complete long RIR, achieving an optimal balance between quality and power consumption.

Inventive Principle:
Principle #16Partial or excessive action

3Adaptability or versatility

If spatial reverberation filter is updated in real-time to reflect changing user viewpoint and location, then the interactive quality is improved, but the computational requirements and processing time increase

Engineering Contradiction:
Improveinteractive qualityVSAvoidprocessing speed
Core Design Contradiction:
Adaptability or versatilityVSProductivity

Solution Approach 1:

The patent pre-computes and stores multiple short RIR filters corresponding to different spatial positions and orientations in the virtual environment. When the user's viewpoint or location changes, the system selects and combines the appropriate pre-computed sub-filters through simple mixing operations rather than performing full real-time convolution. This preliminary preparation enables rapid adaptation to user movements while maintaining high interactive quality.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent segments the spatial environment into discrete zones or positions, each with pre-computed short RIR filters. This segmentation allows the system to handle real-time viewpoint changes by switching between pre-computed segments rather than continuously computing new filters, significantly improving processing speed while maintaining adaptability to user movement.

Inventive Principle:
Principle #1Segmentation

4Manufacturing precision

If all data in the long RIR filter is used for spatial sound reproduction, then the realism of room acoustics is improved, but the memory requirements and data transmission load increase

Engineering Contradiction:
Improveroom acoustic realismVSAvoiddata volume
Core Design Contradiction:
Manufacturing precisionVSQuantity of substance

Solution Approach 1:

The patent divides the long RIR filter data into multiple short sub-filter segments that are stored and processed separately. This segmentation dramatically reduces the memory footprint from storing one large filter to storing multiple small filters. The cascade arrangement of these short filters reconstructs the essential spatial acoustic characteristics without requiring the full data volume of the original long filter, achieving efficient memory usage while maintaining acoustic realism.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS11184727B2Audio signal processing method and device
Publication Date: 2021.11.23 GAUDI AUDIO LAB
  • US11184727B2 patent drawing
  • US11184727B2 patent drawing
  • US11184727B2 patent drawing

AI summary

An audio signal processing apparatus including a receiving unit receiving an input audio signal, a processor generating an output audio signal reproducing a virtual sound source corresponding to the input audio signal in a virtual space, and an output unit outputting an output audio signal generated by the processor is disclosed. The processor may obtain spatial information related to the virtual space including a virtual sound source corresponding to the input audio signal and a listener, filter the input audio signal based on a location of the virtual sound source and the spatial information to generate at least one reflected sound corresponding to each of at least one mirror plane in the virtual space, obtain a relative location of a virtual reflect sound source with respect to a location and a view-point of the listener, based on information of the view-point of the listener and a location of the virtual reflect sound source corresponding to each of the at least one reflected sound, and binaural render the at least one reflected sound, based on the relative location of the virtual reflect sound source corresponding to each of the at least one reflected sound.