Spatial Audio HRTF Selection Using Head-to-Torso Orientation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing virtual audio rendering systems fail to accurately reproduce spatial audio due to the disregard for the orientation of the torso relative to the head, leading to unrealistic sound perceptions when the head and torso move separately.

Innovation Solution

A media system that determines the head-to-torso orientation using head tracking data, either directly measured or inferred, to select an appropriate head-related transfer function (HRTF) for spatial audio reproduction, applying binaural audio filters that account for both head-to-source and head-to-torso orientations.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If existing virtual audio rendering systems use HRTF based on head orientation only, then the system complexity is reduced, but the spatial audio reproduction accuracy deteriorates when head and torso move separately

Engineering Contradiction:
Improvespatial audio reproduction accuracyVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The system segments the orientation determination into two independent components: head orientation and torso orientation. By separately tracking and processing these two orientations, the system can accurately represent the relative position between head and torso, thereby improving spatial audio reproduction accuracy without requiring a completely new complex system architecture.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system adds another dimension to the traditional HRTF parameter set by incorporating torso orientation information. Instead of using only head orientation (azimuth and elevation angles), the system now processes head orientation relative to torso orientation, effectively adding a new dimensional aspect to the spatial audio rendering model.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Reliability

If the system tracks only head orientation, then the tracking complexity is reduced, but the spatial audio realism deteriorates

Engineering Contradiction:
Improvespatial audio realismVSAvoidtracking complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The system uses the relative orientation between head and torso as an intermediary parameter to bridge the gap between tracking complexity and audio realism. By focusing on the relative position rather than absolute positions of both head and torso, the system achieves realistic spatial audio without requiring complex absolute tracking of both components.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The system changes the parameters used in HRTF calculation from traditional head-oriented parameters to head-to-torso relative parameters. This parameter transformation allows the system to maintain audio realism by accounting for torso orientation effects while using processed orientation data that reflects the actual acoustic relevant relationships.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS12513485B2Spatial audio reproduction based on head-to-torso orientation
Publication Date: 2025.12.30 APPLE INC
  • US12513485B2 patent drawing
  • US12513485B2 patent drawing
  • US12513485B2 patent drawing

AI summary

A media system and a method of using the media system to reproduce spatial audio based on head-to-torso orientation, are described. The method includes determining a head-to-source orientation and a head-to-torso orientation based on head orientation data generated by a head tracking device. Determining the head-to-torso orientation includes determining torso movements based on movements of the head. The torso can be determined to move when the head movements meet a head movement condition, such as a predetermined angle of movement or pattern of movement. A binaural audio filter that is based on a head-related transfer function corresponding to both the head-to-source orientation and the head-to-torso orientation is applied to an audio input signal to generate an audio output signal. The audio output signal is played to accurately recreate spatial audio having sounds emitted to the user by a sound source. Other aspects are also described and claimed.