Dynamic Audio Mixing Based on User Pose

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing audio mixing technologies fail to provide an immersive and interactive audio experience for users, as they are often static and do not adapt to the user's pose or environment, limiting the realism and engagement in virtual environments.

Innovation Solution

A method and system that dynamically mix audio based on the user's pose, using sensors to determine the user's position, orientation, and movement, and adjust audio characteristics such as volume and spectral profiles in real-time, creating a customized audio experience tailored to the user's interaction with the media content.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If static audio mixing is used, then device complexity is reduced, but audio immersion and interactivity deteriorate

Engineering Contradiction:
Improveaudio immersionVSAvoidaudio mixing system
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The audio mixing system transitions from static to dynamic operation by continuously tracking user pose parameters (position, orientation, head movements) and adjusting audio characteristics in real-time. The mixer responds to changing user states by modifying volume, panning, and spectral profiles of different audio tracks, creating an adaptive immersive experience that evolves with user interaction.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The system implements feedback loops where sensor data about user pose is continuously fed back to the audio mixer, which then adjusts audio output accordingly. This closed-loop control enables the audio experience to respond to user actions, with the mixer receiving ongoing information about user position and orientation to dynamically recalibrate audio delivery.

Inventive Principle:
Principle #23Feedback

2Adaptability or versatility

If dynamic audio mixing based on user pose is implemented, then audio interactivity and realism are improved, but computational requirements and processing time increase

Engineering Contradiction:
Improveaudio interactivityVSAvoidaudio processing time
Core Design Contradiction:
Adaptability or versatilityVSLoss of time

Solution Approach 1:

The system performs preliminary actions by pre-defining audio tracks and their associated characteristics (volume, panning, spectral profiles) before runtime. When user pose changes are detected, the mixer selectively adjusts only the relevant audio track parameters based on predefined rules and relationships, rather than processing the entire audio signal from scratch, thus reducing real-time computational burden.

Inventive Principle:
Principle #10Preliminary action

3Ease of operation

If multiple audio tracks are mixed dynamically, then audio customization and user engagement are enhanced, but system complexity and computational load increase

Engineering Contradiction:
Improveuser engagementVSAvoidaudio mixing system
Core Design Contradiction:
Ease of operationVSDevice complexity

Solution Approach 1:

The audio system is segmented into multiple independent tracks, each representing a distinct sound source or audio element. The mixer operates on these segmented tracks individually, adjusting parameters such as volume, panning, and spectral content for each track based on user pose. This segmentation enables selective manipulation of audio elements without requiring complex global processing of the entire audio mix.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS11395089B2Mixing audio based on a pose of a user
Publication Date: 2022.07.19 GOOGLE LLC
  • US11395089B2 patent drawing
  • US11395089B2 patent drawing
  • US11395089B2 patent drawing

AI summary

A system, apparatus, and method are disclosed for utilizing a sensed pose of a user to dynamically control the mixing of audio tracks to provide a user with a more realistic, informative, and/or immersive audio experience with a virtual environment, such as a video.