Audio Impulse Responses for Metric 3D Human Pose Lifting

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Recovering a 3D pose of a person in metric scale from a flat 2D image is an ill-posed problem due to the unknown height of the person, and existing methods requiring multiple cameras increase quadratically with the observed environment area.

Innovation Solution

A system and method using a neural network trained with audio and image data to estimate a 3D metric pose of a subject human, employing audio sensors and cameras to lift the 3D pose by processing audio recordings and images, leveraging audio impulse responses to resolve height ambiguity.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If multiple cameras are used to support metric reconstruction, then measurement precision is improved, but device complexity increases quadratically as the area of observed environment increases

Engineering Contradiction:
Improvemetric reconstruction precisionVSAvoidnumber of cameras
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent introduces audio recordings as an intermediary modality to bridge the gap between 2D images and 3D metric reconstruction. Audio data serves as a mediator that provides distance and spatial information, enabling accurate metric reconstruction without requiring multiple cameras. The audio-visual model integrates these intermediary audio cues to resolve the ill-posed nature of monocular 3D pose estimation.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent replaces the mechanical/optical system of multiple cameras with an audio-visual system. Instead of using additional optical sensors (cameras) to gather spatial information, the system substitutes audio recordings that capture acoustic propagation characteristics. This substitution leverages the properties of sound wave propagation to infer spatial relationships, reducing hardware complexity while maintaining measurement precision.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Measurement precision

If additional scene assumptions such as obtaining the height of the person separately are made, then measurement precision is improved, but ease of operation deteriorates

Engineering Contradiction:
Improve3D pose reconstruction precisionVSAvoidoperational simplicity
Core Design Contradiction:
Measurement precisionVSEase of operation

Solution Approach 1:

The patent merges the task of height estimation with the 3D pose reconstruction process by integrating audio-visual information. Instead of separately obtaining height measurements and then performing pose reconstruction, the system combines audio cues (which provide distance and spatial context) with visual data in a unified neural network model. This integration automatically infers height as part of the overall 3D pose estimation, eliminating the need for separate height measurement steps.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent creates a universal audio-visual model that performs multiple functions simultaneously: it estimates 3D pose, infers height, and determines spatial relationships all within a single system. The multi-modal model is designed to handle various aspects of human pose estimation and environmental understanding in one unified framework, making the system more versatile and easier to operate without requiring separate specialized modules for each measurement task.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS12423861B2Metric lifting of 3D human pose using sound
Publication Date: 2025.09.23 SAMSUNG ELECTRONICS CO LTD
  • US12423861B2 patent drawing
  • US12423861B2 patent drawing
  • US12423861B2 patent drawing

AI summary

A pose of a person is estimated using an image and audio impulse responses. The image represents a 2D scene including the person. The audio impulse responses are obtained with the present absent and present in an environment. The pose is reconstructed based on the image and the one or more audio impulse responses. The pose is a metric scale human pose.