TOF and HOA Sensor Fusion for Sound Source Isolation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current audio and video sensor platforms struggle with isolating and identifying sound sources in dynamic environments due to limitations in sound source discrimination and localization, particularly in the presence of external noise and occlusions, leading to inefficiencies in real-time operation and high processing complexity.

Innovation Solution

A system combining a camera array and a microphone array, with DOA estimation, sound source identification, and sensor fusion modules, to accurately locate and isolate sound sources by converting sound vectors and source positions into a common data representation, using beamforming to record audio from active sources while discarding noise.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If multiple sensors tailored to different domains are used to analyze sound sources, then measurement precision and reliability are improved, but device complexity increases

Engineering Contradiction:
Improvesound source identification accuracyVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent combines camera arrays and microphone arrays into an integrated multi-sensor system. The camera array captures visual information while the microphone array captures audio information, and both are processed together through sensor fusion to identify sound sources with high precision. This merging of different sensor types resolves the contradiction by achieving accurate sound source identification through complementary data while managing system complexity through unified processing architecture.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The system employs sensors that serve multiple functions: the camera array not only captures visual data but also helps localize sound sources through visual tracking, while the microphone array provides both audio recording and spatial positioning information. This multi-functionality improves measurement precision without proportionally increasing device complexity, as each sensor contributes to multiple aspects of sound source analysis.

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Reliability

If sensor fusion techniques are used to correct occlusion issues, then reliability is improved, but processing time increases

Engineering Contradiction:
Improvesound source localization accuracyVSAvoidprocessing time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The system performs preliminary tracking of sound sources using the camera array to establish visual positions before audio processing. This preliminary visual localization provides initial position estimates that guide the microphone array processing, reducing the computational search space and enabling faster correction of occlusion issues through sensor fusion without excessive processing time.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent replaces complex mechanical tracking systems with computational sensor fusion algorithms. Instead of using additional mechanical sensors or complex hardware mechanisms to track occluded sound sources, the system uses software-based fusion of camera and microphone data to infer sound source positions, reducing processing time while maintaining reliability.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Measurement precision

If beamforming is used to isolate sound sources, then measurement precision is improved, but device complexity increases

Engineering Contradiction:
Improvesound source isolation accuracyVSAvoidprocessing complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent segments the audio processing into distinct beamforming stages that operate on different spatial regions. The system divides the environment into multiple zones and applies beamforming selectively to each zone, isolating sound sources step-by-step rather than processing all audio data uniformly. This segmentation improves sound source isolation accuracy while reducing processing complexity by focusing computational resources on specific regions of interest.

Inventive Principle:
Principle #1Segmentation

Applied Scientific Principles

This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.

Function Achieved in This Case

The system achieves accurate sound source localization and isolation, maintaining temporal saliency during speech, reducing noise interference, and improving real-time operation by integrating camera and microphone data for enhanced sound analysis.

Implementation Method 1

a microphone array positioned at a second height above the target area, the microphone array including a plurality of microphones oriented to receive sound produced from within the target area

Methodology Applied
Scientific EffectSound: Sound

Data Source

PatentUS20260059251A1Multi-sensor systems and methods for providing immersive virtual environments
Publication Date: 2026.02.26 RENESSELAER POLYTECHNIC INST
  • US20260059251A1 patent drawing
  • US20260059251A1 patent drawing
  • US20260059251A1 patent drawing

AI summary

A time-of-flight (TOF) array including a plurality of cameras is positioned with a coplanar Higher Order Ambisonics (HOA) array including a plurality of microphones above a target area such as a room or space for analyzing speaker activity in that area. The HOA microphone array iteratively samples sound energy level frames captured by the microphones to identify a sound vector associated with the global maximum energy level in the sound energy level frame. This sound vector is fused with the data from the TOF camera array that identified the positions of sound sources in the target area to associate produced sound corresponding with the sound vector with a physical active sound source in the target area. A beamformer can then be used to save audio corresponding to the sound vector and discard audio not associated with the sound vector to produce a sound recording associated with the active sound source.