Voice Biometric Speaker Isolation for Real-Time Noisy Audio
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional methods for extracting target audio signals from multiple audio signals are inefficient in noisy environments, requiring computationally intensive post-production and significant bandwidth, and are impractical for real-time applications in crowded spaces.
Innovation Solution
The method involves extracting acoustic characteristics from audio signals, associating metadata with these characteristics, and creating voice biometric profiles to isolate and differentiate target audio signals using machine learning on edge devices, reducing the need for filters and remote server processing.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If conventional filtering methods are used to extract target audio signals, then frequency bands can be separated, but the methods are insufficient in high noise environments and require computationally intensive post-production processing
Solution Approach 1:
The patent replaces conventional mechanical filtering methods with voice biometric identification systems that use acoustic characteristics (pitch, timbre, formants) and machine learning algorithms to identify and isolate target speakers, achieving better performance in noisy environments without requiring computationally intensive post-production processing
Solution Approach 2:
The system changes the approach from frequency-based filtering to parameter-based voice characterization by extracting and analyzing multiple acoustic parameters (pitch, loudness, timbre, formants, spectral peaks) to differentiate target speakers from background noise and other speakers
2Reliability
If post-production methods are used to filter unwanted audio signals, then signal separation can be achieved, but long latency periods occur and significant bandwidth is required
Solution Approach 1:
The system performs preliminary voice biometric profiling and speaker identification in real-time during audio capture, rather than waiting for post-production processing. This allows immediate separation of target speakers from background noise, eliminating long latency periods while maintaining effective signal separation
Solution Approach 2:
The system processes and identifies target speakers locally using on-device machine learning models, eliminating the need to transmit large amounts of audio data to remote servers for processing, thus reducing both bandwidth requirements and processing latency
3Reliability
If directional microphones are used to focus on individual speakers, then noise filtering can be achieved, but sophisticated planning is required and tracking is limited to controlled situations
Solution Approach 1:
The patent replaces mechanical directional microphone systems with software-based voice biometric identification that automatically identifies and tracks target speakers through acoustic characteristic analysis, eliminating the need for sophisticated planning and manual tracking while maintaining noise filtering capabilities in uncontrolled environments
Data Source
AI summary
The disclosure generally relates to systems and methods for using audio biometrics to isolate a target audio signal from a plurality of audio signals, the methods include determining, isolating, and differentiating metadata associated with the target audio signal and removing or suppressing audio metadata not associated with the target audio signal.


