Voice Biometric Speaker Isolation for Real-Time Noisy Audio

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional methods for extracting target audio signals from multiple audio signals are inefficient in noisy environments, requiring computationally intensive post-production and significant bandwidth, and are impractical for real-time applications in crowded spaces.

Innovation Solution

The method involves extracting acoustic characteristics from audio signals, associating metadata with these characteristics, and creating voice biometric profiles to isolate and differentiate target audio signals using machine learning on edge devices, reducing the need for filters and remote server processing.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If conventional filtering methods are used to extract target audio signals, then frequency bands can be separated, but the methods are insufficient in high noise environments and require computationally intensive post-production processing

Engineering Contradiction:
Improveaudio signal extraction accuracyVSAvoidprocessing complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent replaces conventional mechanical filtering methods with voice biometric identification systems that use acoustic characteristics (pitch, timbre, formants) and machine learning algorithms to identify and isolate target speakers, achieving better performance in noisy environments without requiring computationally intensive post-production processing

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The system changes the approach from frequency-based filtering to parameter-based voice characterization by extracting and analyzing multiple acoustic parameters (pitch, loudness, timbre, formants, spectral peaks) to differentiate target speakers from background noise and other speakers

Inventive Principle:
Principle #35Parameter changes

2Reliability

If post-production methods are used to filter unwanted audio signals, then signal separation can be achieved, but long latency periods occur and significant bandwidth is required

Engineering Contradiction:
Improvesignal separation effectivenessVSAvoidprocessing latency
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The system performs preliminary voice biometric profiling and speaker identification in real-time during audio capture, rather than waiting for post-production processing. This allows immediate separation of target speakers from background noise, eliminating long latency periods while maintaining effective signal separation

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system processes and identifies target speakers locally using on-device machine learning models, eliminating the need to transmit large amounts of audio data to remote servers for processing, thus reducing both bandwidth requirements and processing latency

Inventive Principle:
Principle #25Self-service

3Reliability

If directional microphones are used to focus on individual speakers, then noise filtering can be achieved, but sophisticated planning is required and tracking is limited to controlled situations

Engineering Contradiction:
Improvenoise filtering capabilityVSAvoidsystem operability
Core Design Contradiction:
ReliabilityVSEase of operation

Solution Approach 1:

The patent replaces mechanical directional microphone systems with software-based voice biometric identification that automatically identifies and tracks target speakers through acoustic characteristic analysis, eliminating the need for sophisticated planning and manual tracking while maintaining noise filtering capabilities in uncontrolled environments

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Data Source

PatentUS20240203425A1Effective extraction of voice bio, acoustic and linguistic markers from an audio signal for speaker identification
Publication Date: 2024.06.20 YOBE INC
  • US20240203425A1 patent drawing
  • US20240203425A1 patent drawing
  • US20240203425A1 patent drawing

AI summary

The disclosure generally relates to systems and methods for using audio biometrics to isolate a target audio signal from a plurality of audio signals, the methods include determining, isolating, and differentiating metadata associated with the target audio signal and removing or suppressing audio metadata not associated with the target audio signal.