Speaker Identification Using Spatial Acoustic Features
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Traditional speaker identification methods in multiparty teleconferences, which rely on monaural telephony systems, face challenges when multiple participants share the same endpoint, leading to reduced accuracy and increased computational burden, especially in scenarios with overlapped speech or moving speakers.
Innovation Solution
A method and system for speaker identification using spatial acoustic features and location information across multiple channels, constructing models based on intra-channel and inter-channel shifted delta cepstrum features, and employing generalized linear discriminant sequence (GLDS) kernel functions for improved accuracy and reduced computational complexity.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of manufacture
If traditional monaural speaker identification methods are used, then the system is simple to implement, but the identification accuracy decreases when multiple participants share the same endpoint
Solution Approach 1:
The patent transitions from monaural (single-channel) to multi-channel audio processing by extracting spatial acoustic features from multiple channels. This dimensional expansion allows the system to differentiate speakers based on their spatial positions, thereby improving identification accuracy when multiple participants share the same endpoint while maintaining reasonable system complexity.
Solution Approach 2:
The patent changes the feature extraction parameters from traditional monaural features to spatial acoustic features that incorporate inter-channel time difference of arrival (IC-TDOA) and other spatial characteristics. This parameter transformation enables the system to exploit spatial information for more accurate speaker identification without overly complicating the implementation.
2Measurement precision
If spatial acoustic features across multiple channels are extracted, then speaker identification accuracy is improved, but the computational burden increases
Solution Approach 1:
The patent segments the computational process into distinct stages: first extracting spatial acoustic features from multiple channels, then using IC-TDOA to determine speaker positions, and finally performing speaker identification based on these processed features. This segmentation allows for optimized computation at each stage, improving accuracy while managing computational burden through structured processing.
3Device complexity
If traditional monaural methods are used, then the computational complexity is low, but the system cannot effectively handle overlapped speech from multiple speakers
Solution Approach 1:
The patent introduces spatial dimension information through multi-channel processing, extracting features such as inter-channel time difference of arrival (IC-TDOA) and spatial acoustic characteristics. This additional dimensional information enables the system to separate and identify overlapping speech from multiple speakers by exploiting their different spatial positions, thereby improving robustness without excessively increasing computational complexity.
Data Source
AI summary
Embodiments of the present invention relate to speaker identification using spatial information. A method of speaker identification for audio content being of a format based on multiple channels is disclosed. The method comprises extracting, from a first audio clip in the format, a plurality of spatial acoustic features across the multiple channels and location information, the first audio clip containing voices from a speaker, and constructing a first model for the speaker based on the spatial acoustic features and the location information, the first model indicating a characteristic of the voices from the speaker. The method further comprises identifying whether the audio content contains voices from the speaker based on the first model. Corresponding system and computer program product are also disclosed.


