Speaker Identification Using Spatial Acoustic Features

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Traditional speaker identification methods in multiparty teleconferences, which rely on monaural telephony systems, face challenges when multiple participants share the same endpoint, leading to reduced accuracy and increased computational burden, especially in scenarios with overlapped speech or moving speakers.

Innovation Solution

A method and system for speaker identification using spatial acoustic features and location information across multiple channels, constructing models based on intra-channel and inter-channel shifted delta cepstrum features, and employing generalized linear discriminant sequence (GLDS) kernel functions for improved accuracy and reduced computational complexity.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of manufacture

If traditional monaural speaker identification methods are used, then the system is simple to implement, but the identification accuracy decreases when multiple participants share the same endpoint

Engineering Contradiction:
Improveimplementation simplicityVSAvoidspeaker identification accuracy
Core Design Contradiction:
Ease of manufactureVSMeasurement precision

Solution Approach 1:

The patent transitions from monaural (single-channel) to multi-channel audio processing by extracting spatial acoustic features from multiple channels. This dimensional expansion allows the system to differentiate speakers based on their spatial positions, thereby improving identification accuracy when multiple participants share the same endpoint while maintaining reasonable system complexity.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Solution Approach 2:

The patent changes the feature extraction parameters from traditional monaural features to spatial acoustic features that incorporate inter-channel time difference of arrival (IC-TDOA) and other spatial characteristics. This parameter transformation enables the system to exploit spatial information for more accurate speaker identification without overly complicating the implementation.

Inventive Principle:
Principle #35Parameter changes

2Measurement precision

If spatial acoustic features across multiple channels are extracted, then speaker identification accuracy is improved, but the computational burden increases

Engineering Contradiction:
Improvespeaker identification accuracyVSAvoidcomputational burden
Core Design Contradiction:
Measurement precisionVSPower

Solution Approach 1:

The patent segments the computational process into distinct stages: first extracting spatial acoustic features from multiple channels, then using IC-TDOA to determine speaker positions, and finally performing speaker identification based on these processed features. This segmentation allows for optimized computation at each stage, improving accuracy while managing computational burden through structured processing.

Inventive Principle:
Principle #1Segmentation

3Device complexity

If traditional monaural methods are used, then the computational complexity is low, but the system cannot effectively handle overlapped speech from multiple speakers

Engineering Contradiction:
Improvecomputational complexityVSAvoidrobustness to overlapped speech
Core Design Contradiction:
Device complexityVSReliability

Solution Approach 1:

The patent introduces spatial dimension information through multi-channel processing, extracting features such as inter-channel time difference of arrival (IC-TDOA) and spatial acoustic characteristics. This additional dimensional information enables the system to separate and identify overlapping speech from multiple speakers by exploiting their different spatial positions, thereby improving robustness without excessively increasing computational complexity.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Data Source

PatentUS9626970B2Speaker identification using spatial information
Publication Date: 2017.04.18 DOLBY LABORATORIES LICENSING CORP
  • US9626970B2 patent drawing
  • US9626970B2 patent drawing
  • US9626970B2 patent drawing

AI summary

Embodiments of the present invention relate to speaker identification using spatial information. A method of speaker identification for audio content being of a format based on multiple channels is disclosed. The method comprises extracting, from a first audio clip in the format, a plurality of spatial acoustic features across the multiple channels and location information, the first audio clip containing voices from a speaker, and constructing a first model for the speaker based on the spatial acoustic features and the location information, the first model indicating a characteristic of the voices from the speaker. The method further comprises identifying whether the audio content contains voices from the speaker based on the first model. Corresponding system and computer program product are also disclosed.