Speaker Identification via Acoustic Clustering for Lawful Interception

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current speaker identification methods are inefficient and prone to high error rates, especially in real-time or near-real-time environments with large voice collections, as they require extensive comparisons and degrade in performance with increasing numbers of voices, often failing to identify speakers not present in the collection.

Innovation Solution

The method involves grouping acoustic and non-acoustic models based on various parameters, reducing the number of models to compare a voice sample against, and combining scores to determine identification, allowing for efficient and accurate speaker identification in real-time.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If speaker identification is performed by comparing a voice sample against all stored voice representations in a large collection, then comprehensive identification coverage is achieved, but the identification time becomes excessively long and processing efficiency deteriorates

Engineering Contradiction:
Improveidentification coverageVSAvoididentification time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent segments the large voice collection into multiple clusters or groups based on acoustic characteristics. Instead of comparing a voice sample against all individual voice representations, the system first identifies the relevant cluster and then performs comparison only within that cluster, dramatically reducing computation time while maintaining identification accuracy.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent performs preliminary clustering of voice representations before the actual identification process. Voice samples are pre-processed and organized into clusters based on their acoustic features, so that when identification is needed, the system already has the data structured for efficient retrieval and comparison.

Inventive Principle:
Principle #10Preliminary action

2Quantity of substance

If the number of voices in the collection increases to improve coverage, then more speakers can be identified, but identification performance degrades and statistical significance decreases

Engineering Contradiction:
Improvevoice collection sizeVSAvoididentification performance
Core Design Contradiction:
Quantity of substanceVSMeasurement precision

Solution Approach 1:

By segmenting the large voice collection into smaller acoustic clusters, the system maintains high identification performance within each cluster even as the overall collection size grows. The clustering ensures that comparisons are made between acoustically similar voices, preserving statistical significance.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent changes the parameter of comparison from individual voice representations to cluster-level acoustic characteristics. This allows the system to handle larger collections by comparing against cluster centroids or representative vectors rather than every individual voice sample.

Inventive Principle:
Principle #35Parameter changes

3Measurement precision

If traditional speaker identification methods are used to ensure accurate matching, then identification accuracy is maintained, but the system cannot provide real-time or near-real-time results for large volumes of calls

Engineering Contradiction:
Improveidentification accuracyVSAvoidprocessing throughput
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The patent segments the identification process into two stages: cluster identification followed by within-cluster comparison. This segmentation enables real-time processing by eliminating the need to search through the entire large collection, while the final within-cluster comparison maintains high accuracy.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system performs preliminary organization of voice data into acoustic clusters before identification is needed. This pre-processing creates an efficient data structure that enables rapid retrieval and comparison during real-time operation without sacrificing accuracy.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS8219404B2Method and apparatus for recognizing a speaker in lawful interception systems
Publication Date: 2012.07.10 CYBERBIT
  • US8219404B2 patent drawing
  • US8219404B2 patent drawing
  • US8219404B2 patent drawing

AI summary

A method and apparatus for identifying a speaker within a captured audio signal from a collection of known speakers. The method and apparatus receive or generate voice representations for each known speakers and tag the representations according to meta data related to the known speaker or to the voice. The representations are grouped into one or more groups according to the indices. When a voice to be recognized is introduced, characteristics are determined according to which the groups are prioritized, so that the representations participating only in part of the groups are matched against the voice to be identified, thus reducing identification time and improving the statistical significance.