Speaker Segmentation Using Clustered Speaker Models

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current speaker segmentation and recognition technologies face challenges in accurately identifying speakers in media files, especially when the number of potential speakers is unknown or when the speaker database contains a large number of models, leading to reduced accuracy and increased confusion in the matching process.

Innovation Solution

The proposed media content analytics system estimates a list of potential speakers using applications and social graphs, revises this list based on contextual information, and uses it to constrain speaker segmentation and recognition algorithms, thereby reducing the number of speaker models and improving accuracy.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If speaker segmentation and recognition is performed using a large speaker database, then comprehensive speaker coverage is achieved, but recognition accuracy decreases due to confusion among similar speaker models

Engineering Contradiction:
Improvespeaker coverageVSAvoidrecognition accuracy
Core Design Contradiction:
Adaptability or versatilityVSMeasurement precision

Solution Approach 1:

The patent segments the large speaker database into multiple clusters based on speaker similarity metrics. Each cluster represents a group of acoustically similar speakers, and the system processes each cluster separately. This segmentation reduces the confusion between similar speaker models while maintaining comprehensive coverage of all speakers in the database.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces cluster centroids as intermediary representations between the raw speaker models and the recognition process. Each cluster centroid serves as a representative summary of its cluster, acting as an intermediary that simplifies the comparison process and reduces computational complexity while preserving the essential characteristics of grouped speakers.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Reliability

If all speaker models are considered in the matching process, then complete speaker identification is possible, but processing time increases significantly

Engineering Contradiction:
Improveidentification completenessVSAvoidprocessing time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent performs preliminary clustering of speaker models before the actual recognition process. By organizing speakers into clusters and pre-computing cluster centroids in advance, the system prepares the data structure to enable faster querying during recognition, reducing the time required to process each speaker comparison.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent divides the comprehensive speaker database into smaller, manageable clusters. This segmentation allows the system to process only relevant clusters during recognition rather than evaluating all speaker models, significantly reducing processing time while maintaining complete identification capability through the clustered organization.

Inventive Principle:
Principle #1Segmentation

3Productivity

If the number of speaker models is reduced to improve processing speed, then processing time decreases, but recognition accuracy may be compromised

Engineering Contradiction:
Improveprocessing speedVSAvoidrecognition accuracy
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The patent creates cluster centroids as simplified copies or representations of groups of speaker models. These centroids capture the essential acoustic characteristics of each cluster, serving as efficient proxies that maintain recognition accuracy while reducing the computational burden of processing individual speaker models.

Inventive Principle:
Principle #26Copying

Data Source

PatentUS9058806B2Speaker segmentation and recognition based on list of speakers
Publication Date: 2015.06.16 CISCO TECHNOLOGY INC
  • US9058806B2 patent drawing
  • US9058806B2 patent drawing
  • US9058806B2 patent drawing

AI summary

A method is provided and includes estimating an approximate list of potential speakers in a file from one or more applications. The file (e.g., an audio file, video file, or any suitable combination thereof) includes a recording of a plurality of speakers. The method also includes segmenting the file according to the approximate list of potential speakers such that each segment corresponds to at least one speaker; and recognizing particular speakers in the file based on the approximate list of potential speakers.