Speaker-Cluster ASR for Dialect-Aware Recognition

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Traditional Automated Speech Recognition (ASR) systems, particularly speaker independent systems, struggle to achieve high accuracy without user training and adaptation, as they lack nuanced understanding of individual pronunciation variations.

Innovation Solution

Employing speaker clustering technology to analyze large audio corpora and corresponding transcripts to identify and train for multiple speaker types, allowing for dynamic mapping and transcription of new users based on their dialects, thereby enhancing recognition accuracy without explicit user training.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If speaker independent ASR is used, then ease of operation is improved (no user training required), but measurement precision deteriorates (lower recognition accuracy)

Engineering Contradiction:
Improveease of operationVSAvoidrecognition accuracy
Core Design Contradiction:
Ease of operationVSMeasurement precision

Solution Approach 1:

The patent segments the speaker population into multiple clusters based on acoustic characteristics. Instead of treating all speakers uniformly (speaker independent) or requiring individual training (speaker dependent), the system divides speakers into groups and creates specialized ASR models for each cluster, achieving high accuracy without user training.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system dynamically changes the speaker normalization parameters based on the identified speaker cluster. Different clusters have different acoustic characteristics, and the ASR system adjusts its parameters accordingly to optimize recognition accuracy for each group while maintaining ease of operation.

Inventive Principle:
Principle #35Parameter changes

2Measurement precision

If speaker dependent ASR is used, then measurement precision is improved (higher recognition accuracy), but device complexity worsens (requires user training process)

Engineering Contradiction:
Improverecognition accuracyVSAvoiddevice complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent performs preliminary clustering of speakers into distinct groups before the actual ASR process. By pre-organizing speakers into clusters with similar characteristics, the system eliminates the need for individual user training while maintaining high recognition accuracy, thus reducing device complexity.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

Instead of creating individual ASR models for each speaker (which would be complex), the system creates representative models for each speaker cluster. These cluster-level models capture the essential characteristics of multiple speakers, providing high accuracy without requiring individual training data for each user.

Inventive Principle:
Principle #26Copying

3Ease of operation

If traditional speaker independent ASR with normalization is used, then ease of operation is maintained (no user intervention), but manufacturing precision worsens (inability to capture individual pronunciation nuances)

Engineering Contradiction:
Improveease of operationVSAvoidpronunciation recognition precision
Core Design Contradiction:
Ease of operationVSManufacturing precision

Solution Approach 1:

The patent applies different processing qualities to different speaker groups. Instead of using a single normalization approach for all speakers, the system identifies local acoustic characteristics of each speaker cluster and applies specialized processing tailored to each group's pronunciation patterns, thereby capturing nuanced variations without user intervention.

Inventive Principle:
Principle #3Local quality

Data Source

PatentUS8600750B2Speaker-cluster dependent speaker recognition (speaker-type automated speech recognition)
Publication Date: 2013.12.03 CISCO TECHNOLOGY INC
  • US8600750B2 patent drawing
  • US8600750B2 patent drawing
  • US8600750B2 patent drawing

AI summary

In an example embodiment, there is disclosed herein an automatic speech recognition (ASR) system that employs speaker clustering (or speaker type) for transcribing audio. A large corpus of audio with corresponding transcripts is analyzed to determine a plurality of speaker types (e.g., dialects). The ASR system is trained for each speaker type. Upon encountering a new user, the ASR system attempts to map the user to a speaker type. After the new user is mapped to a speaker type, the ASR employs the speaker type for transcribing audio from the new user.