Speaker Adaptation via Hit Quality Thresholding

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional speaker adaptation methods for automatic speech recognition systems are computationally intensive and time-consuming, making them infeasible for high-speed phonetic wordspotting systems, especially in environments with diverse speech characteristics and structured conversations.

Innovation Solution

A streamlined speaker adaptation system that processes media files from call center agents to identify putative instances of common terms, performs agent-specific acoustic model, pronunciation dictionary, and threshold adaptations, reducing the need for full transcription and leveraging structured conversations to improve recognition accuracy.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If conventional speaker adaptation methods are used, then recognition accuracy is improved, but computational complexity and processing time increase significantly

Engineering Contradiction:
Improverecognition accuracyVSAvoidcomputational complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent segments the speaker adaptation process by identifying only those speakers who need adaptation (those with low hit quality scores) rather than processing all speakers uniformly. This selective segmentation reduces computational complexity while maintaining recognition accuracy for those who need it most.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent applies local quality by customizing adaptation processing only for specific speakers with poor recognition performance rather than applying uniform adaptation to all speakers. This targeted approach optimizes computational resources by focusing on local areas of need.

Inventive Principle:
Principle #3Local quality

2Measurement precision

If conventional speaker adaptation methods are used, then recognition accuracy is improved, but processing time increases

Engineering Contradiction:
Improverecognition accuracyVSAvoidprocessing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent performs partial adaptation by applying speaker adaptation only to the extent necessary - specifically, only to speakers with low hit quality scores and only for the duration needed to improve their recognition performance. This partial action approach reduces processing time while maintaining necessary accuracy.

Inventive Principle:
Principle #16Partial or excessive action

Solution Approach 2:

The system automatically identifies which speakers need adaptation and applies adaptation only to them, eliminating the need for manual selection or exhaustive processing of all speakers. This self-service mechanism reduces processing time by autonomously focusing computational effort where it is most needed.

Inventive Principle:
Principle #25Self-service

3Measurement precision

If exhaustive transcription is performed, then speaker adaptation accuracy is improved, but computational load and time consumption increase

Engineering Contradiction:
Improvespeaker adaptation accuracyVSAvoidprocessing efficiency
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The patent extracts only the essential information needed for speaker adaptation - specifically, putative instances of common terms with associated hit quality scores - rather than performing exhaustive transcription of entire conversations. This extraction approach maintains sufficient adaptation accuracy while dramatically improving processing efficiency.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent applies partial transcription by processing only the portions of speech that contain common terms rather than transcribing entire conversations. This partial action provides sufficient data for speaker adaptation while maintaining high processing efficiency.

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS9001976B2Speaker adaptation
Publication Date: 2015.04.07 NEXIDIA
  • US9001976B2 patent drawing
  • US9001976B2 patent drawing
  • US9001976B2 patent drawing

AI summary

A method for speaker adaptation includes receiving a plurality of media files, each associated with a call center agent of a plurality of call center agents and receiving a plurality of terms. Speech processing is performed on at least some of the media files to identify putative instances of at least some of the plurality of terms. Each putative instance is associated with a hit quality that characterizes a quality of recognition of the corresponding term. One or more call center agents for performing speaker adaptation are determined, including identifying call center agents that are associated with at least one media file that includes one or more putative instances with a hit quality below a predetermined threshold. Speaker adaptation is performed for each identified call center agent based on the media files associated with the identified call center agent and the identified instances of the plurality of terms.