Speaker Category Identification for Speech Recognition Accuracy

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current speech recognition technologies face challenges in achieving accurate and efficient recognition across various speaker types and environments without requiring lengthy enrollment or adaptation processes, often resulting in reduced accuracy and increased computational complexity.

Innovation Solution

A computer-based method that utilizes a speech management server system with multiple speech recognition engines tuned to different speaker types, which associates received speech corpuses with selected speaker types and sends a speaker category identification code to applications, enabling them to select the most accurate engine for improved recognition.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If multiple speech recognition engines are used for different speaker types, then speech recognition accuracy is improved, but device complexity increases

Engineering Contradiction:
Improvespeech recognition accuracyVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent segments the speech recognition system into multiple specialized engines, each tuned to specific speaker types (e.g., male/female, different accents, age groups). Instead of using a single general-purpose engine, the system divides the recognition task across multiple specialized engines, allowing each to excel at recognizing particular speaker categories while maintaining overall system accuracy.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces a speaker type identification module as an intermediary that analyzes incoming speech to determine the speaker's category. This mediator then selects the appropriate specialized recognition engine for processing, avoiding the need to run all engines simultaneously and reducing overall system complexity while maintaining high accuracy for the identified speaker type.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Measurement precision

If speaker-specific enrollment and adaptation processes are implemented, then speech recognition accuracy is improved, but loss of time increases

Engineering Contradiction:
Improvespeech recognition accuracyVSAvoidenrollment time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent performs preliminary classification of speakers into categories during the enrollment process. Instead of requiring extensive adaptation for each individual speaker, the system pre-categorizes speakers based on acoustic features (gender, accent, age group) and assigns them to appropriate pre-trained recognition engines. This preliminary action significantly reduces the time needed for individual speaker adaptation while maintaining high recognition accuracy.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent changes the approach from individual speaker adaptation to category-based parameter selection. By identifying key acoustic parameters that define speaker categories and using these to select appropriate recognition models, the system achieves high accuracy without requiring time-consuming individual enrollment for each user. The parameter change shifts focus from fine-grained individual characteristics to broader categorical features.

Inventive Principle:
Principle #35Parameter changes

3Measurement precision

If multiple speech recognition engines are deployed, then speech recognition accuracy is improved, but use of energy increases

Engineering Contradiction:
Improvespeech recognition accuracyVSAvoidcomputational energy
Core Design Contradiction:
Measurement precisionVSUse of energy by moving object

Solution Approach 1:

The speaker type identification module acts as an energy-efficient intermediary that quickly categorizes incoming speech before routing it to the appropriate specialized engine. This mediation prevents the system from activating multiple heavy computational engines simultaneously, reducing energy consumption while maintaining high recognition accuracy by directing speech to the most suitable engine based on speaker characteristics.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

Instead of deploying and running all speech recognition engines for every input, the patent uses partial action by selecting and activating only the specific engine(s) needed for the identified speaker type. This selective activation significantly reduces computational energy usage while maintaining high accuracy, as the system performs just enough processing (using only the necessary engine) rather than excessive processing with all engines.

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS9305553B2Speech recognition accuracy improvement through speaker categories
Publication Date: 2016.04.05 RPX CORP
  • US9305553B2 patent drawing
  • US9305553B2 patent drawing
  • US9305553B2 patent drawing

AI summary

Methods, systems, and apparatus, including computer programs encoded on a computer storage medium, for speech recognition. In one aspect, a computer-based method includes receiving a speech corpus at a speech management server system that includes multiple speech recognition engines tuned to different speaker types; using the speech recognition engines to associate the received speech corpus with a selected one of multiple different speaker types; and sending a speaker category identification code that corresponds to the associated speaker type from the speech management server system over a network. The speaker category identification code can be used by any one of speech-interactive applications coupled to the network to select one of an appropriate one of multiple application-accessible speech recognition engines tuned to the different speaker types in response to an indication that a user accessing the application is associated with a particular one of the speaker category identification codes.