User Audio Profile Evaluation for Speech Recognition

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current continuous speech recognition systems lack tools to evaluate and improve user audio profiles, leading to poor performance and user abandonment due to phonetic errors and other issues, as they do not provide insights into the quality of audio profiles or offer effective remediation measures.

Innovation Solution

The system evaluates user audio profiles by comparing phoneme sequences from training text and audio pairs, generating phoneme accuracy statistics, and providing remedial measures such as additional training, microphone adjustments, and speech coaching to enhance recognition accuracy.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If user audio profiles are trained using training text and audio pairs, then the speech recognition system can generate phoneme sequences from audio, but there are no good tools to evaluate how good the profile is for the user

Engineering Contradiction:
Improveprofile quality evaluationVSAvoidevaluation system complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The system generates phoneme accuracy statistics by comparing audio-derived phoneme sequences with text-derived phoneme sequences, providing feedback on profile quality. This enables evaluation of how well the audio profile captures the user's speech characteristics without requiring complex external evaluation tools.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The speech recognition system evaluates its own audio profiles using its existing phoneme recognition capabilities. By comparing the phoneme sequences generated from audio against those generated from text, the system performs self-evaluation, eliminating the need for separate complex evaluation infrastructure.

Inventive Principle:
Principle #25Self-service

2Ease of operation

If continuous speech recognition systems use Hidden Markov Models to determine phoneme matching, then they can process natural language speech, but they lack tools to identify causes of poor performing profiles and offer remediation

Engineering Contradiction:
Improveremediation process easeVSAvoiddiagnostic information
Core Design Contradiction:
Ease of operationVSLoss of information

Solution Approach 1:

The system segments the evaluation into phoneme-level analysis, generating statistics for individual phonemes. This detailed segmentation identifies specific phonemes with low accuracy, providing diagnostic information about which sounds are problematic while maintaining ease of operation through automated analysis.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system provides feedback on profile quality through phoneme accuracy statistics and identifies specific causes of poor performance. This enables targeted remediation actions such as additional training with specific phonemes or adjusting recording conditions, making the remediation process both easy to operate and information-rich.

Inventive Principle:
Principle #23Feedback

3Measurement precision

If the system compares audio phoneme sequences with text phoneme sequences to determine accuracy, then it can identify phoneme recognition errors, but it requires processing and storing multiple sequence representations

Engineering Contradiction:
Improvephoneme accuracy measurementVSAvoiddata processing volume
Core Design Contradiction:
Measurement precisionVSQuantity of substance

Solution Approach 1:

The system extracts only the essential comparison needed for evaluation: phoneme sequences from audio and phoneme sequences from text. By focusing on this specific extraction and comparison, it achieves precise phoneme accuracy measurement without unnecessarily processing and storing all intermediate representations.

Inventive Principle:
Principle #2Taking out (Extraction)

Data Source

PatentUS10395640B1Systems and methods evaluating user audio profiles for continuous speech recognition
Publication Date: 2019.08.27 NVOQ INC
  • US10395640B1 patent drawing
  • US10395640B1 patent drawing
  • US10395640B1 patent drawing

AI summary

To attain the advantages and in accordance with the purpose of the technology of the present application, apparatuses, systems, and methods to evaluate a user audio profile are provided. The evaluation of a user audio profile allows for identification of potential causes of poor performing user audio profiles and potential types of remediation to increase the performance of the user audio profile.