Hybrid Speech Recognition Engine for Domain-Specific Accuracy
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing speech recognition systems that combine multiple engines often fail to improve accuracy due to high error rates when handling natural language inputs with specific commands, as they primarily select low-level outputs and lack evaluation of higher-level linguistic features.
Innovation Solution
A hybrid speech recognition method that generates candidate results using both general-purpose and domain-specific engines, aligns common words, and substitutes domain-specific words into general-purpose results, with a pairwise ranker to identify the highest ranked output, improving accuracy by leveraging linguistic significance and confidence scores.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If multiple speech recognition engines are combined by selecting low-level outputs, then the system can process diverse speech inputs, but the recognition accuracy deteriorates due to high error rates in domain-specific terms
Solution Approach 1:
The patent introduces a hybrid speech recognition result generation engine as an intermediary component that bridges general-purpose and domain-specific speech recognition engines. This engine receives outputs from both engine types, aligns them based on common words, and generates hybrid results that combine strengths of both approaches, thereby maintaining versatility while improving accuracy for domain-specific terms
Solution Approach 2:
The patent creates composite speech recognition results by combining outputs from different speech recognition engines. The hybrid results are formed by merging candidate results from general-purpose engines with those from domain-specific engines, creating a composite output that leverages the complementary strengths of each engine type to achieve both versatility and accuracy
2Device complexity
If prior art techniques select outputs from different speech recognition engines based on predetermined ranking, then processing is simplified, but accuracy improves insufficiently because errors from multiple engines compound
Solution Approach 1:
The hybrid speech recognition result generation engine serves as an intermediary that processes outputs from multiple speech recognition engines. Instead of directly selecting from engine outputs, this intermediary aligns candidates based on common words and generates hybrid results, adding a processing layer that improves accuracy without excessive complexity increase
Solution Approach 2:
The patent merges candidate speech recognition results from multiple engines by aligning them on common words and combining their outputs. This merging process allows the system to leverage multiple engine outputs simultaneously rather than selecting one, improving accuracy while maintaining reasonable processing complexity through systematic combination
3Use of energy by moving object
If only low-level outputs such as acoustic model outputs are combined, then computational resources are conserved, but higher-level linguistic features cannot be evaluated
Solution Approach 1:
The patent segments the speech recognition output into different levels: low-level acoustic model outputs and higher-level linguistic features. By processing and aligning candidates at multiple levels rather than only at the acoustic level, the system evaluates linguistic significance while managing computational resources through hierarchical processing
Solution Approach 2:
The patent extends the processing from a single dimension (low-level acoustic outputs) to multiple dimensions by incorporating higher-level linguistic features. The hybrid result generation engine operates in this expanded dimensional space, aligning candidates based on both acoustic and linguistic characteristics, thereby improving evaluation capability without excessive resource consumption
Data Source
AI summary
A method for automated speech recognition includes generating first and second pluralities of candidate speech recognition results corresponding to audio input data using a first general-purpose speech recognition engine and a second domain-specific speech recognition engine, respectively. The method further includes generating a third plurality of candidate speech recognition result including a plurality of words included in one of the first plurality of speech recognition results and at least one word included in another one of the second plurality of speech recognition results, ranking the third plurality of candidate speech recognition results using a pairwise ranker to identify a highest ranked candidate speech recognition result, and operating the automated system using the highest ranked speech recognition result as an input from the user.


