Hybrid Speech Recognition Engine for Domain-Specific Accuracy

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing speech recognition systems that combine multiple engines often fail to improve accuracy due to high error rates when handling natural language inputs with specific commands, as they primarily select low-level outputs and lack evaluation of higher-level linguistic features.

Innovation Solution

A hybrid speech recognition method that generates candidate results using both general-purpose and domain-specific engines, aligns common words, and substitutes domain-specific words into general-purpose results, with a pairwise ranker to identify the highest ranked output, improving accuracy by leveraging linguistic significance and confidence scores.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If multiple speech recognition engines are combined by selecting low-level outputs, then the system can process diverse speech inputs, but the recognition accuracy deteriorates due to high error rates in domain-specific terms

Engineering Contradiction:
Improveability to process diverse speech inputsVSAvoidrecognition accuracy
Core Design Contradiction:
Adaptability or versatilityVSMeasurement precision

Solution Approach 1:

The patent introduces a hybrid speech recognition result generation engine as an intermediary component that bridges general-purpose and domain-specific speech recognition engines. This engine receives outputs from both engine types, aligns them based on common words, and generates hybrid results that combine strengths of both approaches, thereby maintaining versatility while improving accuracy for domain-specific terms

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent creates composite speech recognition results by combining outputs from different speech recognition engines. The hybrid results are formed by merging candidate results from general-purpose engines with those from domain-specific engines, creating a composite output that leverages the complementary strengths of each engine type to achieve both versatility and accuracy

Inventive Principle:
Principle #40Composite materials

2Device complexity

If prior art techniques select outputs from different speech recognition engines based on predetermined ranking, then processing is simplified, but accuracy improves insufficiently because errors from multiple engines compound

Engineering Contradiction:
Improveprocessing simplicityVSAvoidspeech recognition accuracy
Core Design Contradiction:
Device complexityVSMeasurement precision

Solution Approach 1:

The hybrid speech recognition result generation engine serves as an intermediary that processes outputs from multiple speech recognition engines. Instead of directly selecting from engine outputs, this intermediary aligns candidates based on common words and generates hybrid results, adding a processing layer that improves accuracy without excessive complexity increase

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent merges candidate speech recognition results from multiple engines by aligning them on common words and combining their outputs. This merging process allows the system to leverage multiple engine outputs simultaneously rather than selecting one, improving accuracy while maintaining reasonable processing complexity through systematic combination

Inventive Principle:
Principle #5Merging (Combining)

3Use of energy by moving object

If only low-level outputs such as acoustic model outputs are combined, then computational resources are conserved, but higher-level linguistic features cannot be evaluated

Engineering Contradiction:
Improvecomputational resource consumptionVSAvoidevaluation capability
Core Design Contradiction:
Use of energy by moving objectVSMeasurement precision

Solution Approach 1:

The patent segments the speech recognition output into different levels: low-level acoustic model outputs and higher-level linguistic features. By processing and aligning candidates at multiple levels rather than only at the acoustic level, the system evaluates linguistic significance while managing computational resources through hierarchical processing

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent extends the processing from a single dimension (low-level acoustic outputs) to multiple dimensions by incorporating higher-level linguistic features. The hybrid result generation engine operates in this expanded dimensional space, aligning candidates based on both acoustic and linguistic characteristics, thereby improving evaluation capability without excessive resource consumption

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Data Source

PatentUS9959861B2System and method for speech recognition
Publication Date: 2018.05.01 ROBERT BOSCH GMBH
  • US9959861B2 patent drawing
  • US9959861B2 patent drawing
  • US9959861B2 patent drawing

AI summary

A method for automated speech recognition includes generating first and second pluralities of candidate speech recognition results corresponding to audio input data using a first general-purpose speech recognition engine and a second domain-specific speech recognition engine, respectively. The method further includes generating a third plurality of candidate speech recognition result including a plurality of words included in one of the first plurality of speech recognition results and at least one word included in another one of the second plurality of speech recognition results, ranking the third plurality of candidate speech recognition results using a pairwise ranker to identify a highest ranked candidate speech recognition result, and operating the automated system using the highest ranked speech recognition result as an input from the user.