Speech Recognition Arbitration Logic Using Categorical Confidence Levels

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing speech recognition systems face challenges in accurately combining results from different algorithms due to varying confidence score methodologies, leading to potential discarding of low-confidence scores and reduced accuracy in speech recognition tasks.

Innovation Solution

A method that uses confidence levels categorized as high, medium, and low to combine speech recognition results from local and remote algorithms, allowing for the utilization of low-confidence scores and user confirmation or selection to determine the speech topic and slotted values, thereby improving task completion rates.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If multiple speech recognition algorithms are used with numerical confidence score normalization, then the comparison between different algorithms becomes standardized, but the accuracy of speech recognition results deteriorates due to loss of information from low-confidence scores

Engineering Contradiction:
Improvecomparison standardizationVSAvoidspeech recognition accuracy
Core Design Contradiction:
Ease of operationVSMeasurement precision

Solution Approach 1:

The patent segments the confidence assessment into discrete categorical levels (high, medium, low) rather than using continuous numerical scores. This segmentation preserves the qualitative distinctions in confidence levels while enabling standardized comparison across different speech recognition algorithms, resolving the contradiction between standardization and accuracy.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent changes the parameter representation from numerical confidence scores to categorical confidence levels. This parameter transformation maintains the comparative functionality across algorithms while preserving the integrity of low-confidence results, allowing them to be considered rather than discarded.

Inventive Principle:
Principle #35Parameter changes

2Reliability

If low confidence scores are discarded based on normalization thresholds, then the reliability of high-confidence results is improved, but the task completion rate deteriorates due to loss of potentially useful speech data

Engineering Contradiction:
Improveconfidence score reliabilityVSAvoidtask completion rate
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

Instead of discarding low-confidence scores as traditionally done, the patent inverts the approach by actively utilizing them through user confirmation mechanisms. Low-confidence results are presented to users for verification, transforming what was previously discarded data into a source for improving task completion while maintaining reliability through validation.

Inventive Principle:
Principle #13The other way round (Inversion)

Solution Approach 2:

The patent introduces feedback loops where users confirm or correct speech recognition results, particularly for low-confidence scores. This feedback mechanism allows the system to learn from user corrections and improve future recognition accuracy, thereby increasing task completion rates while maintaining high reliability through iterative validation.

Inventive Principle:
Principle #23Feedback

3Measurement precision

If remote speech recognition processing is used to improve accuracy, then the speech recognition performance is enhanced, but the operational cost deteriorates due to remote processing fees

Engineering Contradiction:
Improvespeech recognition accuracyVSAvoidprocessing cost
Core Design Contradiction:
Measurement precisionVSLoss of energy

Solution Approach 1:

The patent applies partial remote processing by sending only speech inputs that require additional verification or are below confidence thresholds to remote servers. High-confidence local recognitions are processed entirely locally, reducing remote processing fees while maintaining accuracy for uncertain cases through selective remote validation.

Inventive Principle:
Principle #16Partial or excessive action

Solution Approach 2:

The local speech recognition system performs self-service for high-confidence recognitions, handling them without remote server intervention. This self-sufficient local processing reduces dependency on expensive remote services while maintaining high accuracy for confident recognitions, reserving remote processing only for edge cases.

Inventive Principle:
Principle #25Self-service

4Measurement precision

If user confirmation is requested for all speech topics, then the accuracy of speech topic determination is improved, but the user interaction complexity deteriorates

Engineering Contradiction:
Improvespeech topic accuracyVSAvoiduser interaction complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent applies different levels of user interaction based on the local quality or confidence level of each speech recognition result. High-confidence results are accepted without user confirmation, while low-confidence results trigger confirmation requests. This differentiated approach improves accuracy for uncertain recognitions without unnecessarily complicating user interaction for confident recognitions.

Inventive Principle:
Principle #3Local quality

Data Source

PatentUS10679620B2Speech recognition arbitration logic
Publication Date: 2020.06.09 GM GLOBAL TECHNOLOGY OPERATIONS LLC
  • US10679620B2 patent drawing
  • US10679620B2 patent drawing
  • US10679620B2 patent drawing

AI summary

A method and associated system for recognizing speech using multiple speech recognition algorithms. The method includes receiving speech at a microphone installed in a vehicle, and determining results for the speech using a first algorithm, e.g., embedded locally at the vehicle. Speech results may also be received at the vehicle for the speech determined using a second algorithm, e.g., as determined by a remote facility. The results for both may include a determined speech topic and a determined speech slotted value, along with corresponding confidence levels for each. The method may further include using at least one of the determined first speech topic and the received second speech topic to determine the topic associated with the received speech, even when the first speech topic confidence level of the first speech topic, and the second speech topic confidence level of the second speech topic are both a low confidence level.