Speech Recognition Arbitration Logic Using Categorical Confidence Levels
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing speech recognition systems face challenges in accurately combining results from different algorithms due to varying confidence score methodologies, leading to potential discarding of low-confidence scores and reduced accuracy in speech recognition tasks.
Innovation Solution
A method that uses confidence levels categorized as high, medium, and low to combine speech recognition results from local and remote algorithms, allowing for the utilization of low-confidence scores and user confirmation or selection to determine the speech topic and slotted values, thereby improving task completion rates.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If multiple speech recognition algorithms are used with numerical confidence score normalization, then the comparison between different algorithms becomes standardized, but the accuracy of speech recognition results deteriorates due to loss of information from low-confidence scores
Solution Approach 1:
The patent segments the confidence assessment into discrete categorical levels (high, medium, low) rather than using continuous numerical scores. This segmentation preserves the qualitative distinctions in confidence levels while enabling standardized comparison across different speech recognition algorithms, resolving the contradiction between standardization and accuracy.
Solution Approach 2:
The patent changes the parameter representation from numerical confidence scores to categorical confidence levels. This parameter transformation maintains the comparative functionality across algorithms while preserving the integrity of low-confidence results, allowing them to be considered rather than discarded.
2Reliability
If low confidence scores are discarded based on normalization thresholds, then the reliability of high-confidence results is improved, but the task completion rate deteriorates due to loss of potentially useful speech data
Solution Approach 1:
Instead of discarding low-confidence scores as traditionally done, the patent inverts the approach by actively utilizing them through user confirmation mechanisms. Low-confidence results are presented to users for verification, transforming what was previously discarded data into a source for improving task completion while maintaining reliability through validation.
Solution Approach 2:
The patent introduces feedback loops where users confirm or correct speech recognition results, particularly for low-confidence scores. This feedback mechanism allows the system to learn from user corrections and improve future recognition accuracy, thereby increasing task completion rates while maintaining high reliability through iterative validation.
3Measurement precision
If remote speech recognition processing is used to improve accuracy, then the speech recognition performance is enhanced, but the operational cost deteriorates due to remote processing fees
Solution Approach 1:
The patent applies partial remote processing by sending only speech inputs that require additional verification or are below confidence thresholds to remote servers. High-confidence local recognitions are processed entirely locally, reducing remote processing fees while maintaining accuracy for uncertain cases through selective remote validation.
Solution Approach 2:
The local speech recognition system performs self-service for high-confidence recognitions, handling them without remote server intervention. This self-sufficient local processing reduces dependency on expensive remote services while maintaining high accuracy for confident recognitions, reserving remote processing only for edge cases.
4Measurement precision
If user confirmation is requested for all speech topics, then the accuracy of speech topic determination is improved, but the user interaction complexity deteriorates
Solution Approach 1:
The patent applies different levels of user interaction based on the local quality or confidence level of each speech recognition result. High-confidence results are accepted without user confirmation, while low-confidence results trigger confirmation requests. This differentiated approach improves accuracy for uncertain recognitions without unnecessarily complicating user interaction for confident recognitions.
Data Source
AI summary
A method and associated system for recognizing speech using multiple speech recognition algorithms. The method includes receiving speech at a microphone installed in a vehicle, and determining results for the speech using a first algorithm, e.g., embedded locally at the vehicle. Speech results may also be received at the vehicle for the speech determined using a second algorithm, e.g., as determined by a remote facility. The results for both may include a determined speech topic and a determined speech slotted value, along with corresponding confidence levels for each. The method may further include using at least one of the determined first speech topic and the received second speech topic to determine the topic associated with the received speech, even when the first speech topic confidence level of the first speech topic, and the second speech topic confidence level of the second speech topic are both a low confidence level.


