Phrase Spotting Confidence via Multi-Feature Classification
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing speech recognition systems face challenges in providing reliable confidence scores for phrase spotting, as posterior probability estimation (PPE) scores can be unreliable or inaccurate, leading to inconsistent phrase spotting results.
Innovation Solution
A system and method that enhance phrase spotting confidence scores by using a classifying model trained with posterior probability features and additional phrase spotting features, selected from decoding process, context-based, and edit probability scores, to improve the identification of phonetic representations in speech.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If posterior probability estimation (PPE) is used to provide confidence scores, then phrase spotting can be performed, but the confidence scores are unreliable or inaccurate
Solution Approach 1:
The patent combines multiple feature types including PPE scores, decoding process features, context-based features, and edit probability scores into a unified feature set that feeds into a trained classifier. This merging of diverse feature sources resolves the contradiction by aggregating multiple weak signals into a more reliable confidence measurement, where the classifier learns optimal weightings of each feature type to produce accurate and reliable phrase spotting results.
Solution Approach 2:
The patent transforms the raw PPE score and other features through a trained classification model that learns optimal parameter transformations during training. The classifier adjusts parameters such as feature weightings, normalization factors, and decision thresholds to convert unreliable individual scores into reliable confidence measurements, effectively changing the parameter space to achieve both accuracy and reliability.
2Reliability
If multiple feature types are incorporated to improve confidence scores, then reliability increases, but system complexity increases
Solution Approach 1:
The trained classifier acts as an intermediary component that receives multiple complex feature types (PPE scores, decoding features, context features, edit probabilities) and transforms them into a simplified confidence score. This intermediary absorbs the complexity of processing multiple feature types while presenting a single reliable output, resolving the contradiction by localizing complexity within the classifier module rather than propagating it throughout the entire system.
Solution Approach 2:
The system performs preliminary extraction and preparation of multiple feature types before they are fed into the classifier. By pre-computing decoding process features, context-based features, and edit probability scores in advance, the system organizes the complexity work beforehand, allowing the classifier to focus on the final integration and decision-making, thus managing overall system complexity while maintaining reliability.
Data Source
AI summary
Embodiments of the disclosed subject matter include a system and method for improving a phrase spotting score. The method may include providing a test speech and a transcription thereof, obtaining an input phrases in a textual form for spotting in the provided test speech, generating a phonetic transcription for the input phrase, and applying a classifying model to the phonetic transcription of the test speech according to a posterior probability feature and to a set of phrase spotting features related to the input phrase, the phrase spotting features selected from features extracted from a decoding process of the phrase, context-based features, and/or a combination thereof, thereby spotting the given phrase with a confidence score. The classifying model is priorly trained for spotting the input phrase according to the posterior probability feature and to the plurality of phrase spotting features.


