Joint Association Score for Speech Recognition and Semantic Classification
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing speech understanding systems face errors in spoken utterance classification due to imperfect automatic speech recognition, which are compounded in semantic classification, leading to inaccurate outputs.
Innovation Solution
A novel system integrates speech recognition and semantic classification by defining a joint association score that incorporates acoustic, language, and semantic model parameters to minimize errors, using discriminative training to revise model parameters and improve classification accuracy.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If traditional two-step speech recognition and semantic classification is used, then system complexity is reduced, but classification accuracy deteriorates due to error propagation
Solution Approach 1:
The patent merges speech recognition and semantic classification into a unified discriminative training framework. The joint association score combines acoustic scores from speech recognition with semantic class probabilities, allowing both tasks to be optimized simultaneously rather than sequentially. This integration prevents error propagation by jointly optimizing both recognition and classification objectives.
Solution Approach 2:
The patent introduces a joint association score that incorporates multiple parameters including acoustic model parameters, language model parameters, and semantic model parameters. By adjusting and optimizing these parameters together through discriminative training, the system achieves better overall performance than separate training approaches.
2Reliability
If separate training of language model and classification model is used, then training complexity is reduced, but overall error rate increases due to independent optimization
Solution Approach 1:
The patent combines language model training and classification model training into a single discriminative training process. The joint association score serves as a unified objective function that simultaneously optimizes both the language model parameters and the classification model parameters, ensuring coordinated optimization rather than independent training.
Solution Approach 2:
The discriminative training framework uses the joint association score to provide feedback for iteratively refining both language model and classification model parameters. The training process repeatedly adjusts parameters to maximize the joint association score, creating a feedback loop that jointly optimizes both components.
Data Source
AI summary
A novel system integrates speech recognition and semantic classification, so that acoustic scores in a speech recognizer that accepts spoken utterances may be taken into account when training both language models and semantic classification models. For example, a joint association score may be defined that is indicative of a correspondence of a semantic class and a word sequence for an acoustic signal. The joint association score may incorporate parameters such as weighting parameters for signal-to-class modeling of the acoustic signal, language model parameters and scores, and acoustic model parameters and scores. The parameters may be revised to raise the joint association score of a target word sequence with a target semantic class relative to the joint association score of a competitor word sequence with the target semantic class. The parameters may be designed so that the semantic classification errors in the training data are minimized.


