Utterance Verification via Multi-Event Detection in Speech Recognition
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing speech recognition systems face challenges in accurately verifying utterances, particularly in natural language contexts, due to misrecognitions related to structure pronunciation, interjections, sound stretching, and other features, which current verification methods fail to adequately address.
Innovation Solution
An apparatus and method that utilize a noise processor, feature extractor, event detector, decoder, and utterance verifier to calculate confidence measurement values based on multi-event detection information, incorporating context-dependent acoustic models, n-gram language models, and support vector machine models to improve utterance verification.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If existing utterance verification methods are used, then the verification process can be performed, but misrecognitions in natural language contexts (structure pronunciation, interjections, sound stretching) cannot be accurately detected
Solution Approach 1:
The utterance verification process is segmented into multiple independent event detectors, each specialized for detecting specific natural language characteristics (interjections, sound stretching, structure pronunciation, hesitations, etc.). This segmentation allows each detector to focus on specific linguistic features, improving overall verification accuracy for natural language contexts.
Solution Approach 2:
The event detection module is designed with multi-functionality to detect various types of linguistic events simultaneously. The system integrates multiple detection capabilities (interjection detection, sound stretching detection, pronunciation structure detection) into a unified verification framework that can handle diverse natural language characteristics.
2Ease of operation
If general information extraction from speech recognition system is used, then verification can be performed, but characteristics for natural language speech recognition cannot be reflected
Solution Approach 1:
Event detectors are introduced as intermediary components between the speech recognition system and the utterance verification process. These detectors extract specific natural language features (interjections, sound stretching, etc.) from the speech signal and provide this information to the verification module, enabling precise detection of natural language characteristics without complicating the overall verification operation.
3Reliability
If multi-event detection information is used, then natural language speech recognition characteristics can be accurately verified, but the system complexity increases
Solution Approach 1:
Multiple event detection functions are merged into a unified event detection module that processes speech signals through multiple specialized detectors simultaneously. The detectors for interjections, sound stretching, pronunciation structures, and other linguistic features are combined in an integrated architecture, reducing overall system complexity while maintaining comprehensive natural language verification capabilities.
Data Source
AI summary
An apparatus and method for verifying an utterance based on multi-event detection information in a natural language speech recognition system. The apparatus includes a noise processor configured to process noise of an input speech signal, a feature extractor configured to extract features of speech data obtained through the noise processing, an event detector configured to detect events of the plurality of speech features occurring in the speech data using the noise-processed data and data of the extracted features, a decoder configured to perform speech recognition using a plurality of preset speech recognition models for the extracted feature data, and an utterance verifier configured to calculate confidence measurement values in units of words and sentences using information on the plurality of events detected by the event detector and a preset utterance verification model and perform utterance verification according to the calculated confidence measurement values.


