Utterance Verification via Multi-Event Detection in Speech Recognition

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing speech recognition systems face challenges in accurately verifying utterances, particularly in natural language contexts, due to misrecognitions related to structure pronunciation, interjections, sound stretching, and other features, which current verification methods fail to adequately address.

Innovation Solution

An apparatus and method that utilize a noise processor, feature extractor, event detector, decoder, and utterance verifier to calculate confidence measurement values based on multi-event detection information, incorporating context-dependent acoustic models, n-gram language models, and support vector machine models to improve utterance verification.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If existing utterance verification methods are used, then the verification process can be performed, but misrecognitions in natural language contexts (structure pronunciation, interjections, sound stretching) cannot be accurately detected

Engineering Contradiction:
Improveutterance verification accuracyVSAvoidnatural language characteristic reflection
Core Design Contradiction:
ReliabilityVSAdaptability or versatility

Solution Approach 1:

The utterance verification process is segmented into multiple independent event detectors, each specialized for detecting specific natural language characteristics (interjections, sound stretching, structure pronunciation, hesitations, etc.). This segmentation allows each detector to focus on specific linguistic features, improving overall verification accuracy for natural language contexts.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The event detection module is designed with multi-functionality to detect various types of linguistic events simultaneously. The system integrates multiple detection capabilities (interjection detection, sound stretching detection, pronunciation structure detection) into a unified verification framework that can handle diverse natural language characteristics.

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Ease of operation

If general information extraction from speech recognition system is used, then verification can be performed, but characteristics for natural language speech recognition cannot be reflected

Engineering Contradiction:
Improveverification process simplicityVSAvoidnatural language feature detection precision
Core Design Contradiction:
Ease of operationVSMeasurement precision

Solution Approach 1:

Event detectors are introduced as intermediary components between the speech recognition system and the utterance verification process. These detectors extract specific natural language features (interjections, sound stretching, etc.) from the speech signal and provide this information to the verification module, enabling precise detection of natural language characteristics without complicating the overall verification operation.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Reliability

If multi-event detection information is used, then natural language speech recognition characteristics can be accurately verified, but the system complexity increases

Engineering Contradiction:
Improvenatural language utterance verification accuracyVSAvoidevent detection system complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

Multiple event detection functions are merged into a unified event detection module that processes speech signals through multiple specialized detectors simultaneously. The detectors for interjections, sound stretching, pronunciation structures, and other linguistic features are combined in an integrated architecture, reducing overall system complexity while maintaining comprehensive natural language verification capabilities.

Inventive Principle:
Principle #5Merging (Combining)

Data Source

PatentUS9799350B2Apparatus and method for verifying utterance in speech recognition system
Publication Date: 2017.10.24 ELECTRONICS & TELECOMM RES INST
  • US9799350B2 patent drawing
  • US9799350B2 patent drawing
  • US9799350B2 patent drawing

AI summary

An apparatus and method for verifying an utterance based on multi-event detection information in a natural language speech recognition system. The apparatus includes a noise processor configured to process noise of an input speech signal, a feature extractor configured to extract features of speech data obtained through the noise processing, an event detector configured to detect events of the plurality of speech features occurring in the speech data using the noise-processed data and data of the extracted features, a decoder configured to perform speech recognition using a plurality of preset speech recognition models for the extracted feature data, and an utterance verifier configured to calculate confidence measurement values in units of words and sentences using information on the plurality of events detected by the event detector and a preset utterance verification model and perform utterance verification according to the calculated confidence measurement values.