Speech Recognition Device Using Segmented Search Models for Disfluency Handling
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional speech recognition systems require high costs to register fillers, disfluencies, and non-speech sounds as words in the recognition dictionary for accurate recognition, limiting their efficiency.
Innovation Solution
A speech recognition device that includes a processor to calculate score vectors from speech signals and search a pre-registered search model to detect paths with likely acoustic scores, allowing for the recognition of phonetic units, fillers, disfluencies, and non-speech sounds with reduced costs by incorporating additional symbols representing these elements.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If fragments including fillers, disfluencies, and non-speech sounds are registered as words in advance in a search model, then recognition accuracy of these elements is improved, but system cost and complexity increase significantly
Solution Approach 1:
The patent segments the search model into two distinct components: a phoneme search model for standard phonetic units and a filler search model for fillers, disfluencies, and non-speech sounds. This segmentation allows each component to be optimized independently, reducing the overall complexity of the search model while maintaining recognition accuracy for all elements.
Solution Approach 2:
The patent extracts filler, disfluency, and non-speech sound recognition from the traditional word-based search model and creates a separate filler search model. This extraction removes the burden of registering numerous filler variants in the main dictionary, simplifying the search model structure while preserving recognition capability for these elements.
2Measurement precision
If a comprehensive search model including all filler and disfluency variants is created, then recognition accuracy is improved, but registration cost and time increase
Solution Approach 1:
The filler search model serves multiple functions simultaneously: it recognizes fillers, disfluencies, and non-speech sounds without requiring separate registration for each type. This multi-functionality reduces the time and effort needed for registration while maintaining comprehensive recognition accuracy across all filler-related elements.
Solution Approach 2:
The system performs preliminary classification by directing filler-related acoustic patterns to the specialized filler search model before attempting standard word matching. This preliminary action prevents the need to pre-register all possible filler variants in the main dictionary, significantly reducing registration time while preserving recognition accuracy.
Data Source
AI summary
A speech recognition device includes one or more processors configured to calculate a score vector sequence on the basis of a speech signal, search a search model to detect a path following the input symbol from which a likely acoustic score in the score vector sequence is obtained, and output an output symbol allocated to the detected path. The symbol set includes a symbol representing a phonetic unit to be recognized, and an additional symbol representing at least one of a filler, a disfluency, and a non-speech sound. A search model includes an input symbol string arranged one or more input symbols, and paths to which output symbols are allocated. When the additional symbol is received as the input symbol from which the likely acoustic score is obtained, the processors start searching for a path associated with a new output symbol from a next score vector.


