Speech Recognition Device Using Segmented Search Models for Disfluency Handling

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional speech recognition systems require high costs to register fillers, disfluencies, and non-speech sounds as words in the recognition dictionary for accurate recognition, limiting their efficiency.

Innovation Solution

A speech recognition device that includes a processor to calculate score vectors from speech signals and search a pre-registered search model to detect paths with likely acoustic scores, allowing for the recognition of phonetic units, fillers, disfluencies, and non-speech sounds with reduced costs by incorporating additional symbols representing these elements.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If fragments including fillers, disfluencies, and non-speech sounds are registered as words in advance in a search model, then recognition accuracy of these elements is improved, but system cost and complexity increase significantly

Engineering Contradiction:
Improverecognition accuracyVSAvoidsearch model complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent segments the search model into two distinct components: a phoneme search model for standard phonetic units and a filler search model for fillers, disfluencies, and non-speech sounds. This segmentation allows each component to be optimized independently, reducing the overall complexity of the search model while maintaining recognition accuracy for all elements.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent extracts filler, disfluency, and non-speech sound recognition from the traditional word-based search model and creates a separate filler search model. This extraction removes the burden of registering numerous filler variants in the main dictionary, simplifying the search model structure while preserving recognition capability for these elements.

Inventive Principle:
Principle #2Taking out (Extraction)

2Measurement precision

If a comprehensive search model including all filler and disfluency variants is created, then recognition accuracy is improved, but registration cost and time increase

Engineering Contradiction:
Improverecognition accuracyVSAvoidregistration time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The filler search model serves multiple functions simultaneously: it recognizes fillers, disfluencies, and non-speech sounds without requiring separate registration for each type. This multi-functionality reduces the time and effort needed for registration while maintaining comprehensive recognition accuracy across all filler-related elements.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The system performs preliminary classification by directing filler-related acoustic patterns to the specialized filler search model before attempting standard word matching. This preliminary action prevents the need to pre-register all possible filler variants in the main dictionary, significantly reducing registration time while preserving recognition accuracy.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS10553205B2Speech recognition device, speech recognition method, and computer program product
Publication Date: 2020.02.04 KK TOSHIBA
  • US10553205B2 patent drawing
  • US10553205B2 patent drawing
  • US10553205B2 patent drawing

AI summary

A speech recognition device includes one or more processors configured to calculate a score vector sequence on the basis of a speech signal, search a search model to detect a path following the input symbol from which a likely acoustic score in the score vector sequence is obtained, and output an output symbol allocated to the detected path. The symbol set includes a symbol representing a phonetic unit to be recognized, and an additional symbol representing at least one of a filler, a disfluency, and a non-speech sound. A search model includes an input symbol string arranged one or more input symbols, and paths to which output symbols are allocated. When the additional symbol is received as the input symbol from which the likely acoustic score is obtained, the processors start searching for a path associated with a new output symbol from a next score vector.