Speech Recognition System Using Unsupervised Phone Generation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional speech recognition systems rely on predetermined phone sets and pronunciation dictionaries, which are inadequate for handling natural continuous speech variations, such as phonetic contractions and omissions, and require transcribed data for acoustic modeling, limiting their effectiveness in recognizing distorted speech.

Innovation Solution

A speech recognition system and method that automatically generates phone sets and acoustic models using unsupervised deep learning from untranscribed speech data, allowing for the allocation of phone sequences to speech data without expert analysis, and generates a pronunciation dictionary based on speech patterns, enabling recognition of distorted speech.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If a predetermined phone set based on expert knowledge is used for acoustic modeling, then the acoustic model can be constructed with established phonetic categories, but it cannot handle natural continuous speech variations such as phonetic contractions and omissions

Engineering Contradiction:
Improveacoustic model constructionVSAvoidhandling speech variations
Core Design Contradiction:
ReliabilityVSAdaptability or versatility

Solution Approach 1:

The patent transforms the static, predetermined phone set into a dynamic system where phones are automatically generated and adapted from continuous speech data. The system dynamically adjusts phone boundaries and categories based on actual speech variations, allowing the acoustic model to adapt to phonetic contractions, omissions, and other natural speech phenomena rather than relying on fixed expert-defined categories

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The system enables self-service by automatically generating phone sets and pronunciation dictionaries from speech data without requiring expert annotation or manual preparation. The acoustic model construction process becomes self-sufficient, extracting phonetic patterns directly from continuous speech and building the necessary linguistic resources autonomously

Inventive Principle:
Principle #25Self-service

2Measurement precision

If transcribed data is used for acoustic model generation, then the acoustic model can be built with accurate phone boundaries, but it cannot be applied to untranscribed or distorted speech recognition

Engineering Contradiction:
Improvephone boundary accuracyVSAvoidrecognition of distorted speech
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The patent performs preliminary action by automatically generating phone sets and pronunciation dictionaries from available speech data before the actual recognition task. This preliminary processing creates adaptive linguistic resources that capture speech variations in advance, enabling the system to handle distorted and untranscribed speech without requiring pre-existing transcriptions or expert annotation for each specific case

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system inverts the conventional approach by not starting with expert-defined phone sets and transcribed data, but rather generating phone boundaries and categories directly from continuous speech data. This inversion allows the system to learn phonetic patterns empirically from actual speech variations rather than imposing predetermined categories, thereby improving both precision and adaptability

Inventive Principle:
Principle #13The other way round (Inversion)

3Reliability

If manual phone set generation by experts is used, then phonetic knowledge can be incorporated into the acoustic model, but the system complexity and time required for preparation increase

Engineering Contradiction:
Improvephonetic knowledge accuracyVSAvoidphone set preparation time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The system eliminates manual expert intervention by enabling automatic generation of phone sets and pronunciation dictionaries from speech data. This self-service approach maintains phonetic accuracy through data-driven pattern learning while completely eliminating the time-consuming manual preparation process, allowing the system to adapt to new speech variations without additional expert annotation time

Inventive Principle:
Principle #25Self-service

4Stability of the object's composition

If a pronunciation dictionary is created using linguistic pronunciation rules, then standard pronunciations can be enforced, but it cannot accommodate variable pronunciations and phonetic contractions

Engineering Contradiction:
Improvepronunciation consistencyVSAvoidhandling variable pronunciations
Core Design Contradiction:
Stability of the object's compositionVSAdaptability or versatility

Solution Approach 1:

The patent transforms the static pronunciation dictionary based on fixed linguistic rules into a dynamic system that adapts to actual speech variations. The system automatically learns and incorporates variable pronunciations, phonetic contractions, and omissions from continuous speech data, maintaining consistency through data-driven patterns rather than rigid rule enforcement, thereby accommodating natural speech variability

Inventive Principle:
Principle #15Dynamics

Data Source

PatentUS10249294B2Speech recognition system and method
Publication Date: 2019.04.02 ELECTRONICS & TELECOMM RES INST
  • US10249294B2 patent drawing
  • US10249294B2 patent drawing
  • US10249294B2 patent drawing

AI summary

A speech recognition method capable of automatic generation of phones according to the present invention includes: unsupervisedly learning a feature vector of speech data; generating a phone set by clustering acoustic features selected based on an unsupervised learning result; allocating a sequence of phones to the speech data on the basis of the generated phone set; and generating an acoustic model on the basis of the sequence of phones and the speech data to which the sequence of phones is allocated.