Adaptive Incremental Learning for Speech Recognition

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current speech recognition systems, particularly for low-resource languages and non-standard speech, face challenges in adapting to user-specific vocalizations and emerging expressions, requiring large amounts of annotated data and struggling with limited economical potential and adaptability to dysarthric speech.

Innovation Solution

An adaptive incremental learning approach for speech recognition systems that learns from demonstrations during usage, using Non-negative Matrix Factorization (NMF) and Maximum A Posteriori (MAP) estimation, allowing for online adaptation and reduced data storage requirements, enabling incremental learning and personalization in vocal user interfaces.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If traditional ASR systems are trained on large amounts of annotated speech data, then recognition accuracy is improved, but data storage requirements and system complexity increase

Engineering Contradiction:
Improverecognition accuracyVSAvoiddata storage requirements
Core Design Contradiction:
Measurement precisionVSQuantity of substance

Solution Approach 1:

The patent extracts only the essential acoustic features and semantic content from speech signals, rather than storing and processing complete annotated speech datasets. This is achieved through acoustic feature extraction and semantic frame generation that capture the core information needed for recognition while discarding redundant data.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The system performs self-training by automatically learning from user demonstrations without requiring manually annotated training data. The incremental learning algorithm enables the system to adapt to user-specific vocalizations and expressions autonomously, eliminating the need for external annotation resources.

Inventive Principle:
Principle #25Self-service

2Adaptability or versatility

If speaker-independent ASR systems are used with speaker adaptation procedures, then versatility is improved, but adaptation speed and performance on non-standard speech deteriorate

Engineering Contradiction:
Improvespeaker independenceVSAvoidadaptation speed
Core Design Contradiction:
Adaptability or versatilityVSProductivity

Solution Approach 1:

The system dynamically adapts its acoustic models through incremental learning, allowing the recognition system to evolve from speaker-independent to speaker-dependent performance over time. The model parameters are continuously updated based on user feedback and demonstrations, enabling fast adaptation to individual speech characteristics including non-standard speech patterns.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The system performs preliminary acoustic feature extraction and semantic frame generation that prepares the data structure for rapid incremental learning. By pre-processing speech into acoustic features and semantic frames, the system enables faster adaptation when user-specific data becomes available.

Inventive Principle:
Principle #10Preliminary action

3Measurement precision

If conventional ASR methods are adapted to dysarthric speech with substantial training material, then recognition accuracy is improved, but training data requirements increase

Engineering Contradiction:
Improverecognition accuracy for dysarthric speechVSAvoidtraining material volume
Core Design Contradiction:
Measurement precisionVSQuantity of substance

Solution Approach 1:

The incremental learning system enables the ASR to self-adapt to dysarthric speech patterns through user demonstrations rather than requiring extensive pre-collected training data. The system learns the specific acoustic characteristics and semantic usage patterns of dysarthric speech autonomously during normal interaction.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system changes its acoustic model parameters through incremental learning to accommodate dysarthric speech characteristics. By continuously updating the acoustic features and semantic frames based on user input, the system adapts to the specific pronunciation patterns, speed, and intonation characteristics of dysarthric speech.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS10573304B2Speech recognition system and method using an adaptive incremental learning approach
Publication Date: 2020.02.25 KATHOLIEKE UNIV LEUVEN
  • US10573304B2 patent drawing
  • US10573304B2 patent drawing
  • US10573304B2 patent drawing

AI summary

The present disclosure relates to speech recognition systems and methods using an adaptive incremental learning approach. More specifically, the present disclosure relates to adaptive incremental learning in a self-taught vocal user interface.