Acoustic Model Adaptation for Speech Recognition

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional speech recognition systems require extensive user voice data for user-dependent acoustic models, leading to inefficiencies in performance improvement and inconvenient online speech recognition due to the need for continuous data collection and speaker authentication.

Innovation Solution

A speech recognition system using incremental device-based acoustic model adaptation, which automatically generates a device key for user devices, allowing for continuous performance improvement without user involvement by categorizing voice data and adapting acoustic models using a multi-model tree structure, including device-independent and device-dependent models, and incrementally updating models based on reliability thresholds.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If extensive user voice data is collected to create user-dependent acoustic models, then speech recognition performance is improved, but data collection time and system complexity increase

Engineering Contradiction:
Improvespeech recognition performanceVSAvoiddata collection time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The system performs preliminary acoustic model adaptation during the initial setup phase by collecting voice data from the user. This preliminary adaptation creates a user-dependent acoustic model that is then applied to subsequent recognition tasks, eliminating the need for continuous data collection and improving recognition performance without requiring extensive ongoing data gathering.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system automatically performs speaker adaptation using voice data collected during normal operation without requiring explicit user involvement or manual configuration. The adaptation process occurs autonomously in the background, allowing the system to improve its acoustic model using its own operational data without user intervention.

Inventive Principle:
Principle #25Self-service

2Reliability

If speaker recognition and authentication are performed to save data for categorizing individual speakers, then user-dependent performance is improved, but ease of operation deteriorates

Engineering Contradiction:
Improveuser-dependent performanceVSAvoidconvenience of online speech recognition
Core Design Contradiction:
ReliabilityVSEase of operation

Solution Approach 1:

The system automatically performs speaker adaptation using voice data collected during normal operation without requiring explicit user involvement or manual configuration. The adaptation process occurs autonomously in the background, allowing the system to improve its acoustic model using its own operational data without user intervention.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system performs preliminary acoustic model adaptation during the initial setup phase by collecting voice data from the user. This preliminary adaptation creates a user-dependent acoustic model that is then applied to subsequent recognition tasks, eliminating the need for continuous data collection and improving recognition performance without requiring extensive ongoing data gathering.

Inventive Principle:
Principle #10Preliminary action

3Reliability

If more voice data is collected continuously, then speech recognition performance is improved, but device complexity and data management requirements increase

Engineering Contradiction:
Improvespeech recognition performanceVSAvoiddata management complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The system segments the acoustic model into a base acoustic model and user-specific adaptation components. This segmentation allows the system to manage data more efficiently by separating universal acoustic information from user-specific information, reducing the complexity of data management while maintaining improved recognition performance through targeted data collection and processing.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS9601112B2Speech recognition system and method using incremental device-based acoustic model adaptation
Publication Date: 2017.03.21 ELECTRONICS & TELECOMM RES INST
  • US9601112B2 patent drawing
  • US9601112B2 patent drawing
  • US9601112B2 patent drawing

AI summary

An embodiment of the present invention relates to a speech recognition system and method using incremental device-based acoustic model adaptation. The speech recognition system comprises a model selection module selecting an acoustic model of multi-model tree by verifying and categorizing a device key transmitted from a user device; a model management module generating and incrementally adapting multi-model tree by categorizing voice data based on a user device; and a speech recognition module performing speech recognition by receiving the acoustic model selected from the model selection module and transmitting data of which reliability exceeds a predetermined threshold value to the model management module.