Domain-Specific Speech Recognition Model Training

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

End-to-end speech recognition technology requires a speech-transcript pair for training, limiting its ability to specialize in specific domains due to insufficient text data and difficulty in collecting domain-specific data, resulting in inferior performance for untrained domains.

Innovation Solution

A method and apparatus generate a domain-specific end-to-end speech recognition model using text data without a speech-transcript pair by collecting domain text, comparing it with a basic transcript database, generating a speech signal through synthesis, and training a neural network to create a specialized model, which can then be applied to an end-to-end speech recognizer.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If end-to-end speech recognition technology is used, then speech recognition performance is improved, but speech recognition performance in specific domains deteriorates due to insufficient text data

Engineering Contradiction:
Improvespeech recognition performanceVSAvoiddomain specialization capability
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The patent applies preliminary action by pre-processing domain text data through text mining and comparison with basic transcript databases to identify domain-specific text before training. This preliminary preparation of domain text data enables the model to adapt to specific domains without requiring extensive domain-specific speech-transcript pairs, thus resolving the contradiction between general performance and domain adaptability

Inventive Principle:
Principle #10Preliminary action

2Measurement precision

If speech-transcript pairs are collected for domain specialization, then domain-specific performance is improved, but data collection difficulty increases

Engineering Contradiction:
Improvedomain-specific speech recognition performanceVSAvoiddata collection ease
Core Design Contradiction:
Measurement precisionVSEase of manufacture

Solution Approach 1:

The patent uses copying by replicating domain text data from easily accessible sources (documents, websites, books) and synthesizing corresponding speech data through text-to-speech conversion. This creates artificial speech-transcript pairs without requiring actual recording and transcription of domain-specific speech, dramatically reducing data collection difficulty while maintaining domain specialization capability

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent introduces text data as an intermediary between domain knowledge and speech recognition training. Instead of directly collecting domain speech data, the system uses domain text as a mediator, which is then converted to speech through TTS and paired with the original text, creating training data through an intermediate textual representation

Inventive Principle:
Principle #24Intermediary (Mediator)

3Ease of manufacture

If domain text data is used for training, then data collection ease is improved, but speech recognition performance may deteriorate without proper speech signal

Engineering Contradiction:
Improvetext data collection easeVSAvoidspeech recognition accuracy
Core Design Contradiction:
Ease of manufactureVSMeasurement precision

Solution Approach 1:

The patent uses text-to-speech synthesis as an intermediary process that converts domain text data into speech signals. This intermediary transformation allows the system to leverage easily collectible text data while generating the necessary speech signals for training, bridging the gap between text availability and speech recognition requirements

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent replaces the mechanical process of recording and transcribing actual domain speech with an automated text-to-speech synthesis system. This substitution eliminates the need for physical speech data collection while maintaining the essential speech signal characteristics needed for training, thereby preserving recognition accuracy while improving data collection ease

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Data Source

PatentUS20230215419A1Method and apparatus for constructing domain-specific speech recognition model and end-to-end speech recognizer using the same
Publication Date: 2023.07.06 ELECTRONICS & TELECOMM RES INST
  • US20230215419A1 patent drawing
  • US20230215419A1 patent drawing
  • US20230215419A1 patent drawing

AI summary

Provided is an end-to-end speech recognition technology capable of improving speech recognition performance in a desired specific domain, which includes collecting domain text data be specialized and comparing the data with a basic transcript text DB to determine domain text that is not included in the basic transcript text DB and requires additional training and constructing a specialization target domain text DB. The end-to-end speech recognition technology generates a speech signal from the domain text of the specialization target domain text DB, and trains a speech recognition neural network with the generated speech signal to generate an end-to-end speech recognition model specialized for the domain to be specialized. The specialized speech recognition model may be applied to the end-to-end speech recognizer to perform the domain-specific end-to-end speech recognition.