Domain-Specific Speech Recognition Model Training
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
End-to-end speech recognition technology requires a speech-transcript pair for training, limiting its ability to specialize in specific domains due to insufficient text data and difficulty in collecting domain-specific data, resulting in inferior performance for untrained domains.
Innovation Solution
A method and apparatus generate a domain-specific end-to-end speech recognition model using text data without a speech-transcript pair by collecting domain text, comparing it with a basic transcript database, generating a speech signal through synthesis, and training a neural network to create a specialized model, which can then be applied to an end-to-end speech recognizer.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If end-to-end speech recognition technology is used, then speech recognition performance is improved, but speech recognition performance in specific domains deteriorates due to insufficient text data
Solution Approach 1:
The patent applies preliminary action by pre-processing domain text data through text mining and comparison with basic transcript databases to identify domain-specific text before training. This preliminary preparation of domain text data enables the model to adapt to specific domains without requiring extensive domain-specific speech-transcript pairs, thus resolving the contradiction between general performance and domain adaptability
2Measurement precision
If speech-transcript pairs are collected for domain specialization, then domain-specific performance is improved, but data collection difficulty increases
Solution Approach 1:
The patent uses copying by replicating domain text data from easily accessible sources (documents, websites, books) and synthesizing corresponding speech data through text-to-speech conversion. This creates artificial speech-transcript pairs without requiring actual recording and transcription of domain-specific speech, dramatically reducing data collection difficulty while maintaining domain specialization capability
Solution Approach 2:
The patent introduces text data as an intermediary between domain knowledge and speech recognition training. Instead of directly collecting domain speech data, the system uses domain text as a mediator, which is then converted to speech through TTS and paired with the original text, creating training data through an intermediate textual representation
3Ease of manufacture
If domain text data is used for training, then data collection ease is improved, but speech recognition performance may deteriorate without proper speech signal
Solution Approach 1:
The patent uses text-to-speech synthesis as an intermediary process that converts domain text data into speech signals. This intermediary transformation allows the system to leverage easily collectible text data while generating the necessary speech signals for training, bridging the gap between text availability and speech recognition requirements
Solution Approach 2:
The patent replaces the mechanical process of recording and transcribing actual domain speech with an automated text-to-speech synthesis system. This substitution eliminates the need for physical speech data collection while maintaining the essential speech signal characteristics needed for training, thereby preserving recognition accuracy while improving data collection ease
Data Source
AI summary
Provided is an end-to-end speech recognition technology capable of improving speech recognition performance in a desired specific domain, which includes collecting domain text data be specialized and comparing the data with a basic transcript text DB to determine domain text that is not included in the basic transcript text DB and requires additional training and constructing a specialization target domain text DB. The end-to-end speech recognition technology generates a speech signal from the domain text of the specialization target domain text DB, and trains a speech recognition neural network with the generated speech signal to generate an end-to-end speech recognition model specialized for the domain to be specialized. The specialized speech recognition model may be applied to the end-to-end speech recognizer to perform the domain-specific end-to-end speech recognition.


