Methods for generating training corpora and methods for training speech recognition models
Patent Information
- Application Number
- TW114112600
- Authority / Receiving Office
- TW · TW
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2025-04-01
- Publication Date
- 2026-09-01
- Estimated Expiration
- 2045-03-31
Smart Images

Figure TWG2TB001908772_001 
Figure TWG2TB001908772_002 
Figure TWG2TB001908772_003
Abstract
Claims
1. A method for generating training corpus, comprising: The input of a special word is received through an input interface, and the special word has a corresponding voice data. A generation module enables a large language model to generate at least one sentence containing the special word received at the input interface; a speech synthesis module synthesizes at least one piece of speech data corresponding to the at least one sentence based on at least the special word, the speech data corresponding to the special word, and the at least one sentence containing the special word; and a corpus processing module converts the at least one sentence and the at least one piece of speech data corresponding to the at least one sentence into a training corpus that can be understood by a neural network; wherein the corpus processing module further establishes an index value in a dictionary based on at least the special word and the speech data corresponding to the special word, the index value being associated with the special word and its corresponding speech data.
2. The method for generating training corpus as described in claim 1, wherein, This special word is composed of Chinese, English, or a combination thereof.
3. The method for generating training corpus as described in claim 1 further includes: receiving input of the special word, wherein the input interface receives the speech data corresponding to the special word.
4. The method for generating training corpus as described in claim 3, wherein, The input interface includes a player configured to play the audio data corresponding to the special word.
5. The method for generating training corpus as described in claim 1, wherein, At least one sentence is generated based on the response of a large language model or a web crawling technique.
6. The method for generating training corpus as described in claim 1 further comprises: establishing an index value in a dictionary by the corpus processing module, at least based on the special word and an English pinyin or a phonetic transcription corresponding to the special word, the index value being associated with the special word and the corresponding speech data.
7. The method for generating training corpus as described in claim 1, wherein, This training corpus is suitable for training a speech recognition model.
8. The method for generating training corpus as described in claim 1 further comprises: expanding, via a copy-and-reproduce module, at least one speech data corresponding to the at least one sentence into a plurality of derived speech data, wherein the derived speech data correspond to the same text data but have different intonation features.
9. A method for training a speech recognition model, comprising: training a speech recognition model through a neural network based on the training corpus as described in any one of claims 1 to 8.
Citation Information
Patent Citations
Portable realtime dialect inter-translationing device and method thereof
CN1645363A
Techniques for machine language model creation
TW202113577A
Integrated Advertising System
US20110258049A1
Method and apparatus for generating speech training data
US20230037892A1