Data amplification method, system and device based on LLM-TTS and storage medium
Through the LLM-TTS-based data amplification method, high-quality speech data are generated and screened, and the problem of low recognition performance of speech recognition systems in scarce data environments is solved, and better speech recognition effect and adaptability are achieved.
Patent Information
- Application Number
- CN202510104378.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-22
- Publication Date
- 2025-05-27
AI Technical Summary
The prior art is difficult to generate high-quality and diverse voice data in a scarce data environment, resulting in poor recognition performance of voice recognition systems in dialect and small language scenarios.
Using the LLM-TTS-based data amplification method, the LLM-based TTS model is trained by collecting and preprocessing the speech data set, and the model is used to generate speech data that conforms to the characteristics of the target dialect and minor languages. Then, the generated data is filtered through the Filter-ASR model to ensure the data quality, and finally train the ASR model using the filtered data.
Effectively generate and filter amplified data to ensure that the speech recognition system can better adapt to the characteristics of dialects and minor languages in a data scarce environment, and improve the system's recognition performance and adaptability.
Smart Images

Figure CN120048256A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of language recognition technology, and particularly to a data augmentation method, system, device, and storage medium based on LLM-TTS. Background Art
[0002] Speech recognition technology has wide applications in fields such as speech human-computer interaction, language translation, and speech assistants. However, with the acceleration of the globalization process, speech recognition technology needs to handle more diverse language environments, especially in scenarios involving scarce dialects and minority languages. The uniqueness of these languages and dialects poses higher requirements for the generalization ability and adaptability of speech recognition systems. However, due to the scarcity of data, the development and optimization process of speech recognition systems face major challenges.
[0003] In the prior art, voice conversion technology (VC) and transfer learning methods based on self-supervised learning have been widely applied to speech data augmentation and speech recognition tasks. Voice conversion technology can generate speech data by converting the input audio into the target timbre to improve the model training effect under scarce data conditions. The strategy of combining self-supervised learning and transfer learning can improve the performance of the model in scarce data scenarios through pre-training on a large amount of unlabeled speech data and fine-tuning on a small amount of labeled data. These methods have alleviated the scarce data problem to a certain extent and provided feasible solutions for dialect and minority language speech recognition.
[0004] However, the prior art still has significant problems. Voice conversion technology has a high dependence on the quality and diversity of training data, and the generated audio content is limited. It can only achieve timbre conversion and cannot change the audio content. The method of combining self-supervised learning and transfer learning is easily affected by the differences between the source task and the target task, resulting in the model performance not meeting expectations. At the same time, the dependence of the fine-tuning process on scarce data may lead to underfitting or overfitting problems, limiting the generalization ability of the system. Therefore, how to generate high-quality and diverse speech data in a scarce data environment and improve the performance of speech recognition systems has become a technical problem to be solved urgently. Summary of the Invention
[0005] This application provides a data augmentation method, system, device, and storage medium based on LLM-TTS. By effectively generating and screening augmented data, it ensures that in a data-scarce environment, the speech recognition system can better adapt to the characteristics of dialects and minority languages, thus solving the problem of low recognition performance of speech recognition systems in scarce languages and dialects in the prior art. This application provides the following technical solutions:
[0006] In the first aspect, this application provides a data augmentation method based on LLM-TTS, and the method includes:
[0007] Collect available speech datasets and preprocess them;
[0008] Train a preset LLM-based TTS model based on the preprocessed speech datasets;
[0009] Use the trained LLM-based TTS model for data augmentation;
[0010] Screen the augmented data;
[0011] Train an ASR automatic speech recognition model using the screened augmented data.
[0012] In a specific feasible implementation, the collecting available speech datasets and preprocessing them includes:
[0013] The speech datasets include mainstream languages, limited target dialects, and minority language data.
[0014] In a specific feasible implementation, the training of a preset LLM-based TTS model based on the preprocessed speech datasets includes:
[0015] Adjust the sampling strategy in each training cycle to make the data volume of each language basically balanced;
[0016] Introduce a chain-of-thought prompting strategy. Before generating each speech token, predict the prosodic features at the syllable level through the model, and use the predicted syllable duration information to guide the self-attention mechanism, so that the model can focus on the parts related to the current phoneme and prosodic features when generating speech.
[0017] In a specific feasible implementation, the training of a preset LLM-based TTS model based on the preprocessed speech datasets further includes:
[0018] The basic architecture used for training the LLM-based TTS model is a Transformer-based decoder-only structure.
[0019] In a specific feasible implementation, the using the trained LLM-based TTS model for data augmentation includes:
[0020] Set specific generation conditions for the target dialects and minority languages;
[0021] During the generation process, use scarce minority language or dialect data as input prompts to guide the LLM-based TTS model to generate speech samples that conform to the characteristics of the target dialects and minority languages through the input prompts.
[0022] In a specific feasible implementation, the screening of the amplified data includes:
[0023] Train a Filter-ASR model and use the audio data generated by the LLM-based TTS model for training; after the training is completed, use the Filter-ASR model to decode the training audio to obtain the decoding result;
[0024] Compare the decoding result with the real text, calculate the word error rate, and according to the word error rate, screen out and remove the audio data with a higher word error rate. After the initial screening, gradually reduce the number of parameters of the Filter-ASR model, and continue to train the Filter-ASR model with the high-quality audio data screened in the previous step;
[0025] After the training is completed, repeat the decoding process, calculate the word error rate, and remove those audio with a higher word error rate again, and keep iterating until the final WER value reaches the preset standard.
[0026] In a specific feasible implementation, the training of the ASR automatic speech recognition model using the screened amplified data includes:
[0027] Use the high-quality speech data set screened by the Filter-ASR model as the training data and input it into the preset ASR model for training;
[0028] After the training is completed, use the test sets of dialects and minority languages to evaluate the ASR model;
[0029] If the performance of the ASR model fails to meet the expectations in the test sets of certain dialects or minority languages, adjust and optimize the generation strategy according to the evaluation results.
[0030] In a second aspect, the present application provides a data augmentation system based on LLM-TTS, adopting the following technical solutions:
[0031] A data augmentation system based on LLM-TTS, comprising:
[0032] A data collection module, configured to collect available speech data sets and preprocess them;
[0033] An LLM-based TTS model training module, configured to train a preset LLM-based TTS model based on the preprocessed speech data set;
[0034] A data augmentation module, configured to perform data augmentation by using the trained LLM-based TTS model;
[0035] A data screening module for screening the amplified data;
[0036] An ASR model training module for training an ASR automatic speech recognition model using the screened amplified data.
[0037] In a third aspect, the present application provides an electronic device, which includes a processor and a memory; a program is stored in the memory, and the program is loaded and executed by the processor to implement a data amplification method based on LLM-TTS as described in the first aspect.
[0038] In a fourth aspect, the present application provides a computer-readable storage medium, in which a program is stored, and when the program is executed by a processor, it is used to implement a data amplification method based on LLM-TTS as described in the first aspect.
[0039] In summary, the beneficial effects of the present application at least include:
[0040] 1) By generating high-quality speech data of scarce dialects and minority languages based on the LLM-based TTS model, the problem of poor generalization ability of the speech recognition system in a scarce data environment is solved. During the data amplification process, self-supervised learning and sampling strategies are adopted to ensure the balance of samples in different languages, enabling scarce minority languages and dialects to obtain more training opportunities. This process, through the diversification and refinement of data, enables the speech recognition system to effectively recognize and process the speech of scarce languages, improving the adaptability and accuracy of the system. Especially in practical applications, it can better cope with the challenges of dialect and minority language speech recognition.
[0041] 2) The present application further improves the quality and diversity of the generated speech through an improved TTS generation strategy, using the chain of thought prompting strategy and syllable-level prosody feature prediction. During the speech generation process, the model first predicts syllable-level prosody features (such as pitch, duration, etc.), and then guides the generation process based on these features to ensure the naturalness and rhythm of the speech. In addition, by setting specific generation conditions (such as speech rate, emotional color, accent, etc.), the LLM-based TTS can generate diverse speech samples covering different speech variants. This not only increases the richness and diversity of the training data but also ensures that the generated speech can more accurately reflect the essential features of speech in different dialect and minority language scenarios, further improving the training effect of the subsequent speech recognition model and solving the problems of unstable voice quality or lack of diversity in the speech generation process.
[0042] By first collecting and preprocessing the available speech datasets to ensure data diversity and quality; then, training the LLM-based TTS model based on the processed data, ensuring the balance of language samples in training, and introducing the chain of thought prompting strategy to optimize the generation quality. Subsequently, using the trained TTS model to generate diverse speech data to further augment the scarce speech datasets. By introducing Filter-ASR for data screening to eliminate low-quality generated data and ensure the quality of the augmented data. Finally, training the ASR model using the screened data and evaluating the model through test sets of dialects and minority languages, continuously optimizing the generation strategy to further improve the recognition ability of the ASR model. This method effectively generates and screens augmented data to ensure that in an environment of scarce data, the speech recognition system can better adapt to the characteristics of dialects and minority languages, thus solving the problem of low recognition performance of the speech recognition system in scarce languages and dialects in the prior art.
[0043] The above description is only an overview of the technical solution of this application. In order to be able to understand the technical means of this application more clearly and implement it in accordance with the content of the specification, the following describes in detail with reference to the preferred embodiments of this application and the accompanying drawings. Brief Description of the Drawings
[0044] Figure 1 is a schematic flowchart of the data augmentation method based on LLM-TTS in an embodiment of this application.
[0045] Figure 2 is a schematic overall flowchart of the data augmentation method based on LLM-TTS in an embodiment of this application.
[0046] Figure 3 is a structural block diagram of the data augmentation system based on LLM-TTS in an embodiment of this application.
[0047] Figure 4 is a block diagram of an electronic device for data augmentation based on LLM-TTS in an embodiment of this application. Detailed Description of the Embodiments
[0048] The following further describes in detail the specific embodiments of this application with reference to the accompanying drawings and embodiments. The following embodiments are used to illustrate this application but are not used to limit the scope of this application.
[0049] Optionally, this application takes the data augmentation method based on LLM-TTS provided in each embodiment and applied to an electronic device as an example for illustration. The electronic device is a terminal or a server, and the terminal can be a mobile phone, a computer, a tablet computer, etc. The type of the electronic device is not limited in this embodiment.
[0050] Refer to Figure 1, which is a schematic flowchart of a data augmentation method based on LLM-TTS provided by an embodiment of this application. The method at least includes the following steps:
[0051] Step S101, collect available speech datasets and preprocess them.
[0052] First, collect available speech datasets. The data sources include but are not limited to public speech databases, data recorded by specific language communities, and enterprise speech data. The data sources need to ensure to a certain extent that they cover data in mainstream languages, target dialects, and minority languages to provide sufficient diversity and representativeness. The speech datasets include mainstream languages, limited target dialects, and minority language data.
[0053] After completing the data collection, perform quality screening on the speech datasets, removing samples with excessive noise, incomplete content, or insufficient duration to ensure the overall quality and consistency of the data. Subsequently, extract discrete acoustic features from the screened data, including information such as pitch, duration, and spectrum. These features can characterize the prosody and content of speech and provide high-quality inputs for subsequent model training.
[0054] Step S102, train a preset LLM-based TTS model based on the preprocessed speech datasets.
[0055] In step S102, first use the speech datasets preprocessed and sampled in step S101 to train the LLM-based TTS model. To ensure the balance of multi-language samples during the training process, adjust the sampling strategy in each training epoch so that the data volumes of each language are basically balanced. This strategy can avoid too much or too little data of a certain language, so that the features of each language can be fully learned during the training process. Especially for data of scarce dialects and minority languages, more training opportunities can be obtained.
[0056] Among them, the basic architecture used for training the LLM-based TTS model is a decoder-only structure based on Transformer, which can efficiently handle sequence generation tasks. The design of the LLM-based TTS model first depends on discrete acoustic features extracted from multi-language dialect data, which are obtained through self-supervised pre-training to adapt to the LLM model structure. These acoustic features include key information such as pitch, duration, and spectrum, ensuring that the model can process and generate rich speech content.
[0057] In addition, to improve the quality of the generated speech, a chain of thought prompting strategy is introduced. Before generating each speech token, the model first predicts syllable-level prosodic features (such as pitch and duration) to ensure the stability of the generation process. By predicting these prosodic features in advance, the model can reasonably plan the pitch, duration, etc. of the speech in advance during speech generation, avoiding unnatural fluctuations during the generation process. The predicted syllable duration information is used to guide the self-attention mechanism, enabling the model to focus on the parts related to the current phoneme and prosodic features when generating speech. The purpose of this design is to enhance the model's fine control over speech during the generation process, ensure that the finally generated speech is more natural and rhythmical, and avoid ignoring some important audio details.
[0058] Step S103: Use the trained LLM-based TTS model for data augmentation.
[0059] In step S103, first, set the specific generation conditions for the target dialect and minority languages to ensure the generation of diverse speech samples during data augmentation. These conditions include but are not limited to various speech variants such as different dialect accents, speech rates, pitches, emotional expressions, etc. Through these settings, more abundant speech data can be generated, covering various possible expressions of the target language and dialects, so as to enhance the diversity and comprehensiveness of the dataset.
[0060] During the generation process, scarce minority language or dialect data is used as input prompts (prompts). These prompts are used to guide the model to generate speech samples that conform to the characteristics of the target dialect and minority languages. During the sampling process, the model will automatically adjust various variants of the speech (such as speech rate, emotional color, pronunciation features, etc.) according to the set generation conditions to ensure that the generated speech not only conforms to the timbre characteristics of the target language but also reflects the performance of the language in different situations. Through the above method, the LLM-based TTS model can effectively augment speech data in an environment with scarce data and ensure that the generated speech samples have diversity, which can better support subsequent speech recognition and other related tasks.
[0061] Step S104: Screen the augmented data.
[0062] In step S104, to ensure that the data generated by LLM-based TTS has high quality, screening of augmented data is carried out. Since the LLM-based TTS model has a certain probability of generating incorrect data during the generation process, an iterative process is needed to filter out this low-quality data. This process is achieved by training an ASR system dedicated to data screening, called Filter-ASR. It should be noted that this Filter-ASR is not the finally optimized ASR system. Its main role is to screen the augmented data, rather than the training model for the final speech recognition task.
[0063] Specifically, first train a Filter-ASR model with a large number of parameters and use the audio data generated by LLM-based TTS for training. After training is completed, use this model to decode the training audio to obtain the decoding result. Then, compare the decoding result with the ground truth text and calculate the word error rate (WER). According to the WER value, screen out the audio data with a high WER and remove them to ensure that only high-quality speech data enters the subsequent training. After the initial screening, gradually reduce the number of parameters of the Filter-ASR model and continue to train this Filter-ASR model using the high-quality audio data screened in the previous step. After training is completed, repeat the decoding process, calculate the WER, and remove the audio with a high WER again. In this way, the model is gradually refined and optimized, thereby improving the accuracy and efficiency of the screening process. Keep iterating until the final WER value reaches the preset standard. In this experiment, the preset standard is that the WER is less than 2%. Once this standard is reached, the screening process ends, and the screened high-quality data will be used for subsequent model training and applications. Through the above process, the screened augmented data will have high quality, ensuring better results in subsequent speech recognition and other tasks.
[0064] Step S105: Train the preset ASR model using the screened augmented data.
[0065] In step S105, use the high-quality TTS augmented data screened in step S104 to further enhance the sparse data of minority languages and dialects, and train the preset ASR (Automatic Speech Recognition) model. The purpose of this step is to utilize the augmented dataset to improve the recognition performance of the ASR model in scarce language and dialect environments.
[0066] First, use the high-quality speech dataset filtered by Filter-ASR as training data and input it into a preset ASR model for training. Since these data have removed low-quality samples through the screening process, they have high accuracy and diversity, which can effectively enhance the ASR model's recognition ability for minority languages and dialects. During the training process, the model will learn richer speech features, especially for the pronunciation characteristics of specific dialects and minority languages, thereby improving its accurate recognition rate for these speeches.
[0067] After training, use the test sets of dialects and minority languages to evaluate the ASR model. The purpose of the test is to examine the actual performance of the trained ASR model in different dialect and minority language environments. Through the test results, evaluate whether the model can achieve the expected recognition accuracy when processing speeches in specific languages. If the performance of the ASR model fails to meet the expectations in the test sets of certain dialects or minority languages, adjust and optimize the generation strategy according to the evaluation results.
[0068] Specifically, optimizing the generation strategy may include adjusting the generation conditions of the LLM-based TTS model, such as changing the pitch, speech rate, emotional color, etc. of the generated speech, to ensure that the generated speech data can better reflect the characteristics of the target dialect or minority language. These adjustments will help improve the diversity and quality of the generated speech data and further improve the performance of the ASR model in dialects and minority languages.
[0069] Through continuous evaluation and optimization, the finally trained ASR model will be able to more accurately recognize the speeches of dialects and minority languages, improving the robustness and accuracy of the speech recognition system in scarce language and dialect scenarios.
[0070] To sum up, combined with Figure 2, this application proposes a data augmentation method based on LLM-based TTS (Large Language Model-based Text-to-Speech generation) technology, aiming to solve the problem of insufficient speech data for scarce dialects and minority languages, thereby improving the performance of speech recognition systems in these language and dialect environments. By first collecting and preprocessing the available speech dataset to ensure data diversity and quality; then, training the LLM-based TTS model based on the processed data to ensure the balance of language samples in training and introducing the chain of thought prompting strategy to optimize the generation quality. Subsequently, using the trained TTS model to generate diverse speech data to further augment the scarce speech dataset. By introducing Filter-ASR for data screening to eliminate low-quality generated data and ensure the quality of the augmented data. Finally, using the screened data to train the ASR model and evaluating the model through test sets of dialects and minority languages, continuously optimizing the generation strategy to further improve the recognition ability of the ASR model. This method effectively generates and screens augmented data, ensuring that in an environment with scarce data, the speech recognition system can better adapt to the characteristics of dialects and minority languages, thus solving the problem of low recognition performance of speech recognition systems in scarce languages and dialects in the prior art.
[0071] Figure 3 FIG. 4 is a structural block diagram of a data augmentation system based on LLM-TTS provided by an embodiment of this application. The system at least includes the following modules:
[0072] A data collection module, configured to collect the available speech dataset and preprocess it;
[0073] An LLM-based TTS model training module, configured to train a preset LLM-based TTS model based on the preprocessed speech dataset;
[0074] A data augmentation module, configured to perform data augmentation using the trained LLM-based TTS model;
[0075] A data screening module, configured to screen the augmented data;
[0076] An ASR model training module, configured to train an ASR automatic speech recognition model using the screened augmented data.
[0077] For relevant details, refer to the method embodiment above.
[0078] Figure 4 FIG. 5 is a block diagram of an electronic device provided by an embodiment of this application. The device at least includes a processor 401 and a memory 402.
[0079] The processor 401 may include one or more processing cores, such as a 4-core processor, an 8-core processor, etc. The processor 401 may be implemented in at least one hardware form of DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), or PLA (Programmable Logic Array). The processor 401 may also include a main processor and a coprocessor. The main processor is a processor for processing data in the wake state, also known as the CPU (Central Processing Unit); the coprocessor is a low-power processor for processing data in the standby state. In some embodiments, the processor 401 may be integrated with a GPU (Graphics Processing Unit), and the GPU is responsible for rendering and drawing the content to be displayed on the display screen. In some embodiments, the processor 401 may further include an AI (Artificial Intelligence) processor, and the AI processor is used to process computational operations related to machine learning.
[0080] The memory 402 may include one or more computer-readable storage media, and the computer-readable storage media may be non-transitory. The memory 402 may further include high-speed random access memory and non-volatile memory, such as one or more disk storage devices and flash storage devices. In some embodiments, the non-transitory computer-readable storage media in the memory 402 is used to store at least one instruction, and the at least one instruction is used to be executed by the processor 401 to implement the data augmentation method based on LLM-TTS provided in the method embodiments of the present application.
[0081] In some embodiments, the electronic device may further optionally include: a peripheral device interface and at least one peripheral device. The processor 401, the memory 402, and the peripheral device interface may be connected through a bus or signal lines. Each peripheral device may be connected to the peripheral device interface through a bus, signal lines, or a circuit board. Schematically, the peripheral devices include but are not limited to: a radio frequency circuit, a touch display screen, an audio circuit, and a power supply, etc.
[0082] Of course, the electronic device may also include fewer or more components, and this embodiment does not limit this.
[0083] Optionally, the present application also provides a computer-readable storage medium, and a program is stored in the computer-readable storage medium, and the program is loaded and executed by the processor to implement the data augmentation method based on LLM-TTS in the above method embodiments.
[0084] Optionally, the present application also provides a computer product, which includes a computer-readable storage medium. A program is stored in the computer-readable storage medium and is loaded and executed by a processor to implement the data augmentation method based on LLM-TTS in the above method embodiment.
[0085] The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.
[0086] The above embodiments only represent several implementation manners of the present application. The description is relatively specific and detailed, but it should not be construed as a limitation on the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application should be subject to the appended claims.
Claims
1. A data augmentation method based on LLM-TTS, characterized in that: The method comprises: Collect available speech datasets and preprocess them; Training a preset LLM-based TTS model based on the preprocessed speech data set; Using the trained LLM-based TTS model to perform data augmentation; Screening the amplification data; The screened augmented data is used to train an ASR automatic speech recognition model.
2. The data augmentation method based on LLM-TTS according to claim 1, characterized in that: The collecting of available speech data sets and preprocessing thereof include: The speech dataset includes mainstream languages, limited target dialects and minority language data.
3. The data augmentation method based on LLM-TTS according to claim 1, characterized in that: The training of the preset LLM-based TTS model based on the preprocessed speech data set includes: Adjust the sampling strategy in each training cycle to ensure that the amount of data for each language is basically balanced; A chain thinking prompt strategy is introduced. Before generating each speech token, the model predicts the prosodic features at the syllable level, and uses the predicted syllable duration information to guide the self-attention mechanism, so that the model can focus on the parts related to the current phonemes and prosodic features when generating speech.
4. The data augmentation method based on LLM-TTS according to claim 3, characterized in that: The training of the preset LLM-based TTS model based on the preprocessed speech data set also includes: The infrastructure used for the LLM-based TTS model training is based on the Transformer-based decoder-only structure.
5. The data augmentation method based on LLM-TTS according to claim 1, characterized in that: The data augmentation using the trained LLM-based TTS model includes: Setting specific production conditions for target dialects and minority languages; During the generation process, scarce minority language or dialect data is used as input prompts, and the LLM-based TTS model is guided by the input prompts to generate speech samples that conform to the characteristics of the target dialect and minority language.
6. The data augmentation method based on LLM-TTS according to claim 1, characterized in that: The screening of the amplification data comprises: Training a Filter-ASR model, and using the audio data generated by the LLM-based TTS model for training; after the training is completed, using the Filter-ASR model to decode the training audio to obtain a decoding result; Compare the decoding results with the real text, calculate the word error rate, and filter out the audio data with higher word error rate based on the word error rate. After the initial screening, gradually reduce the number of parameters of the Filter-ASR model, and continue to train the Filter-ASR model using the high-quality audio data filtered in the previous step. After training is completed, the decoding process is repeated, the word error rate is calculated, and the audio with a high word error rate is eliminated again, and the iteration is continued until the final WER value reaches the preset standard.
7. The data augmentation method based on LLM-TTS according to claim 6, characterized in that: The use of the screened augmented data to train the ASR automatic speech recognition model comprises: The high-quality speech data set filtered by the Filter-ASR model is used as training data and input into the preset ASR model for training; After training, the ASR model is evaluated using test sets of dialects and minority languages; If the performance of the ASR model fails to meet expectations in the test sets of certain dialects or minority languages, the generation strategy will be adjusted and optimized based on the evaluation results.
8. A data augmentation system based on LLM-TTS, characterized in that: include: A data collection module, used to collect available speech data sets and pre-process them; An LLM-based TTS model training module, used for training a preset LLM-based TTS model based on the preprocessed speech data set; A data augmentation module, used for performing data augmentation using the trained LLM-based TTS model; A data screening module, used for screening amplification data; The ASR model training module is used to train the ASR automatic speech recognition model using the screened augmented data.
9. An electronic device, characterized in that: The device includes a processor and a memory; a program is stored in the memory, and the program is loaded and executed by the processor to implement a data augmentation method based on LLM-TTS as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that: The storage medium stores a program, and when the program is executed by the processor, it is used to implement the data augmentation method based on LLM-TTS as described in any one of claims 1 to 7.
Citation Information
Cited By
Chinese dialect English speech processing system and method
CN120748371A