Method for acquiring training data of speech synthesis model

By screening high-frequency syllables from the corpus and using the speech synthesis model to evaluate quality information, a decision table is constructed to simplify attributes and select necessary elements. This solves the problem of suboptimal training data selection for the end-to-end speech synthesis model, achieves efficient and low-cost training data acquisition, and improves the training efficiency and accuracy of the model.

CN120808745AActive Publication Date: 2025-10-17GARZE TIBETAN AUTONOMOUS PREFECTURE INST OF SCI & TECH INFORMATION
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202511065787.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-31
Publication Date
2025-10-17
Estimated Expiration
2045-07-31

AI Technical Summary

Technical Problem

Existing end-to-end speech synthesis models have problems with training data selection, such as small sample size, high professionalism, and high cost. This leads to suboptimal training data selection, wasting time and manpower costs.

Method used

By screening high-frequency syllables from the corpus of the target language, sampling and marking the recordings, using the speech synthesis model to evaluate the quality information of the unlabeled corpus, constructing a decision table and performing attribute simplification, screening out necessary elements, sampling and marking the recordings, and obtaining new training data.

Benefits of technology

While ensuring the accuracy of the speech synthesis model, the time and labor costs of training data are reduced, the efficiency of training data selection is improved, the number of samples is reduced, and costs are saved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120808745A_ABST
    Figure CN120808745A_ABST
Patent Text Reader

Abstract

The invention discloses a method for acquiring training data of a speech synthesis model. The method comprises the following steps: screening a part of corpora with high-frequency syllables from a corpus of a target language; performing recording sampling and marking on the linguistic data obtained by screening to obtain an initial training set, and training the speech synthesis model through the initial training set; performing speech synthesis on unmarked corpora in the corpus according to the speech synthesis model to obtain speech synthesis quality information; determining the decision attribute of the unmarked corpus according to the quality information, and constructing a decision table by taking the occurrence frequency of basic letters in the unmarked corpus as a condition attribute to obtain an advantage rough set. The method effectively solves the problems that in a target language synthesis task, due to the fact that the number of training data samples is small, enrollment and marking of the training data samples are high in professionality, cost is high, and workload is large.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of data processing, in particular to a method for obtaining training data of a speech synthesis model. BACKGROUND

[0002] The end-to-end speech synthesis model not only has a high requirement for the audio quality of the sample, but also needs the training data to cover the common pronunciation of the used language as much as possible from the phoneme level. Therefore, the requirement for the training data of the speech synthesis is much higher than that in other machine learning fields. How to select appropriate samples from the corpus as the training data of the speech synthesis is very important. For the end-to-end speech synthesis model, it is actually a kind of supervised learning, and the mel-frequency spectrum vector of the speech data is a kind of label of the training data. Due to the dialect difference or regional pronunciation characteristics of the target language, it is difficult to obtain the speech synthesis training data with high audio quality and standard pronunciation. In order to make the speech synthesis training data meet the requirements, firstly, the samples most beneficial to the speech synthesis model need to be selected from the corpus as the training data scientifically, and secondly, professional persons need to be hired to use professional equipment to record the audio sampling and label of the selected training data. If the unlabelled samples in the corpus are randomly selected as the training data, the selected training data may not be the optimal training data for the speech synthesis model. In this way, a large amount of time cost and a large amount of human cost are wasted. SUMMARY

[0003] The present application aims to provide a method for obtaining training data of a speech synthesis model, so as to quickly select the samples with the most value as the training data of the speech synthesis model.

[0004] The present application provides a method for obtaining training data of a speech synthesis model, comprising: selecting part of the corpus with high-frequency syllables from the corpus of the target language; recording the sampling and labeling of the selected corpus to obtain an initial training set, and training the speech synthesis model through the initial training set; performing speech synthesis on the unlabelled corpus in the corpus according to the speech synthesis model to obtain quality information of the speech synthesis; determining the decision attribute of the unlabelled corpus according to the quality information, taking the occurrence frequency of the basic letters in the unlabelled corpus as the condition attribute, constructing a decision table, and obtaining a dominant rough set; performing attribute reduction on the dominant rough set to obtain the intersection of the attribute reduction, finding the necessary element from the intersection, and selecting new corpus from the unlabelled corpus through the necessary element to record the sampling and labeling to obtain new training data.

[0005] Furthermore, the quality information is a MOS score, and determining the decision attribute of the unlabeled corpus according to the quality information includes: For any unlabeled corpus, when the corresponding MOS score is greater than a preset threshold, the decision attribute of the unlabeled corpus is determined to be a first-category attribute; when the corresponding MOS score is less than or equal to the preset threshold, the decision attribute of the unlabeled corpus is determined to be a second-category attribute.

[0006] Furthermore, the method of using the number of occurrences of basic letters in the unlabeled corpus as a conditional attribute includes: For any of the unlabeled corpora, convert the corpus into text represented by Willy's Romanization syllable symbols, and determine the number of times each Willy's Romanization syllable symbol appears in the text obtained after the conversion; Each syllable symbol of Willy's Romanization appearing in the text and the number of times the syllable symbol appears are used as conditional attributes of the unlabeled corpus.

[0007] Furthermore, performing attribute simplification on the dominant rough set to obtain an intersection of the reduced attributes, and finding necessary elements from the intersection, includes: By using heuristic attribute reduction, the element algorithm HARCC is used to find the reduced intersection core in the decision table that makes the decision attributes the second type of attributes, and the necessary elements in the reduced intersection core.

[0008] Furthermore, the method of filtering out new corpus from the unlabeled corpus by using the necessary elements for recording sampling and labeling to obtain new training data includes: sorting data in the unlabeled corpus whose decision attributes are the second type of attributes using the necessary elements, and selecting new corpus based on the sorting result, wherein the selected new corpus contains a higher proportion of syllables corresponding to the necessary elements than the unselected corpus; The new corpus is sampled and marked to obtain new training data.

[0009] Furthermore, the target language is any one of Tibetan, Naxi, Hani, Qiang, Miao, and Buyi.

[0010] Furthermore, the new training data is used to update the speech synthesis model, and the method further includes: The speech synthesis model is repeatedly updated and trained according to the quality information of the speech synthesis until the speech synthesis model meets a convergence condition.

[0011] The present application at least includes the following advantages: The present application provides a method for selecting optimal samples from a target language corpus as training data of a target language speech synthesis model based on a small amount of speech synthesis training data. The present application effectively solves the problems of professional and high cost and heavy workload in recording and labeling training data samples due to the small amount of training data samples in the target language synthesis task. The present application selects the most effective samples for the model as training data, which reduces the samples and saves time and labor costs compared with training the language synthesis model by randomly selecting unmarked samples from the target language corpus as training data.

[0012] The present application will be further described below in conjunction with the accompanying drawings and specific embodiments. BRIEF DESCRIPTION OF DRAWINGS

[0013] Figure 1 A flowchart of a method for obtaining training data of a speech synthesis model is provided. Figure 2 A Tibetan alphabet and Wylie transcription comparison diagram is provided. DETAILED DESCRIPTION

[0014] The technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all embodiments.

[0015] The present embodiment is a method for obtaining training data of a speech synthesis model, comprising: Selecting part of the corpus with high frequency syllables from the corpus of the target language; In the present embodiment, the target language is Tibetan. The present application selects optimal samples from a Tibetan corpus as training data of a Tibetan speech synthesis model based on a small amount of Tibetan speech synthesis training data. The present application identifies and extracts high-frequency corpus (such as syllable combinations and common pronunciation fragments) from the Tibetan corpus, focuses on the most commonly used pronunciation patterns in the language, and thus improves the efficiency and accuracy of subsequent model training.

[0016] record and label the screened corpus to obtain an initial training set, and train the speech synthesis model through the initial training set; By recording the high-frequency syllable corpus, obtaining the corresponding audio samples, and labeling the corresponding text content (i.e., the correspondence between speech and text) for each audio sample, a labeled training data is formed. The initial training set is composed of these labeled audio and text data, which is used to preliminarily train the speech synthesis model (such as WaveNet, TTS model, etc.).

[0017] According to the speech synthesis model, the unlabeled corpus in the corpus library is synthesized to obtain the quality information of the speech synthesis; Using the initially trained speech synthesis model, the unlabeled corpus (i.e., unrecorded / annotated text data) in the corpus library is synthesized and the quality information of the speech synthesis is obtained.

[0018] According to the quality information, the decision attribute of the unlabeled corpus is determined, and the occurrence frequency of the basic letters in the unlabeled corpus is used as the condition attribute to construct a decision table to obtain the advantage rough set; According to the quality information, the unlabeled corpus is given a decision label (such as "high quality", "low quality", etc.), which is used as the basis for subsequent classification. And the occurrence frequency of the basic letters in the unlabeled corpus is used as the feature attribute. For example, the frequency of the occurrence of consonant letters, vowel letters, or specific letter combinations in each corpus is counted. Then the decision attribute and the condition attribute are combined into a table form to form a decision system. Each unlabeled corpus is a record containing its letter frequency characteristics and quality label. Then the rough set theory (a mathematical method for handling uncertainty and data classification) is used to analyze the decision table to identify which condition attributes (letter frequency) have a greater impact on the decision (quality label), thereby obtaining the advantage attribute set, i.e., the advantage rough set.

[0019] The advantage rough set is attribute-reduced to obtain the intersection of attribute reduction, and the necessary elements are found from the intersection to screen out new corpora from the unlabeled corpora for recording and labeling to obtain new training data.

[0020] The advantage rough set is simplified through an algorithm (such as a genetic algorithm or a heuristic algorithm), redundant attributes are removed, and the most critical features for decision-making are retained. For example, it is found that the frequency of some letter combinations has a significant impact on quality, while other attributes can be ignored. If there are multiple reduction results, find the common key attributes in these results to obtain the intersection of attribute reduction. Then find the necessary elements from the intersection, that is, the features essential to classification decision-making, and use the necessary elements to further filter a higher-quality subset from the unlabeled corpus. For example, preferentially select the text in which the necessary elements appear frequently. Then record and label the filtered new corpus to obtain new high-quality training data.

[0021] The method effectively solves the problems of few training data samples, high cost and large workload of professional recording and labeling training data samples in the Tibetan language synthesis task. The present application selects the most effective samples for the model as training data, which reduces the samples, saves time and labor costs compared to training the Tibetan language synthesis model by randomly selecting unlabeled samples from the Tibetan corpus as training data for the Tibetan language synthesis model while ensuring the accuracy of the Tibetan speech synthesis model.

[0022] Further, the quality information is a MOS score, and determining the decision attribute of the unlabeled corpus according to the quality information comprises: For any unlabeled corpus, when the corresponding MOS score is greater than a preset threshold, the decision attribute of the unlabeled corpus is determined to be a first type of attribute, and when the corresponding MOS score is less than or equal to the preset threshold, the decision attribute of the unlabeled corpus is determined to be a second type of attribute.

[0023] The decision attribute set D is obtained by performing speech synthesis on the unlabeled sample set U through the speech synthesis model and performing speech synthesis evaluation MOS scoring one by one. The values of the decision attribute set D include: when the speech synthesis evaluation MOS score is greater than or equal to 4, D i =1, and when the speech synthesis evaluation MOS score is less than 4, D i =2.

[0024] Further, the occurrence frequency of the basic letters in the unlabeled corpus is used as the condition attribute, which comprises: For any of the unlabeled corpora, the corpus is converted into a text represented by the syllable symbols of the Wylie Romanization, and the occurrence frequency of each syllable symbol of the Wylie Romanization in the converted text is determined. Each syllable symbol of the Wylie Romanization and the occurrence frequency of the syllable symbol in the text are used as the condition attribute of the unlabeled corpus.

[0025] As Figure 2As shown, the Tibetan language is composed of 30 basic consonant letters and 4 vowels. In order to express conveniently, 34 Tibetan basic letters are transcribed into Wylie Romanization by Wylie transcription method, and the 34 Wylie Romanization are numbered as , and all the 34 letters form a non-empty finite condition attribute set C: , Among them, the value of the condition attribute is the number of times of the corresponding letter appearing in each unlabeled sample .

[0026] Assuming that the existing Tibetan speech synthesis model is trained based on the labeled data, the decision attribute set D is obtained by using the model to estimate each unlabeled data in the unlabeled sample set U, and the value of the decision attribute has only two, when the speech synthesis evaluation MOS score is greater than or equal to 4, D i =1, when the speech synthesis evaluation MOS score is less than 4, D i =2. Thus, a non-empty finite attribute set A is obtained: , From which it can be deduced that: , Among them, is the value range of the condition attribute , and the information function is: , For any condition attribute belongs to the non-empty finite attribute set A and the unlabeled sample data x i belongs to the unlabeled data set U, the following can be obtained: , From the above formula, a four-tuple decision information system (decision table) is obtained: , Assuming that the unlabeled data sample is 5000, the decision information system (decision table) can be represented as follows: Table 1 Decision information system (decision table) of 15000 unlabeled samples

[0027] Further, the attribute reduction of the advantage rough set is performed to obtain the intersection of attribute reduction, and the necessary elements are found from the intersection, including: ​The heuristic attribute reduction is used to find the reduction intersection core of the decision attribute being the second type of attribute in the decision table by the element algorithm HARCC, and the necessary element in the reduction intersection core.

[0028] Further, the new corpus is screened out from the unlabeled corpus by the necessary element for recording sampling and labeling to obtain new training data, comprising: The data with the decision attribute being the second type of attribute in the unlabeled corpus is sorted by the necessary element, and new corpus is screened out based on the sorting result, wherein the syllable proportion corresponding to the necessary element in the screened new corpus is higher than that in the un-screened corpus; The new corpus is recorded, sampled and labeled to obtain new training data.

[0029] The biggest advantage of the dominance rough set compared with the ordinary rough set is that it can solve the ordered decision information system, in the application, the number of times of occurrence of each syllable in the single factor covering information decision system is accumulated in turn, thereby forming an ordered decision information system with different importance degrees, and the strong correlation of the syllable in the single phoneme covering decision information system can be found by the dominance rough set.

[0030] The attribute reduction of the rough set is the key to find the most valuable unlabeled sample, but there is not only one minimal set in a decision information system, a decision information system can have multiple reductions, then the intersection of multiple reductions must be found, that is, the core, and the indispensable criteria, that is, the necessary element in the core is found. Generally, there are two big methods of attribute reduction of rough set, one is attribute reduction based on discernibility matrix, and the other is heuristic attribute reduction. The heuristic attribute reduction is used in the application, and the HARCC (heuristic attribute reduction with computing core) algorithm is used to calculate the core element, In view of the characteristics of the Tibetan phonetic alphabet and the characteristics of the speech synthesis model, in order to cover the Tibetan phonemes as much as possible, the Tibetan speech synthesis model M p The speech synthesis evaluation MOS score is lower than 4, that is, the sample with the decision attribute D i =2 is specially focused on, and data is screened from it. That is, the necessary element syllable of the decision attribute D i =2 in the decision information system is found by the HARCC algorithm, and the unlabeled sample with a large proportion of the necessary element syllable is screened by the ordered characteristics of the dominance rough set, and multiple most valuable training data are screened to obtain new training data N, comprising the following steps: Input: Decision system S (a dataset containing condition attributes and decision attributes), decision system .

[0031] Output: Reduction A of system S, the final set of key attributes.

[0032] Step 1: Initialize the reduction A of system S as an empty set, i.e., set ; Step 2: Calculate the importance of each attribute, for each attribute a, calculate its importance score in classification (e.g., the stronger the attribute a distinguishes samples, the higher the score): ; Step 3: Select core elements (necessary elements), if the score of a certain attribute a meets certain conditions (e.g., the highest score or reaching a certain threshold), add it to the reduction A: if , then select attribute a into the reduction A; Step 4: Initialize the reduction A: ; Steps 5-7: Loop to add attributes (heuristic selection); while the loop condition is that the current reduction A has not yet met the termination condition (e.g., the classification accuracy meets the requirements or the size of the attribute set reaches the limit), select the most suitable attribute a from the remaining attributes (according to heuristic rules, such as selecting the attribute with the highest score or the attribute with the highest dependency), and add attribute a to the reduction A.

[0033] Step 8: Attribute verification and deletion (reduction strategy), delete attribute a from the reduction A, if the classification ability does not decrease after deletion (e.g., by verifying the classification consistency of the reduction A), it means that attribute a is not necessary and can be deleted. This step ensures that each attribute a in the final reduction A is necessary, i.e., the reduced set is minimized.

[0034] Step 9: Output the final reduction set A.

[0035] This algorithm uses the "add-reduce" strategy of core elements, through loop steps 2 and 3, all core elements can be obtained, and the importance of each condition attribute C a is calculated to verify the indispensability of each criterion. Through steps 5 to 7, attributes are added, and the initialized is used as the starting point in each while iteration loop, and the most suitable criterion is selected according to heuristics until the condition is met and the program stops, step 8 ensures that the reduced set does not contain redundant attributes by deleting attributes and checking if the classification ability decreases.

[0036] Suppose there is a more complex decision table with 4 condition attributes (A, B, C, D) and 1 decision attribute (MOS score classification).

[0037] Table 2 Decision table

[0038] Calculate the importance of each attribute: Calculate the importance of each attribute by some algorithm (e.g. HARCC). Suppose we find: The importance of attribute A is 0.6, The importance of attribute B is 0.8, The importance of attribute C is 0.4, The importance of attribute D is 0.2, Select the most important attribute: Select the most important attribute B to join the reduction set A.

[0039] Check if the classification ability is met: Check if all sentences can still be correctly classified when only attribute B is kept. We find: Sentence 1: B=1, classified as 1 (high quality); Sentence 2: B=0, classified as 2 (low quality); Sentence 3: B=1, classified as 1 (high quality); Sentence 4: B=1, classified as 1 (high quality); Sentence 5: B=0, classified as 2 (low quality); We find that when only attribute B is kept, all sentences can still be correctly classified. By using only the currently selected attributes (here only attribute B), rejudge whether the MOS classification of each sentence is consistent with the original table. If all classifications are correct, it means that the attribute is sufficient to distinguish different categories, and other attributes can be further reduced.

[0040] Continue to select other important attributes: Select the second most important attribute A to join the reduction set A.

[0041] Check if all sentences can still be correctly classified when attributes A and B are kept. We find: Sentence 1: A=1, B=1, classified as 1 (high quality); Sentence 2: A=1, B=0, classified as 2 (low quality); Sentence 3: A=0, B=1, classified as 1 (high quality); Sentence 4: A=1, B=1, classified as 1 (high quality); Sentence 5: A=0, B=0, classified as 2 (low quality); When using multiple attribute combinations, it is necessary to ensure that all possible value combinations can uniquely correspond to the correct classification. We found that when retaining attributes A and B, all sentences can still be correctly classified.

[0042] Check if you can remove other attributes: We found that after removing attributes C and D, all sentences can still be correctly classified. Therefore, attributes C and D are not necessary.

[0043] Core: The core is the intersection of all reductions, containing irreplaceable properties. For example, if there are multiple reductions (such as {A, B} or {B, C}), the core is the intersection of their reductions. In this example, the only feasible reduction is {A, B}, so the core is {A, B}.

[0044] Indispensable Criteria: Indispensable criteria are key attributes within the core; removing any one of them will affect classification. Although the core is {A, B}, further analysis reveals that removing A alone (keeping B) still allows for correct classification; however, removing B (keeping A) renders some sentences indistinguishable (for example, sentences 2 and 5 are classified differently when B=0). Therefore, in this example, the essential criteria is {B}. The new training data is designed to find corpus containing the syllable B at a rate greater than a threshold.

[0045] Furthermore, the target language is any one of Tibetan, Naxi, Hani, Qiang, Miao, and Buyi.

[0046] Furthermore, the new training data is used to update the speech synthesis model, and the method further includes: Repeat the steps of performing speech synthesis on the unlabeled corpus in the corpus according to the speech synthesis model to obtain quality information of the speech synthesis, filtering out new corpus from the unlabeled corpus by the necessary elements for recording sampling and labeling to obtain new training data, and updating and training the speech synthesis model according to the new training data, until the speech synthesis model meets the convergence condition.

[0047] After obtaining the most favorable training data, the training data is recorded, sampled, and labeled, and then added to the initial training set for incremental training and updating of the speech synthesis model, thereby strengthening the model and making it more robust.

[0048] Through experiments, we collected a large amount of modern Tibetan text from websites and books to establish a Tibetan corpus, which we used to evaluate our proposed method. We used the well-established end-to-end speech synthesis model, Tacotron-2, as the test model for this experiment. The same experts performed MOS scores on the test set for the model.

[0049] In this experiment, we first selected 3,000 samples from the Tibetan corpus as initial training data through statistical methods, sampled and labeled their recordings, and performed initial training to obtain the initial model M. p 3000 data were selected from the Tibetan corpus as test data, and the initial model M was tested by two methods. p The first method is to use random method to select unlabeled samples from the remaining Tibetan corpus as training data to train the initial model M p The second method is to use the method provided by the present invention to obtain training data for the speech synthesis model and select unlabeled samples from the remaining Tibetan corpus as training data to train the initial model M. p The most favorable training data N value is 50, that is, the first 50 data are selected for recording annotation and added to the initial training set. The experimental results are recorded every time 1000 training data are added and every time 2000 training data are added. The experimental results are shown in Tables 3 and 4.

[0050] Table 3 Comparison of random selection and DRS-AL algorithms when 1000 training data are added

[0051] Table 4 Comparison of random selection and DRS-AL algorithms when 2000 training data are added

[0052] As shown in the experimental results in Tables 3 and 4, regardless of the scenario, using the method provided by the present invention to select unlabeled data as training data and train the model is more efficient than randomly selecting unlabeled data as training data. This method achieved excellent training results after the amount of training data increased to 2,000. Random selection methods, on the other hand, only achieve good results with a large sample size. The proposed method for obtaining training data for a speech synthesis model reduces the labor and time costs of sampling and annotating Tibetan speech synthesis recordings, while significantly improving the model's efficiency.

[0053] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any modifications or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in the present invention should be included in the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be based on the scope of protection of the claims.

Claims

1. A method for obtaining training data for a speech synthesis model, characterized in that: include: Select some corpus with high-frequency syllables from the corpus of the target language; The screened corpus is sampled and labeled to obtain an initial training set, and the speech synthesis model is trained using the initial training set; Performing speech synthesis on unlabeled corpus in the corpus according to the speech synthesis model to obtain quality information of the speech synthesis; Determining a decision attribute of the unlabeled corpus according to the quality information, and constructing a decision table using the number of occurrences of basic letters in the unlabeled corpus as a conditional attribute to obtain a dominant rough set; Attribute simplification is performed on the dominant rough set to obtain an intersection of the attribute simplifications, and necessary elements are found from the intersection to filter out new corpus from unlabeled corpus through the necessary elements for recording sampling and labeling to obtain new training data.

2. The method for obtaining training data for a speech synthesis model according to claim 1, wherein: The quality information is a MOS score, and determining the decision attribute of the unlabeled corpus according to the quality information includes: For any unlabeled corpus, when the corresponding MOS score is greater than a preset threshold, the decision attribute of the unlabeled corpus is determined to be a first-category attribute; when the corresponding MOS score is less than or equal to the preset threshold, the decision attribute of the unlabeled corpus is determined to be a second-category attribute.

3. The method for obtaining training data for a speech synthesis model according to claim 1, wherein: The method of using the number of occurrences of basic letters in the unlabeled corpus as a conditional attribute includes: For any of the unlabeled corpora, convert the corpus into text represented by Willy's Romanization syllable symbols, and determine the number of times each Willy's Romanization syllable symbol appears in the text obtained after the conversion; Each syllable symbol of Willy's Romanization appearing in the text and the number of times the syllable symbol appears are used as conditional attributes of the unlabeled corpus.

4. The method for obtaining training data for a speech synthesis model according to claim 2, wherein: The performing attribute simplification on the dominant rough set to obtain an intersection of the attribute simplifications, and finding necessary elements from the intersection, includes: By using heuristic attribute reduction, the element algorithm HARCC is used to find the reduced intersection core in the decision table that makes the decision attributes the second type of attributes, and the necessary elements in the reduced intersection core.

5. The method for obtaining training data for a speech synthesis model according to claim 4, wherein: The method of filtering out new corpus from unlabeled corpus by using the necessary elements for recording sampling and labeling to obtain new training data includes: sorting data in the unlabeled corpus whose decision attributes are the second type of attributes using the necessary elements, and screening out new corpus based on the sorting results, wherein the proportion of syllables corresponding to the necessary elements in the screened out new corpus is higher than that in the unscreened corpus; The new corpus is sampled and marked to obtain new training data.

6. A method for obtaining training data for a speech synthesis model according to claims 1-5, characterized in that: The target language is any one of Tibetan, Naxi, Hani, Qiang, Miao, and Buyi.

7. A method for obtaining training data for a speech synthesis model according to claims 1-5, characterized in that: The new training data is used to update the speech synthesis model, and the method further includes: Repeat the steps of performing speech synthesis on the unlabeled corpus according to the speech synthesis model to obtain quality information of the speech synthesis, filtering out new corpus from the unlabeled corpus by the necessary elements for recording sampling and labeling to obtain new training data, and updating and training the speech synthesis model according to the new training data, until the speech synthesis model meets the convergence condition.

Citation Information

Patent Citations

  • Corpus system construction method based on rough set

    CN110442729A

  • Large model pre-training corpus construction method and device

    CN119397267A

  • Methods and apparatus for identification and analysis of temporally differing corpora

    US9135243B1

  • Prosody model learning device, prosody model learning method, voice synthesis system, and prosody model learning program

    WO2014061230A1