Method and apparatus for generating a voice dataset
By acquiring speech datasets of standard general-purpose languages and utilizing large language models and retrieval-enhanced generation methods, a speech dataset for the target language is constructed. This solves the problem of low accuracy in low-resource language translation models and achieves the expansion of speech data and the improvement of translation accuracy.
Patent Information
- Application Number
- CN202511340288.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-18
- Publication Date
- 2025-12-23
- Estimated Expiration
- 2045-09-18
AI Technical Summary
Existing multilingual speech datasets do not pay enough attention to low-resource languages, resulting in low translation accuracy of translation models and limited effectiveness due to the lack of corresponding translation annotations and data augmentation methods.
By acquiring a speech dataset of a standard general language, a large language model is used to convert it into text in the target language. A retrieval-enhanced generation method is then used to construct a speech dataset in the target language, ensuring consistency of speech features. This includes multi-level quality control in text generation, speech feature extraction, and evaluation.
The target language speech dataset has been expanded, increasing the amount of speech data and translation accuracy of the translation model. The generated speech data is more in line with actual usage habits, improving the quality and usability of speech synthesis.
Smart Images

Figure CN120877702B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of artificial intelligence, in particular to a voice data set generation method and device. BACKGROUND
[0002] At present, many existing multilingual voice data sets focus on resource-rich languages in the Indo-European language family (such as English and German) and the East Asian language family (such as Chinese and Japanese). Low-resource languages such as domestic dialects lack attention. In addition, these data sets often lack corresponding translation annotations and are difficult to use for training of voice translation models. In recent research, large language models are good at processing general instructions and have shown potential in text generation tasks. Recent research has focused on combining large language models with instructions and examples in raw text training data to prompt them to generate novel and diverse samples. At present, many studies focus on using language theory and machine translation technology to synthesize data enhancement for mixed language text. However, for low-resource languages, the effect of this method is still limited, and usually needs to rely on retrieval enhancement generation methods with additional knowledge bases to improve performance. SUMMARY
[0003] The embodiments of the present application provide a voice data set generation method and device to at least solve the technical problem that the accuracy of a translation model in translating a target language is low due to a small amount of voice data in a voice database of the target language in the related art.
[0004] According to an aspect of an embodiment of the present application, a voice data set generation method is provided, including: obtaining a voice data set of a standard general language, and converting the voice data set of the standard general language into a target language text using a large language model; generating a target language sentence text in a retrieval enhancement generation manner; generating a target language voice according to the target language text and the target language sentence text, and constructing a target voice data set according to the target language voice, wherein a voice feature of the target language voice is consistent with a voice feature of the voice data set of the standard general language.
[0005] Optionally, the target language sentence text is generated according to the search words, including: converting words in a preset target language database into dense vectors, wherein the preset target language database contains a plurality of target language words; obtaining a plurality of sentence topics, and generating a plurality of query texts according to each sentence topic respectively to obtain the plurality of query texts; converting the plurality of query texts into a plurality of query vectors; comparing each query vector with the dense vectors in the preset target language database respectively, screening a plurality of dense vectors related to each query vector, and obtaining a plurality of search words corresponding to the plurality of dense vectors related to each query vector; and generating the target language sentence text according to the plurality of search words corresponding to each query vector respectively.
[0006] Optionally, the target language sentence text is generated according to the plurality of search words corresponding to each query vector respectively, including: generating a target prompt word according to the plurality of target language words corresponding to each query vector and the query text corresponding to each query vector; and analyzing the target prompt word by using a large language model to generate the target language sentence text.
[0007] Optionally, the target language voice is generated according to the target language text and the target language sentence text, including: performing normalization processing on the target language text and the target language sentence text to obtain processed text; marking a pronunciation rule in the processed text to obtain marked text; finding a voice segment related to the marked text from a voice data set of the standard general language as reference audio; extracting the voice feature from the reference audio, wherein the voice feature at least includes tone and intonation; and analyzing the voice feature and the marked text by using a voice generation model to generate the target language voice.
[0008] Optionally, the voice generation model is used to analyze the voice feature and the marked text to generate the target language voice, including: analyzing the voice feature and the marked text by using the voice generation model to generate a preliminary voice; eliminating a mute segment with a time length greater than a preset threshold in the preliminary voice, and adjusting the volume of all time periods in the preliminary voice to be consistent to obtain an adjusted voice; and detecting the adjusted voice by using a voice evaluation model, and determining the adjusted voice as the target language voice in a case where the detection is passed.
[0009] Optionally, the adjusted speech is detected by using a speech evaluation model, including: detecting the adjusted speech by using a speech evaluation model to obtain a naturalness score and a fluency score of the adjusted speech; in a case where the naturalness score is greater than a first threshold and the fluency score is greater than a second threshold, converting the adjusted speech into to-be-verified text; comparing the to-be-verified text with the processed text to obtain a character error rate; and in a case where the character error rate is lower than a preset error rate threshold, determining that the adjusted speech passes the detection.
[0010] Optionally, the method further includes: in a case where the adjusted speech fails the detection, generating a backup speech according to the speech features and the labeled text by using a backup speech generation model; and detecting the backup speech by using a speech evaluation model, and in a case where the detection passes, determining that the backup speech is the target language speech.
[0011] According to another aspect of the embodiments of the present application, a speech data set generation apparatus is further provided, including: an acquisition module configured to acquire a speech data set of a standard universal language, and convert the speech data set of the standard universal language into a target language text by using a large language model; a generation module configured to generate a target language sentence text in a retrieval enhancement generation manner; and a construction module configured to generate a target language speech according to the target language text and the target language sentence text, and construct a target speech data set according to the target language speech, wherein a speech feature of the target language speech is consistent with a speech feature of the standard universal language speech data set.
[0012] According to another aspect of the embodiments of the present application, a computer device is further provided, including: a memory and a processor, wherein the memory is configured to store program instructions; and the processor, connected with the memory, is configured to execute the speech data set generation method.
[0013] According to another aspect of the embodiments of the present application, a computer program product is further provided, including computer instructions, which, when executed by a processor, implement the speech data set generation method.
[0014] In the embodiment of the present application, a speech data set of a standard general language is obtained, and a large language model is used to convert the speech data set of the standard general language into a target language text; a target language sentence text is generated in a retrieval enhancement generation manner; a target language speech is generated according to the target language text and the target language sentence text, and a target speech data set is constructed according to the target language speech, wherein the speech features of the target language speech are consistent with the speech features of the standard general language speech data set, and the target speech data set is constructed by the target language speech converted from the directly converted target language text and the target language text generated in the retrieval enhancement generation manner, so as to achieve the purpose of expanding the target language speech data set, thereby realizing the technical effect of improving the target language speech data amount, and further solving the technical problem that the accuracy of the translation model in translating the target language is low due to the small amount of speech data of the target language speech database in the related art. BRIEF DESCRIPTION OF DRAWINGS
[0015] The accompanying drawings, which are included to provide a further understanding of the present application, constitute a part of this application and help to explain the present application together with the specification. The illustrative embodiments of the present application and their description serve to explain the present application. In the drawings:
[0016] Figure 1 FIG. 1 is a hardware structure block diagram of a computer terminal for implementing a speech data set generation method according to an embodiment of the present application;
[0017] Figure 2 FIG. 2 is a flowchart of a speech data set generation method according to an embodiment of the present application;
[0018] Figure 3 FIG. 3 is a flowchart of another speech data set generation method according to an embodiment of the present application;
[0019] Figure 4 FIG. 4 is a method flowchart for generating a target language sentence based on a retrieval enhancement generation manner according to an embodiment of the present application;
[0020] Figure 5 FIG. 5 is a flowchart of a target language speech generation method according to an embodiment of the present application;
[0021] Figure 6 FIG. 6 is a structural diagram of a speech data set generation device according to an embodiment of the present application. DETAILED DESCRIPTION
[0022] In the following, the technical solutions in the embodiments of the present application will be described clearly and completely in conjunction with the drawings in the embodiments of the present application, so that those skilled in the art can better understand the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, but not all the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative work should be within the scope of protection of the present application.
[0023] It should be noted that the terms "first", "second", and the like in the specification and claims of the present application and the above-described drawings are used to distinguish similar objects, and do not necessarily indicate a specific order or a chronological sequence. It should be understood that the data thus used can be interchanged under appropriate circumstances, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion, for example, a process, method, system, product or device including a series of steps or units does not have to be limited to those steps or units clearly listed, but can include other steps or units not clearly listed or inherent to the process, method, product or device.
[0024] The information collected by the embodiments of the present application is information and data authorized by the user or authorized by all parties, and the collection, storage, use, processing, transmission, provision, disclosure and application of related data comply with relevant laws, regulations and standards in the relevant region, necessary security measures are taken, public order and good customs are not violated, and appropriate operation portals are provided for the user to choose authorization or refuse automatic decision results; if the user chooses to refuse, the expert decision process is entered.
[0025] In order to solve the problems in the related art, the embodiments of the present application provide a method for generating a speech data set, which can be run in Figure 1 The computer terminal is explained and described as follows.
[0026] The method for generating a speech data set provided by the embodiments of the present application can be executed in a mobile terminal, a computer terminal or a similar computing device. Figure 1 A hardware structure block diagram of a computer terminal for implementing the method for generating a speech data set is shown. As Figure 1As shown, the computer terminal 10 can include one or more processors (which can include, but are not limited to, processing devices such as microprocessors (MCU) or programmable logic devices (FPGA)), a memory 104 for storing data, and a transmission module 106 for communication functions connected through wired and / or wireless networks. In addition, it can also include a display, a keyboard, a cursor control device, an input / output interface (I / O interface), a universal serial bus (USB) port (which can be included as one of the ports of the I / O interface), a network interface, a BUS bus. Those skilled in the art can understand that Figure 1 The structure shown is only schematic and does not limit the structure of the above-mentioned electronic device. For example, the computer terminal 10 can include more or fewer components than those shown in Figure 1 or have a different configuration than that shown in Figure 1 .
[0027] It should be noted that the one or more processors and / or other data processing circuits described above can be referred to herein as "data processing circuits" in general. The data processing circuits can be embodied in whole or in part as software, hardware, firmware, or any combination thereof. In addition, the data processing circuits can be a single independent processing module, or any one of the other elements incorporated into the computer terminal 10 in whole or in part. As referred to in the embodiments of the present application, the data processing circuit serves as a processor to control, for example, the selection of the variable resistance terminal path connected to the interface.
[0028] The memory 104 can be used to store software programs and modules of application software, such as program instructions / data storage means corresponding to the voice data set generation method in the embodiments of the present application. The processor executes various functional applications and data processing by running the software programs and modules stored in the memory 104, i.e. implements the voice data set generation method described above. The memory 104 can include a high-speed random access memory, and can also include a non-volatile memory such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some examples, the memory 104 can further include a memory remotely located with respect to the processor, which can be connected to the computer terminal 10 through a network. Examples of the above-mentioned network include but are not limited to the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0029] The transmission module 106 is configured to receive or send data via a network. The network can include, for example, a wireless network provided by a communication provider of the computer terminal 10. In one example, the transmission module 106 includes a network interface controller (NIC) that can connect to other network devices through a base station to communicate with the Internet. In one example, the transmission module 106 can be a radio frequency (RF) module that is configured to communicate with the Internet wirelessly.
[0030] The display can be, for example, a touch screen liquid crystal display (LCD) that enables a user to interact with the user interface of the computer terminal 10.
[0031] It is noted that in some alternative embodiments, the above-described Figure 1 The computer terminal can include hardware elements (including circuitry), software elements (including computer code stored on a computer readable medium), or a combination of both hardware and software elements. It should be noted that in some embodiments, the functions of the computer terminal described above can be provided by one or more of the computer terminals described above. Figure 1 is merely one example of a particular implementation and is intended to illustrate the types of components that can be present in the computer terminal described above.
[0032] In the above-described operating environment, embodiments of the present disclosure provide a method for generating a voice data set. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described herein can be executed in an order different from that shown herein.
[0033] Figure 2 is a flowchart of a method for generating a voice data set according to an embodiment of the present disclosure, as shown in Figure 2 The method includes the following steps:
[0034] In step S202, a voice data set of a standard universal language is obtained, and a large language model is used to convert the voice data set of the standard universal language into a target language text.
[0035] In step S202, the original text of the standard general language voice dataset is extracted from the standard general language voice dataset as the input source of the large language model, and the pre-trained LLM (Large Language Model) is used for direct translation from the standard general language to the target language, for example: translating Mandarin to Cantonese. During translation, specific prompts can be designed to emphasize the use of Cantonese spoken language expressions and localized vocabulary, avoiding mechanical character conversion, for example: "Translate the following Mandarin (standard general language) text into colloquial Cantonese (target language), use local common expressions, avoid direct word-by-word translation", and finally the translated results are preliminarily screened to eliminate sentences that do not obviously conform to the Cantonese expression habits. LLM is a super large deep learning model pre-trained based on a large amount of data. The underlying converter is a set of neural networks composed of encoders and decoders with self-attention functions. The encoder and decoder extract meaning from a series of texts and understand the relationship between the words and phrases in them.
[0036] In step S204, the target language generates target language sentence text in the RAG (Retrieval-Augmented Generation) mode; wherein RAG refers to optimizing the output of the large language model so that it can refer to an authoritative knowledge base outside the training data source before generating the final response. LLM is trained with massive data, containing billions of parameters, which can generate original output for knowledge question and answer, language translation and content generation tasks. Based on the powerful functions of LLM, RAG can use the internal knowledge base of a specific field or organization without retraining the model. This is an economical and efficient way to improve the output of LLM, so that it can maintain relevance, accuracy and practicality in various situations.
[0037] It should be noted that the target language text generated in step S202 is not enhanced by retrieval, but is directly converted by a large language model, while the target language sentence text generated in step S204 is generated in the retrieval enhancement mode.
[0038] In step S204, the pre-set target language database can be constructed by: obtaining a plurality of target language vocabularies from a plurality of data sources, and combining the obtained plurality of target language vocabularies with the vocabularies in the public target language-standard general language dictionary to form the pre-set target language database.
[0039] In step S206, the target language text and the target language sentence text are used to generate a target language voice, and the target language voice is used to construct a target voice dataset, wherein the voice features of the target language voice are consistent with the voice features of the standard general language voice dataset.
[0040] It should be further noted that the target language includes multiple languages, such as Cantonese and other languages.
[0041] Through the above steps S202 to S206, the speech data set of the standard general language is obtained, and the speech data set of the standard general language is converted into the target language text by using a large language model; the target language sentence text is generated in a retrieval enhancement generation manner; the target language speech is generated according to the target language text and the target language sentence text, and the target language speech data set is constructed according to the target language speech, wherein the speech features of the target language speech are consistent with the speech features of the standard general language speech data set, and the target language speech data set is constructed by directly converting the target language text and the target language text generated in the retrieval enhancement generation manner, thereby achieving the purpose of expanding the target language speech data set, realizing the technical effect of improving the target language speech data amount, and further solving the technical problem that the accuracy of the translation model in translating the target language is low due to the small amount of speech data in the target language speech database. The following will be described in detail.
[0042] In some embodiments of the present application, the target language sentence text is generated according to the search words, including: converting the vocabulary in the preset target language database into a dense vector, wherein the preset target language database contains a plurality of target language vocabulary; obtaining a plurality of sentence topics, and generating a query text according to each sentence topic respectively to obtain a plurality of query texts; converting the plurality of query texts into a plurality of query vectors; comparing each query vector with the dense vector in the preset target language database respectively, screening out a plurality of dense vectors related to each query vector, and obtaining a plurality of search words corresponding to the plurality of dense vectors related to each query vector; and generating the target language sentence text according to the plurality of search words corresponding to each query vector respectively.
[0043] Specifically, the natural speech processing model can be used to convert the vocabulary in the preset target language database into a dense vector for semantic retrieval.
[0044] Through the specific scene corresponding to the sentence topic, such as dinner chat and sentence structure, such as interrogative sentence and declarative sentence, the most relevant Cantonese vocabulary and phrases are retrieved from the knowledge base according to the relevance and diversity.
[0045] The specific steps of generating the target language sentence text according to the plurality of search terms corresponding to each query vector include: generating a target prompt word according to the plurality of target language vocabularies corresponding to each query vector and the query text corresponding to each query vector; and analyzing the target prompt word by using a large language model to generate the target language sentence text.
[0046] Specifically, the search results are combined with the original query to generate an enhanced prompt word (target prompt word) which is input into the large language model. An example prompt word is as follows: "Generate a target language sentence in accordance with the content in the following preset database (search results: {search term}) in a daily conversation style". Finally, the large language model generates diversified text in accordance with the expression habits of the target language based on the enhanced prompt word. Post-processing is performed on the generated results to ensure that the sentences are fluent and in accordance with the grammar of the target language.
[0047] The above method extracts original text in a standard general language from a voice data set of the standard general language, uses a pre-trained large language model to directly translate the standard general language into the target language, and designs specific prompt words to emphasize oral expression and localized vocabulary in the target language, thereby avoiding mechanical character conversion, effectively improving the accuracy of translation of the target language text, and making the generated target language text more in line with the actual usage habits of Cantonese. In the search enhancement generation stage, the target language vocabulary knowledge base is constructed based on the target language to standard general language dictionary and supplementary vocabulary, combined with the model for semantic search, the query vector is generated according to the specific scene and sentence structure, the related target language vocabulary and phrases are retrieved from the knowledge base and sorted, combined with the original query to generate an enhanced prompt word input into the large language model, and finally diversified text in accordance with the expression habits of the target language is generated, thereby enriching the expression form of the target language text. In the final stage of text generation, post-processing is performed on the generated results to ensure that the sentences are fluent and in accordance with the grammar of the target language, thereby further improving the quality and readability of the target language text. The entire text generation process is processed and optimized through multiple stages, so that the generated target language text is closer to the actual application scenarios, such as different scenes (such as "dinner chat") and sentence structures (such as interrogative sentences and declarative sentences) in daily conversations, thereby improving the practicality and applicability of the text in actual communication.
[0048] In some embodiments of the present application, the specific steps of generating the target language speech according to the target language text and the target language sentence text include: performing normalization processing on the target language text and the target language sentence text to obtain processed text; marking pronunciation rules in the processed text to obtain marked text; finding a speech segment related to the marked text from the speech data set of the standard universal language as reference audio; extracting the speech features from the reference audio, wherein the speech features at least include tone and intonation; and using a speech generation model to analyze the speech features and the marked text to generate the target language speech.
[0049] In actual application scenarios, the target language text and the target language sentence text output by the text generation stage are standardized, including text normalization (such as number reading method, punctuation processing), target language specific pronunciation marking (such as entering sound words, tone change rules). The speech segment in the speech data set of the standard universal language is used as reference audio, and the speech features (such as tone and intonation) of the speaker are extracted through voiceprint coding. The reference audio features and the marked text are input into the speech generation model to ensure that the synthesized speech is consistent with the original speaker features.
[0050] In some embodiments of the present application, the specific steps of using a speech generation model to analyze the speech features and the marked text to generate the target language speech are as follows: using the speech generation model to analyze the speech features and the marked text to generate a preliminary speech; eliminating the silent segments in the preliminary speech with a time length greater than a preset threshold, and adjusting the volume of all time periods in the preliminary speech to be consistent to obtain an adjusted speech; and using a speech evaluation model to detect the adjusted speech, and determining that the adjusted speech is the target language speech if the detection is passed.
[0051] The speech synthesis process is to input the marked text into the speech generation model to generate a preliminary speech waveform. Basic quality detection is performed on the synthesized speech, including: silent segment detection (to avoid abnormal pauses), and volume equalization (to ensure volume consistency).
[0052] In some embodiments of the present application, the detection of the adjusted speech using a speech evaluation model includes: using a speech evaluation model to detect the adjusted speech to obtain a naturalness score and a fluency score of the adjusted speech; converting the adjusted speech into a to-be-verified text if the naturalness score is greater than a first threshold and the fluency score is greater than a second threshold; comparing the to-be-verified text with the processed text to obtain a character error rate; and determining that the adjusted speech passes the detection if the character error rate is lower than a preset error rate threshold.
[0053] For example, use the DNSMOS model (a speech evaluation model) to evaluate the naturalness and fluency of the speech (threshold > 3.0); calculate the character error rate (CER) of the synthesized speech and the original text through the speech recognition model, and set the threshold to < 15%. Then, artificial verification is performed, and the screened speech data is randomly sampled for listening, focusing on checking the accuracy of Cantonese pronunciation (such as avoiding "lazy sound") and the naturalness of emotional expression (such as the tone of interrogative sentences).
[0054] By the above-mentioned way, the speech quality is accurately evaluated: the naturalness and fluency of the speech are quantitatively evaluated, and the threshold > 3.0 is used as the standard to accurately screen the speech data with high naturalness and fluency, effectively distinguishing the quality of speech synthesis. Ensure the consistency of speech and text: use the speech recognition model to calculate the character error rate (CER) of the synthesized speech and the original text, and set the threshold to < 15%, which can accurately measure the matching degree of speech and text, ensure that the generated speech accurately conveys the original text information, and reduce the speech recognition bias. Improve the accuracy of speech and emotional expression: through artificial verification, randomly sample the screened speech data for listening, and focus on the accuracy of Cantonese pronunciation, avoiding errors such as "lazy sound", etc., to ensure that the speech conforms to the Cantonese pronunciation standard; at the same time, check the naturalness of emotional expression, such as whether the tone of interrogative sentences is appropriate, so that the generated speech is more realistic and infectious, improving the intelligibility and acceptance of speech in actual communication. Build a multi-level evaluation system: combine automatic evaluation model and artificial verification to form a multi-level and comprehensive speech quality evaluation system, fully utilize the efficiency of automatic evaluation and the meticulousness of artificial verification, and improve the overall quality of speech data, providing reliable basis for subsequent optimization and improvement of speech synthesis.
[0055] It should be noted that the naturalness score is mainly used to evaluate whether the synthesized speech sounds similar to a real person speaking, including sound quality, tone, pronunciation accuracy, emotional expression, etc. It can be evaluated from the following aspects, for example: sound quality: evaluate the quality of the audio signal to see if it is close to the clarity and delicacy of natural speech. Tone and rhythm: check the prosodic characteristics of the speech, including the rise and fall of pitch, changes in speech rate, and natural distribution of pauses. Pronunciation accuracy: evaluate the accuracy and fluency of pronunciation. Emotional expression: measure whether emotions and tone can be appropriately conveyed in speech.
[0056] The flow degree score is mainly used for the coherence and fluency of the speech, mainly investigating whether the transition between pronunciation is smooth and whether the speech flow of the whole sentence or paragraph is natural and coherent. For example: pronunciation transition: check whether the connection between syllables and words is natural, and avoid the common phenomenon of discontinuity or abruptness in synthetic speech. Sentence coherence: assess whether the connection between sentences in a paragraph is smooth, and whether it can form a logically complete and comfortable listening experience. Pause and breathing: consider whether the selection of pauses and breathing points is appropriate, which directly relates to the natural fluency and listening experience of the speech.
[0057] In the case where the adjusted speech detection fails, a backup speech generation model is used to generate backup speech according to the speech features and the labeled text; a speech evaluation model is used to detect the backup speech, and in the case where the detection passes, the backup speech is determined as the target language speech.
[0058] If the speech quality output by Cosyvoice2 (a speech generation model) does not meet the standard (such as mechanical sound, sentence error), switch to GPT-SovITS model (a speech synthesis model integrated with a transformer-based pre-training model, a speech cloning and conversion system, a backup speech generation model) to re-synthesize, and finally select the optimal synthesis result through a multi-model voting mechanism.
[0059] The voice data set generation method provided by the embodiment of the application generates natural and fluent target language voice data close to real human pronunciation based on the target language text output by the text generation stage, and improves the quality and availability of target language voice synthesis. The speaker feature consistency is maintained: by using the voice segments in the voice data set of the standard general language as reference audio, the speaker features (such as tone and intonation) are extracted, and these features and the annotated text (target language) are input into the voice generation model, to ensure that the synthesized target language voice can maintain consistency with the original speaker's voice features, and to enhance the personalization and authenticity of voice synthesis. The primary synthesizer uses the Cosyvoice2 model with excellent naturalness and timbre fidelity, and is equipped with a backup synthesizer GPT-SovITS model, which can ensure the stability of the synthesis process in long sentence or complex intonation scenarios, reduce the risk of voice synthesis failure or abnormality, and improve the reliability of voice synthesis. The target language text output by the text generation stage is standardized, including text normalization and Cantonese-specific pronunciation annotation, so that the text input into the voice generation model is more standardized and accurate, which helps to generate voice that conforms to the Cantonese pronunciation rules and language habits, and further improves the accuracy and naturalness of voice synthesis. A multi-model collaboration mechanism is designed, the primary synthesizer model is used to generate preliminary voice waveform, and basic quality detection is performed, if the quality does not meet the standard, the backup model is switched to re-synthesize, and finally the optimal synthesis result is selected through a multi-model voting mechanism, effectively ensuring the high quality of the final output target language voice data, and providing accurate and high-quality voice synthesis services for users.
[0060] Figure 3 Another voice data set generation method is shown, as shown in Figure 3 , including: taking Mandarin and Cantonese as examples, using a large language model to translate Mandarin text into Cantonese text, additionally generating Cantonese text based on a retrieval enhancement generation method, and finally, synthesizing the Cantonese text generated by the two methods, testing the synthesized voice, and constructing a voice data set after passing the test.
[0061] Figure 4 A method for generating target language sentences based on a retrieval enhancement generation method is shown, as shown in Figure 4 , a query vector (question) is generated based on the theme of the sentence, relevant documents are retrieved, and the generated target language sentences are obtained by combining the retrieved documents with the questions.
[0062] Figure 5 A voice generation method flowchart is shown, as shown in Figure 5As shown, the method comprises: generating target language voice based on the labeled text and the reference voice by a main synthesizer, evaluating the generated target language voice by using a voice evaluation model, directly outputting in the case that the evaluation result is passed, and generating new target language voice by using a backup voice generation model in the case that the evaluation result is not passed.
[0063] The method for generating a voice data set proposed in the embodiments of the present application generates a data set, and the embodiments of the present application also provide comparative experiment results with other public data sets:
[0064] Experimental data grouping: control group: CVHK: using a public data set as a benchmark control; DTDT: a Cantonese data set generated only by direct translation. Experimental group: RFDT: a Cantonese data set generated by using retrieval enhancement to generate translation; AUDT: a Cantonese data set constructed by using an automatic generation method. Test set division: ASR test set: AUDT_test (first test set): 500 sentences randomly extracted from AUDT; CVHK_test (second test set): 500 sentences randomly extracted from CVHK. S2TT (Speech-to-Text Translation, speech-to-text translation) test set: 800 samples of the MDT-ASR-D001 data set are used. Baseline model selection in model training: an open source speech recognition model is used as a basic model, and the multilingual recognition and translation capabilities of the original model are maintained. Fine-tuning settings: the training period is 5 rounds, the batch size is 16, and the learning rate is 5e-6. Comparative experiment design: the model is fine-tuned using the DTDT, RFDT and AUDT data sets respectively; experiment 2: the model is fine-tuned using the RFDT+AUDT data sets; experiment 3: performance comparison under different data volumes (20h / 50h / 80h). Evaluation index: the character error rate is used for the speech recognition task, and different evaluation indexes are used for the speech translation task, for example: BLEU-1 / BLEU-2, SacreBLEU, BERTScore.
[0065] The experimental results show that the AUDT data set generated by using the RAG method has a CER of 18.7% in the ASR task, close to the effect of real data; the mixed data set (RFDT+AUDT) has a BERTScore of 0.77 in the translation task, significantly better than a single data set; the data volume experiment shows that the performance improves with the increase of data size, and tends to be stable when the data volume is 80h.
[0066] Table 1 shows the ASR performance comparison (CER / %). As shown in Table 1, the CER of the AUDT data set in the ASR task is reduced to 18.7%, close to the effect of real data.
[0067] Table 1
[0068]
[0069] Table 2 shows the S2TT performance comparison, as shown in Table 2, the mixed data set (RFDT+AUDT) reaches 0.77 in the translation task BERTScore, which is significantly better than the single data set, and it can be understood that the mixed data set is the data set generated by the method of the application.
[0070] Table 2
[0071]
[0072] Figure 6 A voice data set generation device is shown, which comprises:
[0073] The acquisition module 60 is configured to acquire a voice data set of a standard universal language, and convert the voice data set of the standard universal language into a target language text by using a large language model.
[0074] The generation module 62 is configured to generate a target language sentence text in a retrieval enhancement generation manner.
[0075] The construction module 64 is configured to generate a target language voice according to the target language text and the target language sentence text, and construct a target voice data set according to the target language voice, wherein the voice feature of the target language voice is consistent with the voice feature of the standard universal language voice data set.
[0076] The voice data set generation device described above acquires a voice data set of a standard universal language, and converts the voice data set of the standard universal language into a target language text by using a large language model; generates a target language sentence text in a retrieval enhancement generation manner; generates a target language voice according to the target language text and the target language sentence text, and constructs a target voice data set according to the target language voice, wherein the voice feature of the target language voice is consistent with the voice feature of the standard universal language voice data set. The target language voice data set is constructed by the target language voice converted from the target language text directly converted and the target language text generated in the retrieval enhancement generation manner, so as to achieve the purpose of expanding the target language voice data set, realize the technical effect of improving the target language voice data amount, and further solve the technical problem that the accuracy of the translation model in translating the target language is low due to the small amount of voice data in the target language voice database in the related art.
[0077] The generating module 62 comprises a generating submodule for generating the target language sentence text in a retrieval enhancement generation manner, including: converting the vocabulary in the preset target language database into dense vectors, wherein the preset target language database contains a plurality of target language vocabularies; obtaining a plurality of sentence topics, and generating a query text according to each sentence topic respectively to obtain a plurality of query texts; converting the plurality of query texts into a plurality of query vectors; comparing each query vector with the dense vectors in the preset target language database respectively, screening out a plurality of dense vectors related to each query vector, and obtaining a plurality of retrieval words corresponding to the plurality of dense vectors related to each query vector; and generating the target language sentence text according to the plurality of retrieval words corresponding to each query vector respectively.
[0078] The generating submodule comprises a generating unit for generating the target language sentence text according to the plurality of retrieval words corresponding to each query vector respectively, including: generating a target prompt word according to the plurality of target language vocabularies corresponding to each query vector and the query text corresponding to each query vector; and analyzing the target prompt word by using a large language model to generate the target language sentence text.
[0079] The constructing module 64 comprises a voice submodule for generating a target language voice according to the target language text and the target language sentence text, including: performing normalization processing on the target language text and the target language sentence text to obtain processed text; labeling pronunciation rules in the processed text to obtain labeled text; finding out a voice segment related to the labeled text from the voice data set of the standard general language as reference audio; extracting the voice features from the reference audio, wherein the voice features at least include tone and intonation; and analyzing the voice features and the labeled text by using a voice generation model to generate the target language voice.
[0080] The voice submodule comprises a voice unit for analyzing the voice features and the labeled text by using a voice generation model to generate the target language voice, including: analyzing the voice features and the labeled text by using the voice generation model to generate a preliminary voice; eliminating the mute segments with a time length greater than a preset threshold in the preliminary voice, and adjusting the volume of all time periods in the preliminary voice to be consistent to obtain an adjusted voice; and detecting the adjusted voice by using a voice evaluation model, and determining the adjusted voice as the target language voice in the case of passing the detection.
[0081] The voice unit comprises a detection subunit configured to detect the adjusted voice by using a voice evaluation model, including: detecting the adjusted voice by using a voice evaluation model to obtain a naturalness score and a fluency score of the adjusted voice; in a case where the naturalness score is greater than a first threshold and the fluency score is greater than a second threshold, converting the adjusted voice into a text to be verified; comparing the text to be verified with the processed text to obtain a character error rate; and in a case where the character error rate is lower than a preset error rate threshold, determining that the adjusted voice detection is passed.
[0082] The voice unit further comprises an alternative subunit configured to, in a case where the adjusted voice detection is not passed, generate an alternative voice according to the voice feature and the labeled text by using an alternative voice generation model; and detect the alternative voice by using a voice evaluation model, and in a case where the detection is passed, determine that the alternative voice is the target language voice.
[0083] It should be noted that, Figure 6 The voice data set generation apparatus shown is configured to execute the voice data set generation method shown. Figure 2 The voice data set generation method shown, and therefore the above related explanations and descriptions of the voice data set generation method also apply to the voice data set generation apparatus, which will not be described here again.
[0084] The embodiments of the present application also provide a computer device, comprising: a memory and a processor, wherein the memory is configured to store program instructions; the processor is connected with the memory and is configured to execute the voice data set generation method described above.
[0085] The embodiments of the present application also provide a computer program product comprising computer instructions, which, when executed by a processor, implement the steps of the voice data set generation method in the present application.
[0086] The above-mentioned serial numbers of the embodiments of the present application only serve for description, and do not represent the advantages or disadvantages of the embodiments.
[0087] In the above-mentioned embodiments of the present application, the description of each embodiment has its own emphasis, and the parts not described in detail in a certain embodiment can be referred to the relevant description of other embodiments.
[0088] In several embodiments provided in the present application, it should be understood that the disclosed technology can be implemented by other ways. Among them, the above-described device embodiments are only schematic, for example, the division of the units can be a logical function division, and actual implementation can have another division manner, for example, a plurality of units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the displayed or discussed units can be indirect coupling or communication connection through some interfaces, units or modules, and can be electrical or other forms.
[0089] The units described as separate components can or can not be physically separate, and the components shown as units can or can not be physical units, that is, they can be located in one place, or they can be distributed to multiple units. Part or all of the units can be selected according to actual needs to achieve the purpose of the embodiment.
[0090] In addition, the functional units in each embodiment of the present application can be integrated in one processing unit, or each unit can be physically present separately, or two or more units can be integrated in one unit. The integrated unit can be realized in the form of hardware or in the form of a software functional unit.
[0091] The integrated unit, if realized in the form of a software functional unit and sold or used as an independent product, can be stored in a computer readable storage medium. Based on this understanding, the technical solutions of the present application essentially or the part that contributes to the prior art or the whole or part of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium, including a plurality of instructions for causing a computer device (which can be a personal computer, a server or a network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The foregoing storage medium includes: a U disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a mobile hard disk, a magnetic disk or an optical disk, and various media that can store program codes.
[0092] The above is only the preferred embodiment of the present application, and it should be pointed out that for ordinary skilled in the art, without departing from the principles of the present application, a number of improvements and refinements can be made, and these improvements and refinements should be considered as the protection scope of the present application.
Claims
1. A method for generating a speech dataset, characterized in that, include: Obtain a speech dataset of a standard general language and use a large language model to convert the speech dataset of the standard general language into text in the target language; The target language sentence text is generated using a search-enhanced generation method. The target language speech is generated based on the target language text and the target language sentence text, and a target language speech dataset is constructed based on the target language speech, wherein the speech features of the target language speech are consistent with the speech features of the standard general language speech dataset; Generating target language speech based on the target language text and the target language sentence text includes: The target language text and the target language sentence text are normalized to obtain the processed text. The pronunciation rules are annotated in the processed text to obtain the annotated text; Identify the speech segments related to the annotated text from the speech dataset of the standard general language, and use them as reference audio; The speech features are extracted from the reference audio, wherein the speech features include at least: timbre and intonation; The speech features and the annotated text are analyzed using a speech generation model to generate the speech in the target language.
2. The method according to claim 1, characterized in that, The target language sentence text is generated using a search-enhanced generation method, including: The vocabulary in the preset target language database is converted into dense vectors, wherein the preset target language database contains multiple target language vocabulary; Obtain multiple statement topics, and generate query text based on each statement topic to obtain multiple query texts; Convert the multiple query texts into multiple query vectors; Each query vector is compared with the dense vectors in the preset target language database to filter out multiple dense vectors related to each query vector, and multiple search terms corresponding to the multiple dense vectors related to each query vector are obtained. The target language sentence text is generated based on the multiple search terms corresponding to each query vector.
3. The method according to claim 2, characterized in that, Generate the target language sentence text based on multiple search terms corresponding to each query vector, including: Target prompt words are generated based on the multiple target language words corresponding to each query vector and the query text corresponding to each query vector; The target prompt words are analyzed using a large language model to generate the target language sentence text.
4. The method according to claim 1, characterized in that, The speech features and the annotated text are analyzed using a speech generation model to generate the target language speech, including: The speech generation model is used to analyze the speech features and the annotated text to generate preliminary speech. Eliminate silent segments in the initial speech that are longer than a preset threshold, and adjust the volume of all time periods in the initial speech to be consistent to obtain the adjusted speech; The adjusted speech is detected using a speech evaluation model. If the detection is successful, the adjusted speech is determined to be the target language speech.
5. The method according to claim 4, characterized in that, The adjusted speech is detected using a speech evaluation model, including: The adjusted speech is detected using a speech evaluation model to obtain a naturalness score and a fluency score for the adjusted speech; If the naturalness score is greater than the first threshold and the fluency score is greater than the second threshold, the adjusted speech is converted into text to be verified. The text to be verified is compared with the processed text to obtain the character error rate; If the character error rate is lower than a preset error rate threshold, the adjusted speech detection is deemed to have passed.
6. The method according to claim 5, characterized in that, The method further includes: If the adjusted speech detection fails, an alternative speech generation model is used to generate alternative speech based on the speech features and the annotated text. The backup speech is detected using a speech evaluation model. If the detection passes, the backup speech is determined to be the target language speech.
7. A device for generating a speech dataset, characterized in that, include: The acquisition module is used to acquire a speech dataset of a standard general language and use a large language model to convert the speech dataset of the standard general language into text in the target language. The generation module is used to generate target language sentence text using a search-enhanced generation method. A construction module is used to generate target language speech based on the target language text and the target language sentence text, and to construct a target language speech dataset based on the target language speech, wherein the speech features of the target language speech are consistent with the speech features of the standard general language speech dataset; Generating target language speech based on the target language text and the target language sentence text includes: The target language text and the target language sentence text are normalized to obtain the processed text. The pronunciation rules are annotated in the processed text to obtain the annotated text; Identify the speech segments related to the annotated text from the speech dataset of the standard general language, and use them as reference audio; The speech features are extracted from the reference audio, wherein the speech features include at least: timbre and intonation; The speech features and the annotated text are analyzed using a speech generation model to generate the speech in the target language.
8. A computer device, characterized in that, include: A memory and a processor, wherein the memory is used to store program instructions; The processor, connected to the memory, is used to execute the method for generating a speech dataset according to any one of claims 1 to 6.
9. A computer program product comprising computer instructions, characterized in that, When the computer instructions are executed by the processor, they implement the method for generating the speech dataset according to any one of claims 1 to 6.
Citation Information
Patent Citations
Multi-dialect accent mandarin voice recognition model training method and device, and equipment
CN112233653A
Dialect speech recognition method and device based on transfer learning
CN112885351A