A speech synthesis system and method independent of a pronunciation dictionary
The pronunciation representation is extracted through the language-independent speech recognition model, and a pronunciation synthesis system that does not rely on pronunciation dictionary is constructed, and a pronunciation waveform is directly generated from text character sequences, solving the problem of time-consuming and labor-consuming pronunciation dictionary and improving the naturalness and intelligibility of pronunciation synthesis.
Patent Information
- Application Number
- CN202210177013.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-02-24
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2042-02-24
AI Technical Summary
现有语音合成系统依赖于语言专家知识构建发音词典,耗时耗力,且直接将文本字符序列输入后端声学模型会降低合成语音的自然度和可懂度。
The language-independent speech recognition model is used to extract pronunciation representations from the speech waveforms of the target language, and the text-pronunciation representation prediction model and pronunciation representation-acoustic prediction model are trained to directly predict pronunciation representations from text character sequences and generate pronunciation waveforms, avoiding the use of pronunciation dictionary.
It solves the problem of establishing pronunciation dictionary in multilingual pronunciation synthesis system, reduces pronunciation errors, and improves the naturalness and intelligibility of synthetic pronunciation.
Smart Images

Figure CN114495897B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of speech signal processing, and in particular, to a speech synthesis system and method. Background Art
[0002] Speech synthesis aims to enable a machine to speak as smoothly and naturally as a human being, and it has a wide range of applications, such as voice assistants and audiobooks. A speech synthesis system generally consists of two parts: a front end and a back end. The front end focuses on text analysis, which converts a text sequence into linguistic features, and it has a series of functions, such as text normalization, grapheme-to-phoneme conversion, word segmentation, part-of-speech tagging, and prosody prediction, etc. The purpose of grapheme-to-phoneme conversion is to generate a phoneme sequence from a character sequence. A pronunciation dictionary consists of word-pronunciation pairs of a language and is crucial for grapheme-to-phoneme conversion. Since it is impossible for a pronunciation dictionary to cover all the words of a language, a grapheme-to-phoneme conversion model trained on the pronunciation dictionary is usually adopted, which can generate the pronunciations of words not existing in the dictionary. However, a pronunciation dictionary is language-specific, and building a pronunciation dictionary for a new language requires expertise in language and phoneme annotation systems, which is more laborious, time-consuming, and difficult than obtaining the speech recordings of that language. Even though there are some open-source grapheme-to-phoneme conversion tools, considering that there are approximately 7,000 languages globally, the number of languages they cover is still very limited.
[0003] On the other hand, the back end of a speech synthesis system usually consists of an acoustic model that converts language features into acoustic features and a vocoder that reconstructs a speech waveform from the acoustic features. In recent years, neural network-based sequence-to-sequence acoustic modeling has become the mainstream method, which has better performance compared with traditional statistical parametric speech synthesis methods based on hidden Markov models and deep neural networks. Some sequence-to-sequence acoustic models can directly take a character sequence as input, so there is no longer a need for a pronunciation dictionary and a grapheme-to-phoneme conversion model. However, using a character sequence as input generally reduces the naturalness and intelligibility of the synthesized speech compared with using a phoneme sequence.
[0004] A traditional speech synthesis system needs to use a pronunciation dictionary and a grapheme-to-phoneme conversion model to process the input text into a phoneme sequence in the front-end text analysis stage and then send it to the back-end module for acoustic feature prediction and waveform reconstruction. The grapheme-to-phoneme conversion model is usually also trained based on the pronunciation dictionary. The establishment of a pronunciation dictionary relies on language expert knowledge related to the language, and it is time-consuming and laborious to establish a large-capacity and high-precision pronunciation dictionary. However, directly inputting the text character sequence into the back-end acoustic model will reduce the quality of the synthesized speech.
[0005] In view of this, the present invention is specifically proposed. Summary of the Invention
[0006] The object of the present invention is to provide a speech synthesis system and method that do not rely on a pronunciation dictionary, which can perform speech synthesis without relying on a pronunciation dictionary, thereby solving the above-mentioned technical problems existing in the prior art.
[0007] The object of the present invention is achieved by the following technical solutions:
[0008] An embodiment of the present invention provides a speech synthesis system that does not rely on a pronunciation dictionary, including:
[0009] A language-independent speech recognition model, a text-pronunciation representation prediction model, a pronunciation representation-acoustic prediction model, and a neural network vocoder; wherein,
[0010] The language-independent speech recognition model can extract a pronunciation representation from the input speech waveform of the target language during the training phase, and provide the pronunciation representation to the text-pronunciation representation prediction model and the pronunciation representation-acoustic prediction model for training, so as to obtain the trained text-pronunciation representation prediction model and pronunciation representation-acoustic prediction model;
[0011] The text-pronunciation representation prediction model can, after being trained, predict a pronunciation representation according to the character sequence of the input text to be synthesized, and output it to the trained pronunciation representation-acoustic prediction model;
[0012] The pronunciation representation-acoustic prediction model, which is connected to the neural network vocoder, can generate a Mel spectrogram according to the pronunciation representation predicted by the text-pronunciation representation prediction model;
[0013] The neural network vocoder can reconstruct the Mel spectrogram generated by the pronunciation representation-acoustic prediction model into a speech waveform corresponding to the text to be synthesized.
[0014] An embodiment of the present invention also provides a speech synthesis method that does not rely on a pronunciation dictionary. Using the speech synthesis system of the present invention that does not rely on a pronunciation dictionary, first, the language-independent speech recognition model of the speech synthesis system extracts a pronunciation representation from the input speech waveform of the target language, and uses the pronunciation representation to train the text-pronunciation representation prediction model and the pronunciation representation-acoustic prediction model of the speech synthesis system. After the training is completed, the trained text-pronunciation representation prediction model and pronunciation representation-acoustic prediction model are obtained; the synthesis is performed according to the following steps:
[0015] Input the text to be synthesized into the trained text-pronunciation representation prediction model of the speech synthesis system. The text-pronunciation representation prediction model predicts a pronunciation representation according to the character sequence of the text to be synthesized, and outputs it to the pronunciation representation-acoustic prediction model of the speech synthesis system;
[0016] The pronunciation representation - acoustic prediction model predicts and generates a Mel spectrogram based on the pronunciation representation, and outputs the Mel spectrogram to the neural network vocoder of the speech synthesis system;
[0017] The neural network vocoder reconstructs the Mel spectrogram into a speech waveform corresponding to the text to be synthesized.
[0018] Compared with the prior art, the speech synthesis system and method provided by the present invention that do not rely on a pronunciation dictionary have the following beneficial effects:
[0019] By adopting a language - independent automatic speech recognition model, it can automatically extract pronunciation representations from the speech data of the target language, and then use the pronunciation representations to train and construct the text - pronunciation representation prediction model and the pronunciation representation - acoustic prediction model of the speech synthesis system. The constructed speech synthesis system first predicts the pronunciation representation from text characters and then generates speech from the pronunciation representation. This system and method can solve the problem that traditional speech synthesis methods rely on language - related pronunciation dictionaries when constructing a multi - language speech synthesis system, and solve the problem that the establishment of a pronunciation dictionary often requires the participation of language experts, consuming a large amount of manpower and time. Compared with the existing method of directly predicting speech acoustic features from text characters, this method can reduce pronunciation errors in the synthesized speech and improve the naturalness of the synthesized speech. Brief Description of the Drawings
[0020] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0021] Figure 1 It is a schematic diagram of the overall structure of the speech synthesis system that does not rely on a pronunciation dictionary provided by the embodiment of the present invention;
[0022] Figure 2 It is a schematic diagram of the structure of the language - independent speech recognition model of the speech synthesis system that does not rely on a pronunciation dictionary provided by the embodiment of the present invention;
[0023] Figure 3 It is a schematic diagram of the pronunciation representation extraction process of the language - independent speech recognition model of the speech synthesis system that does not rely on a pronunciation dictionary provided by the embodiment of the present invention;
[0024] Figure 4 It is a schematic diagram of the acoustic modeling process of the speech synthesis system that does not rely on a pronunciation dictionary provided by the embodiment of the present invention. Detailed Embodiments
[0025] The following describes the technical solutions in the embodiments of the present invention clearly and completely in combination with the specific content of the present invention; obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments, which does not constitute a limitation to the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the protection scope of the present invention.
[0026] First, the following explanations are given for the terms that may be used in this article:
[0027] The term "and / or" means that either or both of the two can be realized. For example, X and / or Y means that it includes both the case of "X" or "Y" and the three cases of "X and Y".
[0028] The description of terms such as "comprising", "including", "containing", "having" or other similar semantics should be interpreted as non-exclusive inclusion. For example: including a certain technical feature element (such as raw material, component, ingredient, carrier, dosage form, material, size, part, component, mechanism, device, step, process, method, reaction condition, processing condition, parameter, algorithm, signal, data, product or article, etc.) should be interpreted as not only including the clearly listed certain technical feature element, but also including other technical feature elements well-known in the art that are not clearly listed.
[0029] The term "consisting of" means excluding any technical feature element that is not clearly listed. If this term is used in a claim, this term will make the claim a closed type, so that it does not include technical feature elements other than the clearly listed technical feature elements, except for related conventional impurities. If this term only appears in a certain clause of a claim, then it only limits the elements clearly listed in that clause, and the elements recorded in other clauses are not excluded from the overall claim.
[0030] Unless otherwise clearly specified or limited, terms such as "installed", "connected", "joined", "fixed", etc. should be understood in a broad sense. For example: it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be directly connected or indirectly connected through an intermediate medium, and it can be the communication inside two elements. For those of ordinary skill in the art, the specific meanings of the above terms in this article can be understood according to specific situations.
[0031] The terms "center", "longitudinal", "lateral", "length", "width", "thickness", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", "clockwise", "counterclockwise", etc. indicate the orientation or positional relationship based on the orientation or positional relationship shown in the drawings. They are only for convenience of description and simplification of description, rather than explicitly or implicitly indicating that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and thus should not be construed as a limitation to this document.
[0032] The speech synthesis method provided by the present invention that does not rely on a pronunciation dictionary will be described in detail below. The content not described in detail in the embodiments of the present invention belongs to the prior art well-known to those skilled in the art. For the conditions not specified in the embodiments of the present invention, they are carried out according to the conventional conditions in the art or the conditions recommended by the manufacturer. For the reagents or instruments not specified in the embodiments of the present invention for the manufacturer, they are all conventional products that can be obtained by purchasing in the market.
[0033] As Figure 1 shown, the embodiments of the present invention provide a speech synthesis system that does not rely on a pronunciation dictionary, including:
[0034] a language-independent speech recognition model, a text-pronunciation representation prediction model, a pronunciation representation-acoustic prediction model, and a neural network vocoder; wherein,
[0035] the language-independent speech recognition model can extract a pronunciation representation from the input speech waveform of the target language during the training phase, and provide the pronunciation representation to the text-pronunciation representation prediction model and the pronunciation representation-acoustic prediction model for training, so as to obtain the trained text-pronunciation representation prediction model and pronunciation representation-acoustic prediction model;
[0036] the text-pronunciation representation prediction model can predict a pronunciation representation according to the character sequence of the input text to be synthesized after being trained, and output it to the trained pronunciation representation-acoustic prediction model;
[0037] the pronunciation representation-acoustic prediction model is connected to the neural network vocoder, and can generate a Mel spectrogram according to the pronunciation representation predicted by the text-pronunciation representation prediction model;
[0038] the neural network vocoder can reconstruct the Mel spectrogram generated by the pronunciation representation-acoustic prediction model into a speech waveform corresponding to the text to be synthesized.
[0039] See Figure 2 , in the above speech synthesis system, the language-independent speech recognition model includes:
[0040] A sequentially connected wav2vec 2.0 model, a first linear layer, and a second linear layer; wherein,
[0041] The wav2vec 2.0 model uses a wav2vec 2.0 model without a quantization module, and its training input is a multilingual corpus with IPA phoneme transcriptions;
[0042] The first linear layer is a bottleneck layer that can map a 1024-dimensional context representation (C) to a 512-dimensional bottleneck representation (B);
[0043] The second linear layer is a classification layer that can predict the class probability (P) according to the bottleneck representation output by the first linear layer;
[0044] The training objective of this language-independent speech recognition model is the CTC loss between the class probability and the target IPA sequence.
[0045] In the above speech synthesis system, the structure of the text-pronunciation representation prediction model adopts a sequence-to-sequence structure based on Tacotron2;
[0046] The error function for training the text-pronunciation representation prediction model is the mean squared error and mean absolute error between the predicted pronunciation representation and the extracted pronunciation representation, plus the binary cross-entropy of the stop symbol.
[0047] In the above speech synthesis system, the pronunciation representation-acoustic prediction model adopts the structure of Tacotron2;
[0048] The loss function of this pronunciation representation-acoustic prediction model is the mean squared error and mean absolute error between the predicted Mel spectrogram and the true Mel spectrogram, and the binary cross-entropy of the stop symbol.
[0049] An embodiment of the present invention also provides a pronunciation dictionary-independent speech synthesis method. Using the pronunciation dictionary-independent speech synthesis system of the present invention, first, the language-independent speech recognition model of the speech synthesis system extracts the pronunciation representation from the input speech waveform of the target language, and uses the pronunciation representation to train the text-pronunciation representation prediction model and the pronunciation representation-acoustic prediction model of the speech synthesis system. After the training is completed, the trained text-pronunciation representation prediction model and pronunciation representation-acoustic prediction model are obtained; the synthesis is carried out according to the following steps:
[0050] Input the text to be synthesized into the trained text-pronunciation representation prediction model of the speech synthesis system. The text-pronunciation representation prediction model predicts it as a pronunciation representation according to the character sequence of the text to be synthesized and outputs it to the pronunciation representation-acoustic prediction model of the speech synthesis system;
[0051] The pronunciation representation - acoustic prediction model predicts and generates a Mel spectrogram based on the pronunciation representation, and outputs the Mel spectrogram to the neural network vocoder of the speech synthesis system;
[0052] The neural network vocoder reconstructs the Mel spectrogram into a speech waveform corresponding to the text to be synthesized.
[0053] The language - independent speech recognition model extracts pronunciation representations from the input speech waveform of the target language in the following manner, including:
[0054] Calculate the frame - level bottleneck representation B = [b1,..., b T of the input speech waveform of the target language;
[0055] Apply the argmax function to the class probabilities P output by the language - independent speech recognition model to obtain the phoneme symbol class corresponding to each frame;
[0056] Combine the frame - level bottleneck representation B = [b1,..., b T and the phoneme symbol class of each frame to perform a classification operation on the frame - level bottleneck representation, and assign the class of the t - th frame to the bottleneck representation b t ;
[0057] Apply a merging operation to remove the bottleneck representations of the blank class, and merge adjacent bottleneck representations with the same class into a vector R = [r1,..., r N , which is the pronunciation representation. Here, N is the number of pronunciation representations of the speech waveform, and N is less than T, where T is the length of the frame - level bottleneck representation of the speech waveform.
[0058] In summary, in the speech synthesis system and method of the embodiments of the present invention, due to the adoption of a language - independent automatic speech recognition model, pronunciation representations can be automatically extracted from the speech data of the target language, and then the pronunciation representations are used to train and construct the text - pronunciation representation prediction model and the pronunciation representation - acoustic prediction model of the speech synthesis system. The constructed speech synthesis system first predicts the pronunciation representation from text characters and then generates speech from the pronunciation representation. This system and method can solve the problem that traditional speech synthesis methods rely on language - related pronunciation dictionaries when constructing a multi - language speech synthesis system, and solve the problem that the establishment of pronunciation dictionaries often requires the participation of language experts, consuming a large amount of manpower and time. Compared with the existing method of directly predicting speech acoustic features from text characters, this method can reduce pronunciation errors in the synthesized speech and improve the naturalness of the synthesized speech.
[0059] In order to more clearly show the technical solutions provided by the present invention and the technical effects produced, the following uses specific embodiments to describe in detail the pronunciation - dictionary - independent speech synthesis method provided by the embodiments of the present invention.
[0060] Example 1
[0061] As Figure 1 shown, this embodiment provides a speech synthesis system that does not rely on a pronunciation dictionary, including: a language-independent speech recognition model, a text-pronunciation representation prediction model, a pronunciation representation-acoustic prediction model, and a neural network vocoder;
[0062] Among them, the input of the language-independent speech recognition model is the speech waveform of the target language, and the training target is the International Phonetic Alphabet (IPA) phoneme sequence shared by multiple languages. As Figure 2 shown, this language-independent speech recognition model uses wav2vec 2.0 as the basic architecture of the speech recognition model. It adds two linear layers after the context representation and deletes the quantization module in the original wav2vec 2.0 model. The first linear layer is called the bottleneck layer, which maps the 1024-dimensional context representation (C) to a 512-dimensional bottleneck representation (B). The second layer is the classification layer, which predicts the class probability (P) based on the bottleneck representation. The number of classes depends on the size of the phoneme set. To learn the shared pronunciation representation across languages, this speech recognition model is trained using a multi-language corpus with IPA phoneme transcriptions, and the CTC loss between the output class probability and the target IPA sequence is used for training.
[0063] The extraction process of the pronunciation representation of this language-independent speech recognition model is as Figure 3 shown. Here, an audio segment of an English word "work" is taken as an example. φ represents the blank symbol in the CTC loss, and w, o, r, and k are the recognized phoneme symbols;
[0064] Frame-level bottleneck representation B = [b1,..., b T is first calculated using the language-independent automatic speech recognition model, where T is the number of frames of the speech waveform;
[0065] Then, the argmax function is applied to the class probability P output by the speech recognition model to generate the corresponding class (phoneme symbol or blank symbol in CTC) for each frame;
[0066] Then, a classification operation is performed to assign the class of the t-th frame to the bottleneck representation b t ;
[0067] Finally, a merging operation is applied to remove the bottleneck representations of the blank classes, and the adjacent bottleneck representations with the same class are merged into one vector by taking the average, that is, CPR. In Figure 2 it, the extracted pronunciation representation is written as R = [r1,..., r N, where N is the number of pronunciation representations of the waveform. Obviously, N is less than T.
[0068] When training the language-independent speech recognition model, in this embodiment, corpora of 19 languages were used, with a total duration of 2771 hours; first, the texts of all the speeches were transcribed and converted into IPA phoneme sequences using the open-source tool Phonemizer. Finally, the size of the phoneme set was 203, and together with the blank symbol of CTC, the dimension of the class probabilities output by the classification layer in the speech recognition model was 204; the corpora of each language were divided into a training set and a validation set at a ratio of 99 to 1, and the training batch size was 1.2 hours; finally, the model checkpoint with the lowest phoneme recognition error rate of 5.87% on the validation set was selected; the loss function used for training was the CTC error between the output class probabilities and the true IPA sequence.
[0069] Figure 4 Schematically shows the acoustic modeling diagram of the speech synthesis system of this embodiment based on pronunciation representations, which consists of two parts, a text-pronunciation representation prediction model and a pronunciation representation-acoustic prediction model. The text-pronunciation representation model predicts pronunciation representations based on the input text, and the pronunciation representation-acoustic model generates Mel spectrograms based on the pronunciation representations. In the synthesis stage, a neural network vocoder is used to reconstruct the speech waveform from the Mel spectrograms, that is, to complete the speech synthesis of the text to be synthesized. Specifically, the structure of the text-pronunciation representation prediction model adopts a sequence-to-sequence structure based on Tacotron2, and there are three differences compared with Tacotron2; first, the training target changes from an 80-dimensional Mel spectrogram to a 512-dimensional pronunciation representation; second, since the temporal continuity of the pronunciation representation is not as obvious as that of the acoustic features, the post-net part is removed. Finally, dropout in the pre-net is not applied in the synthesis stage.
[0070] The error function for training is the mean square error and mean absolute error between the predicted pronunciation representations and the extracted pronunciation representations, plus the binary cross-entropy of the stop symbol.
[0071] The pronunciation representation-acoustic prediction model is also based on the structure of Tacotron2. The difference is that it takes the pronunciation representation as the input instead of the character or phoneme sequence, so text embedding is not required; its loss function includes the mean square error and mean absolute error between the predicted Mel spectrogram and the true Mel spectrogram, as well as the binary cross-entropy of the stop symbol. In the synthesis stage, the predicted Mel spectrogram is input into the neural network vocoder to reconstruct the speech waveform.
[0072] Embodiment 2
[0073] As Figure 1As shown in the figure, an embodiment of the present invention provides a speech synthesis system that does not rely on a pronunciation dictionary but uses a multilingual general speech recognition model to extract language-independent pronunciation representations. This system can achieve speech synthesis based on continuous phonetic representations (CPR), improving the performance of speech synthesis when there is no pronunciation dictionary for the target language. The language-independent speech recognition model in this system can extract pronunciation representations from the speech waveform of the target language. The pronunciation representation is the output vector of a hidden layer in the speech recognition model and is used as an intermediate representation for acoustic modeling. The acoustic model based on the pronunciation representation consists of two parts: a text-pronunciation representation prediction model and a pronunciation representation-acoustic prediction model. The text-pronunciation representation model predicts the pronunciation representation based on the character sequence, and the pronunciation representation-acoustic model generates the Mel spectrogram based on the pronunciation representation. Since the language-independent speech recognition model is trained using the connectionist temporal classification (CTC) loss, the extracted pronunciation representation is at the segment level and has a similar length to the phoneme sequence. Finally, in the synthesis stage, the system reconstructs the speech waveform from the generated Mel spectrogram through a neural network vocoder, thus completing the speech synthesis of the text to be synthesized.
[0074] The acoustic model based on pronunciation representation of the present invention can automatically predict the pronunciation representation from the character sequence without relying on a pronunciation dictionary, which can improve the naturalness and intelligibility of the synthesized speech compared to directly predicting the Mel spectrogram from the character sequence. In a traditional speech synthesis system, it is necessary to use a pronunciation dictionary and a character-phoneme conversion model to process the input text into a phoneme sequence in the front-end text analysis stage and then send it to the back-end module for acoustic feature prediction and waveform reconstruction. Although some sequence-to-sequence acoustic models can directly input the character sequence, it will reduce the quality of the synthesized speech.
[0075] The effectiveness of the speech synthesis system and method of the present invention is verified in the following ways, including:
[0076] (1) Test settings:
[0077] The present invention conducted experiments in six target languages, namely English (en), Spanish (es), Kazakh (kk), Hindi (hi), Bulgarian (bg), and Malay (ms). The durations of the corpora (i.e., speech waveforms) were 24, 28, 11, 16, 5, and 10 hours respectively, and the corresponding numbers of sentences were 13100, 19351, 6447, 9293, 3006, and 5780; all languages were pronounced by single female speakers; the corpus of each language was divided into a training set, a validation set, and a test set; for English, Spanish, Kazakh, Hindi, Bulgarian, and Malay, the numbers of sentences in the validation set and the test set were 300, 400, 200, 250, 150, and 200 respectively, and the remaining corpus was used for training. In addition, we added another text-only test set for intelligibility assessment based on speech recognition, which contained 1000 sentences for each language. The acoustic model based on pronunciation representation of the present invention (i.e., the speech synthesis system of the present invention) was compared with the three models listed below, including:
[0078] (1) Taco-Char model: The Tacotron2 model that uses character sequences as input, serving as the baseline for the experiment.
[0079] (2) Taco-Phone model: The Tacotron2 model that uses phoneme sequences as input. The open-source text-to-phoneme conversion tool Phonemizer was used to transcribe the text of the dataset into phoneme sequences, except for English which used Festival.
[0080] (3) DPS model: The discrete phoneme symbol results recognized by a language-independent speech recognition model were used as the intermediate representation of the acoustic model to replace the pronunciation representation. Here, the text-pronunciation representation prediction model became a text-to-phoneme conversion model based on a long short-term memory network, and at the same time, the pronunciation representation - acoustic model was changed to take discrete phoneme symbols as input.
[0081] The character error rate (CER) given by speech recognition was used to evaluate the intelligibility of the synthesized speech. In this embodiment, the Google Speech Recognition API was used to transcribe the synthesized speech for CER calculation. As described above, a test set of 1000 sentences was used for CER evaluation for each language. At the same time, subjective listening tests were conducted to evaluate the mean opinion scores (MOS) of the naturalness of the synthesized speech. For each language, 30 sentences were selected from the test set and synthesized using different models. The naturalness of each synthesized speech was rated on a scale of 1 to 5, with an interval of 0.5, and was rated by 4 to 6 native speakers.
[0082] (II) Test results:
[0083] Table 1 below lists the CER results. It can be seen from them that the DPS model has the worst performance because the hard decision affects the derivation of discrete recognition results. The Taco-Phone model is better than the Taco-Char model in three languages, but has the same performance in es, and is slightly worse in kk and bg. There may be two reasons. One is that it is easy to infer the pronunciation of words in the latter three languages from the spelling. The other is that the Phonemizer tool may not provide sufficiently accurate phoneme transcriptions for the latter three languages. The acoustic model based on pronunciation representation of the present invention (i.e., the Proposed model) is better than all other models in all languages, except the Taco-Phone model in hi. This proves the effectiveness of the speech synthesis system and method of the present invention in synthesizing highly intelligible speech without relying on a pronunciation dictionary.
[0084] Table 1: CER (%) of different models on six target languages
[0085]
[0086] Due to the poor CER performance of the DPS model, subjective listening tests were not conducted on it. Table 2 below gives the naturalness MOS results of the subjective listening tests. In this embodiment, the p-value of the paired t-test was used to evaluate the significance of the difference between the two models. Similar to the results in Table 1, the Taco-Phone model achieved better naturalness than the Taco-Char model in en, hi, and ms (p = 5×10 -6 , 0.03, 2×10 -3 ); however, it was comparable to the Taco-Char model in kk (p = 0.14) and bg (p = 1.00), and even worse than the Taco-Char model in es (p = 0.02); the speech synthesis system of this embodiment performed significantly better than the Taco-Char model in en, kk, hi, bg, and ms (p = 2×10 -12 , 7×10 -3 , 8×10 -4 , 7×10 -5 , 1×10 -4 ), and was comparable to the Taco-Char model in es (p = 0.17). This may be because the mapping relationship from spelling to pronunciation in Spanish is very simple. In addition, compared with the Taco-Phone model, the speech synthesis system of the present invention achieved better naturalness in en, es, and bg (p = 3×10 -5 , 3×10 -4 , 1×10 -6) and they are comparable on kk, hi, and ms (p = 0.15, 0.25, 0.18). These results also prove the superiority of the speech synthesis system and method of the present invention.
[0087] Table 2: Naturalness MOS of different models on six target languages, with a confidence interval of 95%
[0088]
[0089] In summary, the speech synthesis system and method of the embodiments of the present invention do not rely on a pronunciation dictionary, but use a language-independent speech recognition model to extract language-independent pronunciation features for text-to-speech prediction and synthesis, with high-quality synthesized speech and improved performance of the speech synthesis system.
[0090] The above are only the preferred specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed by the present invention should be covered by the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims. The information disclosed in the background art section of this article is only intended to deepen the understanding of the overall background art of the present invention and should not be regarded as an admission or any form of implication that this information constitutes prior art known to those skilled in the art.
Claims
1. A speech synthesis system that does not rely on a pronunciation dictionary, characterized in that, Including: A language-independent speech recognition model, a text-pronunciation representation prediction model, a pronunciation representation-acoustic prediction model, and a neural network vocoder; wherein, The language-independent speech recognition model can extract pronunciation representations from the input speech waveform of the target language during the training phase, and provide the pronunciation representations to the text-pronunciation representation prediction model and the pronunciation representation-acoustic prediction model for training, so as to obtain the trained text-pronunciation representation prediction model and pronunciation representation-acoustic prediction model; The text-pronunciation representation prediction model can predict pronunciation representations according to the character sequence of the text to be synthesized after being trained, and output them to the trained pronunciation representation-acoustic prediction model; The pronunciation representation-acoustic prediction model is connected to the neural network vocoder, and can generate a Mel spectrogram according to the pronunciation representation predicted by the text-pronunciation representation prediction model; The neural network vocoder can reconstruct the Mel spectrogram generated by the pronunciation representation-acoustic prediction model into a speech waveform corresponding to the text to be synthesized; The language-independent speech recognition model includes: A wav2vec 2.0 model, a first linear layer, and a second linear layer connected in sequence; wherein, The wav2vec 2.0 model uses a wav2vec 2.0 model without a quantization module, and its training input is a multi-language corpus with IPA phoneme transcriptions; The first linear layer is a bottleneck layer that can map the 1024-dimensional context representation (C) to the 512-dimensional bottleneck representation (B); The second linear layer is a classification layer that can predict the class probability (P) according to the bottleneck representation output by the first linear layer; The training objective of this language-independent speech recognition model is the CTC loss between the class probability and the target IPA sequence.
2. The speech synthesis system independent of a pronunciation dictionary according to claim 1, wherein The structure of the text-pronunciation representation prediction model adopts a sequence-to-sequence structure based on Tacotron2; The error function for training this text-pronunciation representation prediction model is the mean square error and mean absolute error between the predicted pronunciation representation and the extracted pronunciation representation, plus the binary cross-entropy of the stop symbol.
3. The voice synthesis system independent of a pronunciation dictionary according to claim 1, characterized in that The pronunciation representation-acoustic prediction model adopts the structure of Tacotron2; The loss function of this pronunciation representation-acoustic prediction model is the mean square error and mean absolute error between the predicted Mel spectrogram and the true Mel spectrogram, as well as the binary cross-entropy of the stop symbol.
4. A speech synthesis method independent of a pronunciation dictionary, characterized in that, Using the pronunciation dictionary-independent speech synthesis system according to any one of claims 1 to 3, first, the language-independent speech recognition model of the speech synthesis system extracts pronunciation representations from the input speech waveform of the target language, and uses the pronunciation representations to train the text-pronunciation representation prediction model and the pronunciation representation-acoustic prediction model of the speech synthesis system. After the training is completed, the trained text-pronunciation representation prediction model and pronunciation representation-acoustic prediction model are obtained; The language-independent speech recognition model includes: a wav2vec 2.0 model, a first linear layer, and a second linear layer connected in sequence; wherein, the wav2vec 2.0 model adopts a wav2vec 2.0 model without a quantization module, and its training input is a multilingual corpus with IPA phoneme transcriptions; the first linear layer is a bottleneck layer that can map the 1024-dimensional context representation (C) to the 512-dimensional bottleneck representation (B); the second linear layer is a classification layer that can predict the class probability (P) according to the bottleneck representation output by the first linear layer; the training objective of this language-independent speech recognition model is the CTC loss between the class probability and the target IPA sequence; the synthesis is performed according to the following steps: Input the text to be synthesized into the text-pronunciation representation prediction model of the trained speech synthesis system. The text-pronunciation representation prediction model predicts the pronunciation representation according to the character sequence of the text to be synthesized and outputs it to the pronunciation representation-acoustic prediction model of the speech synthesis system; The pronunciation representation-acoustic prediction model predicts and generates a mel spectrum according to the pronunciation representation and outputs the mel spectrum to the neural network vocoder of the speech synthesis system; The neural network vocoder reconstructs the mel spectrum into a speech waveform corresponding to the text to be synthesized.
5. The method for speech synthesis independent of a pronunciation dictionary according to claim 4, characterized in that, The language-independent speech recognition model extracts the pronunciation representation from the input speech waveform of the target language in the following manner, including: Calculate the frame-level bottleneck representation B = [b1,..., b T of the speech waveform of the input target language; Apply the argmax function to the class probability P output by the language-independent speech recognition model to obtain the phoneme symbol category corresponding to each frame; Combined with the frame-level bottleneck representation B = [b1, …, b T , a classification operation is performed on the frame-level bottleneck representation using the phoneme symbol category of each frame, and the category of the t-th frame is assigned to the bottleneck representation b t ; Apply the merging operation to remove the bottleneck representations of blank categories, and merge adjacent bottleneck representations with the same category into one vector R = [r1, …, r N , which is the pronunciation representation. Here, N is the number of pronunciation representations of the speech waveform, N is less than T, and T is the frame-level bottleneck representation length of the speech waveform.
Citation Information
Patent Citations
Improved automated voice synthesis employing enhanced prosodic treatmentof text, spelling of text and rate of annunciation
CA2119397A1
Method, apparatus for synthesizing speech and acoustic model training method for speech synthesis
US20120221339A1