Learning device, data generation device, and program
The learning device and data generation device automate the creation of high-quality training data for speech synthesis by using a speech recognition model and a label data correction model, addressing the cost and efficiency challenges of manual data generation.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2021-08-23
- Publication Date
- 2026-03-30
AI Technical Summary
Conventional speech synthesis technologies face challenges in generating high-quality training data due to the difficulty in creating input data for learning, particularly for Japanese speech synthesis, as Japanese kanji characters have multiple readings, and manual data creation is costly and time-consuming.
A learning device and data generation device that utilizes a speech recognition model to automatically generate label data, including phonemes and prosodic symbols, and a label data correction model to correct errors, reducing the need for manual input data creation.
Enables the generation of large amounts of training data for speech synthesis models, reducing human and time costs while improving the quality of the acoustic feature generation model.
Smart Images

Figure 0007837149000001 
Figure 0007837149000002 
Figure 0007837149000003
Abstract
Description
[Technical Field]
[0001] This invention relates to a learning device, a data generation device, and a program. [Background technology]
[0002] Japanese Seq2seq (sequence-to-sequence) speech synthesis performs Japanese speech synthesis based on input data described using labels representing phonetic readings and prosodic symbols (see, for example, Patent Document 1). In contrast, DNN (Deep Neural Network) speech synthesis performs speech synthesis using full-context labels as input data (see, for example, Non-Patent Document 1). [Prior art documents] [Patent Documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2020-34883 [Non-patent literature]
[0004] [Non-Patent Document 1] Heiga Zen, Andrew Senior, Mike Schuster, "Statistical parametric speech synthesis using deep neural networks," 2013 IEEE International Conference on Acoustics, Speech and Signal Processing [Overview of the project] [Problems that the invention aims to solve]
[0005] To perform speech synthesis using the technologies described in Patent Document 1 and Non-Patent Document 1, learning is required using pairs of input data and correct speech data. The technology in Non-Patent Document 1 uses input data to which various information other than phonemes has been added, taking into account contextual information about what is before and after phonemes. Therefore, it has been difficult to generate input data for learning using speech recognition. This is because conventional phoneme recognition, which is a key technology in speech recognition, can only recognize phonemes and cannot estimate prosody, including accent.
[0006] On the other hand, in the case of the technology described in Patent Document 1, the information used for input data is limited to phonemes and prosody such as accent. Therefore, it is conceivable to generate input data from a script representing the content of the speech data through morphological analysis, etc. However, because Japanese kanji characters have multiple readings, it is not always possible to generate correct input data. Therefore, conventionally, input data for learning was created by manually listening to speech. High-quality speech synthesis requires a large amount of training data, but creating input data manually has the problem of incurring human and time costs.
[0007] This invention has been made in consideration of these circumstances, and provides a learning device, a data generation device, and a program that can generate data for training a speech synthesis model while reducing costs. [Means for solving the problem]
[0008] [1] One aspect of the present invention is a learning device comprising a learning unit that learns a labeling model, which takes speech data of an utterance or feature quantities obtained from the speech data as input and outputs label data of text including characters representing phonemes and prosodic symbols representing accents in the utterance, using pairs of training speech data and correct label data.
[0009] [2] One aspect of the present invention is the learning device described above, characterized in that the learning unit learns a label data correction model, which takes label data estimated using the labeling model as input and outputs label data in which errors in phonemes included in the input label data have been corrected, using pairs of learning label data containing errors and correct label data.
[0010] [3] One aspect of the present invention is a labeling model that takes speech data of an utterance or features obtained from the speech data as input and outputs label data of text including characters representing phonemes and prosodic symbols representing accents, comprising: a speech recognition unit that takes features obtained from speech data to be used for label data estimation as input to the labeling model which has been trained using pairs of training speech data and correct label data, and obtains an estimation result of label data representing the utterance of the speech data to be used for label data estimation.
[0011] [4] One aspect of the present invention is the data generation apparatus described above, wherein the speech recognition unit is a label data correction model that inputs label data estimated based on speech data and outputs label data in which errors in phonemes included in the input label data are corrected, characterized in that the label data correction model, which has been trained using pairs of training label data containing errors and correct label data, is input with the label data estimated using the labeling model to obtain label data in which errors are corrected.
[0012] [5] One aspect of the present invention is the data generation apparatus described above, characterized in that the labeling model includes an encoder that inputs time-series feature quantities obtained from audio data, and a decoder that takes the output from the encoder as input and outputs label data.
[0013] [6]One aspect of the present invention is the data generation device described above, wherein the labeling model corresponds to each of the voice data for each predetermined time interval separated by shifting time, and the voice data of the corresponding time interval is used as the feature amount as an input to a convolutional neural network, and a transformer that takes an output from the convolutional network as an input and outputs label data.
[0014] [7]One aspect of the present invention is the data generation device described above, further comprising a voice processing unit that extracts voice data of each utterance for each sentence from the voice data, and the voice recognition unit inputs the voice data extracted by the voice processing unit into the labeling model as an object for estimating label data.
[0015] [8]One aspect of the present invention is a program for causing a computer to function as any of the learning devices described above.
[0016] [9]One aspect of the present invention is a program for causing a computer to function as any of the data generation devices described above.
Advantages of the Invention
[0017] According to the present invention, it is possible to generate data for learning a model for voice synthesis while reducing costs.
Brief Description of the Drawings
[0018] [Figure 1] It is a diagram showing an outline of the processing of an embodiment of the present invention. [Figure 2] It is a diagram showing a configuration example of a voice synthesis system according to the embodiment. [Figure 3] It is a diagram showing prosodic symbols used for label data according to the embodiment. [Figure 4] It is a diagram showing a voice recognition model according to the embodiment. [Figure 5] It is a diagram showing an acoustic feature quantity generation model for voice synthesis according to the embodiment. [Figure 6] This flowchart illustrates the speech recognition model learning process of the learning data generation device according to the same embodiment. [Figure 7] This flowchart illustrates the learning process for the acoustic feature generation model used in speech synthesis in the speech synthesis system according to the same embodiment. [Figure 8] This is a flowchart illustrating the learning data generation process of the learning data generation device according to the same embodiment. [Figure 9] This figure shows the results of an evaluation experiment of the learning data generation device according to the same embodiment. [Figure 10] This figure shows the results of an evaluation experiment of the learning data generation device according to the same embodiment. [Figure 11] This figure shows a labeling model according to the same embodiment. [Modes for carrying out the invention]
[0019] Embodiments of the present invention will be described in detail below with reference to the drawings. Figure 1 is a diagram illustrating the processing overview of this embodiment. The learning data generation device of this embodiment generates data used for training the speech synthesis acoustic feature generation model M, which is an acoustic feature generation model for speech synthesis. For example, the technology described in Patent Document 1 is used for the speech synthesis acoustic feature generation model M. The speech synthesis acoustic feature generation model M estimates acoustic features by taking readable text data, in which the utterance content is described using phonetic characters and prosodic symbols, as input. This text data described using phonetic characters and prosodic symbols is described as label data. That is, the label data is described by phonetic characters, which are labels representing phonemes, and prosodic symbols, which are labels representing prosody such as accent. Characters other than phonetic characters are used for prosodic symbols. The label data may further include speech style symbols, which represent the features given to the entire utterance as a string of characters. During Japanese speech synthesis, the speech synthesizer inputs label data A2, which is converted from source data A1 (text data of a Japanese sentence containing kanji and kana), into a speech synthesis acoustic feature generation model M to obtain acoustic features A3, such as a Mel spectrogram. The speech synthesizer then generates synthesized speech data A4 from these acoustic features A3 using a vocoder.
[0020] The acoustic feature generation model M for speech synthesis uses a set of training data consisting of pairs of label data and ground truth speech data. The amount of training data directly affects the quality of speech synthesis, so it is desirable to prepare a large amount of training data. Although the amount of information used in label data is less than that used in input data for general speech recognition technology, it is difficult to automatically generate error-free label data from Japanese texts containing both kanji and kana characters because Japanese kanji characters have multiple readings. Therefore, it is necessary to generate label data manually or to manually correct automatically generated label data, making it difficult to prepare a large amount of training data for the acoustic feature generation model M for speech synthesis.
[0021] On the other hand, the techniques described in References 1 and 2 allow for the construction of a speech recognition model that directly converts speech data to text using a small amount of training data. The training data generation device of this embodiment uses a speech recognition model W that applies the techniques of References 1 and 2 to directly generate label data L1 used for training the acoustic feature generation model M for speech synthesis from speech data V1. As a result, the training data generation device of this embodiment can generate a large amount of training data D1, which consists of pairs of speech data V1 and label data L1. The speech synthesizer trains the acoustic feature generation model M for speech synthesis using the training data D1. Note that training with training data D1 may be considered pretraining, followed by fine-tuning. In fine-tuning, the acoustic feature generation model M for speech synthesis is further trained using a small amount of training data D2, which consists of pairs of acoustic features from speech data V2 and manually generated, accurate label data L2.
[0022] By using the speech recognition model W, it is possible to generate label data L1 for each of the large amounts of audio data V1 extracted from, for example, television or radio audio data through speech processing. Therefore, it is possible to reduce the human and time costs required to create the label data used for training the acoustic feature generation model M for speech synthesis, and to improve the quality of the acoustic feature generation model M for speech synthesis by increasing the amount of data.
[0023] (Reference 1) Alexei Baevski, Henry Zhou, Abdelrahman Mohamed, Michael Auli, "wav2vec 2.0: A Framework for Self-Supervised Learning of Speech Representations," 34th Conference on Neural Information Processing Systems (NeurIPS 2020), Vancouver, Canada.
[0024] (Reference 2) Wav2Vec2-XLSR-53, [online],<URL:https: / / huggingface.co / facebook / wav2vec2-large-xlsr-53>
[0025] Figure 2 shows an example configuration of the speech synthesis system 100 according to this embodiment. Figure 2 extracts only the functional blocks related to this embodiment. The speech synthesis system 100 includes a learning data generation device 1 and a speech synthesis device 5. The learning data generation device 1 is an example of a learning device and a data generation device. The learning data generation device 1 and the speech synthesis device 5 may be an integrated device.
[0026] The learning data generation device 1 includes a speech recognition unit 11, a learning data generation unit 13, and a speech processing unit 14. The speech recognition unit 11 estimates label data from speech data using a speech recognition model W. The speech recognition unit 11 includes a speech recognition model learning unit 12. The speech recognition model learning unit 12 learns the speech recognition model W using pairs of speech data V0 and the correct label data L0 of the utterance indicated by the speech data V0. The learning data generation unit 13 generates learning data D1 by associating speech data V1 with label data L1 obtained by the speech recognition unit 11 inputting the speech data V1 into the trained speech recognition model W. The learning data D1 is data for learning an acoustic feature generation model. When the learning data generation unit 13 receives speech data V1', such as speech of multiple sentences or speech containing noise, the speech processing unit 14 extracts speech data V1 for each sentence from the speech data V1'. The audio processing unit 14 uses any existing processing method to extract the audio data V1.
[0027] The speech synthesis device 5 comprises an acoustic feature estimation unit 51, a language processing unit 53, and a vocoder unit 54. The acoustic feature estimation unit 51 estimates acoustic features from label data using an acoustic feature generation model M for speech synthesis. The acoustic feature estimation unit 51 includes an acoustic feature generation model learning unit 52. The acoustic feature generation model learning unit 52 learns the acoustic feature generation model M for speech synthesis using training data for learning the acoustic feature generation model. The training data for learning the acoustic feature generation model includes training data D1 generated by the training data generation device 1, and may also include training data D2. Training data D2 is a pair of speech data V2 and manually generated accurate label data L2. The language processing unit 53 converts the original text data A1, which is a mixed kanji and kana text, into label data A2 using phonetic readings and prosodic symbols. This conversion can be performed using existing technologies such as morphological analysis. The user may modify the generated label data A2 as needed. The speech synthesizer 5 may also receive label data A2 as input. The vocoder unit 54 estimates a speech waveform from the acoustic features A3 estimated by the acoustic feature estimation unit 51 using the label data A2, and outputs synthesized speech data A4 representing the estimated speech waveform. For example, the vocoder unit 54 is a DNN that takes acoustic feature data as input and outputs a speech waveform.
[0028] Figure 3 shows the prosodic symbols used in the label data of this embodiment. The prosodic symbols shown in Figure 3 are modified versions of the prosodic symbols described in Reference 3. Prosodic information includes types such as accent placement, phrase / phrase separation, sentence-ending intonation, and pause placement. Prosodic symbols that indicate accent placement include the rising accent mark "^" and the falling accent mark "!". The rising accent mark "^" indicates that the accent rises on the kana immediately following the mark. The falling accent mark "!" indicates that the accent falls on the kana immediately following the mark. For specifying phrase / phrase separation, the prosodic symbol "#" is used to indicate the separation of accented phrases. The following symbols are used to specify sentence-ending intonation: the punctuation mark "=" for normal sentence endings, the punctuation mark "(" for nominal sentence endings, and the punctuation mark "?" for interrogative sentence endings. The punctuation mark "," is used to specify pauses. Note that these punctuation marks are just examples, and other symbols may be used. Also, some of the above punctuation marks may be used in the label data.
[0029] (Reference 3) Audio Input / Output Method Standardization Committee, "JEITA Standard IT-4006 Symbols for Japanese Text-to-Speech Synthesis", Japan Electronics and Information Technology Industries Association, 2010, pp. 4-10.
[0030] Label data may include speech style symbols that represent features of the entire utterance as a string. These features include, for example, speech style (commentary style, news style), emotion (sad, happy, etc.), and the speaker. Speech style symbols use characters or strings that are different from phonetic readings and also different from characters representing prosodic markers. For example, the utterance tag " <tag>The symbol " can be used as a speech style symbol. <tag>In the "tag" part of "", you can use a string that represents the type of characteristic that applies to the entire utterance. You can also change the number of characters in the string that represents the speech style symbol. For example, if the characteristic that applies to the entire utterance is sad emotion, then " <sad>Use " when speaking in a news-like tone, and " <news>Use " when speaker A, and when speaker A, <spkera>Use "".
[0031] Figure 4 shows a speech recognition model W. The speech recognition model W consists of a labeling model W1 and a label data correction model W2. Labeling model W1 is, for example, Wav2vec2.0 as described in References 1 and 2, or the sequence-to-sequence (Seq2seq) speech recognition method as described in Reference 10. The labeling model W1 shown in Figure 4 is an example using Wav2vec2.0. Generally, the training data for speech synthesis is about 10 hours. Wav2vec2.0 and Seq2seq speech recognition methods learn based on phonemes and prosodic symbols, which have a small number of types as strings, so compared to many other speech recognition technologies, they can achieve high-accuracy speech recognition with a small amount of training data, and their effectiveness has been demonstrated in various languages, including Japanese. In particular, various pre-trained Wav2vec2.0 models have been made publicly available, including a pre-trained model that was trained using 56,000 hours of speech data in 53 languages as pre-training data. However, there are no examples of training a model to output text containing prosodic markers.
[0032] When using Wav2vec2.0, the labeling model W1 is a model that takes raw audio waveform X as input and outputs label data La. The audio data input to the speech recognition unit 11 is used as the audio waveform X of the labeling model W1. The audio data represents sound pressure. The labeling model W1 has multiple CNNs (Convolutional Neural Networks) and a Transformer.
[0033] Multiple CNNs are equivalent to encoders. Each CNN consists of several blocks, including layer normalization and GELU (Gaussian error linear units) activation functions, after temporal convolution. Each CNN corresponds to a different time interval, and each CNN is input with time-series audio data from that time interval. Each CNN outputs acoustic features Z (Latent speech representations) that represent the characteristics of the audio. Acoustic features Z are latent space representations of the audio. In latent space, waveform vectors with similar characteristics become close together.
[0034] The transformer corresponds to the decoder. The transformer is a neural network that outputs context representations C (Contest representations) of time-series acoustic features Z. The time-series acoustic features Z output from each CNN are masked and input to the transformer. That is, a predetermined proportion of the time-series acoustic features Z are randomly selected, and a predetermined number of consecutive acoustic features from the selected acoustic features are replaced with trained features before being input to the transformer. For example, the technique described in Reference 4 is used for the transformer. The context representation C output from the transformer is label data La using phonetic readings and prosodic symbols.
[0035] (Reference 4) Ashish Vaswani, et al., "Attention is all you need," In Proc. of Neural Information Processing Systems (NIPS), 2017.
[0036] Similar to phoneme recognition using acoustic models in general speech recognition, the label data La estimated by the labeling model W1 contains phoneme errors. Therefore, the label data correction model W2 corrects the phoneme errors contained in the label data La. The label data correction model W2 uses a conventional transformer (see, for example, reference 5). This transformer is implemented using a neural network and is configured to include an encoder and a decoder. The encoder accepts the label data La as input data and passes the result of the encoding process to the decoder. Based on the information passed from the encoder, the decoder generates and outputs label data Lb in which the phoneme errors of the label data La have been corrected. In addition to the information passed from the encoder, the decoder also uses the right shift of the previously output label data Lb as input.
[0037] (Reference 5) Colin Raffel, et al., "Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer", Journal of Machine Learning Research 21, 2020, p.1-67
[0038] The speech recognition model learning unit 12 of the speech recognition unit 11 first learns a labeling model W1 using speech data V0 and the correct label data L0 for that speech data V0. The label data L0 is label data in which manually generated or modified phonetic readings and prosodic symbols are described. That is, the speech recognition model learning unit 12 updates the weights, which are the values of each parameter of the labeling model W1, so that the loss L, which represents the difference between the label data La obtained by inputting the speech data V0 for speech recognition model learning and the correct label data L0, becomes small. The loss L is a contrastive loss. In addition, quantized representations Q, which consist of the discretized values q of the acoustic features Z calculated by each CNN, are also used in calculating the loss L.
[0039] After training the labeling model W1, the speech recognition model training unit 12 inputs the label data La obtained by the trained labeling model W1 receiving speech data V0 into the label data correction model W2, and updates the values of each parameter of the label data correction model W2 so that the difference between the obtained label data Lb and the correct label data L0 for the speech data V0 is reduced. The speech recognition model training unit 12 may also modify the label data L0 and use it as label data La. The speech recognition model training unit 12 inputs the label data La generated by modifying the label data L0 into the label data correction model W2, and updates the weights, which are the values of each parameter of the label data correction model W2, so that the difference between the label data Lb produced by the label data correction model W2 and the correct label data L0 is reduced.
[0040] When the speech recognition unit 11 generates label data to be used as training data for the acoustic feature generation model M for speech synthesis, it receives speech data V1 from the training data generation unit 13. The speech recognition unit 11 inputs the speech data V1 into the labeling model W1 to obtain label data La, and then inputs the label data La into the label data correction model W2 to obtain label data Lb. The speech recognition unit 11 outputs the label data Lb to the training data generation unit 13 as label data L1 estimated from the speech data V1. Note that the speech recognition model W does not necessarily have a label data correction model W2. In this case, the speech recognition unit 11 inputs the speech data V1 into the labeling model W1 to obtain label data La, and outputs the label data L1 estimated from the speech data V1 to the training data generation unit 13.
[0041] Figure 5 shows an example of an acoustic feature generation model M for speech synthesis. The acoustic feature generation model M for speech synthesis is a DNN that applies the technique described in Reference 6.
[0042] (Reference 6) Shen et al., [online], February 2018, "Natural TTS Synthesis by Conditioning WaveNet on Mel Spectrogram Predictions", arXiv:1712.05884v2, Internet<URL:https: / / arxiv.org / pdf / 1712.05884.pdf>
[0043] The speech synthesis acoustic feature generation model M includes an encoder 81 and a decoder 85. The encoder 81 uses a CNN and an RNN (Recurrent Neural Network) to generate string features for the utterance content in the sentence indicated by the input label data, taking into account the context before and after the utterance content in the sentence indicated by the label data. The decoder 85 uses an RNN to generate acoustic features for predicting the speech corresponding to the utterance content indicated by the input label data, frame by frame, based on the features generated by the encoder 81 and acoustic features generated in the past.
[0044] The encoder 81 consists of a string conversion processor 811, a convolutional network 812, and a bidirectional LSTM (Long Short-Term Memory) network 813. The string conversion processor 811 converts the phonetic readings and prosodic symbols used in the label data into numerical values, and converts the label data into a vector representation. The convolutional network 812 is a neural network in which multiple (for example, 3) convolutional layers are connected. Each convolutional layer performs convolution on the vector representation of the label data using multiple filters of a size corresponding to a predetermined number of characters, and further performs batch normalization and ReLU (Rectified Linear Units) activation. This models the context of the utterance. For example, the filter size of the 3-layer convolutional network is [5,0,0], and the number of filters is 512. The output of the convolutional network 812 is input to the bidirectional LSTM network 813 in order to generate the feature quantities of the string to be input to the decoder 85. The bidirectional LSTM network 813 is a single bidirectional LSTM with 512 units (256 units in each direction). The bidirectional LSTM network 813 makes it possible to generate string features that take into account the surrounding context within the text data provided.
[0045] Decoder 85 is an autoregressive RNN. Decoder 85 consists of an attention network 851, a preprocessing network 852, an LSTM network 853, a first linear transformation process 854, a postprocessing network 855, an addition process 856, and a second linear transformation process 857.
[0046] The attention network 851 is a network that adds an attention function to an autoregressive RNN, and outputs a fixed-length context vector that summarizes the entire output from encoder 81 for each frame. The attention network 851 receives the output from the bidirectional LSTM network 813 (encoder output) as input. For each frame, the weights used to extract data from the encoder output to generate the summary differ depending on the data position in the encoder output. The attention network 851 generates the context vector (attention network output) for the current frame using data to which features have been added using the context vector generated at the time of the previous decoding, to the data extracted from the encoder output.
[0047] The preprocessing network 852 is input to the data output by the first linear transformation process 854 in the previous time step. The preprocessing network 852 is a neural network containing multiple (e.g., two) fully connected layers, each consisting of 256 hidden ReLU units. Each layer of ReLU units outputs zero if the value of each unit is less than zero, and outputs the value itself if it is greater than zero. The LSTM network 853 is a neural network in which multiple (e.g., two) one-way LSTMs with 1024 units each are connected, and it is input to data that combines the output from the preprocessing network 852 and the output from the attention network 851. Since the acoustic features of a frame are influenced by the acoustic features of the previous frame, the output from the preprocessing network 852 is combined with the features of the current frame output from the attention network 851 to add features based on the acoustic features of the previous frame.
[0048] The first linear transformation process 854 linearly transforms the data output from the LSTM network 853 and generates a context vector, which is the data for one frame of the Mel spectrogram. The first linear transformation process 854 outputs the generated context vector to the preprocessing network 852, the postprocessing network 855, and the addition process 856.
[0049] The post-processing network 855 is a neural network formed by combining multiple (e.g., 5-layer) convolutional networks. For example, a 5-layer convolutional network has a filter size of [5,0,0] and 1024 filters. Each convolutional network performs convolution, batch normalization, and tanh activation except for the last layer. The output from the post-processing network 855 is used to improve the overall quality after wavelength conversion. The summation process 856 adds the context vector generated by the first linear transformation process 854 to the output from the post-processing network 855. The summation process 856 outputs a Mel spectrogram, which is a collection of acoustic features for each frame.
[0050] In parallel with the spectrogram frame prediction described above, the second linear transformation process 857 projects the connection between the output of the LSTM network 853 and the attention context onto a scalar, then performs sigmoid activation to output a stop token used to determine if the output sequence is complete.
[0051] During training, the acoustic feature generation model learning unit 52 of the speech synthesis device 5 updates the parameters of the acoustic feature generation model M for speech synthesis so that the difference between the Mel spectrogram obtained by the acoustic feature estimation unit 51 inputting the label data Ln of the training data Dn into the acoustic feature generation model M for speech synthesis and the Mel spectrogram of the correct speech data Vn for the label data Ln becomes smaller. The pairs of label data Ln and speech data Vn of the training data Dn are the pairs of label data L1 and speech data V1 of the training data D1 generated by the training data generation device 1, and the pairs of label data L2 and speech data V2 of the training data D2 used for fine tuning (i.e., n=1,2).
[0052] During speech synthesis, the acoustic feature estimation unit 51 inputs label data A2 generated from source data A1 to the acoustic feature generation model M for speech synthesis, and outputs the generated Mel spectrogram to the vocoder unit 54. The vocoder unit 54 inputs the Mel spectrogram for each frame to the speech waveform generation model, inversely transforms it into a time-domain waveform to generate speech waveform data, and outputs it as synthesized speech data A4.
[0053] The acoustic feature generation model M for speech synthesis can use not only Tacotron 2, as described in Reference 6, but also Sequence-to-sequence + attention methods such as Deep Voice 3 and Transformer-based TTS. Deep Voice 3 is described, for example, in Reference 7. Transformer-based TTS is described, for example, in Reference 8.
[0054] (Reference 7) Wei Ping et al., [online], February 2018, "Deep Voice 3: Scaling Text-to-Speech with Convolutional Sequence Learning", arXiv:1710.07654v3, Internet<URL:https: / / arxiv.org / pdf / 1710.07654.pdf>
[0055] (Reference 8) Naihan Li et al., [online], January 2019, "Neural Speech Synthesis with Transformer Network", arXiv:1809.08895v3, Internet<URL:https: / / arxiv.org / pdf / 1809.08895.pdf>
[0056] Figure 6 is a flowchart showing the speech recognition model training process of the training data generation device 1. The training data generation device 1 receives speech recognition model training data, which associates speech data V0 of an utterance with the correct label data L0 of that utterance (step S110).
[0057] The speech recognition model learning unit 12 of the speech recognition unit 11 learns the labeling model W1 using the speech recognition model learning data (step S120). Specifically, the speech recognition model learning unit 12 reads pairs of speech data V0 and correct label data L0 from the speech recognition model learning data. The speech recognition unit 11 inputs the speech data V0 read by the speech recognition model learning unit 12 into the labeling model W1 to obtain label data La. The speech recognition model learning unit 12 updates the parameter values of the labeling model W1 so that the difference between the label data La obtained by the speech recognition unit 11 using the speech data V0 as input and the correct label data L0 for that speech data V0 becomes smaller. The speech recognition model learning unit 12 continues to learn the labeling model W1 until a predetermined learning termination condition is met. The learning completion conditions include, for example, completing processing for all input audio data V0 and label data L0 pairs, updating the labeling model W1 a predetermined number of times, or the difference falling below a predetermined level.
[0058] Furthermore, the speech recognition model learning unit 12 may pre-train the labeling model W1 using a large amount of speech recognition model learning data with label data L0 automatically generated from source text data containing a mixture of kanji and kana characters through morphological analysis, and then fine-tune the labeling model W1 using a small amount of speech recognition model learning data with manually generated error-free label data L0.
[0059] Next, the speech recognition model learning unit 12 learns the label data correction model W2 using the speech recognition model learning data (step S130). Specifically, the speech recognition model learning unit 12 reads pairs of speech data V0 and correct label data L0 from the speech recognition model learning data. The speech recognition unit 11 inputs the speech data V0 read by the speech recognition model learning unit 12 into the labeling model W1 to obtain label data La. Furthermore, the speech recognition unit 11 inputs the obtained label data La into the label data correction model W2 to obtain label data Lb with phoneme errors in label data La corrected. The speech recognition model learning unit 12 updates the parameter values of the label data correction model W2 so that the difference between label data Lb and correct label data L0 becomes smaller. The speech recognition model learning unit 12 continues to learn the label data correction model W2 until a predetermined learning termination condition is met. The learning completion conditions include, for example, completing processing for all input audio data V0 and label data L0 pairs, updating the label data correction model W2 a predetermined number of times, or the difference falling below a predetermined level.
[0060] Generally, transformers used in natural language processing require nearly several hundred thousand sentences of training data. Therefore, the training data for the speech recognition model may be expanded using a pair of correct label data L0 and label data La, which is created by randomly deleting characters or swapping consonants from the correct label data L0 to simulate phoneme errors. The speech recognition model training unit 12 inputs the simulated label data La into the label data correction model W2 and updates the parameter values of the label data correction model W2 so that the difference between the obtained label data Lb and the label data L0 becomes small. After pre-training the label data correction model W2 with the expanded training data, the speech recognition model training unit 12 fine-tunes the label data correction model W2 using the speech data V0 and the correct label data L0 as described above.
[0061] Figure 7 is a flowchart showing the acoustic feature generation model learning process of the speech synthesis system 100. The learning data generation unit 13 of the learning data generation device 1 receives multiple audio data V1 and V1' as input (step S210). When audio data V1' such as utterances of multiple sentences or audio containing noise is input, the speech processing unit 14 generates audio data V1 for each sentence from the audio data V1'. The learning data generation unit 13 outputs the audio data V1 to the speech recognition unit 11. The speech recognition unit 11 inputs the audio data V1 to the speech recognition model W to obtain label data L1 and outputs it to the learning data generation unit 13 (step S220). The learning data generation unit 13 generates learning data D1 for learning the acoustic feature generation model, consisting of pairs of audio data V1 and label data L1 output from the speech recognition unit 11 corresponding to the audio data V1 (step S230). The automatically generated learning data generates learning data D1 based on each audio data V1.
[0062] The acoustic feature generation model learning unit 52 of the speech synthesis device 5 acquires a set of training data D1 generated by the training data generation device 1. The acoustic feature generation model learning unit 52 pre-trains the acoustic feature generation model M for speech synthesis using this training data D1 (step S240). That is, the acoustic feature generation model learning unit 52 acquires pairs of speech data V1 and label data L1 from the training data D1. The acoustic feature estimation unit 51 inputs the label data L1 acquired by the acoustic feature generation model learning unit 52 into the acoustic feature generation model M for speech synthesis and obtains the estimation result of the acoustic features. The acoustic feature generation model learning unit 52 updates the acoustic feature generation model M for speech synthesis so that the difference between the acoustic features obtained from the speech data V1 and the estimated acoustic features obtained by the acoustic feature estimation unit 51 becomes smaller. The acoustic feature generation model learning unit 52 continues to train the acoustic feature generation model M for speech synthesis until a predetermined learning termination condition is met. The conditions for ending the learning process include, for example, completing processing for all input training data D1, updating the acoustic feature generation model M for speech synthesis a predetermined number of times, or the difference falling below a predetermined level.
[0063] Next, the acoustic feature generation model learning unit 52 receives training data D2, which includes pairs of speech data V2 and manually generated or modified label data L2. The amount of training data D2 received can be less than the amount of training data D1. The acoustic feature generation model learning unit 52 uses the training data D2 to fine-tune the acoustic feature generation model M for speech synthesis by the same process as in step S240 (step S250).
[0064] Figure 8 is a flowchart showing the learning data generation process of the learning data generation device 1. The learning data generation unit 13 of the learning data generation device 1 in Figure 2 receives audio data V1' as input (step S310). Audio data V1' is, for example, broadcast audio data. The learning data generation unit 13 outputs the audio data V1' to the audio processing unit 14. The audio processing unit 14 performs sound source separation on the audio data V1 (step S320) and then removes noise (step S330). The audio processing unit 14 detects speech and sound effects (SE) in the noise-removed audio data V1' (step S340) and extracts speech data for each sentence based on the detection results (step S350). The audio processing unit 14 outputs the extracted audio data V1 to the learning data generation unit 13. Note that if the learning data generation device 1 receives audio data V1 as input in step S310, it does not perform the processing in steps S320 to S350. The learning data generation device 1 may perform the processing while omitting some of these steps.
[0065] The training data generation unit 13 outputs the audio data V1 to the speech recognition unit 11. The speech recognition unit 11 inputs each audio data V1 into the trained labeling model W1 to obtain label data La. Furthermore, the training data generation unit 13 inputs the label data La into the trained label data correction model W2 to obtain label data Lb and outputs it to the training data generation unit 13 as label data L1 (step S360). If the speech recognition model W does not have a label data correction model W2, the training data generation unit 13 outputs the label data La estimated by the labeling model W1 as label data L1 to the training data generation unit 13. The training data generation unit 13 generates training data D1 consisting of pairs of audio data V1 and label data L1 estimated by the speech recognition unit 11 based on the audio data V1 (step S370).
[0066] To perform speech synthesis, an acoustic feature generation model for speech synthesis must be trained using pairs of training audio data and label data. However, conventionally, when only audio data existed, label data using phonetic readings and prosodic symbols had to be created manually, making it difficult to use such data as training data for an acoustic feature generation model for speech synthesis. According to this embodiment, since prosodic symbols including accents can be estimated from the audio, it becomes possible to use audio-only data as training data for an acoustic feature generation model for speech synthesis. Therefore, it is possible to generate a large amount of training data for an acoustic feature generation model for speech synthesis using audio from a wide range of fields, such as video sharing sites, television and radio audio, conference recordings, audio streaming services, and studio recordings.
[0067] This section describes the evaluation experiment of the training data generation device 1. In the evaluation experiment, the training data for the speech recognition model used to fine-tune the labeling model W1 consisted of audio recorded by NHK announcers in a studio booth and manually corrected label data. Katakana was used for the phonetic transcription. Prosodic markers consisted of accent rise / fall, accent punctuation, pauses, and sentence-ending markers. The experiment used male datasets M001, M002, M003, and M004, and female datasets F001, F002, and F003. The content of each dataset was audio data of news, weather information, and announcements being read aloud, respectively. The sampling frequency of the audio data was 16 kHz (kilohertz), and the bitrate was 16 bits. In addition, 631,014 sentences of news scripts from designated programs broadcast from April 2018 to April 2021 were used as pre-training label data for the label data correction model W2.
[0068] The labeling model W1, which was pre-trained, was trained using approximately 56,000 hours of speech data in 53 languages. Fine-tuning was performed on the pre-trained labeling model W1 using pairs of speech and manually corrected label data as training data for the speech recognition model. The batch size was 16, the gradient accumulation was 2, and the learning rate was 5.0 × 10⁶. -4 The training epoch count was set to 50. Furthermore, for training the label data correction model W2, 631,014 news articles were automatically transcribed using OpenJTalk to generate phonetic readings and prosodic markers. Batch size was 16, gradient accumulation was 1, and learning rate was 5.0 × 10⁶. -4 The number of training epochs was set to 20. Furthermore, the following data augmentation processes (1) and (2) were performed to create training data for pre-training the label data modification model W2.
[0069] (1) Delete characters at a rate of 5% or less. (2) Swap the consonants and prosodic symbols of the phonetic readings at a rate of 10% or less.
[0070] The label data correction model W2, which was pre-trained using the above pre-training data, was fine-tuned using a set of speech recognition model training data consisting of 23,024 manually corrected label data sentences.
[0071] The evaluation targets were label data La obtained by labeling model W1 and label data Lb obtained by labeling model W1 and label data modification model W2. Label data generated using conventional techniques was also used for comparison. The comparison target was label data obtained by Japaneseing speech using a pre-trained Japanese speech synthesis model publicly available on Espnet ASR (see Reference 9), and then automatically converting it into phonetic readings and prosodic symbols using OpenJTalk.
[0072] (Reference 9) Watanabe et al., "ESPnet: End-to-End Speech Processing Toolkit," Interspeech, 2018.
[0073] For fine-tuning the labeling model W1 to obtain label data La and label data Lb, the F003 and M003 speech datasets (2541 sentences, 5.69 hours) were used. For fine-tuning the label data correction model W2 to obtain label data Lb, manually corrected label data (23,024 sentences) was used. For the test set to calculate CER, M002, F002, and M004 (1558 sentences, 3.73 hours) were used. CER was calculated using the label data La estimated by labeling model W1, the label data Lb estimated by labeling model W1 and label data correction model W2, and the label data estimated by conventional techniques (Espnet ASR + OpenJTalk), along with the ground truth label data.
[0074] Figure 9 shows the results of the evaluation experiment. The CER of label data La and the CER of label data Lb were lower than the CER of the conventional technology. Therefore, the effectiveness of this embodiment was confirmed. Furthermore, since the CER of label data Lb was lower than the CER of label data La, the effectiveness of the label data modification model W2 was confirmed.
[0075] Figure 10 shows the results of evaluation experiments with varying amounts of training data. Here, only the labeling model W1 was used, without the label data correction model W2. The corpora M001 and F001 were used as training data for the speech recognition model to fine-tune the labeling model W1. Figure 10 shows the CER when the amount of training data for the speech recognition model was changed. As shown in Figure 10, the highest performance was observed with 5 hours of data.
[0076] From the above experiment, it was confirmed that the learning data generation device 1 of this embodiment can generate label data with high accuracy from speech data alone. Conventional technology cannot accurately estimate phonetic readings and prosodic symbols that reflect acoustic features. This is thought to be because, in conventional technology, when speech recognition is performed, the conversion is done to mixed kanji and kana text, and therefore the speech information cannot be utilized in the conversion from kanji to phonetic readings and the estimation of prosodic information, resulting in errors. On the other hand, in this embodiment, not only accent rise and fall, but also accent punctuation and symbols at the end of sentences can be estimated with high accuracy. Furthermore, it was found that training the labeling model W1 requires a smaller amount of training data compared to conventional speech recognition models.
[0077] The learning data generation device 1 may use the labeling model W1a shown in FIG. 11 instead of the labeling model W1. FIG. 11 is a diagram showing an example of the labeling model W1a using a Seq2seq speech recognition model. The labeling model W1a is, for example, the Seq2seq speech recognition model described in Reference 10. The labeling model W1a using the Seq2seq speech recognition model can be learned with less than several thousand hours of learning data because it learns only limited phonemes and prosodic symbols. The labeling model W1a is a model that takes the acoustic feature amount of speech data as input and outputs label data La. The labeling model W1a has an encoder and a decoder.
[0078] The encoder has a plurality of LSTMs and inputs the feature amount x of speech data. The speech recognition unit 11 generates the feature amount x to be input to the encoder of the labeling model W1a from the speech data. The feature amount x is, for example, a mel spectrogram of a window of a predetermined width (for example, 25 ms) shifted every predetermined time width (for example, 10 ms) smaller than the window. The speech recognition unit 11 downsamples the feature amount x for a predetermined number of frames and inputs it to the encoder. The encoder maps the input feature amount x to a feature representation h enc of another numerical vector and outputs it. The attention determines where to focus on the feature representation h i in order for the decoder to predict the next output y enc and outputs an attention context c i indicating the result. The decoder inputs the attention context c i and the previous output y i-1 and generates the output y i-1 when the previous outputs y i …, y0 and the feature amount x are given. The label data La is generated by arranging the outputs of the decoder.
[0079] (Reference 10) C. Chiu, et al., "State-of-the-Art Speech Recognition with Sequence-to-Sequence Models," 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2018.
[0080] According to the embodiments described above, the learning data generation device 1 of this embodiment can generate data for training a speech synthesis model using speech recognition technology, even from speech alone.
[0081] The aforementioned learning data generation device 1 and speech synthesis device 5 each have an internal computer system. The operation process of the learning data generation device 1 and speech synthesis device 5 is stored in program format on a computer-readable recording medium, and the above processing is performed by reading and executing this program on the computer system. The computer system referred to here includes hardware such as a CPU (Central Processing Unit), various types of memory, an OS (Operating System), and peripheral devices. Furthermore, all or part of the functions of the learning data generation device 1 and speech synthesis device 5 may be implemented using hardware such as an ASIC (Application Specific Integrated Circuit), a PLD (Programmable Logic Device), or an FPGA (Field Programmable Gate Array).
[0082] Furthermore, "computer system" includes the web page provisioning environment (or display environment) if a WWW system is being used. Also, "computer-readable recording medium" refers to portable media such as flexible disks, magneto-optical disks, ROMs, CD-ROMs, and storage devices such as hard disks built into computer systems. Moreover, "computer-readable recording medium" also includes those that dynamically hold programs for a short period of time, such as communication lines used when transmitting programs via networks such as the Internet or communication lines such as telephone lines, and those that hold programs for a certain period of time, such as volatile memory inside computer systems that act as servers or clients in such cases. Furthermore, the above-mentioned program may be intended to implement only a part of the functions described above, and may also be able to implement the above-mentioned functions in combination with programs already recorded in the computer system.
[0083] The learning data generation device 1 and the speech synthesis device 5 can each be implemented by, for example, one or more computer devices. When the learning data generation device 1 and the speech synthesis device 5 are implemented by multiple computer devices, the implementation of each functional unit by each computer device can be arbitrary. For example, the speech recognition unit 11, the learning data generation unit 13, and the speech processing unit 14 of the learning data generation device 1 may be implemented by different computer devices. Alternatively, the speech processing unit 14 may be implemented by an external speech editing device. Furthermore, the learning data generation device 1 that learns the speech recognition model W and the learning data generation device 1 that generates the learning data D1 using the learned speech recognition model W may be different devices. In this case, the learning data generation device 1 that learns the speech recognition model W operates as a learning device, and the learning data generation device 1 that generates the learning data D1 using the learned speech recognition model W operates as a data generation device.
[0084] According to the embodiments described above, the learning device comprises a learning unit. The learning unit is, for example, the speech recognition model learning unit 12 of the embodiment. The learning unit learns a labeling model that takes feature quantities obtained from speech data of an utterance as input and outputs label data of text including characters representing phonemes in the utterance and prosodic symbols representing accents, using pairs of training speech data and correct label data. For example, the learning unit learns the labeling model so that the difference between the label data obtained by inputting feature quantities from the training speech data into the labeling model and the correct label data corresponding to the training speech data becomes small.
[0085] The learning unit may train a label data correction model, which takes label data estimated using a labeling model as input and outputs label data corrected for phoneme errors in the input label data, using pairs of training label data containing errors and correct label data. For example, the learning unit inputs training label data correction data containing errors into the label data correction model and trains the label data correction model so that the difference between the obtained label data and the correct label data corresponding to the training label data becomes small.
[0086] The data generation device also includes a speech recognition unit. The speech recognition unit is a labeling model that takes feature quantities obtained from speech data of an utterance as input and outputs text label data including characters representing phonemes in the utterance and prosodic symbols representing accents. The labeling model is trained using pairs of training speech data and correct label data, and features quantities obtained from the speech data to be labeled are input to obtain the estimation result of label data representing the utterance of the speech data to be labeled. For example, the speech recognition unit uses a labeling model trained by a learning device.
[0087] The speech recognition unit is a label data correction model that takes label data estimated based on speech data as input and outputs label data with errors in the phonemes contained in that label data corrected. The label data correction model is trained using pairs of training label data containing errors and correct label data, and the label data estimated using a labeling model is input to obtain label data with errors corrected. For example, the speech recognition unit uses a label data correction model trained by a learning device.
[0088] The labeling model may include an encoder that takes time-series features obtained from audio data as input, and a decoder that takes the output from the encoder as input and outputs label data for the audio data into which the features have been input.
[0089] Furthermore, the labeling model may include a convolutional neural network that corresponds to each audio data for a predetermined time interval divided by a time shift, and takes the audio data of the corresponding time interval as input as features, and a transformer that takes the output from the convolutional network as input and outputs label data for the audio data in which the features have been input to the convolutional neural network.
[0090] The data generation device may further include a speech processing unit. The speech processing unit extracts speech data for each sentence from the speech data. The speech recognition unit inputs the speech data extracted by the speech processing unit into a labeling model as the target for label data estimation.
[0091] While embodiments of this invention have been described in detail above with reference to the drawings, the specific configuration is not limited to these embodiments and includes designs and the like that do not depart from the spirit of this invention. [Explanation of symbols]
[0092] 1…Training data generation device 5…Speech synthesis device 11…Voice recognition unit 12…Speech Recognition Model Learning Unit 13…Training Data Generation Unit 14…Sound Processing Unit 51…Acoustic feature estimation unit 52…Acoustic Feature Generation Model Learning Unit 53…Language Processing Department 54...Vocoder section 100...Speech synthesis system M... Acoustic feature generation model for speech synthesis W...Speech recognition model W1, W1a… Labeling Models W2…Label data correction model< / spkera> < / news> < / sad> < / tag> < / tag>
Claims
1. A learning unit that trains a labeling model, which takes speech data of an utterance or features obtained from the speech data as input and outputs label data of text described using characters representing phonemes and prosodic symbols representing accents in the utterance, using pairs of training speech data and correct label data. Equipped with, The learning unit takes label data estimated using the labeling model as input and trains a label data correction model that outputs label data corrected for errors in phonemes contained in the input label data, using pairs of training label data containing errors and correct label data. A learning device characterized by the following features.
2. A labeling model that takes speech data of an utterance or features obtained from the speech data as input and outputs label data of text described using characters representing phonemes and prosodic symbols representing accents in the utterance, wherein the labeling model is trained using pairs of training speech data and correct label data, and features obtained from the speech data to be labeled are input to obtain an estimation result of label data representing the utterance of the speech data to be labeled. Equipped with, The speech recognition unit is a label data correction model that inputs label data estimated based on speech data and outputs label data with errors in the phonemes included in the input label data corrected, wherein the label data correction model is trained using pairs of training label data containing errors and correct label data, and the label data estimated using the labeling model is input to obtain label data with errors corrected. A data generation device characterized by the following features.
3. The labeling model includes an encoder that takes time-series features obtained from audio data as input, and a decoder that takes the output from the encoder as input and outputs label data. The data generation apparatus according to feature 2.
4. The labeling model comprises a convolutional neural network that corresponds to each audio data for a predetermined time interval divided by a time shift, and takes the audio data of the corresponding time interval as input as a feature, and a transformer that takes the output from the convolutional network as input and outputs label data. The data generation apparatus according to feature 2.
5. It further includes a speech processing unit that extracts speech data for each sentence from the audio data. The speech recognition unit inputs the speech data extracted by the speech processing unit into the labeling model as the target for label data estimation. A data generation apparatus according to any one of claims 2 to 4.
6. A program for causing a computer to function as the learning device described in claim 1.
7. A program for causing a computer to function as a data generation device according to any one of claims 2 to 5.
Citation Information
Patent Citations
Voice synthesizing device
JP1990238494A
Continuous speech recognizing method
JP1993073094A
Voice communication system
JP2000356995A
Handwritten and voice input with auto-correction
JP2007524949A
Speech recognition device, method thereof, and program
JP2013072922A