Voice conversion device, machine learning method, voice conversion method, and program
Patent Information
- Application Number
- PCT/JP2023/039760
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-11-02
- Publication Date
- 2025-05-08
AI Technical Summary
When existing sound quality conversion technologies use sequence to sequence models, alignment may break, resulting in abnormal speech speed or broken speech content, affecting the practicality of the application.
By introducing machine learning models into sound quality conversion devices, using input and output sound spectrum as training data, combining attention mechanisms and dynamic time scaling technology, the alignment stability of sound quality conversion is achieved.
It effectively suppresses the collapse of voice content, improves the robustness and stability of sound quality conversion, and ensures high-quality output of voice conversion.
Smart Images

Figure JP2023039760_08052025_PF_FP_ABST
Abstract
Description
Voice conversion device, machine learning method, voice conversion method, and program
[0001] The present disclosure relates to a model training technique for voice conversion that converts the speaker characteristics of a voice.
[0002] Voice conversion (VC) is a technology that converts the speaker identity of the input speech to that of the target speaker while preserving the speech content. Statistical model-based voice conversion methods are broadly divided into two categories based on the training data provided. The first is parallel VC, in which the speech content of the source and target speakers is the same due to the correspondence between their respective speech features. The second is non-parallel VC, in which the speech content of the two speakers can be different. The first has been studied for a long time and has the disadvantage of being expensive to collect parallel data, but has the advantage of being easier to model than non-parallel VC and of higher quality. The second is that machine learning is possible with non-parallel data, but has the disadvantage of being difficult to ensure both the intended speech content and the speaker identity of the target speaker during inference.
[0003] Furthermore, the most common method for achieving the first type of parallel VC is to match the number of input and output frames in training data before modeling in order to solve the problem of different speaking rates even when the speech content of the input and output voices is the same. This method uses dynamic time warping (DTW) to stretch or shrink acoustically similar frames to match their lengths, but there is no guarantee that the acoustic features of the input and output will match the same phonemes after stretching or shrinking. Therefore, the accuracy of VC has the drawback of being heavily dependent on the accuracy of DTW.
[0004] To address this drawback, recent advances in sequence-to-sequence (also referred to as seq-to-seq) models have led to the proposal of methods capable of training parallel VC without DTW (Non-Patent Documents 1 and 2). These methods use an attention mechanism to represent the correspondence (alignment) between input and output speech and optimize the model, including the alignment along with the spectral generation error. Non-Patent Document 1 discloses a method that uses an automatic speech recognition (ASR) model as a sequence-to-seq encoder and uses its output posterior probability or intermediate output as input to the decoder to ensure the speech content while reducing the speaker identity of the input speech. On the other hand, Non-Patent Document 2 discloses a method that directly represents the relationship between input and output speech using an attention mechanism. Both methods can achieve high-quality VC.
[0005] Songxiang Liu, Yuewen Cao, Disong Wang, Xixin Wu, Xunying Liu, Helen Meng, "Any-to-Many Voice Conversion with Location-Relative Sequence-to-Sequence Modeling," IEEE. Trans. on Audio, Speech, and Language Processing, vol. 29, pp. 1717-1728, 2021.Hirokazu Kameoka, Kou Tanaka, Damian Kwasny, Takuhiro Kaneko, Nobukatsu Hojo, "ConvS2S-VC: Fully Convolutional Sequence-to-Sequence Voice Conversion," IEEE. Trans. on Audio, Speech, and Language Processing, vol. 28, pp. 1849-1863, 2020.
[0006] However, seq-to-seq based VC alignment can fail if the attention mechanism does not work properly. In this case, not only does the speech rate become abnormally fast or slow, but the speech content also breaks down, with missing utterances and repeated sentences. Because maintaining the speech content of the input speech is essential for VC, this breakdown significantly impairs the practicality of the application.
[0007] The present disclosure has been made in consideration of the above circumstances, and aims to achieve robust voice conversion by suppressing breakdown in speech content.
[0008] In order to achieve the above object, the invention of claim 1 provides a voice conversion device that performs machine learning to generate a voice conversion model using training data consisting of a set of a spectrum of a source speaker and a spectrum of a target speaker whose speech content is identical to that of the source speaker, the device comprising: a spectrum alignment unit that converts at least one of the spectrum of the source speaker and the spectrum of the target speaker based on the spectrum of the source speaker and the spectrum of the target speaker so that the number of frames of the spectrum of the source speaker and the spectrum of the target speaker match, and outputs the aligned spectrum of the source speaker and the aligned spectrum of the target speaker; The voice conversion device includes an input spectrum encoder that receives a spectrum of a source speaker and outputs encoded first features, a decoder that decodes the first features and outputs a predicted spectrum, and an output spectrum encoder that receives the predicted spectrum and outputs encoded second features, wherein the decoder recursively decodes the current first features based on the second features previously encoded by the output spectrum encoder, a voice conversion model, and a pre-training unit that performs pre-training of each parameter of the voice conversion model using the aligned target speaker spectrum as ground truth data for the predicted spectrum.
[0009] As described above, the present disclosure provides the effect of realizing robust voice conversion by suppressing breakdown in speech content.
[0010] FIG. 1 is a functional configuration diagram of a voice conversion device in a main training phase according to the first embodiment. FIG. 2 is a conceptual diagram showing the processing contents of a phoneme label correspondence generation unit. FIG. 3 is a functional configuration diagram of a training unit according to the first embodiment. FIG. 4 is a functional configuration diagram of a voice conversion device in an inference phase according to the first embodiment. FIG. 5 is an electrical hardware configuration diagram of a voice conversion device. FIG. 6 is a flowchart showing processing in a main training phase according to the first embodiment. FIG. 7 is a flowchart showing processing in a main training phase according to the first embodiment. FIG. 8 is a flowchart showing processing in an inference phase according to the first embodiment. FIG. 9 is a diagram showing a comparison between results of voice conversion according to the first embodiment and results of voice conversion according to conventional technology. FIG. 10 is a diagram showing subjective evaluation scores on a five-point scale regarding speech naturalness according to the first embodiment. FIG. 11 is a diagram showing subjective evaluation scores on a five-point scale regarding speaker similarity according to the first embodiment. FIG. 12 is a functional configuration diagram of a voice conversion device in a main training phase according to the second embodiment. FIG. 13 is a functional configuration diagram of a voice conversion device in an inference phase according to the second embodiment. 1 is a flowchart showing processing in the main training phase according to a second embodiment. 2 is a flowchart showing processing in the main training phase according to a second embodiment. 3 is a flowchart showing processing in the inference phase according to a second embodiment. 4 is a functional configuration diagram of a sound quality conversion device in the pre-training phase according to a third embodiment. 5 is a flowchart showing processing in the pre-training phase according to a third embodiment. 6 is a flowchart showing processing in the pre-training phase according to a third embodiment. 7 is a diagram showing the temporal relationship between the number of frames of a correct spectrum (U) and a predicted spectrum (T) obtained using correct data. 8 is a diagram showing the absolute error between a predicted spectrum and a correct spectrum obtained from a model before machine learning of the voice conversion model 13a is started, using random initial values for the voice conversion model 13a according to the first embodiment. 9 is a diagram showing the absolute error between a predicted spectrum and a correct spectrum obtained from a model before machine learning of the voice conversion model 13a is started, using a pre-trained voice conversion model 13p for the voice conversion model 13a according to a third embodiment. 10 is a diagram showing WER and MCD according to the number of pre-training steps according to a third embodiment.10 is a diagram showing subjective evaluation scores on a five-point scale regarding the naturalness of speech according to the third and fourth embodiments; FIG. 11 is a diagram showing subjective evaluation scores on a five-point scale regarding speaker similarity according to the third and fourth embodiments; FIG. 12 is a functional configuration diagram of a learning unit according to the fourth embodiment; and FIG. 13 is a functional configuration diagram of a part of a sound quality conversion device in a pre-learning phase according to the fifth embodiment.
[0011] First Embodiment First, the first embodiment will be described with reference to the drawings.
[0012] This embodiment provides a model that is less likely to fail in alignment in sequence-to-sequence (also referred to as seq-to-seq) based voice conversion (VC) and its machine learning (also simply referred to as "this training") method. This embodiment applies the feedback technology of recurrent neural network transducer (RNN-T) (see Reference 1), which is widely used in automatic speech recognition (ASR), to VC. While ASR outputs a character string from a training model, VC outputs a predicted spectrum of a target speaker from a training model.
[0013] (Reference 1) Alex Graves, "Sequence transduction with recurrent neural networks," Proc. ICML Representation Learning Workshop, 2012. Since ASR requires receiving speech and outputting the recognition result text in real time, it is highly desirable to ensure that alignment does not fail. RNN-T is a seq-to-seq model that is effective in ASR because it is less likely to fail alignment and is capable of streaming processing. This embodiment describes a method for realizing VC based on RNN-T to take advantage of this advantage. Based on this, Modifications 1 and 2 are also described.
[0014] The voice conversion device 10 is configured by one or more computers. When the voice conversion device 10 is configured by multiple computers, it may be referred to as a "voice conversion device" or a "voice conversion system."
[0015] [Functional Configuration of Voice Conversion Device] <Functional Configuration of Main Learning Phase> First, the functional configuration of the voice conversion device 10 in the main learning phase will be described with reference to Fig. 1. Fig. 1 is a functional configuration diagram of the voice conversion device in the main learning phase according to the first embodiment. In the main learning phase, the voice conversion device 10 performs machine learning on a voice conversion model 13a.
[0016] 1, the voice conversion device 10 includes a spectrum calculation unit 12, a voice conversion model 13a, a learning unit 14, and a correspondence generation unit 15. These units each have a function realized by commands from the CPU 101 in FIG.
[0017] The spectrum calculation unit 12 calculates and outputs the spectrum of the source speaker based on the input voice of the source speaker. Note that if the voice conversion device 10 can acquire the spectrum of the source speaker from an external device, the spectrum calculation unit 12 is not necessary.
[0018] The voice conversion model 13a is a machine learning model based on a neural network, and is composed of an input spectrum encoder 131, an output spectrum encoder 132, a decoder 133, a transition probability prediction unit 134, and a spectrum prediction unit 135. These are modules capable of machine learning, each of which is composed of a neural network and has its own model parameters.
[0019] Of these, the input spectrum encoder 131 receives an input spectrum as the spectrum of the source speaker from the spectrum calculation unit 12, encodes this input spectrum, and outputs an input spectrum (an example of a first feature).
[0020] The output spectrum encoder 132 receives, via feedback, the output spectrum (an example of the second feature) as the predicted spectrum of the target speaker output by the spectrum prediction unit 135 (originally the decoder 133), encodes this output spectrum, and outputs it.
[0021] The decoder 133 recursively decodes the current input spectrum (an example of the first feature) input from the input spectrum encoder 131, taking into consideration whether it is possible to output a spectrum of the speech content corresponding to a past input spectrum (an example of the second feature) based on the output spectrum input from the output spectrum encoder 132, and outputs a predicted spectrum.
[0022] The transition probability prediction unit 134 predicts transition probabilities based on the input spectrum of the source speaker input from the decoder 133, and outputs the predicted transition probabilities to the training unit 14. The training unit 14 will be described later.
[0023] The spectrum prediction unit 135 generates and outputs a predicted spectrum of the target speaker based on the input spectrum of the source speaker input from the decoder 133. When the spectrum prediction unit 135 calculates the predicted spectrum of the target speaker at the next time (u+1), it inputs the spectrum at the current time (u) (predicted spectrum of the target speaker) as the output spectrum by feedback to the output spectrum encoder 132. The spectrum prediction unit 135 repeats this feedback process to generate predicted spectra of the target speaker for the number of frames of the input spectrum.
[0024] The decoder 133 may include a spectrum prediction unit 135 and a transition probability prediction unit 134. For example, if the spectrum of a speech is represented by an 80-dimensional vector sequence, the decoder 133 predicts 81 dimensions including a one-dimensional transition probability.
[0025] On the other hand, the correspondence generation unit 15 inputs the phoneme labels with time information of the source speaker ("alignment" in the narrow sense) and the phoneme labels with time information of the target speaker ("alignment" in the narrow sense), generates a correct alignment for machine learning that shows the correspondence between both phoneme labels, and outputs the correct alignment to the learning unit 14.
[0026] (Phoneme Label Correspondence) Here, the generation of phoneme label correspondence will be described with reference to FIG. 2. FIG. 2 is a conceptual diagram illustrating the processing content of the phoneme label correspondence generation unit. FIG. 2(a) shows a source speaker's phoneme label s1 with time information when the source speaker's spectrum ss is generated from the source speaker's speech waveform sw. FIG. 2(b) shows a target speaker's phoneme label t1 with time information when the target speaker's spectrum ts is generated from the target speaker's speech waveform tw. Here, the first half of the utterance content "Thank you" is shown. As shown in FIG. 2, the correspondence generation unit 15 generates and outputs a correct alignment indicating the correspondence of the phoneme label from the source speaker's time-information-added phoneme label s1 to the target speaker's time-information-added phoneme label t1. This correspondence will be described in more detail later.
[0027] (Functional Configuration of Learning Unit) Next, the functional configuration of the learning unit 14 will be described with reference to Fig. 3. Fig. 3 is a diagram showing the functional configuration of the learning unit.
[0028] As shown in FIG. 3, the learning unit 14 includes an alignment calculation unit 141 , an alignment mask unit 142 , a spectral weighting unit 143 , a spectral error calculation unit 144 , and a model optimization unit 145 .
[0029] Of these, the alignment calculation unit 141 calculates and outputs a predicted alignment based on the predicted transition probability input from the transition probability prediction unit 134 .
[0030] The alignment masking unit 142 generates and outputs a masked predicted alignment based on the predicted alignment input from the alignment calculation unit 141 and the correct alignment input from the correspondence generation unit 15 described above.
[0031] The spectral weighting unit 143 generates and outputs a weighted predicted spectrum of the target speaker based on the masked predicted alignment input from the alignment masking unit 142 and the predicted spectrum of the target speaker input from the above-mentioned spectral prediction unit 135.
[0032] The spectral error calculation unit 144 calculates the error between the weighted predicted spectrum of the target speaker input from the spectral weighting unit 143 and the correct spectrum of the target speaker, and outputs the spectral error.
[0033] The model optimization unit 145 updates each model parameter of the voice conversion model 13a (the input spectrum encoder 131, the output spectrum encoder 132, the decoder 133, the spectrum prediction unit 135, and the transition probability prediction unit 134) by error backpropagation so as to minimize the spectrum error input from the spectrum error calculation unit 144. When the model optimization unit 145 treats the spectrum as a continuous value as it is, mean absolute error, mean squared error, etc. can be used as the spectrum error. When the model optimization unit 145 treats the spectrum as a discrete representation by vector quantization, etc., cross entropy may also be used.
[0034] The voice conversion device 10 refines the voice conversion model by repeatedly applying the processing flow configured as above to the training data.
[0035] In the explanation so far, the acoustic features used for input and output have been referred to as spectra, but it is also possible to use spectrograms obtained by short-time Fourier transforms, or mel-spectrograms, which are low-dimensional representations of spectrograms adapted to human hearing. It is also possible to use mel-cepstrum, which is often used in speech synthesis, or a vector that concatenates the fundamental frequency, which is the pitch of the voice.
[0036] <Background Technology Leading to Present Embodiment, Detailed Technology of Present Embodiment> Next, the background technology leading to the present embodiment and detailed technology of the present embodiment will be described.
[0037] In the case of automatic speech recognition (ASR), since the target is a discrete symbol such as a phoneme, the transition probability of an RNN-T can be optimized by minimizing cross-entropy by adding dwell and transition symbols in addition to these symbols and minimizing the cross-entropy of the predictor for whether to transition to the next phoneme. However, in the case of voice conversion (VC), the output is a continuous spectrum, and the scales of the output values of the transition probability and dwell probability predictor (134) and the spectral predictor (135) are significantly different, making it difficult to balance the training of the two. To solve this problem, Reference 2 proposes a lazy forward algorithm for applying RNN-T to speech synthesis tasks where the output is a spectrum, similar to VC.
[0038] (Reference 2) Jiawei Chen, Xu Tan, Yichong Leng, Jin Xu, Guihua Wen, Tao Qin, Tie-Yan Liu, "Speech-T: Transducer for text to speech and beyond," Proc. NeurIPS, vol.34, pp. 6621-6633, 2021. This replaces the transition probability used in the probability calculation of the input / output trellis with the value obtained from the stay probability φ, as shown in the following (Equation 1).
[0039] Here, t and u are the time indices of the input text and output spectrum, respectively, and α(t,u) is the forward probability (= predicted alignment) on the trellis for the t-th input text and the output spectrum of frame u. The transition probability is given by 1-φ(t,u-1). The recurrence formula for this α is calculated for each lattice of the trellis, and the loss function is defined as follows (Equation 2).
[0040] where:
[0041] are the target spectrum and predicted spectrum of the target speaker at time u+1, respectively.
[0042] By designing the loss function in this way, even regression tasks in which the model requires continuous values can be realized within the RNN-T framework. Note that since calculating all trellis elements is unrealistic in terms of both training stability and computational complexity, the voice conversion device 10 masks the predicted alignment using the correct alignment and omits the trellis calculation for impossible paths.
[0043] Furthermore, in this embodiment, focusing on the advantages of the Lazy Forward algorithm, parallel VC is realized by using the spectra of the source speaker and the target speaker as the input and output of Reference 2, respectively. However, unlike speech synthesis, the temporal correspondence between the input spectrum and the output spectrum is not explicitly given. In this embodiment, to realize VC using RNN-T, phoneme alignments of the source speaker and the target speaker are used, and a correct alignment is generated from the correspondence between the alignments. However, even if the phonemes are the same, their sequence lengths are not necessarily the same, so the correct alignment matrix will not be a square matrix and some ingenuity is required.
[0044] Therefore, the voice conversion device 10 of this embodiment has a correspondence generation unit 15 as a function that applies a constraint method (see Reference 3) that forces a non-square alignment matrix to be monotonic. The correspondence generation unit 15 is a means for providing the correct solution for the above α.
[0045] (Reference 3) Hideyuki Tachibana, Katsuya Uenoyama, Shunsuke Aihara, "Efficiently trainable text-to-speech system based on deep convolutional networks with guided attention," Proc. ICASSP, pp.4784-4788, 2018. The technique in Reference 3 designs the following penalty matrix, expressed between 0 and 1, which imposes a small penalty on monotonic alignments and increases the penalty as the alignment deviates from the monotonic alignment.
[0046] where T and U are the number of frames of the input spectrum and the output spectrum, respectively. In other words, when (Equation 5) is a unit matrix, the correct alignment is calculated as argmax for each row of (Equation 6).
[0047]
[0048] However, because this is an approximately calculated correct alignment, the correspondence generation unit 15 can stabilize the convergence of the alignment by considering not only the argmax frames of each row but also a predetermined number of frames before and after them (for example, one frame before and after) as correct when constructing the correspondence. This allows the correspondence generation unit 15 to generate a correct alignment even when, unlike automatic speech recognition (ASR), the speaking time of the source speaker and the speaking time of the target speaker differ, as in voice conversion (VC) (see FIG. 2).
[0049] <Functional Configuration of Inference Phase> Next, the functional configuration of the voice conversion device in the inference phase will be described with reference to Fig. 4. Fig. 4 is a diagram showing the functional configuration of the voice conversion device in the inference phase according to the first embodiment.
[0050] As shown in Fig. 4, the voice conversion device 10 has a spectrum calculation unit 12, a trained voice conversion model 13b, a speech waveform generation unit 16, and an end determination unit 17. These units have functions realized by commands from the CPU 101 in Fig. 5 based on a program. Note that functional components similar to those in the training phase are denoted by the same reference numerals and will not be described again.
[0051] The voice conversion model 13b has the same functional configuration (modules) as the voice conversion model 13a in this learning phase.
[0052] The speech waveform generating unit 16 generates and outputs the speech of the target speaker based on the predicted spectrum of the target speaker input from the spectrum predicting unit 135 .
[0053] The termination determination unit 17 determines whether or not to terminate output of the predicted spectrum of the target speaker to the spectrum prediction unit 135 based on the predicted transition probabilities input from the transition probability prediction unit 134 (see Reference 1). Because the termination determination unit 17 is trained to transition at the timing when the phonemes of the input spectrum change, the termination determination unit 17 causes the spectrum prediction unit 135 to continue outputting the predicted spectrum of the target speaker until it has output a spectrum with the same number of phonemes as the input spectrum. However, unlike ASR, in VC it is not possible to determine termination before the source speaker has finished speaking, so if the number of frames in the output spectrum is equal to or less than the number of frames in the input spectrum, termination will not occur.
[0054] [Hardware Configuration] Next, the electrical hardware configuration of the voice conversion device 10 will be described with reference to Fig. 5. Fig. 5 is a diagram showing the electrical hardware configuration of the voice conversion device.
[0055] As shown in FIG. 5, the voice conversion device 10 is a computer and includes a CPU 101, a ROM 102, a RAM 103, an SSD 104, an external device connection I / F (Interface) 105, a network I / F 106, a display 107, an input device 108, a media I / F 109, and a bus line 110.
[0056] Of these, the CPU 101 controls the overall operation of the voice conversion device 10. The ROM 102 stores programs such as IPL used to drive the CPU 101. The RAM 103 is used as a work area for the CPU 101.
[0057] The SSD 104 reads or writes various data under the control of the CPU 101. Note that instead of the SSD 104, a hard disk drive (HDD) may be used.
[0058] The external device connection I / F 105 is an interface for connecting various external devices, such as a display, a speaker, a keyboard, a mouse, a USB memory, and a printer.
[0059] The network I / F 106 is an interface for performing data communication via the communication network 100 .
[0060] The display 107 is a type of display means such as a liquid crystal display or organic electroluminescence (EL) display that displays various images.
[0061] The input device 108 is a keyboard, a pointing device, or the like, and is an example of an input means for accepting input operations such as characters, numbers, and various instructions.
[0062] The media I / F 109 controls reading and writing (storing) of data from and to a recording medium 109m such as a flash memory, etc. The recording medium 109m includes DVDs and Blu-ray Discs (registered trademarks).
[0063] The bus line 110 is an address bus, a data bus, etc. for electrically connecting the components such as the CPU 101 shown in FIG.
[0064] [Processing or Operation of First Embodiment] Next, the processing or operation of the voice conversion device 10 in the main learning phase and inference phase will be described with reference to FIGS.
[0065] <Processing or Operation in the Main Learning Phase> FIGS. 6 and 7 are flowcharts showing processing in the main learning phase according to the first embodiment.
[0066] S110: First, as shown in FIG. 6, the spectrum calculation unit 12 calculates and outputs the spectrum of the source speaker based on the input speech of the source speaker.
[0067] S111: The input spectrum encoder 131 receives the current input spectrum as the spectrum of the source speaker from the spectrum calculation unit 12, encodes this input spectrum, and outputs it.
[0068] S112: The output spectrum encoder 132 receives, as feedback, the past output spectrum as the predicted spectrum of the target speaker output by the spectrum prediction unit 135 (originally the decoder 133), and encodes and outputs this output spectrum.
[0069] S113: The decoder 133 decodes the current input spectrum input from the input spectrum encoder 131, taking into consideration whether it is possible to output a spectrum of the speech content corresponding to the past input spectrum based on the output spectrum input from the output spectrum encoder 132.
[0070] S114: The transition probability prediction unit 134 predicts and outputs transition probabilities based on the current input spectrum of the source speaker input from the decoder 133. Furthermore, the spectrum prediction unit 135 generates and outputs a predicted spectrum of the target speaker based on the current input spectrum of the source speaker input from the decoder 133.
[0071] S115: As shown in FIG. 7, the alignment calculation unit 141 calculates and outputs a predicted alignment based on the predicted transition probabilities input from the transition probability prediction unit 134.
[0072] S116: The correspondence generation unit 15 receives the phoneme labels with time information of the source speaker and the phoneme labels with time information of the target speaker to generate a correct alignment, and outputs the correct alignment to the training unit 14.
[0073] S117: The alignment masking unit 142 generates and outputs a masked predicted alignment based on the predicted alignment input from the alignment calculation unit 141 and the correct alignment input from the correspondence generation unit 15 described above.
[0074] S118: The spectral weighting unit 143 generates and outputs a weighted predicted spectrum of the target speaker based on the masked predicted alignment input from the alignment masking unit 142 and the predicted spectrum of the target speaker (current input spectrum) input from the above-mentioned spectral prediction unit 135.
[0075] S119: The spectrum error calculation unit 144 calculates the error between the weighted predicted spectrum of the target speaker input from the spectrum weighting unit 143 and the correct spectrum of the target speaker, and outputs the spectrum error.
[0076] S120: The model optimization unit 145 updates each parameter of the voice conversion model 13a by backpropagation so as to minimize the spectral error input from the spectral error calculation unit 144.
[0077] The voice conversion device 10 repeats the above steps S110 to S120 to complete the machine learning of the voice conversion model 13a.
[0078] <Processing or Operation in Inference Phase> FIG. 8 is a flowchart showing processing in the inference phase according to the first embodiment.
[0079] S150: First, as shown in FIG. 8, the spectrum calculation unit 12 calculates and outputs the spectrum of the source speaker based on the input speech of the source speaker.
[0080] S151: The input spectrum encoder 131 receives the current input spectrum as the spectrum of the source speaker from the spectrum calculation unit 12, encodes this input spectrum, and outputs it.
[0081] S152: The output spectrum encoder 132 receives, as feedback, the past output spectrum as the predicted spectrum of the target speaker output by the spectrum prediction unit 135 (originally the decoder 133), and encodes and outputs this output spectrum.
[0082] S153: The decoder 133 decodes the current input spectrum input from the input spectrum encoder 131, taking into consideration whether it is possible to output a spectrum of the speech content corresponding to the past input spectrum based on the output spectrum input from the output spectrum encoder 132.
[0083] S154: The transition probability prediction unit 134 predicts and outputs transition probabilities based on the current input spectrum of the source speaker input from the decoder 133. Furthermore, the spectrum prediction unit 135 generates and outputs a predicted spectrum of the target speaker based on the current input spectrum of the source speaker input from the decoder 133.
[0084] S155: The speech waveform generating unit 16 generates and outputs the speech of the target speaker based on the predicted spectrum of the target speaker input from the spectrum predicting unit 135.
[0085] S156: The termination determination unit 17 determines whether or not to terminate the output of the predicted spectrum of the target speaker based on the predicted transition probabilities input from the transition probability prediction unit 134. If not (NO), the process returns to step S150. On the other hand, if terminated (YES), the inference phase process ends.
[0086] [Experimental Results] Next, experimental results of this embodiment and the prior art will be compared with each other using FIGS.
[0087] FIG. 9 is a diagram showing a comparison between the results of voice conversion according to the first embodiment and the results of voice conversion according to the prior art.
[0088] Figure 9 shows the Mel-Cepstrum Distortion (MCD) and Character Error Rate (CER) for each method. The "Method" column lists the speech input to the speech recognizer: speech recorded in a studio or the like (GT, RESYN), speech obtained using the conventional technique p1 described in Non-Patent Document 1 (BNE-S2SMoL-VC), speech obtained using the conventional technique p2 described in Non-Patent Document 2 (ConvS2S-VC), and speech obtained using the method e1 of this embodiment (VC-T). GT represents the case where the speech recorded in a studio or the like is input as is. RESYN represents the case where the speech waveform recorded in a studio or the like is first converted into a spectrum and then converted back into a speech waveform before input. For conventional technique p2 and this embodiment, the results are also shown for a case where the entire sequence of the source speaker's speech is input to the input spectrum encoder at once, taking future information into account (OFFLINE), and a practical use case where speech is input frame by frame without using future information for streaming operation (STREAM). To obtain these results, Reference 4 was used for the speech recognizer, and Reference 5 was used for the conversion from spectrum to speech waveform generation.
[0089] (Reference 4) Alexei Baevski, Henry Zhou, Abdelrahman Mohamed, Michael Auli, "wav2vec 2.0: A framework for self-supervised learning of speech representations," Proc. NeurIPS, pp. 12449-12460, 2020. (Reference 5) Kong, J. Kim, and J. Bae, "HiFi-GAN: Generative adversarial networks for efficient and high fidelity speech synthesis," Proc. NeurIPS, vol. 33, pp. 17–022–17–033, 2020. CER indicates the character error rate when speech is input to a speech recognizer; a lower value indicates better results. In other words, the closer the CER value is to that of GT, the more easily voice conversion can be performed without disrupting the speech content.
[0090] F2F indicates the case of speech obtained by voice conversion from a female voice to a female voice. M2F indicates the case of speech obtained by voice conversion from a male voice to a female voice. F2M indicates the case of speech obtained by voice conversion from a female voice to a male voice. In general, voice conversion between opposite-sex people is more difficult than voice conversion between same-sex people, and tends to result in a higher character error rate.
[0091] Note that the conventional techniques p1 and p2 and the method of this embodiment require restoring a speech waveform from a voice-converted spectrum. For this reason, the character error rate of the recorded speech (GT) as well as that of RESYN are shown for reference as the upper limit performance when restoring a speech waveform from a spectrum without using voice conversion.
[0092] As shown in Figure 9, the voice conversion method of this embodiment is superior to the conventional techniques p1 and p2 not only in same-sex voice conversion but also in opposite-sex voice conversion. In particular, the results of opposite-sex voice conversion using the conventional techniques p1 and p2 are extremely degraded, while no significant degradation is observed using this embodiment. The reason why this embodiment is superior is that, for example, when the speech content is "yoroshiku," the voice conversion device 10 can take into account both (1) the correspondence between the spectrum of the currently input "Y" and the spectrum of the "Y" it has output, and (2) the transition probability of the phoneme corresponding to the spectrum, such as from "Y" to "O."
[0093] 10 and 11 show subjective evaluation scores on a five-point scale for voice naturalness and speaker similarity, respectively. Because voice conversion involves spectral conversion, the scores are always lower than the evaluation results of RESYN, which does not involve conversion. FIG. 11 shows the evaluation results when comparing each method, excluding RESYN. As shown in the Avg. column in FIG. 11, the evaluation results tend to be worse in the case where no practical future information is used (STREAM) than in the case where operation is not possible until the end of the utterance (OFFLINE). However, the degree of degradation in this embodiment was less than that of conventional technology p2. Furthermore, statistical testing showed no significant difference between STREAM and OFFLINE in this embodiment, indicating no significant degradation.
[0094] [Major Effects of the First Embodiment] As described above, the first embodiment has the effect of realizing a VC that is robust against alignment breakdown by suppressing breakdowns in speech content. Furthermore, since it is desirable in VC to output converted speech immediately after inputting speech, RNN-T is also strong in streaming processing, making it possible to provide a VC with a comfortable response to the user.
[0095] Specifically, to avoid the problem of alignment failure, which was an issue with VCs of conventional inventions based on the seq-to-seq model, the voice conversion device 10 inputs not only the input spectrum but also an output spectrum based on the past (immediately preceding) input spectrum to the decoder via feedback. This allows the decoder 133 to consider whether it is able to output a spectrum of the speech content that corresponds to the input spectrum, which helps to avoid alignment failure.
[0096] Another reason why alignment failure can be avoided is that the voice conversion device 10 learns the transition probability to the next phoneme together with the spectral error. The voice conversion device 10 weights the predicted spectrum based on the ground truth alignment and the predicted alignment derived from the transition probability, and calculates the spectral error based on the trellis of the input and output frames. This makes the predicted alignment monotonic and robust to failure.
[0097] [Modification 1] Next, a description will be given of Modification 1 of the first embodiment. The difference from the first embodiment is that the speaker vector of the target speaker is input to the decoder 133.
[0098] 1, in Modification 1, the decoder 133 receives the speaker vector of the target speaker as well as the input spectrum from the input spectrum encoder 131 and the output spectrum from the output spectrum encoder 132. Also, in the inference phase of FIG. 4, the decoder 133 performs the same processing as in the main training phase.
[0099] Here, speaker vectors can be obtained by converting speaker IDs into one-hot vectors or by passing speaker IDs through a trainable embedding layer. Also, vectors obtained as intermediate representations for speaker recognition tasks (see Reference 6) and vectors obtained from pre-trained models trained on huge data sets (see Reference 4) can also be used as speaker vectors. When using pre-trained models trained on a huge number of speakers for these other tasks, it is assumed that sufficient speaker representations have been acquired, making it possible to realize zero-shot VC, which converts voices of speakers not included in the VC training data as target speakers.
[0100] (Reference 6) David Snyder, Daniel Garcia-Romero, Gregory Sell, Daniel Povey, Sanjeev Khudanpur, "X-vectors: Robust DNN embeddings for speaker recognition," Proc. ICASSP, pp. 5329-5333, 2018. In the first embodiment described above, the source speaker and target speaker were limited to a one-to-one VC. In contrast, Modification 1 uses the speaker vector of the target speaker as conditioning for the decoder 133, making it possible to expand the target speaker to multiple speakers, i.e., the relationship between the source speaker and the target speaker to one-to-many.
[0101] [Modification 2] Next, a description will be given of Modification 2 of the first embodiment. The difference from Modification 1 is that not only the spectrum of the source speaker but also the speaker vector of the source speaker is input to the input spectrum encoder 131.
[0102] 1, in Variation 2, the input spectrum encoder 131 inputs the speaker vector of the source speaker along with the spectrum of the source speaker. Furthermore, as in Variation 1, the decoder 133 inputs the speaker vector of the target speaker along with the input spectrum from the input spectrum encoder 131 and the output spectrum from the output spectrum encoder 132.
[0103] In the inference phase of FIG. 4, the input spectrum encoder 131 and the decoder 133 perform the same processing as in the learning phase.
[0104] In the above-mentioned variant 1, the target speaker can be multiple, but the source speaker is limited to one. In variant 2, the source speaker is expanded to multiple people so that more people can use VC. Specifically, the speaker vector described in variant 1 is also calculated for the source speaker, and this and the source speaker's spectrum are provided as conditioning for the input spectrum encoder 131. This makes it possible to expand the relationship between source speakers and target speakers to many-to-many.
[0105] In the explanation so far, the acoustic features used for input and output have been referred to as spectra, but it is also possible to use spectrograms obtained by short-time Fourier transforms, or mel-spectrograms, which are low-dimensional representations of spectrograms adapted to human hearing. It is also possible to use mel-cepstrum, which is often used in speech synthesis, or a vector that concatenates the fundamental frequency, which is the pitch of the voice.
[0106] Second Embodiment Next, a second embodiment will be described with reference to the drawings. In the first embodiment, the main training and inference are performed using spectra extracted from speech waveforms, but in the second embodiment, an example will be shown in which the main training and inference are performed using speech waveforms as they are.
[0107] [Functional Configuration of Voice Conversion Device] <Functional Configuration of Main Learning Phase> First, the functional configuration of the voice conversion device 20 in the main learning phase will be described with reference to Fig. 12. Fig. 12 is a functional configuration diagram of the voice conversion device in the main learning phase according to the second embodiment. In the main learning phase, the voice conversion device 20 performs machine learning to generate a voice conversion model 23a. Note that the electrical hardware configuration of the voice conversion device 20 is the same as that of the voice conversion device 10 of the first embodiment (see Fig. 5), and therefore description thereof will be omitted.
[0108] 12, the voice conversion device 20 includes a voice conversion model 23a, a learning unit 24, and a correspondence generation unit 25. These units each have a function that is realized by commands from the CPU 101 in FIG. 5 based on a program.
[0109] The voice conversion model 23a is a machine learning model based on a neural network, and is composed of an input voice waveform encoder 231, an output voice waveform encoder 232, a decoder 233, a transition probability prediction unit 234, and a voice waveform prediction unit 235. These are modules capable of machine learning, each of which is composed of a neural network and has its own model parameters.
[0110] Of these, the input speech waveform encoder 231 receives an input speech waveform as the speech waveform of a source speaker from outside the voice conversion device 20, encodes this input speech waveform, and outputs it. During this learning phase, the input speech waveform encoder 231 can predict a speech waveform by having a DNN (Deep Neural Network) mask part of the input speech waveform and perform machine learning. As a result, the input speech waveform encoder 231, unlike the first embodiment, can input a speech waveform and output a speech waveform.
[0111] The output speech waveform encoder 232 receives, via feedback, the output speech waveform as the predicted speech waveform of the target speaker output by the speech waveform prediction unit 235 (originally the decoder 233), encodes this output speech waveform, and outputs it. Like the input speech waveform encoder 231, the output speech waveform encoder 232 can also predict the speech waveform using a DNN.
[0112] The decoder 233 decodes the current input audio waveform input from the input audio waveform encoder 231, taking into consideration whether it is possible to output an audio waveform of the speech content corresponding to the past input audio waveform, based on the output audio waveform input from the output audio waveform encoder 232.
[0113] The transition probability prediction unit 234 predicts transition probabilities based on the input speech waveform of the source speaker input from the decoder 233, and outputs the predicted transition probabilities to the training unit 24. The training unit 24 will be described later.
[0114] The speech waveform prediction unit 235 generates and outputs a predicted speech waveform of the target speaker based on the input speech waveform of the source speaker input from the decoder 233. When the speech waveform prediction unit 235 determines the predicted speech waveform of the target speaker at the next time (u+1), it inputs the speech waveform at the current time (u) (predicted speech waveform of the target speaker) as the output speech waveform by feedback to the output speech waveform encoder 232. The speech waveform prediction unit 235 repeats this feedback process to generate predicted speech waveforms of the target speaker for the number of frames of the input speech waveform.
[0115] The decoder 233 may include a speech waveform prediction unit 235 and a transition probability prediction unit 234. In this case, since the speech sequence is one-dimensional, the decoder 233 predicts two dimensions in total, including the one-dimensional transition probability.
[0116] On the other hand, the correspondence generation unit 25 performs the same processing as the correspondence generation unit 15 of the first embodiment, and therefore a description thereof will be omitted.
[0117] (Functional Configuration of Learning Unit) Next, the functional configuration of the learning unit 24 will be described with reference to Fig. 13. Fig. 13 is a diagram showing the functional configuration of the learning unit.
[0118] As shown in FIG. 13, the learning unit 24 includes an alignment calculation unit 241 , an alignment mask unit 242 , a voice waveform weighting unit 243 , a voice waveform reconstruction error calculation unit 244 , and a model optimization unit 245 .
[0119] Of these, the alignment calculation unit 241 calculates and outputs a predicted alignment based on the predicted transition probability input from the transition probability prediction unit 234 .
[0120] The alignment masking unit 242 generates and outputs a masked predicted alignment based on the predicted alignment input from the alignment calculation unit 241 and the correct alignment input from the correspondence generation unit 25 described above.
[0121] The speech waveform weighting unit 243 generates and outputs a weighted predicted speech waveform of the target speaker based on the masked predicted alignment input from the alignment masking unit 242 and the predicted speech waveform of the target speaker input from the above-mentioned speech waveform prediction unit 235.
[0122] The speech waveform reconstruction error calculation unit 244 calculates the error between the weighted predicted speech waveform of the target speaker input from the speech waveform weighting unit 243 and the correct speech waveform of the target speaker, and outputs the speech waveform reconstruction error.
[0123] The model optimization unit 245 updates each model parameter of the voice conversion model 23a (the input speech waveform encoder 231, the output speech waveform encoder 232, the decoder 233, the speech waveform prediction unit 235, and the transition probability prediction unit 234) by error backpropagation so as to minimize the reconstruction error of the speech waveform input from the speech waveform reconstruction error calculation unit 244. When the model optimization unit 245 treats the speech waveform as a continuous value as it is as the speech waveform error, mean absolute error, mean squared error, etc. can be used. When the model optimization unit 245 treats the speech waveform as a discrete representation by vector quantization, etc., cross entropy may also be used.
[0124] The voice conversion device 20 refines the voice conversion model by repeatedly applying the processing flow according to the above configuration to the training data.
[0125] <Functional Configuration of Inference Phase> Next, the functional configuration of the voice conversion device in the inference phase will be described with reference to Fig. 14. Fig. 14 is a functional configuration diagram of the voice conversion device in the inference phase according to the second embodiment.
[0126] As shown in Fig. 14, the voice conversion device 20 has a trained voice conversion model 23b and an end determination unit 27. These units are functions realized by commands from the CPU 101 in Fig. 5 based on a program. Note that the same functional components as those in the training phase are denoted by the same reference numerals and will not be described again.
[0127] The voice conversion model 23b has the same functional configuration (modules) as the voice conversion model 23a in the main learning phase.
[0128] The end determination unit 27 performs the same processing as the end determination unit 17 of the first embodiment.
[0129] [Processing or Operation of Second Embodiment] Next, the processing or operation of the voice conversion device 20 in the main learning phase and inference phase will be described with reference to FIGS.
[0130] <Processing or Operation in the Main Learning Phase> FIGS. 15 and 16 are flowcharts showing processing in the main learning phase according to the first embodiment.
[0131] S211: The input speech waveform encoder 231 receives a current input speech waveform as the speech waveform of the source speaker from outside the voice quality conversion device 20, encodes this input speech waveform, and outputs it.
[0132] S212: The output speech waveform encoder 232 receives, via feedback, the past output speech waveform as the predicted speech waveform of the target speaker output by the speech waveform prediction unit 235 (originally the decoder 233), and encodes and outputs this output speech waveform.
[0133] S213: The decoder 233 decodes the current input audio waveform input from the input audio waveform encoder 231, taking into consideration whether it is possible to output an audio waveform of the speech content corresponding to the past input audio waveform based on the output audio waveform input from the output audio waveform encoder 232.
[0134] S214: The transition probability prediction unit 234 predicts and outputs a transition probability based on the current input speech waveform of the source speaker input from the decoder 233. In addition, the speech waveform prediction unit 235 generates and outputs a predicted speech waveform of the target speaker based on the current input speech waveform of the source speaker input from the decoder 233.
[0135] S215: As shown in FIG. 16, the alignment calculation unit 241 calculates and outputs a predicted alignment based on the predicted transition probabilities input from the transition probability prediction unit 234.
[0136] S216: The correspondence generation unit 25 receives the phoneme labels with time information of the source speaker and the phoneme labels with time information of the target speaker to generate a correct alignment, and outputs the correct alignment to the training unit 24.
[0137] S217: The alignment masking unit 242 generates and outputs a masked predicted alignment based on the predicted alignment input from the alignment calculation unit 241 and the correct alignment input from the correspondence generation unit 25 described above.
[0138] S218: The speech waveform weighting unit 243 generates and outputs a weighted predicted speech waveform of the target speaker based on the masked predicted alignment input from the alignment masking unit 242 and the predicted speech waveform of the target speaker (current input speech waveform) input from the above-mentioned speech waveform prediction unit 235.
[0139] S219: The speech waveform reconstruction error calculation unit 244 calculates the error between the weighted predicted speech waveform of the target speaker input from the speech waveform weighting unit 243 and the correct speech waveform of the target speaker, and outputs the speech waveform reconstruction error.
[0140] S220: The model optimization unit 245 updates each parameter of the voice conversion model 23a by error backpropagation so as to minimize the speech waveform error input from the speech waveform reconstruction error calculation unit 244.
[0141] The voice conversion device 20 repeats the above steps S211 to S220 to complete the machine learning of the voice conversion model 23a.
[0142] <Processing or Operation in Inference Phase> FIG. 17 is a flowchart showing processing in the inference phase according to the first embodiment.
[0143] S251: The input speech waveform encoder 231 receives a current input speech waveform as the speech waveform of the source speaker from outside the voice quality conversion device 20, encodes this input speech waveform, and outputs it.
[0144] S252: The output speech waveform encoder 232 receives, via feedback, the past output speech waveform as the predicted speech waveform of the target speaker output by the speech waveform prediction unit 235 (originally the decoder 233), and encodes and outputs this output speech waveform.
[0145] S253: The decoder 233 decodes the current input audio waveform input from the input audio waveform encoder 231, taking into consideration whether it is possible to output an audio waveform of the speech content corresponding to the past input audio waveform based on the output audio waveform input from the output audio waveform encoder 232.
[0146] S254: The transition probability prediction unit 234 predicts and outputs a transition probability based on the current input speech waveform of the source speaker input from the decoder 233. In addition, the speech waveform prediction unit 235 generates and outputs a predicted speech waveform of the target speaker based on the current input speech waveform of the source speaker input from the decoder 233.
[0147] S255: The termination determination unit 27 determines whether or not to terminate the output of the predicted speech waveform of the target speaker based on the predicted transition probabilities input from the transition probability prediction unit 234. If not (NO), the process returns to step S251. On the other hand, if it is (YES), the process of the inference phase ends.
[0148] [Major Effects of the Second Embodiment] As described above, the second embodiment provides the same effects as the first embodiment.
[0149] Furthermore, in the second embodiment, a spectrum is not generated from the audio waveform, and machine learning and inference are performed on the audio waveform as is, so there is no mismatch with the audio waveform generation means that can occur in the first embodiment, and the quality is superior.
[0150] Third Embodiment Next, a third embodiment will be described with reference to the drawings. In the first embodiment, the voice conversion model 13a was trained, but in the third embodiment, an example is shown in which pre-training of the voice conversion model 13p is performed before the voice conversion model 13a is trained. That is, the voice conversion device 30 of the third embodiment pre-trains (machine learns) the voice conversion model 13p using training data consisting of a set of a spectrum of a source speaker and a spectrum of a target speaker whose utterance content is identical to that of the source speaker.
[0151] [Functional Configuration of Voice Conversion Device] <Functional Configuration of Pre-Training Phase> First, the functional configuration of the voice conversion device 30 in the pre-training phase will be described using FIG. 18 . FIG. 18 is a functional configuration diagram of a voice conversion device in the pre-training phase according to the third embodiment. In the pre-training phase, the voice conversion device 30 performs machine learning to generate a voice conversion model 23p. For reference, in FIG. 18 , the matrix of each spectrum is shown near each spectrum. T, U, and D are the number of frames of the source speaker's spectrum, the number of frames of the target speaker's spectrum, and the number of dimensions of the spectrum, respectively. For example, T = 1000, U = 1000, and D = 80. U' is the number of aligned frames of the aligned source speaker's or target speaker's spectrum. For example, U' = 1000.
[0152] Furthermore, the electrical hardware configuration of the voice conversion device 30 is the same as that of the voice conversion device 10 of the first embodiment (see FIG. 5), and therefore a description thereof will be omitted.
[0153] 12, the voice conversion device 20 includes a spectrum alignment unit 11, a voice conversion model 23p, and a pre-training unit 64. These units each have a function realized by instructions from the CPU 101 in FIG. 5 based on a program.
[0154] The spectrum alignment unit 11 converts at least one of the source speaker spectrum and the target speaker spectrum based on the source speaker spectrum and the target speaker spectrum so that the number of frames of the source speaker spectrum and the target speaker spectrum match, and outputs the aligned source speaker spectrum and aligned target speaker spectrum. To this end, the spectrum alignment unit 11 includes a path search unit 111 and alignment units 112a and 112b.
[0155] Of these, the path search unit 111 finds the path of the input source speaker and the path of the target speaker. The path search unit 111 executes dynamic programming, typified by the dynamic time warping method described above, to align the number of frames of two input spectra of different lengths. In this embodiment, if the number of frames in the source speaker and target speaker spectra are T and U, respectively, the path search unit 111 finds paths that make both the number of frames U'. A path is an index sequence that satisfies this.
[0156] The alignment units 112a and 112b have the function of applying paths given by the index sequence to the spectrum to expand or contract the spectrum. As a result, the alignment unit 112a outputs an aligned source speaker spectrum based on the source speaker path and source speaker spectrum obtained from the path search unit 111. Similarly, the alignment unit 112b outputs an aligned target speaker spectrum based on the target speaker path and target speaker spectrum obtained from the path search unit 111.
[0157] The voice conversion model 23p is the pre-trained state of the voice conversion model 23a of the first embodiment, but since it has the same model structure as the voice conversion model 23a, its description will be omitted. Note that the input data of the input spectrum encoder 131 is the aligned target speaker spectrum output by the alignment unit 112a.
[0158] The pre-learning unit 64 includes a reconstruction error calculation unit 644 and a model optimization unit 645 .
[0159] Of these, the reconstruction error calculation unit 644 outputs a spectral error by calculating an error using the predicted spectrum of the target speaker input from the spectrum prediction unit 135 and the aligned target spectrum input from the alignment unit 112b as ground truth data for the predicted spectrum. In this case, since the number of frames of the input data to the reconstruction error calculation unit 644 matches, a general loss function used in machine learning for generation tasks, such as squared error or absolute error, can be used.
[0160] The model optimization unit 645 updates each model parameter of the voice conversion model 13p (the input spectrum encoder 131, the output spectrum encoder 132, the decoder 133, the spectrum prediction unit 135, and the transition probability prediction unit 134) by error backpropagation so as to minimize the spectrum error input from the reconstruction error calculation unit 644.
[0161] The voice conversion device 10 refines the voice conversion model by repeatedly applying the processing flow configured as above to the training data.
[0162] In the explanation so far, the acoustic features used for input and output have been referred to as spectra, but it is also possible to use spectrograms obtained by short-time Fourier transforms, or mel-spectrograms, which are low-dimensional representations of spectrograms adapted to human hearing. It is also possible to use mel-cepstrum, which is often used in speech synthesis, or a vector concatenating the fundamental frequency, which is the pitch of the voice.
[0163] <Functional Configuration of Inference Phase> In the third embodiment, the functional configuration of the voice conversion device in the inference phase is the same as that in the first embodiment, and therefore a description thereof will be omitted.
[0164] [Processing or Operation of Third Embodiment] Next, the processing or operation in the pre-training phase of the voice conversion device 20 will be described with reference to Fig. 19. Note that the main training phase and inference phase of the third embodiment are the same as those of the first embodiment, and therefore will not be described here.
[0165] <Processing or Operation in Pre-Learning Phase> FIG. 19 is a flowchart showing processing in the pre-learning phase according to the third embodiment.
[0166] S101: The path search unit 111 finds the input source speaker path and target speaker path, and aligns the number of frames of two different lengths.
[0167] S102: The alignment unit 112a outputs an aligned source speaker spectrum based on the source speaker path and source speaker spectrum obtained from the path search unit 111. Similarly, the alignment unit 112b outputs an aligned target speaker spectrum based on the target speaker path and target speaker spectrum obtained from the path search unit 111.
[0168] S103: The input spectrum encoder 131 receives the current input spectrum as the aligned source speaker spectrum from the alignment unit 112a, encodes this input spectrum, and outputs it.
[0169] S104: The output spectrum encoder 132 receives, as feedback, the past output spectrum as the predicted spectrum of the target speaker output by the spectrum prediction unit 135 (originally the decoder 133), and encodes and outputs this output spectrum.
[0170] S105: The decoder 133 decodes the current input spectrum input from the input spectrum encoder 131, taking into consideration whether it is possible to output a spectrum of the speech content corresponding to the past input spectrum based on the output spectrum input from the output spectrum encoder 132.
[0171] S106: The transition probability prediction unit 134 predicts and outputs transition probabilities based on the current input spectrum of the source speaker input from the decoder 133. Furthermore, the spectrum prediction unit 135 generates and outputs a predicted spectrum of the target speaker based on the current input spectrum of the source speaker input from the decoder 133.
[0172] S107: The reconstruction error calculation unit 644 calculates the error between the predicted spectrum of the target speaker input from the spectrum prediction unit 135 and the aligned target spectrum input from the alignment unit 112b, and outputs the spectrum error.
[0173] S108: The model optimization unit 645 updates each parameter of the voice conversion model 13p by error back propagation so as to minimize the spectrum error input from the reconstruction error calculation unit 644.
[0174] The voice conversion device 30 refines the voice conversion model by repeatedly applying the above-described processing flow to training data. In this training, the initial values of the voice conversion model 13a are not started from random values, but are instead taken from a pre-trained voice conversion model 13p.
[0175] [Major Effects of the Third Embodiment] As described above, the third embodiment provides the following effects in addition to the effects of the first embodiment.
[0176] In the first embodiment, the output of the decoder 133 is a rank 3 tensor
[0177] For example, if T=1000, U=1000, and D=80, this tensor is very large, and the amount of calculation and memory usage during machine learning will also be extremely large. In contrast, in the third embodiment, the output of the decoder 133 is converted into a matrix
[0178] Therefore, the memory usage can be significantly reduced. For example, U' = 1000. In this way, the memory usage for the output of the decoder 133 can be significantly reduced to about (1 / T), so the batch size during learning can be increased without using techniques such as gradient accumulation, leading to a reduction in learning time. As a result, the third embodiment has the effect of significantly reducing the time required for learning compared to the first embodiment.
[0179] [Experimental Results] Next, experimental results of the third embodiment will be described with reference to FIGS.
[0180] Figure 20(a) shows the temporal relationship between the number of frames of the input source speaker spectrum (T) and the output target speaker spectrum (U). This temporal relationship was created using manual ground truth labels.
[0181] FIG. 20(b) shows the correct spectrum.
[0182] and a predicted spectrum before starting machine learning of the voice conversion model 13a using random initial values for the voice conversion model 13a of the first embodiment.
[0183] FIG. 10 is a diagram showing the absolute error of
[0184] FIG. 20(c) shows the correct spectrum obtained by using the trained voice conversion model 13p in addition to the voice conversion model 13a of the third embodiment, similar to FIG. 20(b).
[0185] Predicted Spectrum
[0186] FIG. 10 is a diagram showing the absolute error of
[0187] As shown in Figure 20(b), when machine learning of the voice conversion model 13a is started using random initial values, only an unclear line is output, resulting in a result far removed from the shape of the black line segment in Figure 20(a). On the other hand, when pre-training is performed, as shown in Figure 20(c), it can be confirmed that values become small near the black line segment in Figure 20(a). This indicates that the relationship between the spectra of the source speaker and the target speaker is ensured. Note that in Figure 20(c), there are also areas where values become small other than the components near the diagonal, but these are silent, and since the spectrum of silent speech does not depend on the speaker, the error is always small. This makes it possible to predict that the third embodiment will produce a certain effect even before main training is performed.
[0188] 21 is a graph showing the WER and MCD according to the number of pre-training steps according to the third embodiment. The WER (Word Error Rate) is an index that indicates the degree of insertion errors, deletion errors, and substitution errors when comparing the recognition result obtained by recognizing input speech with the correct text. The speech recognition model used to obtain the recognition results was the model disclosed in Reference 7.
[0189] (Reference 7) Alec Radford, Jong Wook Kim, Tao Xu, Greg Brockman, Christine Mcleavey, and Ilya Sutskever, "Robust speech recognition via large-scale weak supervision," Proc. ICML, vol. 202, pp. 28492-28518, 2023. Here, the WER is 18.4% after 300k steps. Although the WER improves with further training steps, even with as few as 300k training steps, it is possible to preserve a certain degree of linguistic information. The final result is 15.4% after 1000k steps.
[0190] As mentioned above, MCD (Mel-Cepstrum Distortion) is an index of the error between the target speaker's correct spectrum and the predicted spectrum, with a smaller value indicating better performance. Like WER, this also improves as the training steps progress. This embodiment has the advantage of requiring significantly less computational effort than the first embodiment, and does not require alignment. Despite this, a value of MCD = 5.3, comparable to the effect of the first embodiment (average between genders in Figure 9), is obtained.
[0191] Fig. 22 is a diagram showing subjective evaluation scores on a five-point scale regarding the naturalness of speech, according to the third and fourth embodiments. Fig. 22 is the result of an experiment similar to that shown in Fig. 10. Fig. 23 is a diagram showing subjective evaluation scores on a five-point scale regarding speaker similarity, according to the third and fourth embodiments. Fig. 22 is the result of an experiment similar to that shown in Fig. 11. In Figs. 22 and 23, the higher the score, the better the result. Note that, except for analytically synthesized speech (RESYN), all of Figs. 22 and 23 are results obtained by voice conversion using streaming operation.
[0192] 22 and 23, RESYN and ConvS2S-VC are experimental results in the environments described in FIGS. 10 and 11. ORIGINAL VC-T 100k is the result when 100k steps of pre-learning have been performed in the first embodiment. ORIGINAL VC-T 1000k is the result when 1000k steps of pre-learning have been performed in the first embodiment. VC-T PT is the result when the pre-learning is fixed at 1000k steps in the example shown in FIG. 21. VC-T FT W / ALI 100k is the result when 100k steps of pre-learning have been performed in the third embodiment. VC-T FT W / ALI 1000k is the result when 1000k steps of pre-learning have been performed in the third embodiment.
[0193] Note that VC-T FT W / O ALI 100k is the result when pre-learning has progressed 100k steps in the fourth embodiment described below, and VC-T FT W / O ALI 1000k is the result when pre-learning has progressed 1000k steps in the fourth embodiment.
[0194] As shown in Fig. 22, in the first embodiment, a consistent result (score 3.08) is not obtained until main learning has progressed to 1000k steps, but in the third embodiment, a consistent result (score 3.17) can be obtained by progressing main learning to 100k steps. Similarly, in Fig. 23, in the first embodiment, a consistent result (score 3.31) is not obtained until main learning has progressed to 1000k steps, but in the third embodiment, a consistent result (score 3.32) can be obtained by progressing main learning to 100k steps.
[0195] Fourth Embodiment Next, the fourth embodiment will be described with reference to the drawings. Unlike the third embodiment, the fourth embodiment does not mask predicted alignments using correct alignments during actual training. That is, while the voice conversion device 30 of the third embodiment includes the same correspondence generation unit 15 as the first embodiment, the voice conversion device 10 of the fourth embodiment does not include the correspondence generation unit 15. Also, while the voice conversion device 30 of the third embodiment includes the same training unit 14 as the first embodiment, the voice conversion device 30 of the fourth embodiment includes a new training unit 140. However, this training unit 140 is the same as the training unit 14 except that the alignment masking unit 142 has been removed; otherwise, the functional configuration is the same. The electrical hardware configuration of the voice conversion device 40 is the same as that of the voice conversion device 10 of the first embodiment (see FIG. 5 ), and therefore will not be described here.
[0196] [Functional Configuration of Voice Conversion Device] <Functional Configuration of This Training Phase> Fig. 24 is a functional configuration diagram of a training unit according to the fourth embodiment. As shown in Fig. 24, the training unit 140 has the same functional configuration as the training unit 14 shown in Fig. 3 except for the alignment mask unit 142, and therefore a description of these components will be omitted.
[0197] In the voice conversion device 30 of the third embodiment, predicted alignments are masked using a correct alignment during main training after pre-training. Due to the characteristics of the voice conversion task, the correct alignment is not completely diagonal, but each element is located within a few frames near the diagonal components. Thanks to the pre-training assuming diagonal alignment described in the third embodiment, the voice conversion model 13a can predict a reasonable spectrum, thereby reducing the spectral error relative to the weighted spectrum near the correct alignment. In this way, the pre-trained model operates to intensively optimize the area around the diagonal components of the alignment matrix during main training. Therefore, in this embodiment, focusing on this advantage, the voice conversion model 13a can be main trained without a correct alignment, thereby reducing the annotation cost required to prepare a correct alignment.
[0198] <Functional Configuration of Inference Phase> In the fourth embodiment, the functional configuration of the voice conversion device in the inference phase is the same as that in the first embodiment, and therefore a description thereof will be omitted.
[0199] [Processing or Operation of Fourth Embodiment] The processing or operation in the preliminary learning phase and the inference phase of the fourth embodiment is the same as that of the third embodiment, and therefore will not be described. In addition, the processing or operation in the main learning phase of the fourth embodiment is the same as that of the first embodiment, except that the processing by the correspondence generation unit 15 (see S116) and the processing by the alignment mask unit 142 (see S117) in FIG. 8 are omitted.
[0200] That is, in the fourth embodiment, the pre-trained voice conversion model 13a is used as an initial value in the main training, as in the third embodiment, and a predicted alignment is obtained as in the third embodiment. In the fourth embodiment, this predicted alignment and the predicted spectrum of the target speaker are input to the spectrum weighting unit 143. The spectrum weighting unit 143 outputs a weighted predicted spectrum of the target speaker. Other than these processes or operations, the fourth embodiment is the same as the third embodiment. Note that the pre-trained model used as the initial value may be that of the fifth embodiment described below.
[0201] [Major Effects of the Fourth Embodiment] As described above, in addition to the effects of the first and third embodiments, the fourth embodiment has the effect that the voice conversion model 13a can be trained without a correct alignment, thereby reducing the cost of annotation required to prepare a correct alignment.
[0202] [Experimental Results] Next, the experimental results of the fourth embodiment will be described again with reference to Figures 22 and 23. As shown in Figure 22, in the fourth embodiment, by simply progressing the main learning up to 100k steps, a better result (score 3.37) can be obtained than the result of progressing the main learning up to 1000k steps in the third embodiment (score 3.17). Similarly, as shown in Figure 23, in the fourth embodiment, by simply progressing the main learning up to 100k steps, a better result (score 3.74) can be obtained than the result of progressing the main learning up to 1000k steps in the third embodiment (score 3.54).
[0203] Fifth Embodiment Next, the fifth embodiment will be described with reference to the drawings. In the fifth embodiment, unlike the third embodiment, the path search of the spectrum alignment unit 11 is newly calculated based on phoneme posterior probabilities rather than on spectra during pre-training. That is, the voice conversion device 50 of the fifth embodiment has a new speaker-independent spectrum unit 8 in addition to the functional configuration in the pre-training phase of the voice conversion device 30 of the third embodiment. Furthermore, the electrical hardware configuration of the voice conversion device 40 is similar to that of the voice conversion device 10 of the first embodiment (see FIG. 5), and therefore description thereof will be omitted.
[0204] [Functional Configuration of Voice Conversion Device] <Functional Configuration in Pre-Training Phase> Fig. 25 is a diagram showing the functional configuration of a portion of a voice conversion device 50 in the pre-training phase according to the fifth embodiment. As shown in Fig. 25, the voice conversion device 50 newly includes a spectrum speaker-independent unit 8 compared to the voice conversion device 30 shown in Fig. 18. Therefore, only the functional configuration of the spectrum speaker-independent unit 8 will be described below. Note that the pre-training unit 64 shown in Fig. 18 is omitted in Fig. 25 due to limitations on the drawing area, but it actually exists.
[0205] The spectrum speaker-independent unit 8 includes speech recognition units 8a and 8b. The speech recognition units 8a and 8b convert speech or spectra into linguistic information vectors. When the speech recognition units 8a and 8b are configured using machine learning, their acoustic models output posterior probabilities of tokens. Tokens are discrete information such as subwords, such as phonemes or word pieces, or state IDs shared by a decision tree. The posterior probabilities are a series of categorical distributions for the series of discrete information from the speech recognition units 8a and 8b. Taking argmax for the posterior probability at each time point allows the most likely token at that time point to be determined, but token information from the 2nd best onward is lost. Therefore, in this embodiment, the path search unit 111 uses posterior probabilities to perform a path search that retains linguistic information from the 2nd best onward. Note that although phoneme posterior probabilities are shown, posterior probabilities may of course be calculated using other subwords, rather than phonemes, for tokens based on the output of the speech recognition units 8a and 8b.
[0206] With this functional configuration, the speech recognition unit 8a inputs the spectrum of the source speaker and outputs the phoneme posterior probabilities of the source speaker. Similarly, the speech recognition unit 8b inputs the spectrum of the target speaker and outputs the phoneme posterior probabilities of the target speaker. As a result, the path search unit 111 inputs the phoneme posterior probabilities of the source speaker and the phoneme posterior probabilities of the target speaker and outputs the paths of the source speaker and the target speaker.
[0207] <Functional Configuration of Inference Phase> In the fifth embodiment, the functional configuration of the voice conversion device in the inference phase is the same as that in the first embodiment, and therefore a description thereof will be omitted.
[0208] [Processing or Operation of Fifth Embodiment] The processing or operation in the main learning phase and inference phase of the fifth embodiment is the same as that of the third embodiment, and therefore will not be described. Furthermore, in the processing or operation in the preliminary learning phase of the fifth embodiment, the processing of the speech recognition units 8a and 8b is added before the processing of the path search unit 111 (see S101) in Fig. 19 , and the input of the path search unit 111 is changed, but otherwise it is the same as that of the third embodiment.
[0209] That is, the speech recognition unit 8a inputs the spectrum of the source speaker and outputs the phoneme posterior probabilities of the source speaker. Similarly, the speech recognition unit 8b inputs the spectrum of the target speaker and outputs the phoneme posterior probabilities of the target speaker. Next, the path search unit 111 inputs the phoneme posterior probabilities of the source speaker and the phoneme posterior probabilities of the target speaker and outputs the path of the source speaker and the path of the target speaker. The processing from this point onwards is the same as in the third embodiment, so a description thereof will be omitted.
[0210] [Major Effects of Fifth Embodiment] As described above, according to the fifth embodiment, in addition to the effects of the first and third embodiments, the following effects are achieved.
[0211] In the voice conversion device 30 of the third embodiment, the spectrum alignment unit 11 performs path search using the spectra of the source speaker and the target speaker. However, spectra are features that are highly dependent on the speaker, and path search may fail, resulting in inappropriate alignment, particularly in cases of opposite-sex speakers. Training the voice conversion model 13a using incorrectly aligned spectra may result in a deterioration in the conversion performance of the voice conversion model 13p. In contrast, in the fifth embodiment, the output of the speech recognition units 8a and 8b is a linguistic information vector indicating the content of the utterance. Therefore, if the path search unit 111 searches for paths using the linguistic information of the source speaker and the target speaker, comparison can be made without taking into account differences in speaker characteristics, age, or gender between the source speaker and the target speaker. Therefore, performing path search using speaker-independent features has the effect of improving the accuracy of the spectrum alignment unit 11.
[0212] Supplementary Note: The present invention is not limited to the above-described embodiment, and may have the following configurations or processes (operations), for example.
[0213] (1) The voice conversion devices 10, 20, 30, 40, and 50 can be realized by a computer and a program, but this program can also be recorded on a (non-transitory) recording medium or provided via a communication network 100 such as the Internet.
[0214] (2) The CPU 101 as a processor may be a single processor or multiple processors.
[0215] Supplementary Notes The above-described embodiment can also be expressed as the following invention.
[0216] [Supplementary Item 1] A voice conversion device having a processor that performs machine learning to generate a voice conversion model using training data consisting of a set of a source speaker's spectrum and a target speaker's spectrum whose speech content is identical to that of the source speaker's spectrum, wherein the processor performs spectrum alignment processing to convert at least one of the source speaker's spectrum and the target speaker's spectrum based on the source speaker's spectrum and the target speaker's spectrum so that the number of frames of the source speaker's spectrum and the target speaker's spectrum match, and to output aligned source speaker spectra and aligned target speaker spectra; the voice conversion device has an input spectrum encoder that inputs the aligned source speaker's spectrum and outputs encoded first features, a decoder that decodes the first features and outputs a predicted spectrum, and an output spectrum encoder that inputs the predicted spectrum and outputs encoded second features, and the decoder has the voice conversion model that recursively decodes the current first features based on the second features previously encoded by the output spectrum encoder, The voice conversion device wherein the processor executes a pre-training process for pre-training each parameter of the voice conversion model using the aligned target speaker spectrum as ground truth data for the predicted spectrum.
[0217] [Supplementary Item 2] A voice conversion device according to Supplementary Item 2, wherein the processor executes a speaker-independent spectral process that inputs the spectrum of the source speaker and outputs phoneme posterior probabilities of the source speaker, and also inputs the spectrum of the target speaker and outputs phoneme posterior probabilities of the target speaker, and the spectrum alignment process includes a process that inputs the phoneme posterior probabilities of the source speaker and the phoneme posterior probabilities of the target speaker output by the spectrum speaker-independent unit, and outputs the aligned source speaker spectrum and the aligned target speaker spectrum.
[0218] [Supplementary Item 3] The voice conversion device according to Supplementary Item 1 or 2, wherein the processor executes a training process to train the pre-trained voice conversion model so as to minimize an error between the predicted spectrum of the target speaker and the target speaker's correct spectrum.
[0219] [Supplementary Item 4] A voice conversion device according to Supplementary Item 3, wherein the processor executes a correspondence generation processing unit that inputs the phoneme labels with time information of the source speaker and the phoneme labels with time information of the target speaker, and generates and outputs a ground truth alignment that indicates the correspondence between the phoneme labels with time information of the source speaker and the phoneme labels with time information of the target speaker, and the training process includes a process of using the ground truth alignment to train the pre-trained voice conversion model so as to minimize an error between the predicted spectrum of the target speaker and the ground truth spectrum of the target speaker.
[0220] DESCRIPTION OF SYMBOLS 8 Spectral speaker independence unit 8a, 8b Speech recognition unit 10 Voice conversion device 11 Spectral alignment unit 12 Spectral calculation unit 13p Voice conversion model 13a (Pre-trained) voice conversion model 13b (Trained) voice conversion model 14 Training unit 15 Correspondence generation unit 16 Speech waveform generation unit 17 Termination determination unit 23a Voice conversion model 23b (Trained) voice conversion model 24 Training unit 25 Correspondence generation unit 27 Termination determination unit 111 Path search unit 112a, 112b Alignment unit 131 Input spectrum encoder 132 Output spectrum encoder 133 Decoder 134 Transition probability prediction unit 135 Spectral prediction unit 141 Alignment calculation unit 142 Alignment mask unit 143 Spectral weighting unit 144 Spectral error calculation unit 145 Model optimization unit 231 Input speech waveform encoder 232 Output speech waveform encoder 233 Decoder 234 Transition probability prediction unit 235 Speech waveform prediction unit 241 Alignment calculation unit 242 Alignment mask unit 243 Speech waveform weighting unit 244 Speech waveform reconstruction error calculation unit 245 Model optimization unit
Claims
1. A voice conversion device that machine-learns a voice conversion model using training data consisting of a set of a source speaker's spectrum and a target speaker's spectrum having the same speech content as the source speaker's spectrum, the voice conversion device comprising: a spectrum alignment unit that converts at least one of the source speaker's spectrum and the target speaker's spectrum based on the source speaker's spectrum and the target speaker's spectrum so that the number of frames of the source speaker's spectrum and the target speaker's spectrum match, and outputs an aligned source speaker's spectrum and an aligned target speaker's spectrum; an input spectrum encoder that inputs the aligned source speaker's spectrum and outputs an encoded first feature, a decoder that decodes the first feature and outputs a predicted spectrum, and an output spectrum encoder that inputs the predicted spectrum and outputs an encoded second feature, the decoder recursively decodes the current first feature based on the second feature previously encoded in the output spectrum encoder; and a pre-training unit that pre-trains each parameter of the voice conversion model using the aligned target speaker's spectrum as correct answer data for the predicted spectrum.
2. A voice conversion device as described in claim 1, comprising a spectrum speaker independency unit which inputs the spectrum of the source speaker and outputs the phoneme posterior probabilities of the source speaker, and which inputs the spectrum of the target speaker and outputs the phoneme posterior probabilities of the target speaker, wherein the spectrum alignment unit inputs the phoneme posterior probabilities of the source speaker and the phoneme posterior probabilities of the target speaker output by the spectrum speaker independency unit, and outputs the aligned source speaker spectrum and the aligned target speaker spectrum.
3. A voice conversion device as claimed in claim 1, comprising: a learning unit that performs actual training of the pre-trained voice conversion model so as to minimize an error between the predicted spectrum of the target speaker and the correct spectrum of the target speaker.
4. A voice conversion device as described in claim 3, comprising a correspondence generation unit that inputs the phoneme labels with time information of the source speaker and the phoneme labels with time information of the target speaker, and generates and outputs a correct alignment indicating the correspondence between the phoneme labels with time information of the source speaker and the phoneme labels with time information of the target speaker, and the learning unit uses the correct alignment to actually train the pre-trained voice conversion model so as to minimize the error between the predicted spectrum of the target speaker and the correct spectrum of the target speaker.
5. A voice conversion device that converts voice using a machine-learned voice conversion model, wherein the machine-learned voice conversion model has: an input spectrum encoder that inputs a spectrum to be converted and outputs an encoded first feature; a decoder that decodes the first feature and outputs a predicted spectrum; and an output spectrum encoder that inputs the predicted spectrum and outputs an encoded second feature, wherein the decoder recursively decodes the current first feature based on the second feature previously encoded in the output spectrum encoder.
6. A machine learning method executed by a voice conversion device that machine-learns a voice conversion model using training data consisting of a set of a source speaker's spectrum and a target speaker's spectrum having the same speech content as the source speaker's spectrum, wherein the voice conversion device performs a spectrum alignment process that converts at least one of the source speaker's spectrum and the target speaker's spectrum based on the source speaker's spectrum and the target speaker's spectrum so that the number of frames of the source speaker's spectrum and the target speaker's spectrum match, and outputs an aligned source speaker spectrum and an aligned target speaker spectrum; the voice conversion model has an input spectrum encoder that inputs the aligned source speaker's spectrum and outputs an encoded first feature, a decoder that decodes the first feature and outputs a predicted spectrum, and an output spectrum encoder that inputs the predicted spectrum and outputs an encoded second feature, and the decoder recursively decodes the current first feature based on the second feature previously encoded in the output spectrum encoder, the voice conversion device performs a pre-learning process of pre-learning each parameter of the voice conversion model using the aligned target speaker's spectrum as correct answer data for the predicted spectrum; and 7. A machine learning method executed by a voice conversion device that converts voice using a machine-learned voice conversion model, wherein the machine-learned voice conversion model has: an input spectrum encoder that inputs a spectrum to be converted and outputs an encoded first feature; a decoder that decodes the first feature and outputs a predicted spectrum; and an output spectrum encoder that inputs the predicted spectrum and outputs an encoded second feature, wherein the decoder recursively decodes the current first feature based on the second feature previously encoded in the output spectrum encoder.
8. A program for causing a computer to function as the voice quality conversion device according to any one of claims 1 to 5.
Citation Information
Patent Citations
Speech translation method and system using a multilingual text-to-speech synthesis model
JP2022169714A
System and method for streaming end-to-end speech recognition with asynchronous decoder
JP2023504219A