Voice quality conversion device and voice quality conversion method

JP2024018852A5Pending Publication Date: 2025-07-10DOWANGO KK
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2022181983
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2022-11-14
Publication Date
2025-07-10

AI Technical Summary

Technical Problem

Conventional voice quality conversion systems struggle to accurately convert characteristic voices such as whispers, falsettos, and angry voices into target voices, requiring extensive training data for each individual and failing to produce intermediate voice qualities.

Method used

A neural network-based voice quality conversion device that utilizes an encoder to extract speaker-independent latent expressions, removes speaker-specific features, and adds destination speaker characteristics to generate voices reflecting the input voice's manner of speaking, enabling many-to-many voice conversions without needing complete training data for all speakers.

Benefits of technology

The system effectively outputs voices that reflect the input's characteristic, including whispering, falsetto, or angry tones, and can produce intermediate voice qualities, enhancing flexibility and efficiency in voice conversion.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

To output, when a characteristic voice is input during voice quality conversion, the voice that reflects the characteristics.SOLUTION: A voice quality conversion device uses a trained neural network 100 to convert a voice of a conversion source into a voice in accordance with speaker information of a conversion destination. The neural network includes: an encoder 110 for receiving the voice and outputting a latent expression S1; a flow 120 for converting the latent expression S1 into a speaker-independent latent expression from which a speaker property of the conversion source has been removed while leaving characteristics of a method of utterance, and reverse-converting the speaker-independent latent expression into a latent expression S2 by adding the speaker property of the conversion destination; and a vocoder 130 for inputting the latent expression S2 and outputting the voice of the conversion destination. The neural network is trained so that the vocoder can restore the latent expression output by the encoder into an original training voice, and the speaker-independent latent expression obtained through conversion by the flow and an expression created from speaker-independent information output by a text encoder 140 become closer.SELECTED DRAWING: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present invention relates to a voice conversion device, a voice conversion method, a program, and a recording medium. [Background technology]

[0002] Recent advances in deep learning technology have greatly improved the quality of speech synthesis. Non-Patent Document 1 is a technology that can generate speech from text and convert voice quality. Non-Patent Document 2 is a technology based on Non-Patent Document 1 that converts the speech of speakers other than the speaker of the speech used for learning, and can convert the voice quality of any speaker. [Prior art documents] [Non-patent literature]

[0003] [Non-Patent Document 1] Jaehyeon Kim, Jungil Kong, and Juhee Son, "Conditional Variational Autoencoder with Adversarial Learning for End-to-End Text-to-Speech," Proceedings of the 38th International Conference on Machine Learning, 2021, Vol. 139 of PMLR, pp. 5530-5540 [Non-Patent Document 2] “[OV2L Evolving Summit] Session 4 “How to convert VITS to any-to-many VC” presented by kaffelun”, Internet〈 URL:https: / / youtu.be / uRwFHuXw3Qk 〉 Summary of the Invention [Problem to be solved by the invention]

[0004] In conventional voice conversion, even if the source voice is input with a distinctive voice including a whisper, falsetto, angry voice, etc., it is converted to the calm voice (normal voice) of the conversion destination voice used for learning. If whisper, falsetto, angry voice, etc. are learned as the voice of an individual speaker, it is possible to convert to a distinctive voice by specifying whisper, falsetto, or angry voice as the conversion destination voice. However, when converting the voices of multiple people, it is necessary to prepare whisper, falsetto, and angry voices for each person as the training voice. In addition, there was a problem that it was not possible to convert to a voice that was intermediate between a calm voice and a whisper.

[0005] The present invention has been made in view of the above, and has an object to, when a characteristic voice is input during voice conversion, output a voice reflecting the characteristic. [Means for solving the problem]

[0006] A voice conversion device of one embodiment of the present invention includes an input unit for inputting source voice data and meta-information to be manipulated during voice conversion, and a conversion unit for converting the source voice data into voice data corresponding to the meta-information using a trained neural network, the neural network including an encoder for inputting voice data, extracting features from the voice data, and outputting a first latent representation, a flow for converting the first latent representation into a second latent representation by removing features corresponding to the meta-information while retaining predetermined features contained in the voice data, and adding features corresponding to the meta-information to the second latent representation and performing inverse conversion to a third latent representation, and a decoder for inputting the third latent representation and outputting the destination voice data.

[0007] A voice conversion device of one embodiment of the present invention includes an input unit for inputting source voice data and meta-information to be manipulated during voice conversion, and a conversion unit for converting the source voice data into voice data corresponding to the meta-information using a trained neural network, the neural network including a second encoder for inputting voice data and outputting a second latent representation in which predetermined features contained in the voice data are retained while features corresponding to the meta-information are removed, a flow for adding features corresponding to the destination meta-information to the second latent representation and reverse-converting it into a third latent representation, and a decoder for inputting the third latent representation and outputting the destination voice data. Effect of the Invention

[0008] According to the present invention, when a characteristic voice is input during voice conversion, a voice reflecting the characteristic can be output. [Brief description of the drawings]

[0009] [Figure 1] FIG. 1 is a diagram showing an example of the configuration of a voice quality conversion device according to this embodiment. [Diagram 2] FIG. 2 is a diagram illustrating an example of a configuration of a neural network according to the first embodiment. [Diagram 3] FIG. 3 is a flowchart showing an example of a process flow during voice conversion in the first embodiment. [Figure 4] FIG. 4 is a diagram illustrating an example of a configuration of a neural network according to the second embodiment. [Diagram 5] FIG. 5 is a flowchart showing an example of a process flow during voice quality conversion in the second embodiment. [Figure 6] FIG. 6 is a diagram illustrating an example of a learning method for a neural network according to the second embodiment. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0010] [First embodiment] An example of the configuration of a voice conversion device 1 of the first embodiment will be described with reference to Fig. 1. The voice conversion device 1 shown in the figure includes an input unit 11, a conversion unit 12, and a learning unit 13. Each unit included in the voice conversion device 1 may be configured by a computer including an arithmetic processing unit, a storage device, etc., and the processing of each unit may be executed by a program. This program is stored in a storage device included in the voice conversion device 1, and can be recorded on a recording medium such as a magnetic disk, an optical disk, or a semiconductor memory, or can be provided via a network.

[0011] The input unit 11 inputs voice data (hereinafter referred to as voice) and speaker information. Specifically, during learning, the input unit 11 inputs training voice of a speaker to be mutually convertible and speaker information of the voice. The voice conversion device 1 of the first embodiment learns the voice of a speaker to be mutually convertible and enables many-to-many voice conversion of the trained speakers. Speaker information is a speaker identifier. Speaker information is assigned to each speaker of the training voice. During learning, the training voice and speaker information of the training voice input by the input unit 11 are transmitted to the learning unit 13. On the other hand, during inference (during voice conversion), the input unit 11 inputs the source voice, the source speaker information, and the destination speaker information. During inference, the source voice and speaker information input by the input unit 11 are transmitted to the conversion unit 12.

[0012] The conversion unit 12 inputs the source voice, the source speaker information, and the destination speaker information into the trained neural network, and converts the source voice into a voice corresponding to the destination speaker information by reflecting the manner of pronunciation of the source voice. The manner of pronunciation can be a whisper, a falsetto, an angry voice, etc. For example, if the source voice is a whisper, the destination voice is also generated in a whisper. Both the source speaker and the destination speaker are speakers of the voice used for training.

[0013] The neural network of this embodiment includes an encoder that extracts features from the source voice and outputs a latent representation of the source voice, a flow that converts the latent representation of the source voice into a speaker-independent latent representation that removes speaker characteristics (speaker characteristics) while retaining the manner of pronunciation, and adds the speaker characteristics of the destination speaker to the speaker-independent latent representation to convert it back into a latent representation of the destination voice, and a decoder (vocoder) that inputs the latent representation of the destination voice and outputs the destination voice. Details of the model will be described later.

[0014] The learning unit 13 inputs the training speech, speaker information of the training speech, text of the training speech, and pronunciation information of the training speech (hereinafter referred to as conditions), and trains the neural network by constraining the distribution followed by the intermediate representation of the variational autoencoder consisting of an encoder and a vocoder to a distribution created from the text and conditions. In other words, the learning unit 13 trains the neural network so that the latent representation output by the encoder can be restored to the original speech by the vocoder, and the latent representation in which the speaker characteristics are removed while retaining the pronunciation characteristics is close to the representation created from the text and pronunciation information that does not include the speaker characteristics. The text is the phonological information of the training speech. The conditions are flags of 0 and 1 that indicate the pronunciation of the training speech, such as whispering, falsetto, and angry voice. When the training speech is a whisper, information indicating a whisper is input to the learning unit 13 as a condition. The parameters (neural network) learned by the learning unit 13 are stored in a storage device provided in the voice quality conversion device 1.

[0015] (Models and Learning) An example of the neural network and an example of learning according to the first embodiment will be described with reference to Fig. 2. A neural network 100 shown in the figure includes an encoder 110, a flow 120, a vocoder 130, and a text encoder 140.

[0016] The structure consisting of the encoder 110 and the vocoder 130 corresponds to a variational autoencoder. When speech is input to the encoder 110, a latent representation is obtained, and when the latent representation is input to the vocoder 130, speech is output. The latent representation has information about the sound.

[0017] When flow 120 receives latent expressions and speaker information, it outputs speaker-independent latent expressions that have removed as much speaker-specificity as possible from the latent expressions. Flow 120 is also a reversible neural network, and when a speaker-independent latent expression is input in the reverse direction and destination speaker information is added, a latent expression of the destination speaker is obtained. By inputting the latent expressions output by flow 120 to vocoder 130, the voice of the destination speaker can be output.

[0018] The text encoder 140 is a neural network used during learning, and is not necessary during inference. The text encoder 140 inputs the text and conditions of the learning speech, and outputs a latent representation in which the conditions are added to the text. The latent representation output by the text encoder 140 is a representation created from the text and conditions that are not dependent on the speaker, and does not include speaker characteristics.

[0019] During training, training speech and speaker information of the training speech are input to the encoder 110, and text of the training speech and conditions are input to the text encoder 140. The neural network is trained so that the speech input to the encoder 110 and the speech output from the vocoder 130 are the same, and at the same time, the neural network 100 is trained so that the latent representation output from the flow 120 and the representation created from speaker-independent information output from the text encoder 140 are brought closer together. The latent representation output from the encoder 110 is converted and inversely converted in the flow 120 and then input to the vocoder 130. The speaker characteristics are removed during the conversion in the flow 120, and the speaker characteristics are assigned during the inverse conversion. The speaker characteristics assigned during training and inverse conversion are the speaker characteristics of the training speech. Training is performed so that the spectrogram of the speech input to the encoder 110 and the spectrogram of the speech output from the vocoder 130 match. For learning to bring the latent representation output by the flow 120 and the representation output by the text encoder 140 closer together, monotonic alignment search can be used, as in Non-Patent Document 1. The horizontal axis of the latent representation output by the flow 120 is time, and the horizontal axis of the representation created from speaker-independent information output by the text encoder 140 is phonemes. The correspondence between them is determined by monotonic alignment, and constraints are imposed so that the correspondence becomes closer. In this embodiment, in addition to the phonemes, the condition is input to the text encoder 140 as second information. This allows the neural network 100 to learn so that the flow 120 outputs a latent representation that has been speaker-free and includes the characteristics of the way of speaking.

[0020] The training voices are prepared from the voices of people whose voices you want to convert in a many-to-many manner. For example, if you train the voices of three people, A, B, and C, after training, you can convert A's voice to B or C's voice, B's voice to A or C's voice, and C's voice to A or B's voice.

[0021] During learning, it is not necessary to prepare all the training voices of the corresponding conditions. Specifically, if the voice conversion device 1 supports whispering, even if there is no training voice of the whispering voice of person C, there may be training voices of the whispering voice of person A or person B. In other words, it is not necessary to prepare training voices of all variations of conditions supported by the voice conversion device 1 for all speakers to be trained.

[0022] When the training speech includes a manner of pronunciation, the pronunciation information is also input together with the text to text encoder 140. For example, when training the whispering speech of person A as training speech, the training speech of person A and speaker information indicating person A are input to encoder 110, and the text of the training speech and a flag indicating the whispering speech are input to text encoder 140.

[0023] The speaker information input to the neural network 100 can be considered as meta-information to be manipulated during voice conversion. As described above, when it is desired to control the speaker characteristics, the speaker information is input as meta-information. If pitch or intonation is used as the speaker information, the pitch or intonation can be controlled to convert the voice. By specifying the pitch or intonation, it is possible to output a voice in which the high or low voice and intonation of the converted speaker are controlled. On the other hand, the text and condition input to the text encoder 140 are information that remains unchanged during conversion. In other words, they are features contained in the voice that are desired to remain after conversion.

[0024] When intonation is input as a condition to be input to the text encoder 140 together with the text, that is, when intonation is treated as invariant information during voice conversion, each phoneme obtained from the text has a time length, but the condition does not, so it is advisable to devise a way to match the time length of the phoneme information and the time length of the condition in monotonic alignment. For example, intonation information is extracted from training speech, and the time length of the intonation is matched to the time length of the speech information.

[0025] In order to take into account differences in the microphone of the training voice and the environment such as space, training voice with added noise may be input to the encoder 110, and training may be performed so that clean voice is output from the vocoder .

[0026] (Voice conversion processing) The flow of processing during voice conversion will be described with reference to FIG.

[0027] In step S11, the input unit 11 inputs the source voice, source speaker information, and destination speaker information, and transmits them to the conversion unit 12. The voice conversion device 1 processes the voice in units of a predetermined number of samples (slice). When the source voice is input in real time, it is processed in slice units in real time, and voice quality can be converted in real time. Both the source speaker and the destination speaker are one of the speakers of the training voice.

[0028] In step S12, the conversion unit 12 inputs the source speech and the source speaker information to the encoder 110, and obtains a latent representation S1 from the encoder 110. The latent representation S1 is a latent representation including the speaker characteristics of the source speech.

[0029] In step S13, the conversion unit 12 inputs the latent expression S1 and the source speaker information to the flow 120 to obtain a speaker-independent latent expression. The speaker-independent latent expression includes features of the pronunciation of the source voice.

[0030] In step S14, the conversion unit 12 adds speaker information of the converted speech and performs inverse conversion on the speaker-independent latent expression in flow 120 to obtain a latent expression S2 of the converted speech.

[0031] In step S15, the conversion unit 12 inputs the latent expression S2 and the speaker information of the conversion destination to the vocoder 130, and outputs the conversion destination voice that reflects the manner of pronunciation of the conversion source voice.

[0032] As described above, the voice conversion device 1 of the present embodiment includes an input unit 11 for inputting the source voice, the source speaker information, and the destination speaker information, and a conversion unit 12 for converting the source voice into a voice corresponding to the destination speaker information by using a trained neural network 100. The neural network 100 includes an encoder 110 for inputting a voice, extracting features from the voice, and outputting a latent representation S1, a flow 120 for converting the latent representation S1 into a speaker-independent latent representation from which the source speaker characteristics are removed while leaving the vocalization characteristics included in the voice, and for inversely converting the speaker-independent latent representation into a latent representation S2 by adding the destination speaker characteristics, and a vocoder 130 for inputting the latent representation S2 and outputting the destination voice. As a result, the voice conversion device 1 can convert the input voice into a voice quality of the destination speaker that reflects the vocalization manner of the input voice, such as a whisper, falsetto, or angry voice. The voice quality conversion device 1 does not specify the manner in which the converted voice is to be pronounced, but rather the encoder 110 and the flow 120 output a latent expression that includes the manner in which the original voice is pronounced. For example, if the original voice is somewhere between a calm voice and a whisper, a voice that reflects the intermediate manner of pronunciation is output.

[0033] The voice quality conversion device 1 of this embodiment inputs training speech to the encoder 110, inputs the text of the training speech and the condition indicating the manner of pronunciation included in the training speech data to the text encoder 140, and includes a learning unit 13 that trains the neural network 100 so that the vocoder 130 can restore the latent representation output by the encoder 110 to the original training speech, and the speaker-independent latent representation obtained by the conversion by the flow 120 and the representation output by the text encoder 140, created from speaker-independent information, are close to each other. As a result, the conversion by the flow 120 removes the speaker characteristics, and a latent representation including the manner of pronunciation is obtained. By assigning the speaker characteristics of the conversion destination speaker to this latent representation and performing inverse conversion, a latent representation including the speaker characteristics and manner of pronunciation of the conversion destination speaker is obtained.

[0034] [Second embodiment] The voice conversion device of the second embodiment additionally learns the neural network 100 of the first embodiment and converts the voice of any speaker. The first embodiment is a voice conversion device that performs many-to-many voice conversion. In the second embodiment, after generating the neural network of the first embodiment, learning is performed with the task of obtaining speaker-independent latent expressions without correct speaker information. The configuration of the voice conversion device of the second embodiment is the same as that of the first embodiment, so a description thereof will be omitted here.

[0035] (Models and Learning) An example of a neural network and a learning method according to the second embodiment will be described with reference to Fig. 4. The neural network 100 shown in the figure includes an encoder 110, a flow 120, a vocoder 130, and an any encoder 150. The encoder 110, the flow 120, and the vocoder 130 used are those already learned in the first embodiment. The text encoder 140 is not necessary during learning in the second embodiment.

[0036] The any encoder 150 is a neural network that receives speech without speaker information and outputs a speaker-independent latent expression. In the second embodiment, the neural network is trained so that the output of the any encoder 150, to which training speech is input without speaker information of the training speech to be converted, approaches a speaker-independent latent expression.

[0037] During learning, the learning voice and speaker information of the learning voice are input to the encoder 110, and the learning voice is input to the any encoder 150. The learning voice used in the first embodiment is also used in the second embodiment. The learning voice is input to the encoder 110 and the any encoder 150, and the neural network is trained so that the latent representation obtained by converting the output of the encoder 110 in the flow 120 and the output of the any encoder 150 are close to each other. The latent representation converted in the flow 120 is a latent representation in which the speaker characteristics are removed from the learning voice and which includes the characteristics of the way of speaking. The any encoder 150 is trained to output a latent representation in which the speaker characteristics are removed from the input voice and which includes the characteristics of the way of speaking. It is considered that generality is obtained by training with the learning voices of many speakers, about tens to 100 people, and even if the voice of any speaker other than the speaker of the learning voice is input to the any encoder 150, a latent representation in which the speaker characteristics are removed and which includes the characteristics of the way of speaking is obtained.

[0038] The latent representation output by the any encoder 150 is inversely converted by the flow 120 and the speaker information of the conversion destination is added, so that the voice input to the any encoder 150 can be converted into the voice of the conversion destination speaker.

[0039] (Voice conversion processing) The flow of processing during voice conversion in the second embodiment will be described with reference to FIG.

[0040] In step S21, the input unit 11 inputs the source voice and the speaker information of the conversion destination, and transmits them to the conversion unit 12. The speaker of the source voice does not have to be the speaker of the learning voice. In other words, the voice of any speaker may be input.

[0041] In step S22, the conversion unit 12 inputs the source speech to the any encoder 150, and obtains a speaker-independent latent representation from the any encoder 150. The speaker-independent latent representation includes features of how the source speech is uttered.

[0042] In step S23, the conversion unit 12 adds speaker information of the converted speech and performs inverse conversion on the speaker-independent latent expression in flow 120 to obtain a latent expression S2 of the converted speech.

[0043] In step S24, the conversion unit 12 inputs the latent expression S2 and the speaker information of the conversion destination to the vocoder 130, and outputs the conversion destination voice that reflects the manner of pronunciation of the conversion source voice.

[0044] (Another learning example) Another example of the learning method of the neural network of the second embodiment will be described with reference to Fig. 6. The configuration of the neural network in Fig. 6 is the same as the configuration of the neural network in Fig. 4.

[0045] In the learning example of Fig. 6, the neural network is trained so that the latent representation obtained by converting and inversely converting the output of encoder 110 in flow 120 and the latent representation obtained by inversely converting the output of encoder 150 for any in flow 120 become similar. Training voice is input to encoder 150 for any. When inversely converting in flow 120, speaker information of the destination is added. In this way, learning may be performed so that the latent representation of the destination speaker's voice obtained by inverse conversion in flow 120 becomes similar.

[0046] Furthermore, the latent representation obtained by the inverse transformation in flow 120 may be input to a vocoder 130 to train a neural network so that the waveforms or spectrograms are close to each other.

[0047] 6, learning speech and the converted speaker information S2 may be input to the any encoder 150, and the any encoder 150 may output the latent representation S2 without passing through the flow 120. In this case, the degree of freedom of the network configuration, such as the presence or absence of the flow 120, can be increased.

[0048] The learning method shown in FIG. 4 and the learning method shown in FIG. 6 may be combined.

[0049] As described above, the voice conversion device 1 of this embodiment includes an input unit 11 for inputting the source voice and the speaker information of the destination, and a conversion unit 12 for converting the source voice into a voice corresponding to the speaker information of the destination by using a trained neural network 100. The neural network 100 includes an encoder for any 150 for inputting a voice and outputting a speaker-independent latent expression in which the speaker characteristics of the source voice are removed while leaving the characteristics of the manner of pronunciation contained in the voice, a flow 120 for inversely converting the speaker-independent latent expression into a latent expression S2 by adding the speaker characteristics of the destination voice, and a vocoder 130 for inputting the latent expression S2 and outputting the destination voice. As a result, the voice conversion device 1 can convert any voice into the voice quality of the destination speaker that reflects the manner of pronunciation of the input voice.

[0050] The voice quality conversion device 1 of this embodiment is equipped with a learning unit that, after training the neural network 100 of the first embodiment, inputs training voice data to the encoder 110 and the any encoder 150, and trains the neural network 100 so that a speaker-independent latent representation (teacher) obtained by conversion using the flow 120 and a latent representation output by the any encoder 150 become close to each other. As a result, when the any encoder 150 receives input of a voice of an arbitrary speaker, the speaker characteristics are removed, and the encoder 150 becomes able to output a latent representation including the manner of pronunciation. By assigning the speaker characteristics of the conversion destination speaker to this latent representation and performing inverse conversion, a latent representation including the speaker characteristics and manner of pronunciation of the conversion destination speaker is obtained.

[0051] The voice quality conversion device 1 may train the neural network 100 so that the latent representation S2 (teacher) obtained by inverse conversion after conversion in flow 120 becomes close to the latent representation S2 obtained by inverse conversion of the speaker-independent latent representation output by the any encoder 150 in flow 120. [Explanation of symbols]

[0052] 1. Voice conversion device 11 Input section 12 Conversion section 13 Learning Department 100 Neural Networks 110 Encoder 120 Flow 130 Vocoder 140 Text Encoder 150 any encoder

Claims

1. An input unit to which information on a target speaker for voice quality conversion and voice data of a source are input; A conversion unit that uses a learned neural network to convert the voice data of the source into a voice quality according to the information on the target speaker, wherein the conversion unit reflects the way of utterance of the voice data of the source and converts it into a voice quality according to the information on the target speaker, A voice quality conversion device.

2. The voice quality conversion device according to Claim 1, wherein the neural network includes an encoder that inputs voice data and outputs a first latent representation in which the characteristics of the speaker of the voice data are removed while retaining the characteristics of the way of generation included in the voice data, a flow that adds characteristics corresponding to the information on the target speaker to the first latent representation and performs inverse conversion to a second latent representation, and a decoder that inputs the second latent representation and outputs voice data of a destination; A voice quality conversion device.

3. A computer inputs information on a target speaker for voice quality conversion and voice data of a source, and uses a learned neural network to convert the voice data of the source into a voice quality according to the information on the target speaker, reflecting the way of utterance of the voice data of the source, A voice quality conversion method.