Cross-lingual any-to-one voice conversion

The voice conversion system addresses the challenge of cross-lingual voice conversion by using a bottleneck structure and fine-tuning to disentangle content and speaker information, enabling high-quality speech output in the target speaker's voice without extensive training data.

WO2025184148A1PCT designated stage Publication Date: 2025-09-04CERENCE OPERATING CO
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
PCT/US2025/017301
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-02-29
Filing Date
2025-02-26
Publication Date
2025-09-04

AI Technical Summary

Technical Problem

Existing cross-lingual any-to-one voice conversion systems struggle to preserve source accents and emotional cues without requiring multilingual data from the target speaker, especially when few or no recordings are available for the target speaker in the target language.

Method used

A voice conversion system utilizing a bottleneck structure and fine-tuning approach that disentangles content and speaker information, allowing conversion of utterances from any speaker into the voice of a target speaker, even if the target speaker does not speak the input language, by using a content encoder and acoustic model trained with limited target speaker data.

Benefits of technology

The system effectively converts speech to retain emotional content and source accents while reducing the need for extensive high-quality training data, achieving high-quality speech output in the target speaker's voice.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025017301_04092025_PF_FP_ABST
    Figure US2025017301_04092025_PF_FP_ABST
Patent Text Reader

Abstract

A voice converter includes a first stage that receives an input signal and a second stage that provides an output signal. The input signal comprises content and first speaker-information. The output signal comprises the same content but with second speaker-information having supplanted the first speaker-information. The stages are trained independently of each other with the first having been trained using a training dataset that comprises utterances from other than the target speaker and the second having been trained using a tuning dataset that comprises utterances that consist of the second speaker-information.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] CROSS-LINGUAL ANY-TO-ONE VOICE CONVERSION

[0002] Cross-Reference to Related Applications

[0003]

[0001] This application claims the benefit of U.S. Provisional Application 63 / 559,411, filed on February 29, 2024, which is incorporated herein by reference.

[0004] Background

[0005]

[0002] This disclosure relates to generation of artificial spoken utterances, and more particularly relates to cross-lingual voice conversion.

[0006]

[0003] The generation of spoken utterances in artificial voices remains a challenging task, and this is particularly true of cross-lingual any-to-one voice conversion systems that are able to preserve a source accent without the need for multilingual data from a target speaker in whose voice the output is generated. In such conversion, the goal is to create an artificial spoken utterance, referred to as the target utterance, in a target language such that it sounds like it has been spoken by the target speaker, even though the target speaker does not speak that target language, or at least recordings of the target speaker are not available in that target language. In voice conversion, a source speaker who can speak the target languages speaks the utterance in the target language, and the conversion process transforms that source utterance into the target utterance. Preferably, characteristics of the source speaker, such as the speaker’s gender, are not manifested in the target utterance, while speaking style characteristics such as emotional cues or language accents are preserved and transferred from the source utterance to the target utterance. An “any-to-one voice conversion” (A2O VC) process can accept a source utterance from any speaker, even one not previously known, and nevertheless can convert the utterance to the target speaker’s voice. A “polyglot” process can accept input utterances in any of a set of multiple languages, at least some of which are not spoken by the target speaker, and nevertheless convert the utterance to the target speaker’s voice, which retaining the speaking style characteristics.

[0007] Summary

[0008]

[0004] In a general aspect, approaches described herein address the technical problem of configuring an any-to-one voice conversion system when few if any recordings are available for a particular target speaker in a language in which the voice conversion system will be used and for which it is to be configured. The voice conversion system, when presented with an input utterance from any speaker the system transforms the utterance into the voice of a specific target speaker. This transformation is achieved even if the target speaker does not speak the language of the input utterance or has a different gender than the speaker of the input utterance. Consequently, the output utterance will have the same linguistic content and the same style (e g., emotion or language accent) as in the input but with a different voice. The approaches described herein address goals of (a) enhancing the speech intelligibility and overall quality of the converted speech, especially in cross-lingual scenarios, (b) disentangling content and speaker information so that speaker characteristics are not transferred to the output while content including characteristics such as emotional cues or language accents are transferred, and (c) removing a need for training recordings of the target speaker in the target language, and more generally, reducing the amount of training records needed from the target speaker in any language. The approaches make use of one or both of a structural aspect of the transformation system, referred to herein as a “bottleneck”, that helps with the process of disentangling the content and the speaker information, and a training approach that involves “fine-tuning” the system using a relatively limited number of the target speaker’s utterances in a language other than the target language.

[0009]

[0005] The concept disclosed herein includes a way to convert a source utterance by any source speaker into a corresponding target utterance spoken by a target speaker and to do so in a way that is agnostic to gender and language and in a way that causes the target utterance to retain the style of the source utterance in the target utterance. Such a system makes it possible to create polyglot voices, to give those voices new styles and emotions, and to do so even if specific data for the style or language are not available for the target speaker.

[0010]

[0006] An object of the disclosure is that of achieving intra-lingual voice conversion and cross-lingual voice conversion that retains emotional content in the source utterance and that also circumvents the need for an extensive amount of high-quality training data.

[0011]

[0007] Another object is that of receiving an input utterance from an input speaker in a first language and having that input be transformed into an output utterance that sounds as if it had been spoken, instead, by a specific target speaker, even if that target speaker does not speak the first language and to do so in a way that preserves, in the output utterance, the same semantic content, style, and emotion as that used by the input speaker.

[0012]

[0008] In one aspect, in general, voice conversion to have voice characteristics of a target speaker of utterances spoken in a first language of a set of target languages by any source speaker other than the target speaker uses a voice converter that includes a content encoder and an acoustic model. The method of voice conversion includes configuring the voice converter, including pre-training the acoustic model using a pre-training dataset using pre-training utterances spoken in the first language. The pre-training includes processing an input signal representation for each utterance of the pre-training utterances including using the content encoder to determine respective vector sequences and computing output signal representations of said utterances. The pre-training further includes determining values of parameters of the acoustic model such that processing vector sequences for utterances of the pre-training dataset yields outputs that match output signal representations of said utterances. The configuring of the voice converter further includes fine tuning the acoustic model using a fine-tuning dataset including finetuning utterances spoken by the target speaker in a first language not in the set of target languages. The fine tuning includes processing an input signal representation each utterance of the fine-tuning utterances including using the content encoder to determine respective vector sequences and computing output signal representations of said utterances, and updating the values of the parameters of the acoustic model such that processing vector sequences for utterances of the fine-tuning dataset yields outputs improve their match to output signal representations of said utterances. The amount of data in the fine-tuning dataset is substantially smaller than the amount of data in the pretraining dataset.

[0013]

[0009] Aspects can include one or more of the following features.

[0014]

[0010] The input signal representation comprises a waveform representation, and wherein the output signal representation comprises a spectrogram representation.

[0015] [OU] The voice converter further comprises a vocoder that accepts output signal representation from the acoustic model and generates a waveform representation corresponding to output signal representation

[0016]

[0012] The configuring of the voice converter further comprises training the vocoder using utterances spoken by the target speaker, including determining an output signal representation of each utterance of said utterances spoken by the target speaker, and determining values of parameters of the vocoder such that processing output signal representations for said utterances yields outputs that match waveform signal representations of said utterances.

[0017]

[0013] Determining the output signal representation of an utterance comprises using an output of the content encoder as said output signal representation, or alternatively comprises processing a waveform signal representation of said utterance.

[0018]

[0014] The training of the vocoder is performed independently of the training of the acoustic model.

[0019]

[0015] The acoustic model comprises a dimension reducing stage that accepts a vector sequence from the content encoder and produces a reduced dimension vector sequence for processing by further stages of the acoustic model, and wherein one or both of the pre-training and the fine-tuning of the acoustic model include determining values of parameters of the dimension reducing stage.

[0020]

[0016] The configuring of the voice converter further comprises self-supervised training of the content encoder using waveform signals including utterances in the first language.

[0021]

[0017] The method further comprising using the voice converter, including processing a first waveform signal of an utterance spoken in the first language by a speaker other than the target speaker including providing the first waveform signal to the content encoder, passing an output of the content encoder to the acoustic model, and using an output of the acoustic model to generate a second waveform signal to have voice characteristics of the target speaker.

[0022]

[0018] In another aspect, in general, the disclosure features an apparatus that receives an input signal that includes content and first speaker-information and that produces an output signal that includes the content and second speaker-information, the second speaker-information having been derived from utterances by a target speaker and having supplanted the first speaker-information. Such an apparatus includes a first stage that receives the input signal and a second stage that receives an output of the first stage and provides the output signal, the first stage and the second stage being constituents of a voice converter and having been trained independently of each other, the first stage having been trained using a training dataset that includes utterances from other than the target speaker, and the second stage having been trained using a tuning dataset that includes utterances that consist of the second speaker-information.

[0023]

[0019] In some embodiments, the first stage includes an encoder and the output signal is a semantic vector that carries the first speaker-information. In these embodiments, the second stage includes a decoder that further transforms the semantic vector by incorporating the second speaker-information therein.

[0024]

[0020] In other embodiments, the output of the first stage is a semantic vector having a first dimensionality and the second stage includes a reducer that projects the semantic vector onto a subspace of dimensionality that is lower than the first dimensionality.

[0025]

[0021] In still other embodiments, the output of the first stage is a semantic vector having a first dimensionality and the second stage includes a reducer that suppresses noise and speaker information from the output of the first stage.

[0026]

[0022] Embodiments include those in which the training dataset is a monoglot dataset and those in which it is a polyglot dataset.

[0027]

[0023] Among the embodiments in which the second stage includes a decoder are those that further include a tuner that receives parameters that are used for the first stage and trains the second stage by fine tuning the parameters to reduce a difference between the content of the output signal and the content of the input signal.

[0028]

[0024] Also among those embodiments are those that include a tuner and a tuning dataset that is used by the tuner to train the second stage. In such embodiments, the tuning dataset includes utterances in a language that differs from that used for utterances in a training dataset used to train the first stage.

[0029]

[0025] In other embodiments, the first stage includes layers of nodes, the layers including a first layer, a last layer, and intermediate layers therebetween. In such embodiments, the output of the first stage that is received by the second stage is an output of one of the intermediate layers. Description of Drawings

[0030]

[0026] FIG. 1 shows an overview of a voice-converter according to one embodiment.

[0031]

[0027] FIGS. 2-4 shows details of three phases of operation of the voice converter of

[0032] FIG. 1.

[0033] Detailed Description

[0034]

[0028] Referring to FIG. 1, a voice converter 10 receives an input signal 12 and provides an output signal 14. The input signal 12 is spoken by an arbitrary input speaker 16. The output signal 14 is spoken using the voice of a target speaker 18.

[0035]

[0029] For ease of exposition, the term “signal” shall be used herein to refer to information that represents the acoustic manifestation of spoken utterance rather than the content spoken utterance itself. Examples of a “signal” includes samples of a time domain waveform (e.g., sampled at 16kHz), which represents an amplitude as a function of time, and a spectrogram, which represents energy of such a waveform as a function of time (e.g., in steps of 20ms) and frequency (e.g., in 10-100 mel-spaced frequency bins). In an example described below, the input signal 12 and the output signal 14 are waveforms, and may be referred to as the input waveform 12 and the output waveform 14, respectively. But it should be recognized that most aspects of approaches described below are applicable to other forms of signals.

[0036]

[0030] For further ease of exposition, voice conversion to generate an output signal 14 to sound as if a target speaker 18 had uttered the content of the input signal 12 can be referred to as “channeling.” Thus, the target speaker 18 can be viewed as “channeling” the input speaker 16 by delivering the same content as that input speaker 16 but with a different voice and with no attempt to impersonate the input speaker 16. In effect, the target speaker 18 can be viewed as speaking on behalf of the input speaker 16.

[0037]

[0031] Continuing to refer to FIG. 1, the voice converter 10 consists of two stages. A first stage 20 includes a content encoder 26 that extracts a speech representation 24 from an input the time-domain waveform 12 of any speaker. In the discussion below, this speech representation is described as consisting of a sequence of “semantic vector”, where the term “semantic” is used to connote that the vectors at least contain the content of the input signal, although they may inadvertently also include other information such as the characteristics of the input speaker. The second stage 22 includes an acoustic model 23, and a vocoder 29. The acoustic model 23 in combination with the vocoder 29 translates the speech representation 24, or semantic vector representation 24, into a timedomain waveform 14. In particular, the acoustic model 23 translates the representation 24 into a mel-spectrogram 13, and the vocoder 29 converts mel-spectogram 13 into the output waveform 14.

[0038]

[0032] A more precise problem definition can be expressed as follows. Consider a dataset D of a target speaker containing M utterances in the time domain. Let us denote the ithutterance 14 of the target speaker as its features 24 extracted by the content encoded 26 as vte IRS, and its mel-spectrogram 13 as x where i E [1, M].

[0039]

[0033] The content encoder 26 (whose function is denoted C ) takes utas its input and generates vtas vt= C(ui; wc'), where wcrepresents the content encoder parameters. The acoustic model 23 includes two parts. A first part 32, which may be referred to as a “reducer,” which removes speaker information in Vt while minimizing its impact on other components (e g., linguistic content). This part is implemented as a neural network (a “pre-network”), whose function is denoted P, which reduces the dimensionality of the semantic vectors (e.g., “projects”) into a spaced, where d « s. The acoustic model includes a second part, whose function is denoted A , which uses the output of the pre-net P. This dimensionality gap imposes an information bottleneck, which allows the acoustic model to discard content-irrelevant information such as noise and speaker identity. Thus, the acoustic model predicts the mel-spectrogram xtas xt= A(P(viwP); vy . That is, the acoustic model has two sets of parameters, wPand wA, where wPrepresents the parameters of the pre-net, while wArepresents those of the second part of the acoustic model. Finally, the vocoder V, given x;, generates utas vt= V(xt; wv) where wvrepresents the vocoder parameters. Therefore, the entire transformation from the input to the second stage to the waveform output is parameterized by wP, wA, and wv.

[0040]

[0034] A particular advantage of the method described herein is that the target speaker 18 need not actually know the second language. The voice converter 10 configured as described herein outputs audio in a second language as if it were spoken by the target speaker 18 even though the target speaker 18 knows nothing of that second language. Moreover, the voice converter 10 does so in such a way that one who hears the output signal 14 might reasonably believe that the target speaker 18 actually had some fluency in that second language.

[0035] A further advantage of the method described herein is that speech characteristics, such as manifestations of emotion (e.g., anger, empathy, etc.), present in the input signal 12 are passed through to the output signal 14.

[0041]

[0036] Configuration of the voice converter involves determining the numerical values for the parameters wP, wA, wv, and wc, a process referred to as “training”. This training process is performed in several distinct steps, which address different parts of the converter in each step.

[0042]

[0037] Turning first to the content encoder 26, which transforms an input waveform into a more concise semantic vector representation 24 (also referred to as a selfsupervised learning (SSL) representation). The content encoder is training to determine the values of the parameters wcin a self-supervised manner, which means that input waveforms are used for the training, but annotations of those waveforms, for example, with the words spoken or other content or style information is not needed. The content encoder 26 is a neural network that uses a transformer model as its backbone, and convolutional network that processes the input waveform to yield the inputs to the transformer model. Those inputs each represent about 25ms of audio with a stride of about 20ms. The output of the content encoder comprises one semantic vector for each 20ms of input. The hidden layers of the transformer have roughly 1024 units in each layer. The encoder is trained in an unsupervised manner based on a masked prediction loss, where masking of a section of input is done by corrupting the section of waveform by adding noise or additional speech. In particular, an example of the content encoder was trained using 94,000 hours of speech. Preferably, the encoder is trained on recordings in multiple languages (i.e., a “polyglot” encoder), and preferably includes recordings in the target language or languages. However, the encoder does not necessarily have to be trained with recordings from the target language as other steps in training can compensate for some such omissions.

[0043]

[0038] In one example of the content encoder 26, a WavLM pretrained transformer is used as described in Sanyuan Chen et al., “WavLM: Large-scale self-supervised pretraining for full stack speech processing,” IEEE Journal of Selected Topics in Signal Processing, 16: 1505-1518 (2021). In particular, publicly available trained parameters of WavLM-Large were used. WavLM-Large has about 24 layers in the transformer model. For the content encoder 26, the output of the 15thlayer is used, thereby yielding 1024- dimensional semantic vectors vtat a rate of one every 20ms. The part of the encoder up to the 15thlayer may be referred to as the encoder “core”. In the example described below, the parameters W of the content encoder are fixed for the remaining training steps. In alternative examples, other encoders can be used, for example, HuBERT (Wei- Ning Hsu etal., “HuBERT: Self-supervised speech representation learning by masked prediction of hidden units.” IEEE / ACM Transactions on Audio, Speech, and Language Processing, 29:3451-3460 (2021)), or Wav2vec2 (Alexei Baevski et al., “Wav2vec 2.0: A framework for self-supervised learning of speech representations.” ArXiv, abs'2006.11477 (2020)) can be used instead of WavLM in a similar manner.

[0044] The vocoder V 29 is trained only on target speaker data of paired spectrograms and waveforms (x£, uz). In some examples of this training, the spectrograms x£are true mel-spectrograms computed using standard techniques directly from the corresponding waveforms u(. As an alternative, the spectrograms are the outputs x£of the encoder with the corresponding waveform being input to the encoder. In some examples, the vocoder is initially trained on the true spectrograms, and then later retrained on the output of the encoder after other training steps. Note that the vocoder is trained using recording of the target speaker in a language other than the target language because such recordings are not assumed to be available. The criterion for training the vocoder is either or minWvLv(ui, V(_xi; wv'))>where Lvis a loss function in the time domain. A variety of computational techniques may be used for training. One such technique makes use of gradient descent using back- propagation calculations.

[0045]

[0039] By training only on target speaker data, the loss makes the vocoder target speaker specific so that it generates waveforms that have many of the character of the target speaker.

[0046]

[0040] In some examples, the vocoder is a HiFi-GAN model, as described in Jungil Kong et al. “Hifi-GAN: Generative adversarial networks for efficient and high fidelity speech synthesis.” In Advances in Neural Information Processing Systems, volume 33, pages 17022-17033. Curran Associates, Inc. (2020).

[0041] The next step of training determines the initial parameters of the acoustic model, wPand wA. The acoustic model is a neural network that takes as input the sequence of semantic vectors and outputs a mel-spectrogram In an example of the acoustic model, the architecture features an encoder and an autoregressive decoder. Both the encoder and decoder are preceded by a feed-forward pre-net, while a final linear layer with n -MELs units follows the decoder. No attention module is used. The encoder prenet is a feed-forward neural net composed of a stack of two linear layers with 256 units, ReLU activations, and dropout. The encoder consists of a stack of three ID-convolutional layers with 512 units, kernel size 5, stride 1, padding 2, and ReLU activations. An Instance Normalization (IN) layer is added after each ConvlD layer to help to further reduce speaker identity while preserving content information. The decoder predicts each spectrogram frame from the output of the encoder and the previously generated frames. First, a decoder pre-net, similar to the encoder one, is applied and then three LSTMs with 768 units each are followed by a final linear layer with n-MELs units are used.

[0047]

[0042] Recall above that a dataset D of a target speaker contains M utterances as the time domain waveforms of speech in a language lang2. The goal of the voice conversion is to generate waveforms in a language langl for which training utterances are not available in that language for the target speaker. The acoustic model, P and A, is preferably pretrained on a dataset D' consisting of N speakers, each of whom has M’ utterances in langl. Let us denote the ithutterance of the jthspeaker as Uji, semantic vector as 1^7, and its mel-spectrogram as xjwhere j E [1, IV] and i E The pretraining of the model employs pairs for each Ujtfrom the dataset D' . That is, the Vji are computed using the content encoder and the Xjtare computed from the waveforms Ujt. The training loss becomes

[0048]

[0043] After pre-training, the acoustic model A is “fine tuned” using the dataset D (or a subset thereof) from the target speaker. This procedure involves pairs (i?i, xt~) for each U[ of the dataset D. The acoustic model A is fine tuned by optimizing the following objective function: minWp iWALAxi, A(P c, wP, w ) where LAis a loss function in the time-frequency domain such as the distance between target and predicted mel-spectrograms.

[0049]

[0044] In some examples, the acoustic model is pretrained in multiple languages langla, langlb, etc. , and then fine-tuned on lang2. This results in a “polyglot” system in which outputs can be generated in any language langla, langlb, etc. as well as lang2. In some examples, the content encoder is trained on multiple languages, for instance including langla, langlb, etc. as well as lang2. Preferably, both datasets D and D' include utterances spoken in a variety of speaking styles, including with various emotional content.

[0050]

[0045] The pre-training followed by fine tuning allows the acoustic model to acquire phonetic coverage of langl during the pre-training, which can be leveraged during the fine-tuning on lang2. Consequently, the lang2 target speaker will exhibit improved langl linguistic characteristics after the conversion.

[0051]

[0046] The cross-lingual fine-tuning technique offers several advantages. It enhances the linguistic content and pronunciation in cross-lingual voice conversion and reduces the amount of training data needed from the target speaker. While training the acoustic model without fine-tuning typically requires at least 10-12h of high-quality audio from the target speaker, the fine-tuning only requires, for example, 2h.

[0052]

[0047] Referring now to FIGS. 2-4, the overall operation of the voice converter 10 includes three distinct phases: a training phase 36, a tuning phase 38, and a runtime phase 40. In the examples shown, FIG. 2 represents an example training phase 36, FIG. 3 illustrates an example tuning phase 38, and FIG. 4 illustrates an example runtime phase 40.

[0053]

[0048] Referring first to FIG. 2, the training phase 36 is carried out in the first language using a training dataset 42. The training dataset 42 is either a single-speaker dataset or a multi-speaker dataset.

[0054]

[0049] Referring next to FIG. 3, the tuning phase 38 carries out further training to accommodate the target speaker 18 on a second language. The training phase 36 and the tuning phase 38 cooperate to cause the output signal 14 corresponding to a particular input signal 12 to both have the same non-linguistic content (e.g., emotional content) as the input signal 12 and to be spoken in the voice of the target speaker 18, in effect to channel the input speaker 16.

[0050] Referring finally to FIG. 4, the runtime phase 40 is the simplest of them all. The voice converter 10 simply receives an input signal 12 and uses the parameters that result from having carried out the training phase 36 and the tuning phase 38 to process that input signal 12.

[0055]

[0051] Returning now to FIG. 2, the training dataset 42 includes utterances 44 generated by one or more training speakers 46 in the first language. The goal of the training phase 36 it to cause the output signal 14 corresponding to a particular input signal 12 to have the same non-linguistic content (e.g., emotional content) as the input signal 12 and to be spoken in the voice of the target speaker 18, in effect to channel the input speaker 16.

[0056]

[0052] The training phase 36 uses a trainer 48 whose purpose is to adjust parameters 64 for the decoder 34 so that the output signal 14 matches the input signal 12 whose semantic vector representation 24 is given as input to the second stage 22. For reasons that will be apparent in the tuning phase 38, these parameters 64 will henceforth be referred to as the “pre-tuning parameters 64.”

[0057]

[0053] To carry out this adjustment, the trainer 48 executes a training loop in which it receives the output signal 14 and, based on a difference between the output signal 14 and the input signal 12, incrementally adjusts the pre-trained parameters 50 in an effort to make the output signal 14 sound like the desired signal. The trainer 48 repeats this process multiple times. This results in pre-tuning parameters 64 that have been optimized to cause the output signal 14 to sound as much as possible like the target speaker(s) 18 speaking in the first language.

[0058]

[0054] However, even after having completed the training phase 36, the voice converter 10 still lacks competence in channeling speech as the target speaker 18. In addition, the voice converter 10, although competent in a first language, lacks usable competence in other languages, and in particular, in a second language that is spoken by the target speaker 18. To acquire such competence, something more is required: the tuning phase 38.

[0059]

[0055] Referring next to FIG. 3, the tuning phase 38 makes use of a tuning dataset 52 that comprises utterances 54 by the target speaker 18 in a second language. The utterances 54 are ground-truth utterances. In either case, training the decoder 34 only using utterances by the target speaker 18 forces the decoder 34 to incorporate speaker information for the target speaker 18. In the case in which a reducer 32 has removed speaker information from the input speaker 16, the speaker information from the target speaker 18 fills in the gap without changing the remaining content, namely the semantic content and the emotion content. In the case in which no reducer 32 is used, the same effect occurs, albeit more slowly, with the speaker information for the target speaker 18 diluting that from the input speaker 16. In either case, the technical effect is the same, namely that of modifying speaker information independently of all other content in the input signal 12.

[0060]

[0056] The tuning dataset 52 and a tuner 58 cooperate to train the decoder 34 to enable the target speaker 18 to both channel the input speaker 16 and to also improve the target speaker’s pronunciation in the first language. Since the target speaker 18 may not be fluent in the first language, no utterances of the target speaker 18 in the first language are available in the tuning dataset 52.

[0061]

[0057] The tuner 58 uses the tuning dataset 52 to tune, or adjust, the pre-trained parameters 50 that were developed by the trainer 48 in an effort to accommodate the second language. The procedure is similar to that carried out by the trainer 48. In this case, the tuner 58 fine tunes the pre-tuning parameters 64 in an effort to make the voice converter’s output signal 14 sound as much as possible as if it had been uttered by the target speaker 18 instead of being uttered by the input speaker 16. The resulting parameters that are provided by the tuner 58 will be referred to herein as the “tuned parameters 60.”

[0062]

[0058] Referring now to FIG. 4, in the runtime phase 40, the voice converter 18 receives an input signal 12 and carries out a transformation using these tuned parameters 60 instead of the pre-tuning parameters 64 developed by the trainer 48.

[0063]

[0059] The foregoing process has been described in terms of two languages. However, it is not limited to two languages. It is possible to repeat the tuning phase 38 in more than one language, thus endowing the decoder 34 with the ability to channel in more than two languages.

[0064]

[0060] As shown in FIG. 4, when using the tuned parameters 60 on an input signal 12 in the first language, the result is an output signal 14 that sounds as if it had been spoken by the target speaker 18, even though the target speaker 18 may never have heard the second language in his life. The resulting linguistic content of the output signal 14 is substantially similar to what would have resulted had the pre-tuning parameters 64 been used instead of the tuned parameters 60 except for the target speaker identity. When using the tuned parameters 60 on an input signal 12 in the second language, the result is an output signal 14 that sounds as if it had been spoken by the target speaker 18 in the second language.

[0065]

[0061] In view of the above, the acoustic model is able to influence the output signal 14 to an extent that is disproportionate to the amount of training data used to train the decoder 34. Whereas the training data required by the content encoder core 28 is quite voluminous, the decoder 34 can exert a significant influence after having been trained by only a modest amount of training data. This makes it possible to leverage the existence of a pre-trained content encoder with immutable encoder parameters by customizing it with an acoustic model that has been trained using only a small amount of additional training data.

[0066]

[0062] Implementations of a voice converter as described above can be used in a variety of applications. For instance, configuration of a voice user interface may require generation of voice output that sounds like a particular person, but speaking a language that the particular person may not be fluent in. The voice conversion may be used to precompute output prompts, or may be used to augment training data that is used to train a text-to-speech system, for example, by replacing the content encoder with automated conversion of text to a sequence of semantic vectors. As an example, a voice user interface for an automotive system may need to be customized with different languages for different countries of sale, and may further be customized to support output in a number of languages for some countries. Other examples of use of the voice converter include in entertainment, for example, when dubbing a movie into a language not spoken by the actors, thereby maintaining the actor’s voice characteristic transferring the emotional cues present in the voice talent that would otherwise have been directly recorded for the dubbed movie.

[0067]

[0063] Implementations of the approaches described above, including the training steps as well as the runtime steps, may make use of software, which is stored on a non- transitory machine-readable medium (e.g., a solid state storage). The instructions, when executed on a processor cause the processor to perform the steps of the described approaches. The processor (or processors) can include general-purpose processors (e.g., CPU of a general-purpose computer) and can include special-purpose processors suitable for the numerical computations of the approaches (e.g., graphics processing units, GPUs). In addition to such circuit-implemented processors, virtual (i.e., software-implemented) processors can be used. The approaches may also make use of hardware (circuitry) that performs special purpose functions instead of or in addition to software-based components. An apparatus for performing the approaches described above may include such a processor or processors and storage for the instructions and / or circuitry for special purpose functions.

[0068]

[0064] Having described the disclosure and a preferred example thereof, what is claimed as new and secured by letters patent is:

Claims

CLAIMS1. A method for voice conversion to have voice characteristics of a target speaker of utterances spoken in a first language of a set of target languages by any source speaker other than the target speaker, the voice conversion using a voice converter that includes a content encoder (26) and an acoustic model (23), the method comprising configuring the voice converter, including: pre-training the acoustic model using a pre-training dataset (£>') including pretraining utterances spoken in the first language, the pre-training including processing an input signal representation (u) for each utterance of the pretraining utterances including using the content encoder to determine respective vector sequences (v) and computing output signal representations (x) of the utterances, and determining values of parameters of the acoustic model such that processing vector sequences for utterances of the pre-training dataset yields outputs that match output signal representations of the utterances; fine tuning the acoustic model using a fine-tuning dataset (D) including fine- tuning utterances spoken by the target speaker in a first language not in the set of target languages, the fine tuning including processing an input signal representation each utterance of the fine-tuning utterances including using the content encoder to determine respective vector sequences and computing output signal representations of the utterances, and updating the values of the parameters of the acoustic model to such that processing vector sequences for utterances of the fine-tuning dataset yields outputs improves their match to output signal representations of the utterances, wherein the amount of data in the fine-tuning dataset is substantially smaller than the amount of data in the pre-training dataset.

2. The method of claim 1, wherein the input signal representation comprises a waveform representation, and wherein the output signal representation comprises a spectrogram representation.

3. The method of any of claims 1 and 2, wherein the voice converter further comprises a vocoder (29) that accepts the output signal representation from the acoustic model and generates a waveform representation corresponding to the output signal representation, and wherein the configuring of the voice converter further comprises: training the vocoder using utterances spoken by the target speaker, including determining an output signal representation of each utterance of the utterances spoken by the target speaker, and determining values of parameters of the vocoder such that processing output signal representations for the utterances yields outputs that match waveform signal representations of the utterances.

4. The method of claim 3, wherein determining the output signal representation of an utterance comprises using an output of the content encoder as the output signal representation.

5. The method of claim 3, wherein determining the output signal representation of an utterance comprises processing a waveform signal representation of the utterance.

6. The method of claim 3, wherein the training of the vocoder is performed independently of the training of the acoustic model.

7. The method of any of the preceding claims wherein the acoustic model comprises a dimension reducing stage that accepts a vector sequence from the content encoder and produces a reduced dimension vector sequence for processing by further stages of the acoustic model, and wherein one or both of the pre-training and the finetuning of the acoustic model include determining values of parameters of the dimension reducing stage.

8. The method of any of the preceding claims wherein configuring the voice converter further comprises self-supervised training of the content encoder using waveform signals including utterances in the first language.

9. The method of any of the preceding claims further comprising using the voice converter, including: processing a first waveform signal of an utterance spoken in the first language by a speaker other than the target speaker including providing the first waveform signal to the content encoder, passing an output of the content encoder to the acoustic model, and using an output of the acoustic model to generate a second waveform signal to have voice characteristics of the target speaker.

10. A non-transitory machine-readable medium comprising instructions stored thereon, the instructions when executed by a data processing system cause the system to perform all the steps of any one of the preceding claims.

11. An apparatus that receives an input signal that comprises content and first speaker-information and that produces an output signal that comprises the content and second speaker-information, the second speaker-information having been derived from utterances by a target speaker and having supplanted the first speaker-information, wherein the apparatus comprises a first stage that receives the input signal, and a second stage that receives an output of the first stage and provides the output signal, the first stage and the second stage being constituents of a voice converter and having been trained independently of each other, the first stage having been trained using a training dataset that comprises utterances from other than the target speaker, and the second stage having been trained using a tuning dataset that comprises utterances that consist of the second speaker-information.

12. The apparatus of claim 11, wherein the first stage comprises an encoder and the output signal is a semantic vector that carries the first speaker-information and wherein the second stage comprises a decoder that further transforms said semantic vector by incorporating the second speaker-information therein.

13. The apparatus of claim 11, wherein the output of the first stage is a semantic vector having a first dimensionality and the second stage comprises a reducer that projects the semantic vector onto a subspace of dimensionality that is lower the said first dimensionality.

14. The apparatus of claim 11, wherein the output of the first stage is a semantic vector having a first dimensionality and the second stage comprises a reducer that suppresses noise and speaker information from the output of the first stage.

15. The apparatus of claim 11, wherein the training dataset is a monoglot dataset.

16. The apparatus of claim 11, wherein the training dataset is a polyglot dataset.

17. The apparatus of claim 11, wherein the second stage comprises a decoder and wherein the apparatus further comprises a tuner that receives parameters that are used for the first stage and trains the second stage by fine tuning the parameters to reduce a difference between the content of the output signal and the content of the input signal.

18. The apparatus of claim 11, further comprising a tuner and a tuning dataset that is used by the tuner to train the second stage, wherein the tuning dataset comprises utterances in a language that differs from that used for utterances in a training dataset used to train the first stage.

19. The apparatus of claim 11, wherein the first stage comprises layers the nodes, the layers comprising a first layer, a last layer, and intermediate layers therebetween and wherein the output of the first stage that is received by the second stage is an output of one of the intermediate layers.

Citation Information

Patent Citations

  • Multilingual speech synthesis and cross-language voice cloning

    US11580952B2