A method for residual adapters for few-shot text-to-speech
speaker adaptation includes obtaining a text-to-speech (TTS) model configured to convert text into representations of synthetic speech, the TTS model pre-trained on an initial training
data set. The method further includes augmenting the TTS model with a stack of residual adapters. The method includes receiving an adaption training
data set including one or more spoken utterances spoken by a target speaker, each spoken
utterance in the
adaptation training
data set paired with corresponding input text associated with a transcription of the spoken
utterance. The method also includes adapting, using the adaption training data set, the TTS model augmented with the stack of residual adapters to learn how to synthesize speech in a voice of the target speaker by optimizing the stack of residual adapters while parameters of the TTS model are frozen.