Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

6 results about "Speaker adaptation" patented technology

Video translation method and system based on artificial intelligence

The invention discloses a video translation method and system based on artificial intelligence. The method relates to the technical field of video translation and comprises the following steps of original sound track extraction, target AI speaker adaptation, AI dubbing generation and mouth shape synchronization and video synthesis. According to the method, independent audio and video streams are obtained by adopting an audio and video separation technology, and multiple original sound tracks are extracted through a voice separation model; matching or generating an adaptive target AI speaker module in a preset tone library; converting the original language voice into a text, translating the text into a target language text, and synthesizing an AI dubbing audio track in combination with a target AI speaker module; and finally, the independent video stream and the multi-AI dubbing audio track are input into the mouth shape synchronization model to output a translated video, so that the timbre fitting degree, the voice quality and the voice consistency of the same speaker of AI dubbing are improved, and meanwhile, the resource utilization rate of video translation and the processing efficiency under batch tasks are improved. The problem that in the prior art, video translation is low in quality and efficiency is solved.
Owner:BEIJING DEEP LOGIC INTELLIGENT TECHNOLOGY CO LTD

IMPROVING THE NATURALNESS OF SPEAKER-ADAPTED SPEECH SYNTHESIS

Techniques for improving the naturalness of synthetic speech are revealed. A speaker-adaptation model for a speech synthesis pipeline is presented, achieved by training an acoustic model, along with a post-processing model for modifying speech features output by the acoustic model. Additional reference training examples based on simulated input-output datasets with limited resources for reference speakers are provided for this training.
Owner:FRAUNHOFER GESELLSCHAFT ZUR FORDERUNG DER ANGEWANDTEN FORSCHUNG EV

Speaker-adaptive speech end detection for conversational AI applications

In various instances, an end of speech (EOS) of an audio signal is determined based at least in part on a speech rate of a speaker, and for a segment of the audio signal, an EOS is indicated based at least in part on an EOS threshold determined based at least in part on the speech rate of the speaker.
Owner:NVIDIA CORP

Naturalness of speaker-adapted speech synthesis

Techniques of improving the naturalness of synthetic speech are disclosed. Speaker-adaption of a speech synthesis pipeline by training of an acoustic model and a post-processing model for modifying speech features output by the acoustic model are disclosed. For this training, additional reference training samples are obtained based on simulated low-resource input-output datasets for reference speakers.
Owner:FRAUNHOFER GESELLSCHAFT ZUR FORDERUNG DER ANGEWANDTEN FORSCHUNG EV

Residual adapters for few-shot text-to-speech speaker adaptation

A method for residual adapters for few-shot text-to-speech speaker adaptation includes obtaining a text-to-speech (TTS) model configured to convert text into representations of synthetic speech, the TTS model pre-trained on an initial training data set. The method further includes augmenting the TTS model with a stack of residual adapters. The method includes receiving an adaption training data set including one or more spoken utterances spoken by a target speaker, each spoken utterance in the adaptation training data set paired with corresponding input text associated with a transcription of the spoken utterance. The method also includes adapting, using the adaption training data set, the TTS model augmented with the stack of residual adapters to learn how to synthesize speech in a voice of the target speaker by optimizing the stack of residual adapters while parameters of the TTS model are frozen.
Owner:GOOGLE LLC

Residual adapters for few-shot text-to-speech speaker adaptation

A method for residual adapters for few-shot text-to-speech speaker adaptation includes obtaining a text-to-speech (TTS) model configured to convert text into representations of synthetic speech, the TTS model pre-trained on an initial training data set. The method further includes augmenting the TTS model with a stack of residual adapters. The method includes receiving an adaption training data set including one or more spoken utterances spoken by a target speaker, each spoken utterance in the adaptation training data set paired with corresponding input text associated with a transcription of the spoken utterance. The method also includes adapting, using the adaption training data set, the TTS model augmented with the stack of residual adapters to learn how to synthesize speech in a voice of the target speaker by optimizing the stack of residual adapters while parameters of the TTS model are frozen.
Owner:GOOGLE LLC