Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

9 results about "Speaker adaptation" patented technology

Video translation method and system based on artificial intelligence

The invention discloses a video translation method and system based on artificial intelligence. The method relates to the technical field of video translation and comprises the following steps of original sound track extraction, target AI speaker adaptation, AI dubbing generation and mouth shape synchronization and video synthesis. According to the method, independent audio and video streams are obtained by adopting an audio and video separation technology, and multiple original sound tracks are extracted through a voice separation model; matching or generating an adaptive target AI speaker module in a preset tone library; converting the original language voice into a text, translating the text into a target language text, and synthesizing an AI dubbing audio track in combination with a target AI speaker module; and finally, the independent video stream and the multi-AI dubbing audio track are input into the mouth shape synchronization model to output a translated video, so that the timbre fitting degree, the voice quality and the voice consistency of the same speaker of AI dubbing are improved, and meanwhile, the resource utilization rate of video translation and the processing efficiency under batch tasks are improved. The problem that in the prior art, video translation is low in quality and efficiency is solved.
Owner:BEIJING DEEP LOGIC INTELLIGENT TECHNOLOGY CO LTD

Speaker embedding-based speaker adaptation method and system generated by using global style tokens and prediction model

Disclosed are a speaker embedding-based speaker adaptation method and system generated by using global style tokens and a prediction model. The speaker adaptation method performed by the speaker adaptation system, according to an embodiment, may comprise the steps of: generating a plurality of speaker embeddings representing the tone of a speaker from a speaker embedding by using a voice transformation model including a global style token mechanism; and predicting the final speaker embedding representing a new speaker through similarity comparison between a new speaker embedding predicted by using a prediction model for predicting a speaker embedding and the plurality of generated speaker embeddings.
Owner:INDUSTRY UNIVERSITY COOPERATION FOUNDATION HANYANG UNIVERSITY

IMPROVING THE NATURALNESS OF SPEAKER-ADAPTED SPEECH SYNTHESIS

Techniques for improving the naturalness of synthetic speech are revealed. A speaker-adaptation model for a speech synthesis pipeline is presented, achieved by training an acoustic model, along with a post-processing model for modifying speech features output by the acoustic model. Additional reference training examples based on simulated input-output datasets with limited resources for reference speakers are provided for this training.
Owner:FRAUNHOFER GESELLSCHAFT ZUR FORDERUNG DER ANGEWANDTEN FORSCHUNG EV

Speaker-adaptive speech end detection for conversational AI applications

In various instances, an end of speech (EOS) of an audio signal is determined based at least in part on a speech rate of a speaker, and for a segment of the audio signal, an EOS is indicated based at least in part on an EOS threshold determined based at least in part on the speech rate of the speaker.
Owner:NVIDIA CORP

Naturalness of speaker-adapted speech synthesis

Techniques of improving the naturalness of synthetic speech are disclosed. Speaker-adaption of a speech synthesis pipeline by training of an acoustic model and a post-processing model for modifying speech features output by the acoustic model are disclosed. For this training, additional reference training samples are obtained based on simulated low-resource input-output datasets for reference speakers.
Owner:FRAUNHOFER GESELLSCHAFT ZUR FORDERUNG DER ANGEWANDTEN FORSCHUNG EV

Residual adapters for few-shot text-to-speech speaker adaptation

A method for residual adapters for few-shot text-to-speech speaker adaptation includes obtaining a text-to-speech (TTS) model configured to convert text into representations of synthetic speech, the TTS model pre-trained on an initial training data set. The method further includes augmenting the TTS model with a stack of residual adapters. The method includes receiving an adaption training data set including one or more spoken utterances spoken by a target speaker, each spoken utterance in the adaptation training data set paired with corresponding input text associated with a transcription of the spoken utterance. The method also includes adapting, using the adaption training data set, the TTS model augmented with the stack of residual adapters to learn how to synthesize speech in a voice of the target speaker by optimizing the stack of residual adapters while parameters of the TTS model are frozen.
Owner:GOOGLE LLC

Residual adapters for few-shot text-to-speech speaker adaptation

A method for residual adapters for few-shot text-to-speech speaker adaptation includes obtaining a text-to-speech (TTS) model configured to convert text into representations of synthetic speech, the TTS model pre-trained on an initial training data set. The method further includes augmenting the TTS model with a stack of residual adapters. The method includes receiving an adaption training data set including one or more spoken utterances spoken by a target speaker, each spoken utterance in the adaptation training data set paired with corresponding input text associated with a transcription of the spoken utterance. The method also includes adapting, using the adaption training data set, the TTS model augmented with the stack of residual adapters to learn how to synthesize speech in a voice of the target speaker by optimizing the stack of residual adapters while parameters of the TTS model are frozen.
Owner:GOOGLE LLC

Conference room sound system based on single-channel target spectrum adaptive FIR (Finite Impulse Response) filter

The invention provides a conference room sound system based on a single-channel target spectrum adaptive FIR filter, and relates to the technical field of audio signal processing. According to the conference room sound system based on the single-channel target spectrum adaptive FIR filter, only one microphone is used for collection, deployment is simplified, a multi-microphone array and single-channel multi-speaker adaptation are not needed, and a user can import any'good hearing 'target spectrum envelope curve; when the system operates, voice collected by a single-channel microphone is used as a driving signal, errors between an output frequency spectrum and a target curve are calculated in real time, FIR filter coefficients are updated online, spectral errors between output and a target are calculated in each frame, FIR frequency response is adaptively updated in a frequency domain through NLMS / LMS, meanwhile, only spectral envelope is corrected, an original phase / F0 is not affected, and the frequency response of the FIR is corrected. And the suppression, pause and personality of the speaker are kept, so that the output spectrum is continuously close to the target curve and the rhythm and fundamental frequency of the speaker are kept.
Owner:SHENZHEN TONGCHUANG AUDIO TECH CO LTD

Zero-sample speech synthesis method based on denoising non-parametric speaker adaptation

PendingCN120708592ASpeech synthesisSpeech synthesisNearest neighbor classifier
The invention relates to a zero-sample speech synthesis method based on denoising non-parametric speaker adaptation, and belongs to the technical field of natural language processing. The method is used for a high-fidelity zero-sample speaker to adapt to speech without a speaker, and zero-sample speech synthesis faces the problems that a model is usually trained on a small data set, and the generalization ability is poor; when the condition depends on noise reference, unseen speaker identities are difficult to model; according to the zero-sample speech synthesis method based on denoising non-parametric speaker adaptation, speaker embedding is predicted by using a nearest neighbor classifier on a database of a large number of cache examples, similarity search is carried out by using representation of a speaker encoder, and the method does not need extra training and is high in robustness. Therefore, a highly robust model is formed, and the performance can be improved when the method depends on noise reference.
Owner:小语智能信息科技(云南)有限公司