Voice style extraction model training method, voice synthesis method, device and medium

The speech style extraction model trained with data augmentation and loss function solves the problems of high sample cost and labeling errors in object style transfer tasks, and improves the reliability and accuracy of the model.

CN115985284BActive Publication Date: 2026-07-24BEIJING DAJIA INTERNET INFORMATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
BEIJING DAJIA INTERNET INFORMATION TECH CO LTD
Filing Date
2022-12-09
Publication Date
2026-07-24

Smart Images

  • Figure CN115985284B_ABST
    Figure CN115985284B_ABST
Patent Text Reader

Abstract

The present disclosure relates to the technical field of computer, in particular to a voice style extraction model training method, a voice synthesis method, a voice style extraction model training device, a voice synthesis device, a computer readable storage medium and an electronic device. The voice style extraction model training method comprises: obtaining a reference voice sample; performing data enhancement processing to obtain an adversarial voice sample; obtaining a synthesized voice sample; inputting the reference voice sample, the adversarial voice sample and the synthesized voice sample into a voice style extraction model to be trained respectively for style coding processing to obtain a predicted reference style feature, a predicted adversarial style feature and a predicted synthesized style feature; determining an adversarial loss function and a consistency loss function; and updating parameters of the voice style extraction model to be trained. Through the technical scheme of the embodiment of the present disclosure, the problem of low accuracy in realizing object style transfer task in the prior art can be solved.
Need to check novelty before this filing date? Find Prior Art