Voice conversion method and system for non-parallel data
A non-parallel data and voice conversion technology, which is applied in the voice conversion method and system field of non-parallel data, can solve problems such as difficulty in obtaining parallel data and poor conversion effect, and achieve the effect of ensuring conversion quality
Patent Information
- Authority / Receiving Office
- CN · China
- Current Assignee / Owner
- Publication Date
- 2020-11-20
Smart Images

Figure 1 
Figure 2
Abstract
Description
technical field
[0001] The invention relates to the technical field of voice conversion, in particular to a voice conversion method and system for non-parallel data. Background technique
[0002] Speech conversion is a technique for modifying a source speaker's speech signal to match a target speaker's speech signal, so that it has the speech characteristics of the target speaker while keeping the speech information unchanged. The main task of speech conversion includes extracting and converting the characteristic parameters representing the personality of the speaker, and then reconstructing the converted parameters into speech. This process must not only ensure the clarity of the converted speech, but also ensure the similarity of the converted speech features.
[0003] In the existing speech conversion technology, most of the methods require two speakers to have parallel data (the text content corresponding to the speech is consistent). The main disadvantage of this meth...
Examples
Embodiment Construction
[0051] The preferred embodiments of the present invention will be described below in conjunction with the accompanying drawings. It should be understood that the preferred embodiments described here are only used to illustrate and explain the present invention, and are not intended to limit the present invention.
[0052]The embodiment of the present invention provides a voice conversion method for non-parallel data, such as figure 1 As shown, the method performs the following steps:
[0053] Step 1: using large-scale synthetic sound database data other than the source speaker and the target speaker to train the speech synthesis model of the target speaker, wherein the large-scale synthetic sound database data includes text data and speech pair data;
[0054] Step 2: Based on the text corresponding to the speech data of the source speaker, generate parallel data corresponding to the target speaker according to the speech synthesis model of the target speaker, wherein the paral...