Embodiments of the present application provide a phoneme-based speech domain migration method,
system and electronic device. The method comprises: performing
grapheme-to-phoneme conversion on target domain text to obtain a target domain phoneme sequence; converting the target domain phoneme sequence into a plurality of target domain phoneme N-
gram sequences according to a phoneme N-
gram dictionary; generating target domain speech segments with speaker diversity using the plurality of target domain phoneme N-
gram sequences; and generating synthesized audio for the target domain based on the target domain speech segments. Embodiments of the present application use a dictionary constructed from speech segments of basic phoneme n-grams to generate speech for the text of the target domain, the splicing and synthesis method guided by phonemes has the ability to model the connected reading between words, and because the audio segments for splicing and synthesis come from a large amount of real speech, the synthesized speech has better speaker diversity, and the required computing resources are reduced, avoiding
overfitting of the ASR model on the synthesized data.