The invention relates to the field of
artificial intelligence, is applied to financial and medical scenes, and discloses a text-to-voice method and device,
computer equipment and a storage medium, and the method comprises the steps: receiving a
target text, a voice prompt and an emotion
label; performing semantic coding on the
target text to extract text semantic features; performing acoustic coding on the voice prompt to extract voice acoustic features; based on the emotion
label and the text
semantic feature, generating a dynamic emotion feature through a
time sequence emotion model; performing cross-
modal fusion on the voice acoustic features and the dynamic emotion features to obtain cross-
modal alignment features, and inputting the cross-
modal alignment features and the text semantic features into a pre-training
language model for
feature fusion and voice token prediction to obtain an initial voice token sequence; and performing
semantic alignment and spectrum conversion on the initial voice token sequence to generate a target voice waveform. According to the invention, high-
quality voice with natural emotional circulation, controllable
timbre and high conformity with professional contexts can be synthesized.