The invention relates to the technical field of voice
semantics, can be applied to business scenes of financial science and technology,
medical health and the like, and discloses a voice synthesis method, device, equipment and medium based on acoustic
branch weight, which comprises the following steps: acquiring a voice prompt and a
target text, extracting acoustic embedding of the voice prompt and text embedding of the
target text; the method comprises the steps of analyzing acoustic features of voice prompts to determine acoustic
branch weights, adjusting acoustic embedding according to the acoustic
branch weights to obtain weighted acoustic embedding, fusing the weighted acoustic embedding and text embedding to generate fused features, inputting the fused features into a
language model to generate a voice marking sequence, and decoding the voice marking sequence into a target voice waveform. According to the method, acoustic embedding and text embedding can adaptively reflect the
rhythm, tone and
semantic information of the voice through acoustic feature analysis and a weighted
fusion mechanism, so that the naturalness and expressive power of the voice are improved while the
semantic consistency is kept, and the method is suitable for high-fidelity voice generation of dialects and complex acoustic scenes.