The invention relates to the technical field of
speech synthesis, can be applied to business scenes of
medical health, financial science and technology, debate and the like, and discloses a
speech generation method and device based on speech style
adaptation, equipment and a medium, and the method comprises the steps: obtaining a
target text, a target speaker speech and a reference style speech, extracting a phoneme feature sequence, acoustic features and
rhythm coding information, and performing
feature fusion to generate a fusion coding vector; and
processing the fusion coding vector through a style
adaptation module, generating an acoustic code containing a target style feature, generating a Mel-
frequency spectrum feature based on the acoustic code, and inputting the Mel-
frequency spectrum feature into a pre-training vocoder to generate a voice waveform. The voice style matching and
rhythm control capability is improved through multi-
feature fusion, the generalization capability of unseen speakers is enhanced by optimizing the style
adaptation module, high-naturalness voice is generated in combination with the Mel spectrum and the vocoder, and the voice synthesis definition and adaptability are improved.