The invention relates to the technical field of
artificial intelligence, can be applied to business scenes of
medical health, financial science and technology, cultural transmission and the like, and discloses a voice generation method based on multi-
modal fusion, which comprises the following steps: collecting audio data to extract
timbre features, and training a field feature
timbre generation model; analyzing text semantic recognition emotion information, adjusting
speech synthesis parameters, combining personalized information to construct a parameter mapping table, fusing to generate a synthesis control parameter sequence, aligning the synthesis control parameter sequence with character labels, visual elements and background music data, driving
a domain feature
timbre generation model, and generating synthesis data of synchronous speech, text, vision and music. Domain timbres are generated through timbre feature training, speech expression is optimized by combining semantic analysis and
emotion recognition, user requirements are matched based on personalized information, and
time alignment is performed by fusing text, vision and music data, so that the synthesized speech has domain features, emotion adaptability and
personalization, and the
speech quality is improved. And the voice immersion and the
information transmission capability are improved.