The invention discloses a text-driven
digital human audio and video generation method, relates to the technical field of intelligent virtual
digital human generation, and provides the following scheme: obtaining a driving text and a reference picture, preprocessing the driving text to generate a phoneme sequence containing pronunciation
time sequence information, and sending the phoneme sequence to the reference picture; and extracting key face feature points in the reference picture through a face feature point
positioning technology, extracting tone reference feature vectors associated with tones in the reference picture through a
deep learning model, and inputting the phoneme sequence into a text-to-
speech model to generate an original audio. According to the method, the
timbre feature vectors and the face feature points are synchronously extracted from the single reference picture, and a multi-
modal feature cross fusion technology is combined, so that the problem of
timbre and visual attribute splitting in a traditional method is solved, the unified
digital human identity of personalized
timbre and visual image is driven by only one picture, extra sound recording is not needed, and the user experience is improved. The use threshold is reduced.