The invention discloses a real-time voice-driven
human body posture generation method based on a variational auto-
encoder, and belongs to the technical field of
digital human interaction, and the method comprises the steps: S1, carrying out the preprocessing and alignment of multi-
modal data, and constructing a three-dimensional training sample comprising audio features, emotion labels and posture data; s2, constructing a cross-
modal generation model based on a variational auto-
encoder, wherein the cross-
modal generation model comprises an
encoder and a decoder; s3, designing a multi-objective
loss function to
train and optimize the cross-modal generation model; and S4, inputting the preprocessed audio to be processed into the trained cross-modal generation model, and obtaining a
human body posture sequence with a coherent
time sequence through real-time reasoning. A variational auto-encoder is applied to a voice-
human body posture cross-modal generation scene, the traditional application boundary of the variational auto-encoder is broken, the complex mapping relation between audio features and human body postures is encoded into low-dimensional probability distribution through the probability modeling capability of a
potential space of the variational auto-encoder, and a new technical path is provided for posture generation.