Embodiments of the present application provide a method and device for generating a dynamic
image based on audio, an apparatus, and a storage medium, relating to the field of natural human-computer interaction. The method comprises: first obtaining a
reference image and a reference audio input by a user; then, based on the
reference image and a trained generation
network model, determining a target head action feature and a
target expression coefficient feature, and adjusting the trained generation
network model based on the target head action feature and the
target expression coefficient feature to obtain a target generation
network model; finally, based on the reference audio, the
reference image, and the target generation network model,
processing a to-be-processed image to obtain a target dynamic image; wherein the to-be-processed image is the same as an
image object in the reference image; in this way, a corresponding digital person can be obtained based on a single picture of a target person; in this way, video acquisition work and data cleaning work are not required, the production cost of the digital person can be reduced, and the
production cycle of the digital person is shortened.