The application discloses a
digital human modeling method, comprising the following steps: S1, three-dimensional reconstruction data about each frame of image is extracted from
monocular video, the three-dimensional reconstruction data comprises
original data, high-precision human
mask, original
audio signal, optimized SMPL parameters and camera parameters corresponding to each frame of image; S2, a three-dimensional structure
human body grid in a "T" posture is generated based on the optimized SMPL parameters, after being discretized into a 3D
Gaussian point set, the initialization of the 3D
Gaussian point set is completed in combination with the
original data, the high-precision human
mask and the camera parameters corresponding to each frame of image; S3, a facial
animation sequence capable of representing
mouth shape and emotional dynamics is obtained based on the original
audio signal, a text description of an expected facial morphology and a
reference image; S4, a
Gaussian physical model is added in a 3DGS process to render the initialized 3D Gaussian
point set, and the 3DGS rendering result is mixed and rendered with the facial
animation sequence to generate a
digital human model with facial expressions.