The present invention provides a multi-
modal portrait
video editing method, which can be applied to the technical field of
video editing. The method comprises: given a portrait video, preprocessing same to obtain camera parameters, a human identity coefficient, a human expression coefficient, a human
pose coefficient, a human semantic segmentation map, and a two-dimensional portrait
mask; on the basis of a neural
Gaussian texture mechanism, embedding a learnable three-dimensional
Gaussian feature into a parameterized human geometric surface, using a neural renderer to convert a three-dimensional
Gaussian splatting feature map into an image, and optimizing the reconstruction of a three-dimensional portrait on the basis of RGB and segmentation map information of the video; and using an iterative dataset update technique to distill knowledge of a multi-
modal two-dimensional
image generation model into three-dimensional portrait editing, and using expression similarity guidance and a face-aware portrait editing model to improve editing quality. The method elevates a two-dimensional editing task to three-dimensional space, and ensures good three-dimensional consistency and
temporal consistency. By means of the knowledge of the multi-
modal generation model, high-quality portrait
video editing functions can be achieved.