The application provides a 4D person-object interaction generation method based on a text description, which comprises the following steps: in stage one, 3D person-object interaction
key frame recovery: a
human body motion sequence is obtained through a
human body motion model and uniformly down-sampled to extract a
key frame; a
human body grid is reconstructed through an SMPL-X model for each
key frame to extract a vertex position and form a human body
point cloud; an object position anchoring network takes the human body
point cloud, an object template
point cloud and a text prompt as input to predict an object position and generate a 3D person-object interaction key frame; in stage two, 4D person-object interaction sequence generation: a contact
perception diffusion model is constructed to take a sparse 3D person-object interaction key frame as input, extract a conditional
signal through a contact
perception encoder; based on the conditional
signal, the 3D person-object interaction key frame is subjected to
time series interpolation through the contact
perception diffusion model to generate a
time series coherent dense 4D person-object interaction sequence. The application realizes natural and realistic 4D person-object interaction synthesis of an unseen object.