The invention relates to a method and a device for generating a video with a body and
electronic equipment. The method comprises the following steps: analyzing a task instruction and initial environment
observation data, generating a key operation step sequence and an associated object set thereof, and constructing a three-dimensional physical
constraint graph; generating an initial action
video sequence based on the coding condition; in response to single-frame action execution, calculating a spatial error index, and if the object contact distance deviation, the trajectory
collision probability or the physical constraint violation value exceeds a preset threshold, triggering the space-time
diffusion model to generate a correction
frame based on the current physical
constraint graph; in response to the completion of the key operation steps, verifying the
semantic matching degree of the generated result and the task target, and if the key
object attribute loss or the operation
logic error is detected, regenerating the task operation step sequence and the physical
constraint graph; and updating the object position, the
distance threshold and the motion feasible region in the physical constraint graph, and inputting the space-time
diffusion model. According to the method, the problems of physical consistency deficiency, error accumulation and semantic understanding splitting are solved.