This application provides a
natural language-driven video generation method based on intent deconstruction, relating to the field of video generation technology, including the following steps: receiving a user's
natural language intent description, thereby obtaining a structured
event sequence containing multiple event nodes arranged in chronological order, each event node containing structured attribute labels for scene, subject, and behavior, and for logically related event node pairs, they are marked as at least one of temporal relationship or causal dependency relationship; based on the structured
event sequence and the
logical relationship between its nodes, generating a multi-condition
control signal group, integrating it with a video generation model, and driving the video generation model to generate a video
stream corresponding to the structured
event sequence; wherein, the first type of condition
signal is encoded as a text prompt embedding in the denoising process of the video generation model, and the second type of condition
signal is encoded as a spatiotemporal attention
mask or
motion dynamics prior vector acting on the latent feature space of the video generation model.