The present invention discloses a full-link text-to-video generation method and
system based on a multimodal
large model, which belongs to the field of
artificial intelligence content generation technology. Through the collaborative work of multiple intelligent agents,
user input text is analyzed, a cross-
modal memory
library is constructed, and the unified video and audio of the generated
storyboard is ensured based on the memory
library content, thereby realizing the full-process automatic generation from text to video. The implementation of this method includes the following steps: obtaining user text input; text analysis, through collaborative agents, dynamically extracting, analyzing, generating, associating, and storing multimodal information of images, texts, and audios from the input text to construct a multimodal memory
library; generating storyboards, generating
storyboard videos and audios according to the memory library; audio and video synthesis, and forming the final video after synchronous alignment of audio and video. The present invention can achieve narrative coherence in long video generation, improve the feature consistency of storyboards, enhance the consistency of cross-
modal emotions, reduce manual intervention, and improve the efficiency of
video production.