The application provides a video intelligent synthesis method for multi-
modal content conversion, comprising: acquiring video demand data, performing semantic analysis on the video demand data, generating a video intelligent synthesis
data set, performing multi-path material retrieval based on the video intelligent synthesis
data set,
pruning in combination with a preset dynamic
pruning strategy, generating a video synthesis material candidate set and a
key frame demand set, generating an intelligent synthesis video segment based on the
key frame demand set, performing quality screening on the intelligent synthesis video segment, generating a candidate video synthesis segment set, performing video intelligent synthesis based on
beam search guidance according to the candidate video synthesis segment set and the video synthesis material candidate set, generating a candidate intelligent synthesis video set, performing quality evaluation on the candidate intelligent synthesis video set based on a preset
video quality evaluation mechanism, and performing secondary synthesis according to the
evaluation result to generate an optimal synthesis video, thereby improving the efficiency, accuracy and reliability of automatic
video production.