The present application relates to the technical field of multi-
modal abstract generation, in particular to a multi-
modal abstract generation method and
system based on keyword prediction enhancement, comprising the following steps: 1) obtaining multi-
modal abstract data, wherein the multi-modal abstract data comprises text, images corresponding to the text, and text keywords, the text keywords are obtained by calculating entities in the intersection of the original text and the reference abstract, and provide reliable training data for subsequent visual key
information extraction; the present application effectively uses a small model as an auxiliary model to extract keywords to assist a
large model in performing a multi-modal abstract generation task, more effectively guides the
large model to generate a multi-modal abstract by using this framework, can lock core information in the text, thereby obtaining more robust and stronger fact consistency results, effectively solves problems such as focus shift, visual redundancy, and
semantic gap between different modalities, and improves the quality of abstract generation.