The invention belongs to the technical field of
artificial intelligence and
office automation, and particularly relates to a PPT generated picture
insertion method based on multi-
modal text and graph retrieval, which comprises the following steps of: structurally analyzing a PPT page text by fusing
artificial intelligence and a multi-
modal representation learning technology, extracting multi-dimensional information such as themes, abstracts, keywords, semantic hierarchies, contexts and the like; constructing a precise semantic expression model; a text and an image are mapped to a unified
semantic space through a multi-
modal depth coding model, cross-page
semantic enhancement features are obtained by means of an attention mechanism, and image-text
semantic consistency modeling and evaluation are completed. On the basis, factors such as page structure
layout, text density, image proportion, style matching and the like are synthesized, and picture display position and size suggestions are generated by adopting a content driving strategy, so that dual optimization of semantic adaptability and visual coordination of image
insertion is realized; the problems that in a current automatic PPT generation
system, picture selection is not accurate, image-text
semantics are not matched, and page
layout is not coordinated are solved.