The invention discloses a
feature fusion optimization-based multi-
modal map enhancement retrieval method and a dialogue
system. The method comprises the following steps of: respectively carrying out pre-training and
fine tuning on a visual model and a
language model by utilizing
a domain image and text data; constructing a
knowledge graph based on the text data in the
knowledge base and constructing a vector
database containing
associated image data; performing semantic analysis and optimization on the original query of the user by using the
language model and forming a structured retrieval intention; searching related sub-graphs, text
semantic vector information and
associated image data based on the search intention; encoding the sub-images into knowledge contexts, inputting the knowledge contexts into a dynamic prompt generator to generate visual prompts, and extracting enhanced visual features from the
associated image data through a visual model; and inputting the subgraph, the text
semantic vector information and the enhanced visual features into a
language model for collaborative reasoning, and generating and outputting a final answer. According to the method, deep fusion and accurate retrieval of multi-
modal knowledge can be realized, and the accuracy and efficiency are remarkably improved.