The invention discloses a multi-
modal natural language understanding and generating
system and method. The method comprises the following steps: constructing a cross-
modal pre-training module, training a multi-
modal encoder, and establishing a cross-modal
association mapping space; mixing prompt
fine tuning is carried out, and a complete blank filling template is constructed; according to the intention reasoning network, extracting multi-round dialogue intention representation of the user, and retrieving an external
knowledge base for fine-grained reasoning; constructing a unified
semantic representation framework, embedding the text, the image and the voice into a unified space, and generating a query vector of multi-modal intention
perception; and the knowledge query module based on key value memory generates entity-level multi-modal replies and optimizes the semantic comprehension and generation capability of the dialogue model. According to the method, the multi-modal information understanding and
generating capacity is improved, deep association and understanding of image and text information are achieved, downstream task adaptability is enhanced,
task completion accuracy and efficiency are improved, unified
semantic representation of the multi-modal information is achieved, and support is provided for
information retrieval and utilization.