The invention belongs to the technical field of
natural language processing and
computer vision crossing, and discloses a multi-
modal topic modeling method based on
semantic consistency driving, which comprises the following steps: step 1, constructing a text-image multi-
modal corpus and preprocessing; step 2, obtaining vector representation; step 3, creating a theme
inference network; and step 4, training the topic
inference network by taking the Dirichlet prior loss of the comparative learning loss, the multi-
modal topic alignment loss and the
energy loss on the multi-modal topic distribution as a joint optimization target, and completing cross-modal topic alignment and joint topic modeling. According to the method, the shared combined topic space is constructed, the topic distribution centrality of the same semantic samples is improved, the topic distribution overlapping degree of different semantic samples is reduced, it is ensured that the topic distribution of the text and the topic distribution of the image keep the corresponding relation statistically, and combined modeling and
semantic association of the text and the image in the unified topic space are achieved.