The invention discloses a method and a
system for identifying a multi-
modal named entity, and belongs to the technical field of
digital data processing. In order to solve the technical problems that in the prior art, when images and text information are processed, shared information and private information are sequentially connected in series, and the shared information and the private information of
visual objects in the images are directly connected in series, so that feature information
confusion is caused, fine-grained alignment in visual
modes is influenced, and cross-
modal understanding of a GMNER
system is influenced. The shared visual features and the private visual features of the image are extracted respectively, the features of the
visual objects in the image and the relation features between the
visual objects are distinguished, and then the images are dynamically integrated and projected to the text embedding space, so that the corresponding relation between the visual object entities and the text entities is clearer, and the text embedding efficiency is improved. And the accuracy of fine
granularity alignment is improved, so that the comprehensive cross-
modal understanding capability of the GMNER
system is improved. The method is mainly used for multi-modal
named entity recognition.