The present application belongs to the field of multi-
modal named entity recognition, and particularly relates to a
power equipment multi-
modal named entity recognition method,
system, device and medium. The present application adopts a co-occurrence entity method to obtain entities appearing in both text and images, which can ensure text and
image matching and reduce
image noise. By extracting entity triples and then adopting an entity auxiliary method,
semantic information can be supplemented and semantic disambiguation can be reduced. The present application adopts the co-occurrence entity and entity auxiliary method of extracting triples, which helps to improve the multi-
modal named entity recognition accuracy, accurately recognize the entities existing in the image or text, and especially performs better in data with higher image-text correlation. Meanwhile, the present application is not limited to the field of
power equipment, and can be constructed into a
knowledge base of other fields to expand to other fields for multi-modal
named entity recognition.