This invention discloses an event-visual-inertial semantic simultaneous localization and map building method, which is applicable to navigation and localization of mobile robots, unmanned vehicles and drones. This method acquires
event data streams from event cameras and inertial data from IMUs, and optionally standard camera images. Based on the
event data, it constructs an event activity surface, performs coarse-to-fine event
corner detection on it, constructs an event representation for event
optical flow estimation, and estimates the event
optical flow. It then uses the event
optical flow to perform temporal tracking of event corners to obtain multi-time-stack related observations. Based on the event corners and event representations, it extracts descriptive information, performs
loop closure detection and relocalization, and introduces
loop closure constraints as additional residual terms into a sliding window
graph optimization. Within the sliding window, it combines residuals from events, images, IMUs,
edge detection, and
loop closure relocalization, and sets adaptive weights for multi-source fusion to obtain a six-degree-of-freedom
pose sequence. It performs target detection on the
event data stream, outputting detection boxes, categories, and confidence scores. Based on the event optical flow, it implements missed detection compensation and motion consistency checks to identify dynamic target regions, eliminates dynamic corners and their observation constraints, and generates an object-level
semantic map based on
pose estimation. This improves the robustness of localization mapping and environmental understanding in complex lighting, high-speed motion, and dynamic interference scenarios.