The invention provides a semantic aerial view visual
relocation method and device in a non-exposed scene,
electronic equipment, a storage medium and a
computer program product. The method comprises the following steps: acquiring a multi-view
image sequence under a non-exposed scene (such as a tunnel, an underground
pipe gallery or an underground
parking lot); semantic recognition is carried out based on a pre-trained semantic target detection model, and
spatial consistency semantic features are extracted through a semantic-geometric dual-channel
fusion mechanism combining a semantic
mask and geometric constraints; the method comprises the following steps of: realizing three-dimensional reconstruction by using a
voxel micro-renderable modeling method (VGGT), and generating a dense three-dimensional semantic
point cloud fusing
semantics and a geometric structure; two-dimensional
semantics are mapped to a three-dimensional space through a projection and
back projection relation, and
point cloud semantics are endowed; main structure planes such as the ground, the
left wall surface and the right wall surface are extracted, and a two-dimensional semantic aerial view with
semantic annotation is generated; and
pose estimation is carried out based on a reciprocal matching strategy guided by a semantic
mask, so that visual repositioning with high precision, high robustness and semantic
interpretability is realized. The method breaks through the problems of low precision, sparse features and poor
semantic consistency of traditional visual repositioning in a non-exposed environment, and can be widely applied to the fields of intelligent transportation, underground inspection and unmanned
system positioning.