This invention discloses a cross-view positioning method based on semantic coarse localization and BEV fine registration, aiming to solve the problem of large-scale, high-precision global positioning under conditions of no initial position prior. The invention employs a two-stage
processing architecture: First, hierarchical
semantic information is extracted from multi-view images from the vehicle using a visual
language model. Combined with heading constraints and a
backtracking retrieval strategy, candidate regions are quickly screened in a
standard map to achieve semantic coarse localization. Then, within the candidate regions, the vehicle-mounted BEV
perception results are finely geometrically registered with the map template, and a dynamic weight
fusion mechanism based on road and building responses is introduced. The fusion weights are adaptively adjusted according to the degree of support from building structures to complete the fine registration. This invention organically combines high-level
semantic information with low-level geometric information, achieving high-precision and robust cross-view
geolocation even in the absence of coarse position prior. It effectively solves the problems of easy
confusion in pure
geometric matching, low accuracy in pure semantic localization, and poor adaptability of fixed fusion methods, making it suitable for global positioning of unmanned systems in complex urban environments.