A visual positioning model training method and device based on hierarchical space refinement, equipment and medium
By employing a hierarchical spatial refinement method for visual positioning model training, the bounding box coordinates of the visual positioning training data are split into hundreds, tens, and units digit sequences. A geometric perception reward mechanism is introduced, which solves the problem of insufficient visual positioning accuracy in existing models, achieves pixel-level high-precision positioning, and improves the accuracy of multimodal understanding and human-computer interaction.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- UNIV OF CHINESE ACAD OF SCI
- Filing Date
- 2026-04-14
- Publication Date
- 2026-07-21
AI Technical Summary
Existing computer vision models suffer from insufficient high-precision performance in visual localization tasks, especially when locating objects in images based on text descriptions, where high-precision bounding box localization is difficult to achieve.
A visual localization model training method based on hierarchical spatial refinement is adopted. The original coordinates of the target true bounding box in the visual localization training data are mapped to a set coordinate range and then split into a hierarchical sequence of hundreds, tens and units digits. An axis-decoupled hierarchical coordinate representation is established and the model is trained in combination with a verifiable geometric perception reward mechanism.
It significantly improves the model's localization accuracy in visual localization tasks, outperforming existing models with fewer training samples and achieving pixel-level high-precision alignment. It is suitable for multimodal understanding and human-computer interaction scenarios, and improves the response accuracy of robot operation, autonomous driving perception, and fine-grained image retrieval.
Smart Images

Figure CN122435233A_ABST