A visual positioning model training method and device based on hierarchical space refinement, equipment and medium

By employing a hierarchical spatial refinement method for visual positioning model training, the bounding box coordinates of the visual positioning training data are split into hundreds, tens, and units digit sequences. A geometric perception reward mechanism is introduced, which solves the problem of insufficient visual positioning accuracy in existing models, achieves pixel-level high-precision positioning, and improves the accuracy of multimodal understanding and human-computer interaction.

CN122435233APending Publication Date: 2026-07-21UNIV OF CHINESE ACAD OF SCI
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
UNIV OF CHINESE ACAD OF SCI
Filing Date
2026-04-14
Publication Date
2026-07-21

AI Technical Summary

Technical Problem

Existing computer vision models suffer from insufficient high-precision performance in visual localization tasks, especially when locating objects in images based on text descriptions, where high-precision bounding box localization is difficult to achieve.

Method used

A visual localization model training method based on hierarchical spatial refinement is adopted. The original coordinates of the target true bounding box in the visual localization training data are mapped to a set coordinate range and then split into a hierarchical sequence of hundreds, tens and units digits. An axis-decoupled hierarchical coordinate representation is established and the model is trained in combination with a verifiable geometric perception reward mechanism.

Benefits of technology

It significantly improves the model's localization accuracy in visual localization tasks, outperforming existing models with fewer training samples and achieving pixel-level high-precision alignment. It is suitable for multimodal understanding and human-computer interaction scenarios, and improves the response accuracy of robot operation, autonomous driving perception, and fine-grained image retrieval.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122435233A_ABST
    Figure CN122435233A_ABST
Patent Text Reader

Abstract

The present application relates to the field of computer vision, and discloses a visual positioning model training method and device based on hierarchical space refinement, equipment and medium, the original coordinates of the target real boundary box corresponding to the visual positioning training data are mapped to the set coordinate interval, the mapped coordinates of the target real boundary box are obtained, the mapped coordinates are split and arranged, and the real hierarchical coordinate sequence corresponding to the visual positioning training data is obtained. The visual positioning training data is input into the to-be-trained model for visual positioning, and a plurality of candidate hierarchical coordinate sequences corresponding to the visual positioning training data are obtained. Based on the difference between each candidate hierarchical coordinate sequence and the real hierarchical coordinate sequence, the to-be-trained model is updated to obtain a trained visual positioning model. The present application can effectively improve the high-precision performance of the model in the visual positioning task, and realize high-precision positioning of the target object in the image by the visual positioning model.
Need to check novelty before this filing date? Find Prior Art