Multi-Frame Visual Positioning Using Inertial Graph Constraints
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing visual positioning technologies based on machine vision face instability in positioning results, especially when using a single image frame or when GPS positioning is lost, leading to inaccurate and unreliable positioning for unmanned devices and smart devices.
Innovation Solution
A visual positioning method that acquires a video from an image sensor, determines visual positioning information for key image frames, establishes capture pose transformation relationships using inertial navigation data, and performs graph optimization using these relationships as edge constraints to achieve accurate positioning.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If visual positioning is performed using a single image frame, then the positioning process is simple and fast, but the positioning result is unstable and inaccurate
Solution Approach 1:
The video sequence is segmented into multiple key frames selected based on feature point quantity and image quality metrics. By processing only these key frames rather than every frame, the system achieves accurate positioning while controlling computational complexity through selective frame sampling.
Solution Approach 2:
The system transitions from single-frame positioning to multi-frame positioning by adding the time dimension. Multiple frames are processed together with temporal relationships established through graph optimization, improving positioning accuracy through temporal information integration.
2Measurement precision
If multiple image frames are processed for visual positioning, then positioning accuracy improves, but computational complexity and processing time increase
Solution Approach 1:
The video sequence is segmented into multiple key frames selected based on feature point quantity and image quality metrics. By processing only these key frames rather than every frame, the system achieves accurate positioning while controlling computational complexity through selective frame sampling.
Solution Approach 2:
Instead of processing all frames or a fixed number of frames, the system selectively processes only those frames that meet specific criteria (feature point threshold, image quality metrics). This partial action approach processes fewer frames than exhaustive methods while achieving better accuracy than single-frame methods.
3Reliability
If key frames are selected based on content repeatability and image quality, then positioning stability improves, but the selection process becomes more complex
Solution Approach 1:
The system evaluates frames using multiple quantifiable parameters including feature point quantity, image quality metrics, and content repeatability measures. By changing from simple frame selection to multi-parameter evaluation, the system achieves more stable positioning through objective, measurable criteria.
Solution Approach 2:
The frame selection process is self-regulating through automated evaluation of content repeatability and image quality. The system automatically identifies and selects key frames based on predefined thresholds without manual intervention, making the complexity manageable through algorithmic automation.
Data Source
AI summary
A visual positioning method and apparatus are provided. In some embodiments, the method includes: acquiring a video captured by an image sensor; determining visual positioning information respectively corresponding to a plurality of key image frames in the video; determining a capture pose transformation relationship between each of the plurality of key image frames according to inertial navigation information of the image sensor recorded when taking the video; performing, according to the visual positioning information corresponding to each of the plurality of key image frames, graph optimization processing on the visual positioning information corresponding to each of the plurality of key image frames by using the capture pose transformation relationship between each of the plurality of key image frames as an edge constraint; and determining, according to a result of the graph optimization processing, a visual positioning result of the image sensor when taking the video.


