Shot-Processing Device Pose Estimation Ellipse-Ellipsoid Matching
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current visual positioning and odometry methods face challenges in robustness, especially in large environments with repeated patterns and varying illumination, and require additional sensors or markers, limiting their adaptability and applicability in confined spaces.
Innovation Solution
A photo processing device and method that uses three-dimensional object doublets and ellipsoids to determine pose without external sensors, by detecting two-dimensional object doublets in images and calculating a pose through ellipse-ellipsoid correspondences, allowing for robust localization in a common reference frame.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If visual positioning uses local descriptors (e.g., SIFT) for 2D-3D point matching, then the method can operate without predefined markers, but it fails in large environments with repeating patterns and is sensitive to illumination changes and viewpoint variations
Solution Approach 1:
The patent transitions from 2D image space to 3D object space by using 3D object models with known geometries. Instead of matching 2D descriptors directly, the system projects 3D object points onto the 2D image plane and matches these projections with detected 2D points, adding the dimension of known 3D object structure to the matching process.
Solution Approach 2:
The system performs preliminary action by pre-defining 3D object models with known geometries and control points before the actual positioning task. These pre-established 3D models serve as reference frameworks that guide the matching process, eliminating the need for real-time descriptor calculation and comparison in uncertain environments.
2Reliability
If visual positioning uses CNN-based learning techniques for direct inference of control point projections, then robustness to viewpoint and illumination changes improves, but the system requires large numbers of training images for each object and is difficult to adapt to new environments
Solution Approach 1:
The patent uses copying by projecting known 3D object model points onto the 2D image plane to generate synthetic control point projections. Instead of learning from大量 training images, the system copies the geometric relationships from the 3D model and applies them to the current image through projection mathematics, eliminating the need for extensive training data.
Solution Approach 2:
The system achieves universality by using a single 3D object model that can be applied to multiple instances and viewpoints. The same 3D model serves as the reference for all matching operations regardless of illumination conditions, camera angles, or specific object instances, making the system universally applicable without retraining.
3Reliability
If visual positioning uses marker detection with known geometry, then robustness to lighting and viewpoint changes improves and repeating pattern problems are avoided, but markers must be positioned in the environment and their relative positions must be known, making the system cumbersome to implement
Solution Approach 1:
The system applies self-service by using the scene's own 3D objects as positioning references. Instead of requiring external markers to be placed in the environment, the system utilizes the objects already present in the scene, whose 3D geometries are either known from models or can be reconstructed, making the positioning system self-sufficient without environmental modifications.
4Adaptability or versatility
If visual odometry reconstructs a 3D point cloud on the fly from 2D points matched between consecutive images, then no predefined map is required, but the reconstructed map is initially devoid of useful external information for augmented reality or navigation assistance
Solution Approach 1:
The system performs preliminary action by incorporating or associating external information about 3D objects (such as semantic labels, models, or pre-acquired geometric data) into the on-the-fly reconstruction process. This allows the dynamically built map to include not just geometric structure but also semantic information useful for augmented reality and navigation applications.
Data Source
Figure 1

AI summary
The invention relates to a shot-processing device which comprises a memory (10), a detector (20), a preparer (30), a combiner (40), an estimator (50) and a selector (60). The memory (10) is arranged to receive, on the one hand, scene data (12) comprising three-dimensional object pairs each associating an object identifier, and ellipsoid data defining an ellipsoid and its orientation and a position of its centre in a common frame of reference and, on the other hand, shot data (14) defining a two-dimensional image of the scene associated with the scene data (12), from a viewpoint corresponding to a desired pose. The detector (20) is arranged to receive shot data (14) and to return one or more two-dimensional object pairs (22) each comprising an object identifier present in the scene data, and a shot region associated with this object identifier. The preparer (30) is arranged to determine, for at least some of the two-dimensional object pairs (22) from the detector (20), a set of positioning elements (32) whose number is less than or equal to the number of three-dimensional object pairs in the scene data (12) that comprise the object identifier of the two-dimensional object pair (22) in question, each positioning element (32) associating the object identifier and the ellipsoid data of a three-dimensional object pair comprising this object identifier, and ellipse data which define an ellipse approximating the shot region of the two-dimensional object pair in question and its orientation as well as a position of its centre in the two-dimensional image. The combiner (40) is arranged to generate a list of candidates (42) each associating one or more positioning elements (32) and a shot orientation, and/or the combination of at least two positioning elements (32), the positioning elements (32) of a single candidate (42) being taken from separate two-dimensional object pairs (22) and not relating to the same three-dimensional object pair. The estimator (50) is arranged to calculate, for at least one of the candidates, a pose (52) comprising a position and an orientation in the common frame of reference from the ellipse data and the ellipsoid data of the positioning elements, or from the ellipse data and the ellipsoid data of the one or more positioning elements and the shot orientation. The selector (60) is arranged, for at least some of the poses, to project all of the ellipsoid data of the scene data onto the shot data from the pose, to determine a measurement of similarity between each projection of the ellipsoid data and each ellipse defined by a two-dimensional object pair (22) coming from the detector (20), and to calculate a likelihood value by using, for each projection of the ellipsoid data, the largest measurement of similarity determined, and to select the pose that has the highest likelihood value.