Shot-Processing Device Pose Estimation Ellipse-Ellipsoid Matching

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current visual positioning and odometry methods face challenges in robustness, especially in large environments with repeated patterns and varying illumination, and require additional sensors or markers, limiting their adaptability and applicability in confined spaces.

Innovation Solution

A photo processing device and method that uses three-dimensional object doublets and ellipsoids to determine pose without external sensors, by detecting two-dimensional object doublets in images and calculating a pose through ellipse-ellipsoid correspondences, allowing for robust localization in a common reference frame.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If visual positioning uses local descriptors (e.g., SIFT) for 2D-3D point matching, then the method can operate without predefined markers, but it fails in large environments with repeating patterns and is sensitive to illumination changes and viewpoint variations

Engineering Contradiction:
Improveability to operate without predefined markersVSAvoidrobustness to illumination changes, viewpoint changes, and repeating patterns
Core Design Contradiction:
Adaptability or versatilityVSReliability

Solution Approach 1:

The patent transitions from 2D image space to 3D object space by using 3D object models with known geometries. Instead of matching 2D descriptors directly, the system projects 3D object points onto the 2D image plane and matches these projections with detected 2D points, adding the dimension of known 3D object structure to the matching process.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Solution Approach 2:

The system performs preliminary action by pre-defining 3D object models with known geometries and control points before the actual positioning task. These pre-established 3D models serve as reference frameworks that guide the matching process, eliminating the need for real-time descriptor calculation and comparison in uncertain environments.

Inventive Principle:
Principle #10Preliminary action

2Reliability

If visual positioning uses CNN-based learning techniques for direct inference of control point projections, then robustness to viewpoint and illumination changes improves, but the system requires large numbers of training images for each object and is difficult to adapt to new environments

Engineering Contradiction:
Improverobustness to viewpoint and illumination changesVSAvoidtraining data requirements and adaptability to new environments
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent uses copying by projecting known 3D object model points onto the 2D image plane to generate synthetic control point projections. Instead of learning from大量 training images, the system copies the geometric relationships from the 3D model and applies them to the current image through projection mathematics, eliminating the need for extensive training data.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The system achieves universality by using a single 3D object model that can be applied to multiple instances and viewpoints. The same 3D model serves as the reference for all matching operations regardless of illumination conditions, camera angles, or specific object instances, making the system universally applicable without retraining.

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Reliability

If visual positioning uses marker detection with known geometry, then robustness to lighting and viewpoint changes improves and repeating pattern problems are avoided, but markers must be positioned in the environment and their relative positions must be known, making the system cumbersome to implement

Engineering Contradiction:
Improverobustness to lighting, viewpoint changes, and repeating patternsVSAvoidease of implementation without environmental modifications
Core Design Contradiction:
ReliabilityVSEase of operation

Solution Approach 1:

The system applies self-service by using the scene's own 3D objects as positioning references. Instead of requiring external markers to be placed in the environment, the system utilizes the objects already present in the scene, whose 3D geometries are either known from models or can be reconstructed, making the positioning system self-sufficient without environmental modifications.

Inventive Principle:
Principle #25Self-service

4Adaptability or versatility

If visual odometry reconstructs a 3D point cloud on the fly from 2D points matched between consecutive images, then no predefined map is required, but the reconstructed map is initially devoid of useful external information for augmented reality or navigation assistance

Engineering Contradiction:
Improveability to operate without predefined mapsVSAvoidlack of external information for augmented reality and navigation
Core Design Contradiction:
Adaptability or versatilityVSLoss of information

Solution Approach 1:

The system performs preliminary action by incorporating or associating external information about 3D objects (such as semantic labels, models, or pre-acquired geometric data) into the on-the-fly reconstruction process. This allows the dynamically built map to include not just geometric structure but also semantic information useful for augmented reality and navigation applications.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentEP3977411B1Shot-processing device
Publication Date: 2024.11.20 INRIA INSTITUT NATIONAL DE RECHERCHE EN INFORMATIQUE ET EN AUTOMATIQUE
  • EP3977411B1 patent drawingFigure 1
  • EP3977411B1 patent drawing
  • EP3977411B1 patent drawing

AI summary

The invention relates to a shot-processing device which comprises a memory (10), a detector (20), a preparer (30), a combiner (40), an estimator (50) and a selector (60). The memory (10) is arranged to receive, on the one hand, scene data (12) comprising three-dimensional object pairs each associating an object identifier, and ellipsoid data defining an ellipsoid and its orientation and a position of its centre in a common frame of reference and, on the other hand, shot data (14) defining a two-dimensional image of the scene associated with the scene data (12), from a viewpoint corresponding to a desired pose. The detector (20) is arranged to receive shot data (14) and to return one or more two-dimensional object pairs (22) each comprising an object identifier present in the scene data, and a shot region associated with this object identifier. The preparer (30) is arranged to determine, for at least some of the two-dimensional object pairs (22) from the detector (20), a set of positioning elements (32) whose number is less than or equal to the number of three-dimensional object pairs in the scene data (12) that comprise the object identifier of the two-dimensional object pair (22) in question, each positioning element (32) associating the object identifier and the ellipsoid data of a three-dimensional object pair comprising this object identifier, and ellipse data which define an ellipse approximating the shot region of the two-dimensional object pair in question and its orientation as well as a position of its centre in the two-dimensional image. The combiner (40) is arranged to generate a list of candidates (42) each associating one or more positioning elements (32) and a shot orientation, and/or the combination of at least two positioning elements (32), the positioning elements (32) of a single candidate (42) being taken from separate two-dimensional object pairs (22) and not relating to the same three-dimensional object pair. The estimator (50) is arranged to calculate, for at least one of the candidates, a pose (52) comprising a position and an orientation in the common frame of reference from the ellipse data and the ellipsoid data of the positioning elements, or from the ellipse data and the ellipsoid data of the one or more positioning elements and the shot orientation. The selector (60) is arranged, for at least some of the poses, to project all of the ellipsoid data of the scene data onto the shot data from the pose, to determine a measurement of similarity between each projection of the ellipsoid data and each ellipse defined by a two-dimensional object pair (22) coming from the detector (20), and to calculate a likelihood value by using, for each projection of the ellipsoid data, the largest measurement of similarity determined, and to select the pose that has the highest likelihood value.