3D Object Localization with Virtual Cameras for Distance Accuracy

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional methods for estimating three-dimensional positions of objects suffer from reduced accuracy, particularly in determining the distance to an object, and existing gaze estimation systems lack precision in object positioning.

Innovation Solution

A method involving the use of a physical camera to capture an image, detect target objects, generate virtual cameras based on object coordinates, and apply a trained model to estimate three-dimensional positions using probability distributions, constrained by predetermined parameters and calibration settings to enhance accuracy.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If conventional methods are used to estimate three-dimensional positions from images, then the process is simple, but the accuracy of distance estimation is reduced

Engineering Contradiction:
Improvedistance estimation accuracyVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent introduces virtual cameras as intermediary computational models between the physical camera and the target object. These virtual cameras simulate different viewing angles and positions, allowing the system to estimate three-dimensional position with higher accuracy by combining multiple virtual perspectives without requiring multiple physical cameras

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The system changes parameters by creating virtual cameras with varying positions, orientations, and focal lengths. By adjusting these camera parameters computationally, the system can estimate depth and distance more accurately from a single physical image, transforming a 2D image problem into a multi-perspective 3D estimation problem

Inventive Principle:
Principle #35Parameter changes

2Measurement precision

If multiple physical cameras are used to improve three-dimensional position accuracy, then measurement precision improves, but device complexity and cost increase

Engineering Contradiction:
Improvethree-dimensional position accuracyVSAvoidcamera system complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent creates virtual copies of cameras (virtual cameras) that simulate multiple viewing perspectives. Instead of deploying multiple physical cameras, the system generates computational copies with different positions and orientations, processing a single physical image through multiple virtual camera models to achieve three-dimensional position estimation with the accuracy of multi-camera systems

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The system replaces the mechanical approach of using multiple physical cameras with a computational approach using virtual cameras. The physical camera captures a single image, and then computational algorithms generate multiple virtual views and estimate three-dimensional position, substituting mechanical complexity with software-based processing

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Data Source

PatentEP4287123B1Method of estimating a three-dimensional position of an object
Publication Date: 2025.11.12 TOBII TECH AB
  • EP4287123B1 patent drawingFigure 1
  • EP4287123B1 patent drawingFigure 2
  • EP4287123B1 patent drawingFigure 3

AI summary

The disclosure relates to a method (300) of estimating a three-dimensional position (150) of at least one target object (111) performed by a computer, the method comprising obtaining a first image (130), detecting at least one target object (111), generating a set of images depicting subsets of a captured scene, and estimating a three-dimensional position of at least one target object using a probability distribution (P) indicative of a likelihood that the at least one target object is in a particular position (x, y, z) in the scene.