Monocular 3D Perception for Accurate Object Counting

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing methods for accurate object counting in 3D scenes, such as people in a specific volume, require additional sensors like stereo cameras or LIDAR, which are costly and complex to calibrate, leading to inefficiencies and inaccuracies, especially when objects appear in mirrors or glass.

Innovation Solution

A method utilizing monocular 3D perception with a single RGB camera to generate reference and current depth maps, comparing depth changes to determine if objects are within a defined volume of interest, eliminating the need for additional sensors and complex calibration.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If additional sensors (stereo camera, LIDAR, ToF) are used for 3D scene understanding, then measurement precision of object location is improved, but device complexity and cost increase

Engineering Contradiction:
Improveobject location accuracyVSAvoidsensor system complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent creates a virtual 3D model (copy) of the environment using data from existing sensors, eliminating the need for additional depth sensors. The 3D model serves as a digital replica that enables accurate object location measurement without requiring physical stereo cameras or LIDAR systems.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent replaces mechanical/optical depth sensing systems (stereo cameras, LIDAR, ToF sensors) with a computational approach using monocular video frames and AI-based 3D reconstruction algorithms, substituting physical sensing mechanisms with information processing.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Measurement precision

If additional sensors and 3D modeling are used for accurate object counting, then object count accuracy is improved, but use of energy and computational resources increase

Engineering Contradiction:
Improveobject count accuracyVSAvoidcomputational energy consumption
Core Design Contradiction:
Measurement precisionVSUse of energy by moving object

Solution Approach 1:

The patent performs preliminary 3D environmental mapping and object detection when the environment is relatively static, storing this information for later use. This preliminary action reduces computational burden during actual object counting tasks, as the system can compare against pre-established 3D models rather than performing full 3D analysis in real-time.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent applies AI-based 3D perception selectively - using full 3D reconstruction only when necessary (e.g., when objects are occluded or depth information is ambiguous), and relying on simpler 2D analysis for routine counting tasks, thus avoiding excessive computational energy consumption.

Inventive Principle:
Principle #16Partial or excessive action

3Productivity

If traditional object counting methods are used, then processing speed is maintained, but measurement precision decreases due to false counts from mirrors and glass

Engineering Contradiction:
Improveobject counting speedVSAvoidobject count accuracy
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The patent transitions from 2D image analysis to 3D spatial reasoning by reconstructing the environment in three dimensions. This additional dimension enables the system to distinguish between real objects and reflections by analyzing depth relationships and spatial consistency, eliminating false counts from mirrors and glass surfaces.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Solution Approach 2:

The patent introduces a 3D environmental model as an intermediary between the raw video frames and the object counting decision. This intermediate representation provides contextual information about the scene geometry, helping to distinguish real objects from reflections without requiring additional sensors.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS20240281990A1Object count using monocular three-dimensional (3D) perception
Publication Date: 2024.08.22 QUALCOMM INC
  • US20240281990A1 patent drawing
  • US20240281990A1 patent drawing
  • US20240281990A1 patent drawing

AI summary

Systems and techniques are provided for performing an accurate object count using monocular three-dimensional (3D) perception. In some examples, a computing device can generate a reference depth map based on a reference frame depicting a volume of interest. The computing device can generate a current depth map based on a current frame depicting the volume of interest and one or more objects. The computing device can compare the current depth map to the reference depth map to determine a respective change in depth for each of the one or more objects. The computing device can further compare the respective change in depth for each object to a threshold. The computing device can determine whether each object is located within the volume of interest based on comparing the respective change in depth for each object to the threshold.