Tight 2D Bounding Box Generation for Autonomous Driving Perception

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional motion planning and control for autonomous vehicles do not accurately consider differences in vehicle types, and existing synthetic datasets for training perception modules lack tight 2D bounding boxes that cover both visible and occluded or truncated parts of objects, making them inadequate for effective autonomous driving.

Innovation Solution

A method for generating tight 2D bounding boxes in a 3D scene by rendering a 2D segmentation image, identifying visible objects, and creating amodal segmentation images to generate accurate 2D bounding boxes that cover all parts of visible objects, even if occluded, which can be used to train the perception module of autonomous vehicles.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If 2D bounding boxes are generated from 3D bounding boxes in existing synthetic datasets, then bounding boxes are provided for all objects including occluded parts, but the bounding boxes are bigger than the objects themselves and not tight

Engineering Contradiction:
Improvebounding box accuracyVSAvoidbounding box tightness
Core Design Contradiction:
Measurement precisionVSManufacturing precision

Solution Approach 1:

The patent segments the 3D scene into individual object models with their respective 3D bounding boxes, then projects each object's 3D bounding box separately onto the 2D image plane. This segmentation allows for precise calculation of each object's 2D bounding box dimensions based on its specific orientation and position, rather than using a uniform approach that results in oversized bounding boxes.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent transforms 3D bounding box information into 2D bounding box information through projection mathematics. By using the 3D object models and their bounding boxes in three-dimensional space, the system calculates the corresponding two-dimensional bounding boxes on the image plane, taking into account camera parameters, object positions, and orientations to achieve tight fitting.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Measurement precision

If manual labeling is used for training perception modules, then accurate labels can be obtained, but the process is both time-consuming and costly

Engineering Contradiction:
Improvelabeling accuracyVSAvoidlabeling efficiency
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The patent uses synthetic 3D object models as copies or representations of real objects to generate training data. By rendering these digital models with known ground truth information (positions, orientations, bounding boxes), the system creates labeled training images without requiring manual annotation, thus maintaining accuracy while dramatically improving productivity.

Inventive Principle:
Principle #26Copying

3Ease of manufacture

If 2D bounding boxes are generated only for visible pixels, then the labeling process is simpler, but occluded or truncated parts of objects are not covered

Engineering Contradiction:
Improvelabeling simplicityVSAvoidobject detection completeness
Core Design Contradiction:
Ease of manufactureVSReliability

Solution Approach 1:

The patent performs preliminary action by using 3D object models that inherently contain information about the complete object geometry and dimensions. Before generating 2D bounding boxes, the system has already established the full 3D extent of each object, allowing it to calculate 2D bounding boxes that encompass all parts of the object including occluded portions, based on the projected 3D boundaries.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS11699235B2Way to generate tight 2D bounding boxes for autonomous driving labeling
Publication Date: 2023.07.11 BAIDU USA LLC
  • US11699235B2 patent drawing
  • US11699235B2 patent drawing
  • US11699235B2 patent drawing

AI summary

A method, apparatus, and system for generating tight two-dimensional (2D) bounding boxes for visible objects in a three-dimensional (3D) scene is disclosed. A two-dimensional (2D) segmentation image of a three-dimensional (3D) scene comprising one or more objects is generated by rendering the 3D scene with a segmentation camera. Each of the objects is rendered in a single respective different color. Next, one or more visible objects in the 3D scene are identified among the one or more objects based on the segmentation image. Next, a 2D amodal segmentation image for each of the visible objects in the 3D scene is generated separately. Each amodal segmentation image comprises only the single visible object for which it is generated. Thereafter, a 2D bounding box is generated for each of the visible objects in the 3D scene based on the amodal segmentation image for the visible object.