3D Vehicle Box Correction Using Camera Frame Detection

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing automated driving technologies face challenges in accurately acquiring the location of surrounding vehicles using three-dimensional detection information, as this information is often treated as two-dimensional, leading to potential misidentification and false detection of vehicle positions.

Innovation Solution

An image processing apparatus and method that captures images of surroundings, detects vehicle frames, generates three-dimensional boxes based on three-dimensional detection information, and corrects these boxes using frame recognition to accurately determine vehicle locations, thereby reducing false detections and improving positional accuracy.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Device complexity

If three-dimensional detection information is treated as two-dimensional information for combining with image data, then the processing complexity is reduced, but the location accuracy of surrounding vehicles deteriorates

Engineering Contradiction:
Improveprocessing complexityVSAvoidlocation accuracy
Core Design Contradiction:
Device complexityVSMeasurement precision

Solution Approach 1:

The patent applies dimensionality change by converting two-dimensional image frame data into three-dimensional box representations. The frame detector identifies vehicle regions in 2D images, and the three-dimensional box generator creates corresponding 3D boxes with depth information. This allows the system to maintain accurate 3D location information while still utilizing 2D image data, resolving the contradiction between processing simplicity and location accuracy.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Measurement precision

If three-dimensional detection information is used to generate three-dimensional boxes, then the location accuracy of surrounding vehicles is improved, but the possibility of false detection and ghost phenomena increases

Engineering Contradiction:
Improvelocation accuracyVSAvoiddetection reliability
Core Design Contradiction:
Measurement precisionVSReliability

Solution Approach 1:

The patent merges two different detection approaches: frame detection from 2D images and three-dimensional detection information from sensors like LiDAR or radar. The three-dimensional box correction section integrates these two sources, using the frame detector's results to validate and correct the three-dimensional boxes. This combination reduces false detections and ghost phenomena while maintaining high location accuracy, as the system cross-validates data from multiple detection modalities.

Inventive Principle:
Principle #5Merging (Combining)

3Device complexity

If three-dimensional detection information is used without correction, then the detection process is simplified, but the positional accuracy of vehicle locations deteriorates

Engineering Contradiction:
Improvedetection process complexityVSAvoidpositional accuracy
Core Design Contradiction:
Device complexityVSMeasurement precision

Solution Approach 1:

The patent implements a feedback mechanism where the frame detector provides correction information to adjust the three-dimensional boxes generated from three-dimensional detection data. The three-dimensional box correction section uses the frame detection results as feedback to refine the positional accuracy of the 3D boxes. This feedback loop maintains detection process simplicity while significantly improving positional accuracy through iterative correction.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS11978261B2Information processing apparatus and information processing method
Publication Date: 2024.05.07 SONY SEMICON SOLUTIONS CORP
  • US11978261B2 patent drawing
  • US11978261B2 patent drawing
  • US11978261B2 patent drawing

AI summary

An information processing apparatus and an information processing method to properly acquire a location of a surrounding vehicle using three-dimensional detection information regarding an object around an own vehicle. A camera captures an image of surroundings of an own automobile, and a region of a vehicle in the captured image is detected as a frame, the vehicle being in the surroundings of the own automobile. Three-dimensional information regarding an object in the surroundings of the own automobile is detected, and a three-dimensional box that indicates a location of the vehicle in the surroundings of the own automobile is generated on the basis of the three-dimensional information. Correction is performed on the three-dimensional box on the basis of the frame, and the three-dimensional box is arranged to generate surrounding information.