Monocular Vehicle Detection via Virtual Point Cloud Generation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current vehicle information detection in road and driving scenes relies heavily on laser radar or millimeter wave radar for point cloud data, which is costly and may not provide accurate detection results, especially when using monocular images.

Innovation Solution

A vehicle information detection method and apparatus that performs multiple stages of target detection operations based on images, including a first target detection, error detection, and a second target detection, using models like Densely Connected Networks and lightweight error detection models to improve precision and robustness, reducing reliance on radar data.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If laser radar or millimeter wave radar is used to detect point cloud data for vehicle information, then detection accuracy is improved, but detection cost increases

Engineering Contradiction:
Improvedetection accuracyVSAvoiddetection cost
Core Design Contradiction:
Measurement precisionVSQuantity of substance

Solution Approach 1:

The patent creates a virtual point cloud copy from monocular image data through depth estimation algorithms, replacing the need for physical radar point cloud data. This virtual copy contains sufficient 3D spatial information for vehicle detection while being generated computationally from inexpensive 2D images, thus achieving accurate detection without expensive radar hardware

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent replaces the mechanical radar detection system with a computational vision system. Instead of using physical radar waves to capture point cloud data, the system uses monocular images processed through deep learning models and depth estimation algorithms to generate virtual 3D point cloud representations, substituting mechanical detection with information processing

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Quantity of substance

If monocular images are used for vehicle detection, then detection cost is reduced, but detection precision deteriorates

Engineering Contradiction:
Improvedetection costVSAvoiddetection precision
Core Design Contradiction:
Quantity of substanceVSMeasurement precision

Solution Approach 1:

The patent transforms 2D monocular image data into 3D spatial information through depth estimation and virtual point cloud generation. By adding the depth dimension computationally, the system recovers 3D geometric characteristics from 2D images, enabling accurate vehicle detection without requiring 3D sensing hardware

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Solution Approach 2:

The patent performs preliminary depth estimation and virtual point cloud generation from monocular images before the actual vehicle detection process. This preprocessing step creates enriched 3D spatial data structures that contain depth and distance information, which are then used by the detection model to achieve high precision results

Inventive Principle:
Principle #10Preliminary action

3Productivity

If single-stage target detection is performed on monocular images, then detection speed is improved, but detection precision deteriorates

Engineering Contradiction:
Improvedetection speedVSAvoiddetection precision
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The patent divides the detection process into two distinct stages: first generating virtual point cloud data from monocular images through depth estimation, then performing vehicle detection on the generated 3D data. This segmentation allows each stage to be optimized independently, maintaining speed while improving precision through the addition of depth information

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces virtual point cloud data as an intermediary between monocular images and the vehicle detection model. This intermediate representation contains both 2D image information and estimated 3D spatial characteristics, serving as a bridge that enriches the input data without requiring direct radar hardware

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS11867801B2Vehicle information detection method, method for training detection model, electronic device and storage medium
Publication Date: 2024.01.09 BEIJING BAIDU NETCOM SCI & TECH CO LTD
  • US11867801B2 patent drawing
  • US11867801B2 patent drawing
  • US11867801B2 patent drawing

AI summary

A vehicle information detection method, a method for training a detection model, an electronic device and a storage medium are provided, and relates to the technical field of artificial intelligence, in particular to the technical field of computer vision and deep learning. The method includes: performing a first target detection operation based on an image of a target vehicle, to obtain a first detection result for target information of the target vehicle; performing an error detection operation based on the first detection result, to obtain error information; and performing a second target detection operation based on the first detection result and the error information, to obtain a second detection result for the target information.