Multihead Deep Learning for 3D Object Detection

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current systems require multiple sensors and complex processing to generate a 3D model of a vehicle's environment, which increases processing power and time, and may not accurately identify objects without multiple inputs.

Innovation Solution

A method using a monocular camera to capture 2D images and process them through multi-head deep learning to differentiate between objects and traversable space, reducing the need for multiple sensors and enhancing object detection accuracy by assigning values for heading, depth, and motion, thereby generating a 3D model for driver assistance features.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If multiple sensors and complex processing systems are used to generate 3D models, then measurement precision and reliability improve, but device complexity and processing power requirements increase

Engineering Contradiction:
Improveobject detection accuracyVSAvoidsensor network complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent combines multiple detection functions (object detection, traversable space identification, depth estimation, and 3D model generation) into a single integrated deep learning model that processes monocular camera images, eliminating the need for separate sensors and processing systems for each function

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The deep learning model performs multiple functions simultaneously including detecting objects, identifying traversable space, estimating depth, and generating 3D models from a single monocular camera input, making the system multi-functional without requiring additional sensors

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Reliability

If multiple sensors and complex processing are used, then object detection reliability improves, but processing time and computational power increase

Engineering Contradiction:
Improveobject detection reliabilityVSAvoidprocessing time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent merges sequential processing steps (object detection, space identification, depth estimation) into a single parallel deep learning inference process that operates on monocular images, significantly reducing total processing time while maintaining reliability

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The system performs preliminary learning and feature extraction during the training phase, enabling the model to make accurate predictions with minimal computational effort during real-time operation, thus reducing processing time without sacrificing reliability

Inventive Principle:
Principle #10Preliminary action

3Measurement precision

If multiple sensors are deployed to capture environmental data, then measurement precision improves, but use of energy and device complexity increase

Engineering Contradiction:
Improveenvironmental characterization accuracyVSAvoidsensor network energy consumption
Core Design Contradiction:
Measurement precisionVSUse of energy by moving object

Solution Approach 1:

The patent extracts depth and 3D spatial information from 2D monocular images through deep learning processing, eliminating the need for energy-consuming additional sensors like stereocameras or LIDAR while maintaining measurement precision

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The system replaces physical sensor arrays (mechanical/optical systems) with a computational approach using deep learning models that process images algorithmically, reducing energy consumption associated with multiple physical sensors

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Data Source

PatentUS20240331288A1Multihead deep learning model for objects in 3D space
Publication Date: 2024.10.03 RIVIAN HOLDINGS LLC
  • US20240331288A1 patent drawing
  • US20240331288A1 patent drawing
  • US20240331288A1 patent drawing

AI summary

Systems and methods are presented herein for generating a three-dimensional model based on data from one or more two-dimensional images to identify a traversable space for a vehicle and objects surrounding the vehicle. A bounding area is generated around an object identified in a two-dimensional image captured by one or more sensors of a vehicle. Semantic segmentation of the two-dimensional image is performed based on the bounding area to differentiate between the object and a traversable space. The three-dimensional model of an environment comprised of the object and the traversable space is generated based on the semantic segmentation. The three-dimensional model is used for one or more of processing or transmitting instructions useable by one or more driver assistance features of the vehicle.