Bird's Eye View Feature Map Augmented with Semantic Information

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current computer vision systems for autonomous mobile agents face challenges in accurately detecting objects in environments, particularly when objects are occluded or have limited data points, as they often rely on either two-dimensional images with limited distance information or point cloud data with limited visual appearance information.

Innovation Solution

A system that uses a bird's eye view feature map augmented with semantic information, obtained from point cloud data sets, to enhance object detection capabilities by extracting features, producing an initial bird's eye view feature map, and then augmenting it with semantic information to improve distinction between objects and determine their distances.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Loss of information

If two-dimensional images are used for object detection, then visual appearance information is abundant, but distance information between the camera and objects is limited

Engineering Contradiction:
Improvevisual appearance informationVSAvoiddistance information
Core Design Contradiction:
Loss of informationVSMeasurement precision

Solution Approach 1:

The patent combines two-dimensional image data with three-dimensional point cloud data into a unified feature map. The two-dimensional image provides rich visual appearance information while the point cloud contributes precise distance information, merging both data sources to resolve the contradiction between visual detail and spatial accuracy.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent transforms two-dimensional image data into a three-dimensional bird's eye view feature map by integrating depth information from point clouds. This dimensional transformation allows the system to preserve visual appearance details while adding accurate distance measurements, effectively resolving the information loss versus measurement precision contradiction.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Measurement precision

If point cloud data sets are used for object detection, then precise distance information is available, but visual appearance information is limited

Engineering Contradiction:
Improvedistance informationVSAvoidvisual appearance information
Core Design Contradiction:
Measurement precisionVSLoss of information

Solution Approach 1:

The patent merges point cloud data with two-dimensional image data, where the point cloud provides accurate distance measurements and the image supplies rich visual appearance information. This combination resolves the contradiction by ensuring both precise spatial data and detailed visual characteristics are available for object detection.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent introduces a bird's eye view feature map as an intermediary representation that integrates both point cloud and image data. This feature map serves as a mediator that preserves distance information from the point cloud while incorporating visual appearance information from the image, resolving the information limitation.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Device complexity

If traditional object detection methods are used, then processing simplicity is maintained, but detection accuracy for small or occluded objects is poor

Engineering Contradiction:
Improveprocessing simplicityVSAvoidobject detection accuracy
Core Design Contradiction:
Device complexityVSMeasurement precision

Solution Approach 1:

The patent segments the detection process into distinct modules: a feature extraction module that processes both image and point cloud data, a feature map production module that generates the bird's eye view feature map, and an object detection module that identifies objects. This segmentation improves detection accuracy for small and occluded objects while maintaining processing organization and efficiency.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent creates a bird's eye view feature map that adds a top-down spatial dimension to traditional detection approaches. This dimensional change provides superior spatial context and object relationships, significantly improving detection accuracy for small and occluded objects while the modular architecture maintains processing efficiency.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

4Measurement precision

If semantic information is integrated into the feature map, then object distinction capability is improved, but data processing complexity increases

Engineering Contradiction:
Improveobject distinction capabilityVSAvoiddata processing complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent segments semantic information processing into a dedicated feature map production module that handles semantic augmentation separately from object detection. This modular approach improves object distinction capability through enriched semantic features while managing processing complexity through organized, separated functionality.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent performs semantic information integration and feature map production before the actual object detection process. By preliminarily enriching the feature map with semantic information, the system improves object distinction capability upfront, allowing the subsequent detection phase to operate more efficiently with pre-processed, information-rich data.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS11430218B2Using a bird's eye view feature map, augmented with semantic information, to detect an object in an environment
Publication Date: 2022.08.30 TOYOTA JIDOSHA KK
  • US11430218B2 patent drawing
  • US11430218B2 patent drawing
  • US11430218B2 patent drawing

AI summary

A bird's eye view feature map, augmented with semantic information, can be used to detect an object in an environment. A point cloud data set augmented with the semantic information that is associated with identities of classes of objects can be obtained. Features can be extracted from the point cloud data set. Based on the features, an initial bird's eye view feature map can be produced. Because operations performed on the point cloud data set to extract the features or to produce the initial bird's eye view feature map can have an effect of diminishing an ability to distinguish the semantic information in the initial bird's eye view feature map, the initial bird's eye view feature map can be augmented with the semantic information to produce an augmented bird's eye view feature map. Based on the augmented bird's eye view feature map, the object in the environment can be detected.