Warp Structure Viewpoint Conversion for Object Identification

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing object identification methods for advanced driving support systems struggle to accurately grasp the positional relationship between multiple objects in an image, particularly when objects overlap, as they recognize objects only from a specific capture viewpoint, making it difficult to understand depth and positional relationships.

Innovation Solution

An object identification apparatus and method that utilizes a convolutional neural network to generate a viewpoint conversion map by warping feature maps from a capture coordinate system to a different coordinate system, allowing objects to be identified and positioned regardless of the capture viewpoint, thereby improving generalization performance and understanding of object relationships.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If objects are recognized only from a specific capture viewpoint using conventional image recognition methods, then the recognition process is simple, but the ability to grasp depth and positional relationships between objects deteriorates

Engineering Contradiction:
Improveaccuracy of positional relationship recognitionVSAvoidcomplexity of neural network structure
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent transforms the feature map from the capture coordinate system to a different coordinate system (e.g., bird's-eye view coordinate system) through warping operations. This dimensional transformation enables the system to grasp depth and positional relationships between objects by viewing them from a different perspective, thereby improving measurement precision without requiring an overly complex neural network structure.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Solution Approach 2:

The patent introduces a warp structure as an intermediary component between the convolutional neural network and the object identification process. This warp structure transforms feature maps to different coordinate systems, serving as a mediator that enables accurate positional relationship recognition while keeping the neural network structure manageable in complexity.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Measurement precision

If the neural network structure is made more complex to improve object identification accuracy, then identification precision improves, but the device complexity and computational load increase

Engineering Contradiction:
Improveobject identification accuracyVSAvoidneural network structure complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The warp structure acts as an intermediary that performs coordinate transformation on feature maps before they are used for object identification. This intermediary component enables the system to achieve high identification accuracy by transforming features to different coordinate systems without requiring the neural network itself to become overly complex.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent separates the coordinate transformation function from the object identification function by introducing a dedicated warp structure. This segmentation allows the neural network to focus on feature extraction and object identification while the warp structure handles coordinate transformations, thereby maintaining neural network simplicity while achieving high accuracy.

Inventive Principle:
Principle #1Segmentation

3Adaptability or versatility

If multiple capture viewpoints are used to improve understanding of object relationships, then the comprehensiveness of object recognition improves, but the quantity of data and processing requirements increase

Engineering Contradiction:
Improveunderstanding of object relationshipsVSAvoidamount of image data
Core Design Contradiction:
Adaptability or versatilityVSQuantity of substance

Solution Approach 1:

Instead of capturing images from multiple physical viewpoints, the patent transforms the single captured image's feature map to different coordinate systems through warping operations. This dimensional transformation provides comprehensive object relationship understanding by virtually presenting the same scene from different perspectives, thereby achieving versatility without increasing the quantity of image data.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Solution Approach 2:

The patent creates virtual copies of the feature map in different coordinate systems through warping operations. These copied and transformed feature maps provide multiple viewpoint information without requiring multiple physical captures, thereby reducing the quantity of actual image data while maintaining comprehensive object relationship understanding.

Inventive Principle:
Principle #26Copying

Data Source

PatentUS11734918B2Object identification apparatus, moving body system, object identification method, object identification model learning method, and object identification model learning apparatus
Publication Date: 2023.08.22 DENSO CORP
  • US11734918B2 patent drawing
  • US11734918B2 patent drawing
  • US11734918B2 patent drawing

AI summary

An object model learning method includes: in an object identification model forming a convolutional neural network and a warp structure warping a feature map extracted in the convolutional neural network to a different coordinate system, preparing, in the warp structure, a warp parameter for relating a position in the different coordinate system to a position in a coordinate system before warp; and learning the warp parameter to input a capture image in which an object is captured to the object identification model and output a viewpoint conversion map in which the object is identified in the different coordinate system.