Object Recognition via Component Decomposition for Partial Appearances

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Image recognition systems face difficulties in identifying objects with partial appearances, orientation variations, and occlusions due to challenges in decomposing objects into components and estimating pose information effectively.

Innovation Solution

A system and method that decompose objects into components using a learning module, trained with a component-based approach, incorporating overall objectness scores, component scores, and pose information, allowing for recognition and reconstruction of objects with partial appearances and orientation variations by using a deep neural network and generative adversarial networks.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If object decomposition into components is performed to improve recognition accuracy, then recognition precision improves, but device complexity increases

Engineering Contradiction:
Improverecognition accuracyVSAvoiddevice complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent applies segmentation by decomposing objects into multiple components or parts. The system divides a complex object recognition task into smaller sub-tasks, where each component is detected and recognized separately. This allows the recognition system to handle partial appearances and occlusions more effectively by focusing on detectable components rather than requiring the complete object to be visible.

Inventive Principle:
Principle #1Segmentation

2Reliability

If component-based training is used to handle partial appearances, then recognition reliability improves, but training complexity increases

Engineering Contradiction:
Improverecognition reliabilityVSAvoidtraining complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent applies preliminary action by performing component decomposition and labeling during the training phase. Training images are pre-processed to identify and label individual components of objects, creating a structured dataset that teaches the system to recognize parts separately. This preliminary structuring of training data enables the system to reliably handle partial appearances during deployment, as it has already learned component relationships during training.

Inventive Principle:
Principle #10Preliminary action

3Adaptability or versatility

If pose estimation for each component is performed to handle orientation variations, then adaptability improves, but measurement complexity increases

Engineering Contradiction:
ImproveadaptabilityVSAvoidmeasurement complexity
Core Design Contradiction:
Adaptability or versatilityVSDifficulty of detecting and measuring

Solution Approach 1:

The patent applies segmentation by estimating pose for individual components rather than treating the object as a single unit. Each component can have its own pose parameters (position, orientation, scale), allowing the system to handle objects in various orientations and configurations. This component-level pose estimation provides greater adaptability to orientation variations while managing complexity through modular processing.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS10460470B2Recognition and reconstruction of objects with partial appearance
Publication Date: 2019.10.29 FUTUREWEI TECHNOLOGIES INC
  • US10460470B2 patent drawing
  • US10460470B2 patent drawing
  • US10460470B2 patent drawing

AI summary

Various embodiments include systems and methods structured to provide recognition of an object in an image using a learning module trained using decomposition of the object into components in a number of training images. The training can be based on an overall objectness score of the object, an objectness score of each component of the object, a pose of the object, and a pose of each component of the object for each training image input. Additional systems and methods can be implemented in a variety of applications.