Object Recognition Using Multi-Resolution Depth Image Segmentation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional depth cameras, such as those using Time of Flight or Light Coding techniques, are not widely available and produce imprecise depth images, making it difficult to recognize hand gestures, especially when the hand is slightly away from the camera, limiting their effectiveness in recognizing individual fingers.

Innovation Solution

A method utilizing two cameras with an overlapping field of view to capture and process images, where the resolution is reduced to generate low-level and high-level depth images, allowing for precise object recognition and depth information acquisition with reduced computational complexity by calculating shift amounts between pixel blocks and determining object areas within these images.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If conventional depth cameras using Time of Flight or Light Coding techniques are used, then depth information can be obtained, but the depth image precision is insufficient and computational complexity is high

Engineering Contradiction:
Improvedepth image precisionVSAvoidcomputational complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent divides the depth calculation process into two levels: first calculating a low-level depth image from original images at full resolution, then calculating a high-level depth image from sub-images at reduced resolution. This segmentation allows the system to obtain precise depth information for hand gestures while reducing overall computational complexity by processing only relevant regions at high resolution.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent applies partial action by calculating depth information only for specific regions of interest (hand gestures) rather than processing the entire image at high resolution. By identifying hand regions in the low-level depth image and then calculating high-level depth information only for those regions, the system achieves precise hand gesture recognition with reduced computational burden.

Inventive Principle:
Principle #16Partial or excessive action

2Measurement precision

If depth camera is used to capture images, then depth information can be acquired, but the precision is not sufficient for recognizing hand gestures when hand is away from camera

Engineering Contradiction:
Improvehand gesture recognition precisionVSAvoiddepth camera availability and performance
Core Design Contradiction:
Measurement precisionVSReliability

Solution Approach 1:

The patent transitions from relying solely on depth camera data to using a combination of color images and calculated depth information. By processing color images through multi-resolution analysis and calculating depth from pixel correspondences between different resolutions, the system achieves reliable hand gesture recognition without being limited by depth camera availability or performance.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Solution Approach 2:

The patent introduces an intermediary processing stage that calculates a low-level depth image from color images at full resolution. This intermediate depth map serves as a basis for identifying hand regions, which then guides the calculation of high-level depth information. This intermediary step enables reliable hand gesture recognition without requiring a physical depth camera.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Measurement precision

If original images are processed directly to generate depth image, then precision is maintained, but computational complexity increases significantly

Engineering Contradiction:
Improvedepth information precisionVSAvoidprocessing efficiency
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The patent segments the image processing into multiple resolution levels. First, a low-level depth image is calculated from original images at full resolution to identify regions of interest. Then, sub-images are extracted and processed at reduced resolution to generate the final high-level depth image. This segmentation maintains precision for hand gestures while significantly improving processing efficiency by reducing the amount of data processed at high resolution.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent applies partial action by calculating high-level depth information only for sub-images corresponding to hand gesture regions rather than processing the entire original image at high resolution. This approach maintains the precision needed for hand gesture recognition while dramatically reducing computational complexity by limiting high-resolution processing to relevant regions only.

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS8948493B2Method and electronic device for object recognition, and method for acquiring depth information of an object
Publication Date: 2015.02.03 WISTRON CORP
  • US8948493B2 patent drawing
  • US8948493B2 patent drawing
  • US8948493B2 patent drawing

AI summary

A method for recognizing an object from two original images, includes the steps of accessing the two original images, reducing resolutions of the two original images so as to generate two resolution-reduced images, respectively, calculating a plurality of shift amounts, each of which is between two corresponding pixels in pixel blocks that have similar content and that are respectively in the two resolution-reduced images and generating a low-level depth image based on the shift amounts, determining an object area of the low-level depth image containing the object therein, and obtaining a sub-image, from one of the original images, corresponding to the object area of the low-level depth image, thereby recognizing the object based on the sub-image.