Neural Network Viewpoint Estimation for Rectangular Objects
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing methods for estimating another viewpoint image from two-viewpoint images with lateral disparity, particularly for rectangular-shaped objects, face challenges in accurately determining depth and corresponding points, leading to reduced estimation accuracy.
Innovation Solution
The use of a neural network-based image processing system that generates feature maps from two-viewpoint images, compares them to find corresponding points, and estimates another viewpoint image by deepening the network to expand the receptive field, facilitating the detection of rectangular-shaped objects with high accuracy.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Extent of automation
If a machine learning model is used to estimate another viewpoint image from two-viewpoint images with lateral disparity, then the estimation process can be automated, but the estimation accuracy for rectangular-shaped objects is lowered
Solution Approach 1:
The patent segments the image processing into multiple stages: generating candidate images at multiple viewpoints, selecting candidate images based on selection criteria, and performing final image generation. This segmentation allows the system to handle rectangular-shaped objects more accurately by breaking down the complex estimation process into manageable steps where each stage can optimize for specific object characteristics.
Solution Approach 2:
The patent introduces an additional dimension by generating candidate images at multiple viewpoints beyond the original two viewpoints. Instead of directly estimating from two viewpoints, the system creates intermediate candidate images at various viewpoints and selects the most appropriate ones, effectively adding a viewpoint dimension to the processing pipeline.
2Measurement precision
If the machine learning model deepens the network to expand the receptive field, then the detection of rectangular-shaped objects improves, but the device complexity increases
Solution Approach 1:
The patent divides the deep network into multiple processing stages or modules, each handling specific aspects of image analysis. Rather than using a single monolithic deep network, the system segments the processing into candidate image generation, selection, and refinement stages, making the overall system more manageable and interpretable.
Solution Approach 2:
The patent introduces intermediary candidate images as mediators between the input two-viewpoint images and the final estimated viewpoint image. These candidate images serve as intermediate representations that facilitate the detection of rectangular-shaped objects without requiring the final network to directly process all complexity in a single step.
3Measurement precision
If corresponding points are difficult to find between two-viewpoint images of rectangular-shaped objects, then depth determination becomes inaccurate, but using alternative methods increases processing complexity
Solution Approach 1:
The patent performs preliminary actions by generating candidate images at multiple viewpoints before final selection. This preliminary generation of candidate images with potentially better corresponding points allows the system to establish depth relationships more accurately before committing to the final image estimation, avoiding the need for complex alternative methods.
Solution Approach 2:
The patent creates copies of the original two-viewpoint images at multiple intermediate viewpoints as candidate images. These copied and transformed images provide additional geometric information and corresponding point relationships that facilitate more accurate depth determination without requiring fundamentally different processing approaches.
Data Source
AI summary
An image processing apparatus includes at least one processor or circuit configured to execute a plurality of tasks including an acquisition task configured to acquire two first images made by capturing the same object at two different viewpoints, and an image processing task configured to input the two first images into a machine learning model and to estimate a second image at one or more viewpoints different from the two viewpoints.


