Multi-view Image Learning via Geometric Model Synthesis

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current learning image processing techniques face difficulties in handling multiple views and background-foreground separation, especially when the number of views increases or when background is close to foreground, leading to increased computational complexity and labeling burdens.

Innovation Solution

An image processing apparatus and method that estimates foreground and background images using geometric transforms on respective view models, synthesizes views, and updates model parameters based on evaluation values, employing stochastic generation models and posterior probability calculations to reduce computational load and labeling requirements.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If learning is executed for each view separately in multi-view learning, then learning accuracy for each view is improved, but the overall learning efficiency deteriorates when the number of views increases

Engineering Contradiction:
Improvelearning accuracyVSAvoidlearning efficiency
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The patent combines multiple view learning processes into a unified learning framework that processes multiple views simultaneously. The learning unit integrates image features from multiple views and executes learning in a coordinated manner, rather than processing each view separately. This merging approach maintains learning accuracy while significantly improving efficiency when handling multi-view data.

Inventive Principle:
Principle #5Merging (Combining)

2Measurement precision

If geometric relation between multiple views is precisely modeled, then learning accuracy is improved, but device complexity and calculation amount increase

Engineering Contradiction:
Improvelearning accuracyVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent extracts and utilizes only the necessary geometric relations for learning purposes, rather than modeling all possible geometric relationships between views. The system identifies and processes key geometric features that are sufficient for accurate learning, eliminating unnecessary complexity in the geometric modeling process while maintaining learning precision.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent changes the parameter representation of geometric relations from precise but complex models to simplified parameter sets that capture essential geometric information. By transforming the geometric relation parameters into a more efficient representation, the system achieves accurate learning with reduced computational complexity and simpler system requirements.

Inventive Principle:
Principle #35Parameter changes

3Measurement precision

If geometric relation between multiple views is precisely modeled, then learning accuracy is improved, but calculation amount increases

Engineering Contradiction:
Improvelearning accuracyVSAvoidcalculation amount
Core Design Contradiction:
Measurement precisionVSUse of energy by moving object

Solution Approach 1:

The patent extracts only the essential geometric information needed for learning, discarding redundant geometric details. This selective extraction reduces the volume of data that requires processing while preserving the critical geometric relationships necessary for accurate learning, thereby reducing overall calculation amount.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent transforms geometric relation parameters into a compressed or optimized representation that requires fewer computational operations. By changing how geometric parameters are stored and processed, the system maintains learning accuracy while significantly reducing the calculation amount and energy consumption required for processing multi-view geometric relationships.

Inventive Principle:
Principle #35Parameter changes

4Ease of manufacture

If foreground and background are not clearly disaggregated, then processing simplicity is maintained, but learning accuracy deteriorates when background is close to foreground

Engineering Contradiction:
Improveprocessing simplicityVSAvoidlearning accuracy
Core Design Contradiction:
Ease of manufactureVSMeasurement precision

Solution Approach 1:

The patent applies segmentation to separate foreground and background regions in the image data. By dividing the image into distinct foreground and background segments, the system can process each region with appropriate learning parameters while maintaining overall processing simplicity. This segmentation enables accurate learning even when foreground and background are closely positioned.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent applies different processing qualities and parameters to different regions of the image based on whether they are foreground or background. By assigning local quality characteristics to specific regions, the system can enhance learning accuracy in critical foreground areas while using simpler processing for background regions, balancing accuracy requirements with processing simplicity.

Inventive Principle:
Principle #3Local quality

Data Source

PatentUS8849017B2Image processing apparatus, image processing method, program, and recording medium for learning from moving images
Publication Date: 2014.09.30 SONY GROUP CORP
  • US8849017B2 patent drawing
  • US8849017B2 patent drawing
  • US8849017B2 patent drawing

AI summary

An image processing apparatus includes: an image feature outputting unit that outputs each of image features in correspondence with a time of the frame; a foreground estimating unit that estimates a foreground image at a time s by executing a view transform as a geometric transform on a foreground view model and outputs an estimated foreground view; a background estimating unit that estimates a background image at the time s by executing a view transform as a geometric transform on a background view model and outputs an estimated background view; a synthesized view generating unit that generates a synthesized view by synthesizing the estimated foreground and background views; a foreground learning unit that learns the foreground view model based on an evaluation value; and a background learning unit that learns the background view model based on the evaluation value by updating the parameter of the foreground view model.