3D Virtual Space Rendering for Neural Network Depth Data

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current methods for improving the generalization performance of neural networks (NNs) in image recognition tasks, particularly in capturing depth information, are inefficient and lack effective means to generate learning data with accurate depth context.

Innovation Solution

An information processing apparatus and method that arrange three-dimensional models of subjects and cameras in a virtual space, generate depth maps based on rendered images and distance information, and create defocus maps using depth values and camera parameters, thereby efficiently acquiring learning data with depth information.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If data augmentation methods (blurring, shaking, rotation, etc.) are used to improve generalization performance, then the NN can handle varied inputs better, but the depth information and three-dimensional context are lost or distorted

Engineering Contradiction:
Improvegeneralization performanceVSAvoiddepth information
Core Design Contradiction:
Adaptability or versatilityVSLoss of information

Solution Approach 1:

The patent transitions from 2D image processing to 3D virtual space processing. By arranging 3D models of subjects and cameras in a virtual three-dimensional space and rendering images with depth maps and defocus maps, the system preserves depth information while generating training data that maintains three-dimensional context throughout augmentation processes.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Solution Approach 2:

The patent creates synthetic training data by copying and rendering 3D models in virtual space. Instead of augmenting real 2D images which loses depth information, the system generates new training examples by rendering 3D models with various camera parameters, lighting conditions, and positions, thereby preserving depth information while providing diverse training data.

Inventive Principle:
Principle #26Copying

2Loss of information

If 3DCG rendering methods are used to generate learning data with depth information, then depth context is preserved, but the complexity of generating and processing three-dimensional models increases

Engineering Contradiction:
Improvedepth information preservationVSAvoidcomplexity of generating and processing 3D models
Core Design Contradiction:
Loss of informationVSDevice complexity

Solution Approach 1:

The patent creates a universal 3D model database that can serve multiple purposes. Once 3D models are created and stored in the database, they can be rendered repeatedly with different camera parameters, lighting conditions, and positions to generate diverse training data without requiring separate modeling efforts for each training example.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent performs preliminary actions by pre-creating and storing 3D models in a database before the actual training data generation. This preliminary modeling work is done once, and then the stored models can be efficiently rendered and augmented multiple times without repeating the complex modeling process.

Inventive Principle:
Principle #10Preliminary action

3Speed

If phase difference AF method is used to calculate shift amount of focal plane, then focusing can be performed more quickly, but the system requires two images with parallax which increases data requirements

Engineering Contradiction:
Improvefocusing speedVSAvoidnumber of images required
Core Design Contradiction:
SpeedVSQuantity of substance

Solution Approach 1:

The patent extracts depth information directly from rendered 3D models and distance information in the virtual space, rather than calculating it from multiple images. The system generates depth maps and defocus maps by directly measuring distances in the 3D virtual space, eliminating the need for complex multi-image processing while providing accurate depth data for training.

Inventive Principle:
Principle #2Taking out (Extraction)

Data Source

PatentUS20250139799A1Information processing apparatus, image capturing apparatus, method, and non-transitory computer readable storage medium
Publication Date: 2025.05.01 CANON KK
  • US20250139799A1 patent drawing
  • US20250139799A1 patent drawing
  • US20250139799A1 patent drawing

AI summary

The at least one processor arranges a three-dimensional model of a subject and a camera in a virtual three-dimensional space. The at least one processor generates a depth map including at least a depth value of a partial region of a region around the subject, based on an image in which the subject appears, the image being rendered based on a photographing field of view of the camera, and distance information corresponding to the photographing field of view. The at least one processor generate a defocus map including a defocus amount of the partial region based on a depth value of the partial region and a photographing parameter of the camera.