3D Model Generation Using Multi-Camera Depth Mapping

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional methods for creating 3D models of products for online marketplaces face challenges in generating photorealistic renderings that allow users to interactively view products from various angles, as they often require extensive calibration, long capture times, and result in sparse structures unsuitable for visualization and interaction.

Innovation Solution

A system that simultaneously captures real-time depth and color data using pre-calibrated sensors, combined with high-resolution cameras, to generate dense 3D models with photo-realistic texture, allowing for interactive product visualization by mapping high-resolution image data to depth data and distributing projection errors for accurate 3D model generation.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If Structure from Motion (SFM), Visual SLAM, or Bundle Adjustment (BA) techniques are used to match image features and estimate camera viewpoints, then relative viewpoints and sparse structure can be obtained, but the resulting sparse structure is not suitable for creating photorealistic renderings needed for visualization and interaction

Engineering Contradiction:
Improvecamera viewpoint estimation accuracyVSAvoid3D model density quality
Core Design Contradiction:
Measurement precisionVSManufacturing precision

Solution Approach 1:

The patent combines multiple data sources including color images from RGB cameras, depth data from time-of-flight sensors, and normal maps to create a dense 3D model. This merging of multiple data types transforms the sparse structure from traditional SFM/BA methods into a dense, photorealistic model suitable for visualization and interaction while maintaining accurate camera viewpoint estimation.

Inventive Principle:
Principle #5Merging (Combining)

2Manufacturing precision

If 3D time-of-flight sensors (e.g., LIDAR) are augmented with cameras to generate high quality 3D models, then high quality dense 3D models can be created, but extensive calibration and long capture times are required

Engineering Contradiction:
Improve3D model qualityVSAvoidcapture time
Core Design Contradiction:
Manufacturing precisionVSLoss of time

Solution Approach 1:

The patent implements continuous capture of color images and depth data simultaneously using synchronized RGB and time-of-flight sensors. This continuous simultaneous capture eliminates the need for sequential calibration procedures, maintaining high 3D model quality while significantly reducing capture time by performing both data acquisition and calibration in an integrated continuous process.

Inventive Principle:
Principle #20Continuity of useful action

3Manufacturing precision

If 3D time-of-flight sensors (e.g., LIDAR) are augmented with cameras to generate high quality 3D models, then high quality dense 3D models can be created, but extensive calibration is required

Engineering Contradiction:
Improve3D model qualityVSAvoidcalibration complexity
Core Design Contradiction:
Manufacturing precisionVSDevice complexity

Solution Approach 1:

The system performs self-calibration by automatically determining transformation parameters between the RGB camera and time-of-flight sensor using feature matching and optimization algorithms. This self-service calibration approach eliminates the need for manual extensive calibration procedures, reducing device complexity while maintaining high 3D model quality through automated alignment of multiple sensor data streams.

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS10574974B23-D model generation using multiple cameras
Publication Date: 2020.02.25 AMAZON TECH INC
  • US10574974B2 patent drawing
  • US10574974B2 patent drawing
  • US10574974B2 patent drawing

AI summary

Various embodiments provide for the generation of 3D models of objects. For example, depth data and color image data can be captured from viewpoints around an object using a sensor. A camera having a higher resolution can simultaneously capture image data of the object. Features between images captured by the image sensor and the camera can be extracted and compared to determine a mapping between the camera and the image. Once the mapping between the camera and the image sensor is determined, a second mapping between adjacent viewpoints can be determined for each image around the object. In this example, each viewpoint overlaps with an adjacent viewpoint and features extracted from two overlapping viewpoints are matched to determine their relative alignment. Accordingly, a 3D point cloud can be generated and the images captured by the camera can be projected on the surface of the 3D point cloud to generate the 3D model.