3D Scene Modeling from 2D Images Using Depth Prediction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing technologies for producing three-dimensional models from two-dimensional images are either too expensive or too complicated for small contractors to use effectively, limiting their application in residential remodeling and other built environments.
Innovation Solution
A method involving a computing device that receives two-dimensional images, applies scene generator and depth prediction services to generate metadata for a three-dimensional model, using convolutional neural networks and conditional random fields to infer pixel depths and place objects within a spatial shell, enabling the creation of 3D models from a single image.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Manufacturing precision
If existing technologies for producing three-dimensional models from two-dimensional images are used, then manufacturing precision is improved, but device complexity and cost increase
Solution Approach 1:
The patent replaces complex mechanical 3D scanning systems with computer vision algorithms. Specifically, it uses convolutional neural networks and conditional random fields to infer depth information from standard 2D images, substituting physical depth sensors and mechanical scanning apparatus with software-based processing that analyzes image patterns and predictions to generate 3D models
Solution Approach 2:
The patent creates simplified copies of depth information by generating depth predictions from 2D image data rather than directly measuring it. The system produces metadata that includes predicted depth values for pixels, which are then used to construct 3D models, effectively copying spatial information from 2D representations without requiring complex 3D sensing hardware
2Manufacturing precision
If existing technologies for producing three-dimensional models from two-dimensional images are used, then manufacturing precision is improved, but ease of manufacture deteriorates
Solution Approach 1:
The patent enables the system to automatically perform all 3D modeling operations without requiring expert intervention. The convolutional neural network automatically identifies objects, estimates their properties, and generates depth predictions, while the conditional random field automatically refines these predictions by considering spatial relationships. This self-service capability allows small contractors to generate accurate 3D models using simple 2D images without needing specialized knowledge or equipment
Solution Approach 2:
The patent introduces metadata as an intermediary layer between the 2D image and the final 3D model. This metadata includes depth predictions, object labels, and spatial relationships that bridge the gap between simple 2D input and complex 3D output, making the transformation process more manageable and accessible to users without requiring them to understand the underlying complexity
3Productivity
If existing technologies for producing three-dimensional models from two-dimensional images are used, then productivity is improved, but device complexity increases
Solution Approach 1:
The patent performs preliminary processing by generating depth predictions using convolutional neural networks before finalizing the 3D model construction. The system pre-processes the 2D image to extract depth information, object labels, and spatial relationships, which are then used to efficiently build the 3D model. This preliminary action accelerates the overall process by preparing essential data structures in advance
Solution Approach 2:
The patent segments the complex 3D modeling task into distinct processing stages: first, the convolutional neural network segments the image processing to generate depth predictions; second, the conditional random field segments the refinement process to incorporate spatial relationships; third, the 3D model construction segments the final assembly. This segmentation allows each component to be optimized independently while maintaining high overall productivity
Data Source
AI summary
Described are techniques for producing a three-dimensional model of a scene from one or more two dimensional images. The techniques include receiving by a computing device one or more two dimensional digital images of a scene, the image including plural pixels, applying the received image data to scene generator/scene understanding engine that produces from the one or more digital images a metadata output that includes depth prediction data for at least some of the plural pixels in the two dimensional image and that produces metadata for a controlling a three-dimensional computer model engine, and outputting the metadata to a three-dimensional computer model engine to produce a three-dimensional digital computer model of the scene depicted in the two dimensional image.


