Deep Learning Network for 3D Indoor Scene Structural Estimation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional deep learning methods are inefficient in estimating and outputting 3D indoor scenes directly, requiring additional processes and lacking the capability to generate 3D indoor scene images effectively.

Innovation Solution

An indoor scene structural estimation system based on a deep learning network, comprising a 2D encoder, 2D plane decoder, 2D edge decoder, 2D corner decoder, and 3D encoder, which decodes input images into 3D parameters to generate 3D indoor scene images through a two-phase training process, eliminating the need for additional processes and optimizing the operation efficiency.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If conventional deep learning methods are used to estimate 3D indoor scenes, then the system can process 2D images, but it requires additional processes and cannot directly output 3D scenes, resulting in increased operational complexity and reduced efficiency

Engineering Contradiction:
Improveoperational efficiencyVSAvoidprocess complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent merges multiple separate modules (2D encoder, 2D plane decoder, 2D edge decoder, 2D corner decoder, and 3D encoder) into a unified deep learning network architecture. This integration allows the system to process 2D images and directly output 3D indoor scenes through a single coordinated network, eliminating the need for additional separate processes and reducing operational complexity while improving productivity

Inventive Principle:
Principle #5Merging (Combining)

2Productivity

If conventional deep learning methods are used, then the network can be trained with standard approaches, but it cannot directly generate 3D indoor scene images, requiring post-processing steps that reduce operational efficiency

Engineering Contradiction:
Improveoperational efficiencyVSAvoiddirect 3D output capability
Core Design Contradiction:
ProductivityVSEase of operation

Solution Approach 1:

The patent extends the network's output capability from 2D image processing to 3D scene generation by introducing a 3D encoder module. This dimensional transition enables the network to directly output 3D indoor scene images with depth information, eliminating the need for post-processing steps and improving both operational efficiency and ease of operation

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

3Measurement precision

If a single-phase training approach is used, then the training process is simpler, but it cannot achieve accurate 3D parameter estimation and scene generation

Engineering Contradiction:
Improve3D parameter estimation accuracyVSAvoidtraining process complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent segments the training process into two distinct phases: a first training phase that trains the 2D encoder and decoders separately, and a second training phase that trains the 3D encoder with the pre-trained 2D components. This segmented approach enables accurate 3D parameter estimation by allowing each module to be optimized independently before integration, while keeping the training process manageable through systematic breakdown

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS10839606B2Indoor scene structural estimation system and estimation method thereof based on deep learning network
Publication Date: 2020.11.17 NATIONAL TSING HUA UNIVERSITY
  • US10839606B2 patent drawing
  • US10839606B2 patent drawing
  • US10839606B2 patent drawing

AI summary

An indoor scene structural estimation system and an estimation method based on deep learning network are provided. An indoor scene structural estimation system based on deep learning network includes a 2D encoder, a 2D plane decoder, a 2D edge decoder, a 2D corner decoder, and a 3D encoder. The 2D encoder receives an input image and encodes the input image. The 2D plane decoder is connected to the 2D encoder, decodes the encoded input image, and generates a 2D plane segment layout image. The 2D plane decoder is connected to the 2D encoder, decodes the encoded input image, and generates a 2D plane segment layout image. The 2D corner decoder is connected to 2D encoder, decodes the encoded input image, and generates a 2D corner layout image.