3D Point Cloud Generation Using Deep Learning Semantic Segmentation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional structure-from-motion techniques for generating 3D models from 2D images often discard metadata and colorspace information, leading to complex programming tasks and limited analysis capabilities.

Innovation Solution

The use of deep learning and structure-from-motion techniques to analyze 2D images and generate semantically-segmented 3D point clouds, where labeled points are identified and combined using a voting algorithm to preserve meaningful information.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If conventional structure-from-motion techniques are used to generate 3D models from 2D images, then the 3D model can be created, but metadata and colorspace information are discarded leading to complex programming tasks and limited analysis capabilities

Engineering Contradiction:
Improveprogramming complexityVSAvoidmetadata and colorspace information
Core Design Contradiction:
Ease of operationVSLoss of information

Solution Approach 1:

The patent applies preliminary action by performing semantic segmentation on 2D images before the structure-from-motion process. Deep learning models analyze and label pixels in advance, identifying objects, surfaces, and features. This preprocessing ensures that meaningful information is preserved and organized before 3D reconstruction begins, eliminating the need for complex post-processing programming and enabling richer analysis capabilities in the final 3D model.

Inventive Principle:
Principle #10Preliminary action

2Loss of information

If conventional SFM techniques retain colorspace information, then color data is preserved, but other useful information is discarded and programming remains complex

Engineering Contradiction:
Improvecolorspace informationVSAvoidprogramming complexity
Core Design Contradiction:
Loss of informationVSDevice complexity

Solution Approach 1:

The patent applies segmentation by dividing the image data into meaningful semantic categories using deep learning. Instead of treating all pixels uniformly or discarding non-color information, the system segments pixels into labeled categories (objects, surfaces, features) while preserving colorspace information. This semantic segmentation organizes information structurally, making it accessible without complex programming while retaining both color and semantic meaning for enhanced analysis.

Inventive Principle:
Principle #1Segmentation

3Quantity of substance

If 3D data is stored contiguously in memory for efficient access, then memory usage is optimized, but programming tasks become more complicated

Engineering Contradiction:
Improvememory efficiencyVSAvoidprogramming complexity
Core Design Contradiction:
Quantity of substanceVSEase of operation

Solution Approach 1:

The patent applies dimensionality change by organizing 3D point cloud data with additional semantic dimensions. Each point is not only stored with its spatial coordinates (x, y, z) but also enriched with semantic labels, object categories, and surface properties derived from deep learning analysis. This multi-dimensional organization allows efficient memory access while providing structured, interpretable information that simplifies programming tasks through meaningful data categorization rather than raw coordinate manipulation.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Data Source

PatentUS20250148632A1Using deep learning and structure from motion techniques to generate 3D point clouds from 2d data
Publication Date: 2025.05.08 STATE FARM MUTAL AUTOMOBILE INSURANCE COMPANY
  • US20250148632A1 patent drawing
  • US20250148632A1 patent drawing
  • US20250148632A1 patent drawing

AI summary

A server includes a processor and a memory storing instructions that, when executed by the processor, cause the server to receive two-dimensional (2D) images, analyze the images using a trained deep network to generate points, process the labeled points to identify tie points, and combine the 2D dimensional images into a three-dimensional (3D) point cloud using structure-from-motion. A method for generating a semantically-segmented 3D point cloud from 2D data includes receiving 2D images, analyzing the images using a trained deep network to generate labeled points, processing the points to identify tie points, and combining the 2D images into a 3D point cloud using structure-from-motion. A non-transitory computer readable storage medium stores executable instructions that, when executed by a processor, cause a computer to receive 2D images, analyze the images using a trained deep network to generate labeled points, process the points to identify and combine tie points using structure-from-motion.