3D Environment Mapping Using Semantic Segmentation and Depth Fusion

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Depth sensors often fail to provide depth data for non-edge portions of objects, leading to incomplete and inaccurate environment maps, and existing systems that rely solely on depth data or image data suffer from resolution and shape representation issues.

Innovation Solution

A system that combines image data with semantic segmentation to generate a voxel-based three-dimensional map, using depth data to identify object edges and image data to fill in non-edge portions, ensuring complete and accurate object representation.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If depth data is used to identify object edges, then location precision is improved, but non-edge portions of objects are not captured leading to incomplete shape representation

Engineering Contradiction:
Improvelocation precisionVSAvoidshape representation
Core Design Contradiction:
Measurement precisionVSShape

Solution Approach 1:

The patent combines depth data with semantic segmentation results to create a complete three-dimensional map. The depth data provides accurate edge location information, while the semantic segmentation fills in the non-edge portions of objects, merging both data sources to achieve complete and accurate shape representation.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent uses semantic segmentation as an intermediary to bridge the gap between depth data edges and complete object shapes. The segmentation process identifies and fills non-edge portions that depth sensors miss, acting as a mediator to complete the shape information while preserving the precise edge locations from depth data.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Shape

If image data is used to capture complete objects, then shape completeness is improved, but resolution and shape precision are reduced

Engineering Contradiction:
Improveshape completenessVSAvoidshape precision
Core Design Contradiction:
ShapeVSMeasurement precision

Solution Approach 1:

The patent merges image data with depth data to achieve both complete shape representation and high precision. The image data ensures all portions of objects are captured including non-edge areas, while the depth data provides precise edge location and shape information, combining both advantages in the final three-dimensional map.

Inventive Principle:
Principle #5Merging (Combining)

3Loss of information

If depth sensors are used to obtain depth data, then depth information is obtained, but non-edge portions of surfaces are not identified

Engineering Contradiction:
Improvedepth informationVSAvoidsurface coverage
Core Design Contradiction:
Loss of informationVSShape

Solution Approach 1:

The patent employs semantic segmentation as an intermediary mechanism to supplement depth sensor limitations. The segmentation process identifies and fills non-edge surface portions that depth sensors fail to capture, while preserving the accurate depth information where available, thereby achieving complete surface coverage without losing depth precision.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS20260073631A1Systems and methods for environment mapping based on multi-domain sensor data
Publication Date: 2026.03.12 QUALCOMM INC
  • US20260073631A1 patent drawing
  • US20260073631A1 patent drawing
  • US20260073631A1 patent drawing

AI summary

Systems and techniques for environment mapping are described. In some examples, a system receives image data and depth data captured using at least one sensor. The image data and the depth data both include respective representations of an environment. The system processes the image data using semantic segmentation to identify segments of the environment that represent different types of objects in the environment in the image data. The system combines the depth data with the semantic segmentation to generate a voxel-based three-dimensional map of the environment.