RGB-D and LiDAR Map Alignment for Semantic 3D Indoor Mapping

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing 2D LiDAR maps lack semantic details about objects and walls, making it difficult to generate full 3D maps of environments, and merging scans across rooms is error-prone due to the lack of overlapping features in image data.

Innovation Solution

Align RGB-D image sequences with 2D LiDAR maps to augment them with 3D geometry, texture, and semantics, using techniques like SLAM, ray casting, and inpainting to generate composite maps that include semantic information.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Loss of information

If 2D LiDAR is used to generate spatial maps, then the robot can capture floor layout and navigate efficiently, but the maps lack semantic details and qualitative information about objects and walls

Engineering Contradiction:
Improvesemantic informationVSAvoidmapping system complexity
Core Design Contradiction:
Loss of informationVSDevice complexity

Solution Approach 1:

The patent merges 2D LiDAR spatial maps with RGB-D image sequences to create composite maps that contain both geometric layout information and semantic object information. The 2D LiDAR provides accurate floor plan and navigation data, while RGB-D images provide semantic labels and 3D geometry, combining their strengths to resolve the information loss problem without requiring completely new hardware systems

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent transitions from 2D LiDAR point clouds to 3D reconstruction by processing RGB-D image sequences. This dimensional upgrade allows the system to capture semantic information and object properties that are invisible in 2D representations, while the 2D LiDAR map serves as a coordinate reference frame for aligning the enriched 3D data

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Loss of information

If commercial grade 3D mapping solutions are used, then full 3D maps with semantic information can be generated, but expensive hardware and meticulous scanning procedures are required

Engineering Contradiction:
Improve3D geometry and semantics informationVSAvoidimplementation ease
Core Design Contradiction:
Loss of informationVSEase of manufacture

Solution Approach 1:

The patent uses RGB-D images to create 3D reconstructions that copy and augment the environmental data captured by 2D LiDAR. Instead of requiring expensive commercial 3D scanners, the system creates digital copies of the environment from consumer-grade camera data, aligning these copies with the existing 2D LiDAR map coordinate system to achieve comprehensive 3D mapping

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent makes the 2D LiDAR map serve multiple functions: it acts as both a navigation guide for the robot and a common coordinate reference frame for aligning RGB-D image sequences. This multi-functionality allows the same hardware to support both navigation and 3D mapping tasks, reducing the need for specialized expensive equipment

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Ease of manufacture

If consumer grade cameras are used to reconstruct local areas, then 3D maps can be generated without expensive hardware, but scanning an entire house is extremely labor intensive and impractical

Engineering Contradiction:
Improvehardware costVSAvoidscanning time
Core Design Contradiction:
Ease of manufactureVSLoss of time

Solution Approach 1:

The patent performs preliminary alignment by using the 2D LiDAR map as a pre-established coordinate reference frame. This allows RGB-D image sequences to be directly registered to the global map without requiring time-consuming manual alignment or extensive overlapping captures, significantly reducing the scanning time required for entire environments

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The 2D LiDAR map serves as an intermediary that bridges local RGB-D reconstructions with the global environment model. By providing a common coordinate system and spatial context, it enables efficient merging of multiple local scans into a comprehensive 3D map without requiring labor-intensive manual registration procedures

Inventive Principle:
Principle #24Intermediary (Mediator)

4Area of stationary object

If multiple video sequences are captured from different locations to scan entire spaces, then complete coverage can be achieved, but merging scans is error-prone due to lack of overlapping features

Engineering Contradiction:
Improvemapped area coverageVSAvoidmerge accuracy
Core Design Contradiction:
Area of stationary objectVSReliability

Solution Approach 1:

The patent creates an equipotential reference frame by using the 2D LiDAR map as a common coordinate system for all RGB-D sequences. This eliminates the accumulation of alignment errors that occurs when sequences are merged directly with each other, as all data is registered to the same stable reference, ensuring consistent and reliable merging across the entire environment

Inventive Principle:
Principle #12Equipotentiality

Data Source

PatentUS12518407B2Aligning image data and map data
Publication Date: 2026.01.06 SAMSUNG ELECTRONICS CO LTD
  • US12518407B2 patent drawing
  • US12518407B2 patent drawing
  • US12518407B2 patent drawing

AI summary

A method of generating a composite map from image data and spatial map data may include: acquiring a spatial map of an environment; acquiring a plurality of images of a portion of the environment; generating a three-dimensional (3D) image of the portion of the environment using the plurality of images; identifying a floor in the 3D image; generating a synthetic spatial map of the portion of the environment based on the floor in the 3D image; determining a location of the portion of the environment within the spatial map by identifying a region of the spatial map that corresponds to the synthetic spatial map; and generating a composite map by associating the 3D image with the location within the spatial map of the portion of the environment.