Cylindrical Transformer for 3D Depth Estimation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing methods for accurately detecting 3D spatial information in autonomous driving systems are limited by the high cost and density limitations of LiDAR sensors, while image sensors offer denser and more affordable alternatives for estimating 3D spatial information.

Innovation Solution

A learning device and method that utilize a plurality of image sensors to estimate 3D spatial information by processing main and peripheral images through an encoder, cylindrical feature mapping, and a cylindrical transformer to produce accurate depth information.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If LiDAR sensor is used to obtain 3D spatial information, then measurement precision is improved, but device cost increases excessively and device complexity increases

Engineering Contradiction:
Improve3D spatial information detection accuracyVSAvoidsensor system complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent uses image sensors to capture 2D images as a substitute (copy) for LiDAR's direct 3D measurement capability. By processing these image copies through neural networks and feature mapping, the system reconstructs 3D spatial information without requiring the expensive LiDAR hardware, thus reducing device complexity while maintaining measurement precision

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent replaces the mechanical/optical LiDAR system with an image sensor-based computational system. Instead of using physical light emission and detection mechanisms, the system uses camera images processed through encoder, cylindrical feature mapping, and transformer modules to achieve 3D reconstruction, substituting mechanical measurement with computational inference

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Measurement precision

If LiDAR sensor is used to obtain 3D spatial information, then measurement precision is improved, but the number of channels is limited resulting in sparse 3D information

Engineering Contradiction:
Improve3D spatial information densityVSAvoidnumber of measurement channels
Core Design Contradiction:
Measurement precisionVSQuantity of substance

Solution Approach 1:

The patent segments the image processing task into multiple components: encoder for feature extraction, cylindrical feature mapping for spatial transformation, and transformer for integration. This segmentation allows each component to process and contribute to different aspects of 3D information, enabling dense reconstruction from standard image sensors without requiring multiple LiDAR channels

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent transforms 2D image data into 3D spatial information by adding depth dimension through neural network inference. The cylindrical feature mapping creates multiple depth layers (first depth value to n-th depth value) from single 2D images, effectively adding dimensional information without increasing the number of physical sensors

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

3Device complexity

If image sensor is used to estimate 3D spatial information, then device cost is reduced and density is improved, but measurement precision deteriorates

Engineering Contradiction:
Improvesensor system costVSAvoid3D spatial information accuracy
Core Design Contradiction:
Device complexityVSMeasurement precision

Solution Approach 1:

The patent incorporates loss calculation and parameter update mechanisms that provide feedback to the neural network. The parameter update device adjusts encoder, cylindrical feature mapping, and transformer parameters based on prediction errors, continuously improving measurement precision through iterative learning while using cost-effective image sensors

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The patent creates a composite processing system combining multiple techniques: traditional image processing, cylindrical coordinate transformation, neural network inference, and temporal-spatial fusion. This composite approach leverages the strengths of each component to achieve high precision 3D reconstruction from standard image sensors

Inventive Principle:
Principle #40Composite materials

4Measurement precision

If 3D spatial information is estimated for each pixel of the image, then measurement precision is improved, but processing time increases

Engineering Contradiction:
Improvedepth information densityVSAvoidprocessing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent performs preliminary feature extraction and cylindrical feature mapping to create pre-processed depth information structures. By preparing these intermediate representations in advance through efficient neural network forward passes, the system reduces the computational burden during final depth estimation, maintaining high precision while minimizing processing time

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS20250173552A1Learning device, learning method, and test device and test method using same
Publication Date: 2025.05.29 HYUNDAI MOTOR CO LTD
  • US20250173552A1 patent drawing
  • US20250173552A1 patent drawing
  • US20250173552A1 patent drawing

AI summary

In a learning device and learning method, and a test device and a test method using the same, the learning device includes an encoder that outputs a main encoding feature, and a peripheral encoding, a cylindrical feature mapping device that maps the main encoding feature and the peripheral encoding feature to a first cylindrical shell to a n-th cylindrical shell, and outputs an integrated shell feature including a main feature and a peripheral feature, a cylindrical transformer that updates the integrated shell feature by modifying a value of the main feature with reference to the peripheral feature, a decoder that outputs predicted main depth information, and a parameter update device, that is configured to determine a first loss, and updates at least some of parameters of the encoder, the cylindrical feature mapping device, the cylindrical transformer, and the-decoder by use of the first loss.