Learned Stereo Architecture with Dynamic Disparity Ranges

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing stereo depth estimation systems are limited by predefined baselines, require extensive real-world data collection and labeling, and lack flexibility in adapting to different stereo camera systems, leading to high training costs and inefficiencies.

Innovation Solution

A learned stereo architecture utilizing fully differentiable 3D convolutions and trained with fully synthetic labeled data, enabling flexibility across various stereo camera systems and allowing dynamic adjustments to range and resolution.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If stereo depth estimation systems use image-based sensors only, then system versatility and sensor package requirements improve, but depth sensing accuracy and reliability deteriorate

Engineering Contradiction:
Improvesystem versatilityVSAvoiddepth sensing accuracy
Core Design Contradiction:
Adaptability or versatilityVSReliability

Solution Approach 1:

The system dynamically adjusts the disparity range parameter to adapt to different depth sensing requirements. By changing the parameter that defines the disparity search range, the system can optimize performance for different scenarios while maintaining accurate depth estimation using only image-based sensors

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The learned stereo architecture implements dynamic adjustments to the disparity range based on input conditions. The system can modify its operational parameters in real-time, allowing it to adapt to varying depth requirements and maintain reliability across different environments without additional sensors

Inventive Principle:
Principle #15Dynamics

2Loss of time

If learned stereo architecture is trained with fully synthetic labeled data, then training cost and data collection time improve, but measurement precision and reliability of depth estimation deteriorate

Engineering Contradiction:
Improvedata collection timeVSAvoiddepth estimation accuracy
Core Design Contradiction:
Loss of timeVSMeasurement precision

Solution Approach 1:

The system uses synthetic rendered images as copies or simulations of real-world scenes for training. These synthetic images with known ground truth disparity values serve as training data, eliminating the need for expensive and time-consuming real-world data collection while still enabling the network to learn accurate depth estimation patterns

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The system performs preliminary rendering of synthetic training data in advance, creating a comprehensive training dataset through computer graphics simulations before deployment. This preliminary action of generating synthetic training pairs with known disparities prepares the model for accurate real-world performance without requiring real-world data collection

Inventive Principle:
Principle #10Preliminary action

3Productivity

If learned stereo architecture is designed for a predefined range of disparities, then computational efficiency and processing speed improve, but adaptability to different stereo systems and ranges deteriorates

Engineering Contradiction:
Improveprocessing speedVSAvoidflexibility across stereo systems
Core Design Contradiction:
ProductivityVSAdaptability or versatility

Solution Approach 1:

The learned stereo architecture is designed with universal adaptability to work across multiple stereo camera systems with different baselines and disparity ranges. The network can be configured with different disparity range parameters to suit various hardware configurations, making it a multi-functional solution that maintains processing efficiency while adapting to diverse systems

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The system dynamically adjusts the disparity range parameter based on the specific stereo system configuration and application requirements. This dynamic reconfigurability allows the same base architecture to efficiently process different disparity ranges without requiring multiple specialized models, balancing speed and adaptability

Inventive Principle:
Principle #15Dynamics

Data Source

PatentUS20250238946A1Learned Stereo Flexibility
Publication Date: 2025.07.24 TOYOTA RESEARCH INSTITUTE INC
  • US20250238946A1 patent drawing
  • US20250238946A1 patent drawing
  • US20250238946A1 patent drawing

AI summary

A method for controlling a learned stereo architecture includes implementing a learned stereo architecture capable of generating refined disparity estimates for a predefined range of disparities based on training with fully synthetic image data; inputting a first range parameter into the learned stereo architecture, where the first range parameter is a subset of the predefined range of disparities; receiving, from a first stereo system having a first baseline, a first stereo image pair; generating a first disparity estimate, wherein disparities estimated in the first disparity estimate correspond to the first range parameter; upsampling the first disparity estimate to a resolution corresponding to a resolution of the first stereo image pair to form a first full resolution disparity estimate; refining the first full resolution disparity estimate with a first disparity residual thereby generating a first refined full resolution disparity estimate; and outputting the first refined full resolution disparity estimate.