Interleaved Video Stream Distance Estimation for Mobile Robots

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current methods for determining distance to objects in video streams from moving robotic platforms are inefficient in terms of energy consumption and computational complexity, particularly when dealing with differential motion and changing backgrounds.

Innovation Solution

The use of a non-transitory computer-readable storage medium with instructions to produce a video stream by interleaving images from multiple cameras, evaluating the stream for binocular disparity, and encoding using motion estimation to determine distance, leveraging specialized hardware encoders like MPEG-4, H.262, H.263, H.264, and H.265 for efficient motion and depth analysis.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Use of energy by moving object

If traditional methods are used to determine distance to objects in video streams from moving robotic platforms, then distance measurement can be achieved, but energy consumption and computational complexity are high

Engineering Contradiction:
Improveenergy consumptionVSAvoiddistance measurement accuracy
Core Design Contradiction:
Use of energy by moving objectVSMeasurement precision

Solution Approach 1:

The video stream is segmented into individual frames from multiple cameras, which are then processed independently through interleaving. This segmentation allows the system to analyze only relevant portions of the video data rather than processing entire continuous streams, reducing computational complexity and energy consumption while maintaining distance measurement accuracy through selective frame analysis

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces an intermediary processing stage that interleaves frames from multiple cameras and uses motion estimation algorithms as a mediator between raw video data and final distance measurements. This intermediary layer filters and prepares data before final analysis, reducing the computational burden on the main processing system while preserving measurement precision

Inventive Principle:
Principle #24Intermediary (Mediator)

2Measurement precision

If complex motion estimation and binocular disparity evaluation are performed on all video frames, then accurate depth and distance information can be obtained, but computational complexity increases

Engineering Contradiction:
Improvedepth estimation accuracyVSAvoidcomputational complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The system performs preliminary actions by pre-interleaving frames from multiple cameras and pre-processing video streams before main distance calculation. Motion estimation is performed on selected frames in advance, creating prepared data structures that reduce the computational complexity of subsequent depth estimation operations while maintaining measurement precision

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

Instead of performing full motion estimation and binocular disparity evaluation on all video frames, the patent applies these computationally intensive operations selectively to key frames or regions of interest. This partial action approach reduces computational complexity by processing only necessary portions of the data while still achieving accurate depth estimation through strategic selection of analysis targets

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS10989521B2Apparatus and methods for distance estimation using multiple image sensors
Publication Date: 2021.04.27 BRAIN CORP
  • US10989521B2 patent drawing
  • US10989521B2 patent drawing
  • US10989521B2 patent drawing

AI summary

Data streams from multiple image sensors may be combined in order to form, for example, an interleaved video stream, which can be used to determine distance to an object. The video stream may be encoded using a motion estimation encoder. Output of the video encoder may be processed (e.g., parsed) in order to extract motion information present in the encoded video. The motion information may be utilized in order to determine a depth of visual scene, such as by using binocular disparity between two or more images by an adaptive controller in order to detect one or more objects salient to a given task. In one variant, depth information is utilized during control and operation of mobile robotic devices.