Monocular Distance Estimation for Real-Time AGV Visual Navigation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing vehicle distance estimation schemes for automated terminals are complex, expensive, difficult to deploy in real-time, and provide inaccurate predictions, leading to high costs and limited deployment in various working scenarios.

Innovation Solution

A lightweight attention mechanism distance estimation method using a depth monocular camera, enhanced with a modified Squeeze Former and WT Bins module, for accurate and efficient visual navigation in container terminals.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If a binocular or multi-purpose camera is used for distance estimation, then measurement precision is improved, but device complexity and cost increase

Engineering Contradiction:
Improvedistance estimation accuracyVSAvoidcamera system complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent replaces expensive binocular cameras with a single monocular camera, using a cheaper, simpler device that achieves comparable performance through algorithmic enhancement rather than hardware complexity

Inventive Principle:
Principle #27Cheap short-living objects (Disposable)

Solution Approach 2:

The patent transforms the monocular depth estimation problem by changing parameters in the attention mechanism and network architecture, optimizing the model to achieve accurate distance estimation without requiring multiple cameras

Inventive Principle:
Principle #35Parameter changes

2Measurement precision

If a Transformer-based backbone network with self-attention mechanism is used, then measurement precision is improved, but computing power and memory requirements increase

Engineering Contradiction:
Improvedepth estimation accuracyVSAvoidcomputational resource consumption
Core Design Contradiction:
Measurement precisionVSUse of energy by moving object

Solution Approach 1:

The patent extracts and removes the computationally intensive self-attention mechanism from the Transformer backbone, retaining only the essential feature extraction capabilities while eliminating the heavy computational burden

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent replaces the expensive Transformer-based backbone with a lighter alternative that achieves comparable performance with significantly reduced computational requirements, making real-time deployment feasible

Inventive Principle:
Principle #27Cheap short-living objects (Disposable)

3Measurement precision

If a complex network model with large number of parameters is used, then measurement precision is improved, but ease of operation and real-time deployment become difficult

Engineering Contradiction:
Improvedistance prediction accuracyVSAvoidreal-time deployment capability
Core Design Contradiction:
Measurement precisionVSEase of operation

Solution Approach 1:

The patent extracts and removes unnecessary complex components and excessive parameters from the network model, retaining only the essential elements needed for accurate depth estimation while enabling real-time operation

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent segments the network into modular components with optimized parameter counts, allowing efficient deployment while maintaining accuracy through structured design rather than brute-force parameter increases

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS20250342604A1Lightweight attention mechanism distance estimation method for assisting visual navigation of a vehicle at a container terminal
Publication Date: 2025.11.06 SHANGHAI SHIP & SHIPPING RES INST CO LTD
  • US20250342604A1 patent drawing
  • US20250342604A1 patent drawing
  • US20250342604A1 patent drawing

AI summary

The present discloses a lightweight attention mechanism distance estimation method for assisting visual navigation of a vehicle at a container terminal. Firstly, using a depth monocular camera calibrated with the imaging parameters of the planar checkerboard tool to collect RGB-Depth image pairs in the working scenario of the automatic guided vehicle. Secondly, performing depth completion and manual annotation processing on the collected depth image. Thirdly, inputting image pairs into a lightweight monocular metric depth estimation framework which uses an improved lightweight attention mechanism Squeeze Former as the token mixer for training. Finally, fusing the results of relative depth estimation and absolute depth estimation to obtain a prediction of an actual distance between an object in the RGB image and the camera in the real world. The method and model provided by the present invention feature simple equipment, low cost, high timeliness of prediction and accurate results.