AGV Vision Switching Between Stereo and Monocular Detection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Automated logistics systems face challenges in navigation and object detection due to limitations in existing sensors, particularly when stereo vision is impaired by obstructions or unsuitable image data, leading to reduced accuracy and increased pick failure rates in densely packed and dynamically varying storage environments.
Innovation Solution
The implementation of an autonomous guided vehicle equipped with a vision system using a combination of stereo and monocular vision, employing artificial neural networks and machine learning models to select the most confident detection protocol, allowing for robust object detection and localization even in constrained environments with dense spacing, deformities, and dynamic case size distributions.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If stereo vision systems are used for object detection and localization, then depth perception and spatial understanding are improved, but the system becomes vulnerable to impairment from obstructions and unsuitable image data
Solution Approach 1:
The system dynamically switches between stereo vision mode and monocular vision mode based on real-time assessment of image quality and availability. When obstructions are detected or image data becomes unsuitable, the system transitions from relying on both stereo cameras to using a single monocular camera, ensuring continuous operation without failure.
Solution Approach 2:
The system changes operational parameters by switching detection protocols between stereo-based depth estimation and monocular-based depth estimation. This parameter change allows the system to adapt to varying environmental conditions, maintaining measurement precision across different scenarios by selecting the appropriate detection mode.
2Reliability
If monocular vision is used as backup when stereo vision fails, then system availability is improved, but measurement precision and depth determination capability are reduced
Solution Approach 1:
The system replaces the mechanical stereo vision mechanism with an artificial intelligence-based monocular depth estimation system. Instead of relying on physical camera pairs, the system uses neural networks to infer depth information from single images, maintaining measurement capability through intelligent algorithms rather than mechanical redundancy.
Solution Approach 2:
The artificial neural network acts as an intermediary that translates monocular image data into depth information. This intermediary component bridges the gap between limited monocular input and the depth perception requirements of the logistics system, enabling accurate depth determination even with a single camera.
3Adaptability or versatility
If multiple detection protocols are implemented to handle various scenarios, then adaptability is improved, but system complexity and processing requirements increase
Solution Approach 1:
The artificial neural network serves multiple functions: it processes monocular images for depth estimation, assesses image quality, determines when to switch between detection modes, and adapts to various lighting and obstruction conditions. This multi-functional approach consolidates what would otherwise require separate specialized systems into a single versatile component.
Solution Approach 2:
The system implements feedback mechanisms where the AI model continuously evaluates image quality and detection confidence, then adjusts its operational mode accordingly. This feedback loop allows the system to automatically adapt to changing conditions without external intervention, managing complexity through intelligent self-regulation rather than rigid multi-protocol architecture.
Data Source
AI summary
An autonomous guided vehicle including a frame, a drive section, a payload handler, a vision system, and a controller. The vision system has a camera disposed to generate video stream data imaging of an object. The controller being communicably connected to register the video stream data imaging from the at least one camera and communicably connected to at least one or more of a time of flight sensor and a distance sensor that detects a distance of the object. The controller is configured so to effect, from the video stream data imaging, robust object detection and localization within a predetermined reference frame via alternately both binocular vision and monocular vision from the video stream data imaging, the detection determined via monocular vision having confidence commensurate with detection determined via the binocular vision.


