Warehouse Picking Robot Using Transformer-Guided Grasp Poses
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing hybrid order-picking systems (HOPS) require extensive training to handle complex picking scenarios and are brittle, limiting their effectiveness in real-world warehouse operations.
Innovation Solution
A robot system utilizing a perceiver transformer and goal generator to autonomously navigate and pick objects in a warehouse, employing a perceiver transformer to generate latent vectors for end effector positioning and a goal generator to determine poses, combined with a planner to avoid collisions, all supported by machine-learning tools and compute devices.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If extensive training is applied to HOPS systems, then handling capability of complex picking scenarios improves, but system complexity and training time increase
Solution Approach 1:
The robot system performs self-training by autonomously observing and learning from human worker actions in the warehouse environment. The observation buffer stores human demonstrations, and the reinforcement learning agent uses these observations to automatically improve its picking skills without requiring manual retraining, thereby reducing training complexity while maintaining high adaptability
Solution Approach 2:
The system implements feedback mechanisms where the robot observes human worker performance, compares its own actions with observed successful actions, and uses this feedback to adjust its reinforcement learning model. This continuous feedback loop enables the system to improve handling capability through automated learning rather than extensive manual training
2Productivity
If HOPS systems are deployed in real-world warehouses, then productivity increases, but reliability decreases due to brittleness
Solution Approach 1:
The robot system transitions from static, pre-programmed picking operations to dynamic, adaptive decision-making using reinforcement learning. The agent continuously learns from observations of human workers and adjusts its picking strategies in real-time, enabling it to handle unexpected situations and maintain reliability while achieving high productivity in real-world warehouse environments
Solution Approach 2:
The system performs self-training by autonomously observing human workers and learning from their actions. This self-service learning capability allows the robot to adapt to real-world variations and maintain reliable operation without requiring extensive manual training or intervention, thereby improving both productivity and reliability
3Adaptability or versatility
If human workers perform manual order picking, then flexibility and adaptability are maintained, but productivity and cost efficiency decrease
Solution Approach 1:
The robot system performs self-training by autonomously observing human workers and learning from their actions. This self-service learning enables the robot to develop human-like adaptability and flexibility in handling diverse picking scenarios while maintaining higher productivity and cost efficiency through automated operation, effectively replacing manual workers without sacrificing adaptability
Solution Approach 2:
The system uses feedback from observing human worker actions to continuously improve its picking strategies. By learning from human demonstrations and adjusting its behavior accordingly, the robot achieves both the flexibility needed for diverse scenarios and the productivity benefits of automation, resolving the trade-off between adaptability and throughput
Data Source
AI summary
Technologies for generative AI for warehouse picking are disclosed. In an illustrative embodiment, a robot can partially or fully autonomously perform object picking in a warehouse. A description of an object to be picked can be sent to a compute device controlling the robot. The compute device can direct the robot to move to where the object is located. The robot can then take a picture that can be analyzed. A text description as well as the image is encoded and provided to a transformer. The transformer generates an output vector indicating, e.g., where an end effector should grab hold of an object. A goal generator can then determine a pose for the end effector based on the output latent vector. A planner can move the end effector to the determined pose. A new picture can be taken, and the cycle can repeat until the item is picked.


