Robotic Meal Assembly With Real-Time Pose Estimation for Similar Food Items
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing robotic systems face challenges in high-speed assembly of meals with high-mix and high-speed requirements due to the high resemblance and unpredictable poses of food items, leading to failed picking and imperfect assembly, which is exacerbated by the reliance on human labor in commercial kitchens.
Innovation Solution
A robotic system utilizing a robotic device with a gripper, imaging device, and computing device for real-time object pose estimation and automated visual inspection, employing a Fast Image to Pose Detection (FI2PD) method based on convolutional neural networks (CNN) to handle high-resemblance random food items, with a dataset like FdIngred328, and an item-arrangement verifier (IAV) algorithm for verifying meal assembly.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If traditional pattern matching methods with CAD models are used for pose estimation, then angular or richly-textured items can be estimated accurately, but food items with high resemblance and random poses cannot be distinguished
Solution Approach 1:
The patent replaces traditional mechanical pattern matching methods with a vision-based CNN system. The mechanical/CAD model-based approach is substituted by deep learning neural networks that process visual data directly, enabling the system to handle high-resemblance food items without requiring precise CAD models or texture-dependent features.
Solution Approach 2:
The patent changes the fundamental parameters of pose estimation by transitioning from 2D image processing to 6D pose estimation (three positional coordinates and three rotational angles). This parameter expansion allows the system to capture the full spatial configuration of food items, including their unpredictable poses, enabling accurate identification and localization for robotic manipulation.
2Measurement precision
If CNN-based methods are used for 6D pose estimation, then accuracy improves, but huge annotated datasets with 6D pose labels are required which are labor-consuming to create
Solution Approach 1:
The patent performs preliminary actions by pre-processing images to extract key features and pre-computing potential pose hypotheses before the main pose estimation step. This preliminary processing reduces the complexity of the subsequent pose refinement stage, enabling accurate 6D pose estimation without requiring exhaustive annotated datasets for all possible poses.
Solution Approach 2:
The patent introduces an intermediary approach by using 2D detection results as a intermediate step to guide 6D pose estimation. Rather than directly estimating 6D poses from raw images (which would require huge datasets), the system first performs 2D object detection to locate items, then uses these detections to constrain and guide the 6D pose estimation process, significantly reducing annotation requirements.
3Productivity
If high-speed robotic assembly is implemented for meals, then productivity increases, but the high-mix and high-speed requirements with unpredictable food item poses lead to failed picking and imperfect assembly
Solution Approach 1:
The patent implements a feedback loop where the vision system continuously monitors food item positions and poses, provides real-time feedback to the robotic manipulator, and allows for dynamic adjustment of picking parameters. This closed-loop control enables the system to adapt to unpredictable food item configurations at high speeds while maintaining high picking success rates through real-time correction of positioning errors.
Solution Approach 2:
The patent introduces dynamics by making the robotic system adaptive and flexible rather than rigid and predetermined. The system dynamically adjusts its pose estimation and picking parameters based on real-time visual feedback, allowing it to handle the variability of high-mix meal components at high speeds. The dynamic nature of the vision-guided control enables the system to respond to unpredictable food item poses without sacrificing productivity or reliability.
Data Source
AI summary
Methods, systems and computer readable media are provided for automatic kitting of items. The system for automatic kitting of items includes a robotic device, a first imaging device, a computing device and a controller. The robotic device includes an arm with a robotic gripper at one end. The first imaging device is focused on a device conveying kitted items. The computing device is coupled to the first imaging device and is configured to process image data from the first imaging device. The computing device includes item arrangement verification software configured to determine whether each item desired to be in the kitted items is present or absent in the kitted items in response to the processed image data from the first imaging device and generates data based on whether an item desired to be in the kitted items is absent from the kitted items. The controller is coupled to the computing device to receive the data from the computer representing whether an item desired to be in the kitted items is absent from the kitted items. The controller is also coupled to the robotic device for providing instructions to the robotic device to control movement of the arm and the robotic gripper, wherein at least some of the instructions provided to the robotic device are generated in response to the item desired to be in the kitted items being absent from the kitted items.


