Synthetic Image Training for Robotic Multi-Pick Detection

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Machine learning services for autonomous robotic control face challenges in training due to the scarcity of data for infrequently occurring events, such as multi-pick errors, which are difficult to source sufficiently.

Innovation Solution

A method is developed to synthesize training datasets by selecting and generating image data sequences from unlabelled data to simulate multi-pick events, using first, intermediate, and final images of items being picked by a robotic manipulator to teach the machine learning service to recognize and detect accidental multi-item grasps.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If training data is sourced from actual robotic operations, then the training data reflects real-world scenarios, but sufficient data for infrequently occurring events (multi-pick errors) cannot be obtained

Engineering Contradiction:
Improvetraining data qualityVSAvoidtraining data quantity
Core Design Contradiction:
ReliabilityVSQuantity of substance

Solution Approach 1:

The patent creates synthetic training data by copying and transforming existing single-pick operation data to simulate multi-pick error scenarios. Image sequences from normal operations are processed to generate artificial multi-pick error cases, allowing the machine learning service to be trained on sufficient data without requiring actual multi-pick errors to occur during data collection

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent performs preliminary data synthesis before actual robotic operations begin. By pre-generating synthetic multi-pick error training data from single-pick operation data, the system prepares comprehensive training datasets in advance, eliminating the need to wait for rare multi-pick errors to occur naturally during operation

Inventive Principle:
Principle #10Preliminary action

2Productivity

If the machine learning service is trained on limited data, then training time and computational resources are reduced, but detection accuracy for multi-pick events deteriorates

Engineering Contradiction:
Improvetraining efficiencyVSAvoidmulti-pick event detection accuracy
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The patent synthesizes additional training data by copying and transforming existing single-pick operation image sequences into multi-pick error scenarios. This multiplication of effective training data improves detection accuracy without requiring proportionally more computational resources or training time

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent transforms existing data by changing key parameters to create synthetic error cases. By modifying image sequences to represent multi-pick scenarios (changing the state from normal single-pick to error multi-pick), the system generates diverse training examples from limited source data, improving model generalization and detection accuracy

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS12515345B2Method of synthesising training datasets for autonomous robotic control
Publication Date: 2026.01.06 OCADO INNOVATION LTD
  • US12515345B2 patent drawing
  • US12515345B2 patent drawing
  • US12515345B2 patent drawing

AI summary

The present disclosure relates to a computer-implemented method of synthesising a training dataset for training a comparison function of a machine learning service in the detection of a multi-pick event in which a robotic manipulator erroneously picks two or more items concurrently, the method comprising: selecting first image data representative of an image of a plurality of items to be picked by the robotic manipulator; selecting intermediate image data representative of an image of the plurality of items following the removal, by the robotic manipulator, of an item from the plurality of items; selecting final image data representative of an image of the plurality of items following the removal, by the robotic manipulator, of another item from the plurality of items; and, generating the training dataset representative of a notional multi-pick event based on a sequence of unlabelled image data consisting of the first and final image data.