Robot End-Effector Vision Control for Multi-Step Insertion

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing robot control methods for insertion tasks, such as peg-in-hole operations, are inefficient and slow, particularly when dealing with complex shapes and varying locations, and often require additional sensors beyond visual techniques.

Innovation Solution

A method utilizing a machine learning model trained with image data from dual cameras on the robot's end-effector to derive delta movements for efficient robot control, incorporating contrastive learning and one-shot learning techniques to minimize data requirements and sensor reliance.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Speed

If traditional visual techniques are used for robot insertion tasks, then the robot can perform the task, but the speed is about three times slower than human operators

Engineering Contradiction:
Improverobot operation speedVSAvoidtask completion time
Core Design Contradiction:
SpeedVSLoss of time

Solution Approach 1:

The patent replaces traditional mechanical visual processing systems with a machine learning-based controller that processes image data more efficiently. The ML model directly maps image inputs to motor control outputs, eliminating the need for complex intermediate visual processing steps that slow down traditional systems.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The patent changes the control parameters from traditional multi-step visual processing to direct image-to-movement mapping. By training the ML model on pairs of origin and target images with associated movement vectors, the system learns to directly translate visual information into actionable motor commands, significantly reducing processing time.

Inventive Principle:
Principle #35Parameter changes

2Adaptability or versatility

If existing robot control methods are used for insertion tasks with complex shapes and varying locations, then the robot can perform simple tasks, but it is not applicable to small subsets of problems involving simple shapes in fixed locations

Engineering Contradiction:
Improvetask applicability rangeVSAvoidtask execution reliability
Core Design Contradiction:
Adaptability or versatilityVSReliability

Solution Approach 1:

The patent creates a universal control method that can handle various insertion tasks regardless of object complexity or location. The ML model is trained on diverse image pairs representing different shapes, sizes, and positions, enabling it to generalize across task variations without requiring task-specific programming or additional sensors.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent performs preliminary training of the ML model using collected image data and movement vectors before actual task execution. This pre-training phase allows the robot to learn from examples of successful insertions, building a knowledge base that enables reliable performance across different task scenarios without requiring real-time complex calculations.

Inventive Principle:
Principle #10Preliminary action

3Measurement precision

If additional sensors beyond visual techniques are used, then measurement precision may improve, but device complexity increases

Engineering Contradiction:
Improveposition detection accuracyVSAvoidsensor system complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent makes the visual sensor system universal by training the ML model to extract all necessary positional and orientational information from standard images. The same camera system used for basic visual detection is also trained to provide precise positioning data for insertion tasks, eliminating the need for specialized sensors while maintaining measurement precision.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent enables the existing visual sensor system to serve multiple functions through ML processing. The same camera that captures images for basic object detection is also used for precise position measurement and movement calculation, with the ML model automatically extracting all necessary information from the image data without requiring additional sensing capabilities.

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS12370679B2Device and method for controlling a robot to perform a task
Publication Date: 2025.07.29 ROBERT BOSCH GMBH
  • US12370679B2 patent drawing
  • US12370679B2 patent drawing
  • US12370679B2 patent drawing

AI summary

A method for controlling a robot to perform a task. The method includes acquiring a target image data element comprising at least one target image from a perspective of an end-effector of the robot at a target position of the robot in which the robot has performed the task, acquiring an origin image data element comprising at least one origin image from the perspective of the end-effector of the robot at an origin position of the robot, supplying the origin image data element and the target image data element to a machine learning model configured to derive a delta movement between the origin current position and the target position and controlling the robot to move according to the delta movement to perform the task.