Robot Insertion Control Using Image-Based Delta Movement Learning

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing robot control methods for insertion tasks, such as peg-in-hole operations, are inefficient and slow, particularly when dealing with complex shapes and varying locations, and often require additional sensors beyond visual techniques.

Innovation Solution

A method utilizing a machine learning model trained with image data from multiple cameras and force sensors to derive delta movements for robot control, enabling efficient multi-step tasks without the need for additional sensors.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Speed

If traditional visual techniques are used for robot insertion tasks, then the robot can perform simple insertion tasks, but the speed is about three times slower than human operators

Engineering Contradiction:
Improveinsertion task execution speedVSAvoidoverall task efficiency
Core Design Contradiction:
SpeedVSProductivity

Solution Approach 1:

The patent replaces traditional visual-based control systems with a machine learning model that processes image data more efficiently. The ML model predicts delta movements directly from image pairs, substituting the slower traditional visual processing pipeline and achieving speeds comparable to or exceeding human operators while maintaining accuracy in insertion tasks.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Adaptability or versatility

If existing robot control schemes are used for insertion tasks, then the robot can handle simple shapes in fixed locations, but it cannot handle complex shapes and varying locations

Engineering Contradiction:
Improveability to handle complex shapes and varying locationsVSAvoidcontrol system complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent changes the control parameters from fixed-location simple shape handling to variable location and complex shape handling by training the machine learning model on diverse image data. The model learns to generalize across different shapes, locations, and orientations by processing image pairs and predicting delta movements, enabling adaptable insertion tasks without increasing physical device complexity.

Inventive Principle:
Principle #35Parameter changes

3Measurement precision

If additional sensors beyond visual techniques are used, then measurement precision may improve, but device complexity and cost increase

Engineering Contradiction:
Improveposition detection precisionVSAvoidsensor system complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent uses image data as a copy or representation of the physical scene, processing visual information through machine learning to achieve precise position detection without adding physical sensors. The ML model extracts spatial relationships and movement cues from image pairs, creating a virtual model of the insertion task that replaces the need for additional tactile or force sensors.

Inventive Principle:
Principle #26Copying

Data Source

PatentUS12479108B2Device and control method using machine learning for a robot to perform an insertion task
Publication Date: 2025.11.25 ROBERT BOSCH GMBH
  • US12479108B2 patent drawing
  • US12479108B2 patent drawing
  • US12479108B2 patent drawing

AI summary

A method for controlling a robot to perform a task. The method includes acquiring, for each target of a sequence of targets comprising at least one intermediate target of the task and a final target of a task, a target image data element comprising at least one target image from a perspective of an end-effector of the robot at a respective target position of the robot and successively according to the sequence of targets, for each target in the sequence, acquiring, for the target, an origin image data element, supplying the origin image data element and the target image data to a machine learning model configured to derive a delta movement between the origin current position and the target position and controlling the robot to move according to the delta movement.