Training Image Generation Using Difference Masks for Object Segmentation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing machine learning models for object segmentation in robotic systems require large amounts of training data and are sensitive to visual context, making them inefficient for handling unseen objects and densely stacked scenarios, and often necessitate human intervention.

Innovation Solution

A method and apparatus for generating labelled training images by computing depth and visual difference masks from before-and-after images of robotic or manual manipulation of stackable objects, enabling automated data collection and model training without human intervention.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If a large amount of training data is acquired for object segmentation, then the machine learning model performance is improved, but the time consumption and human intervention required increase

Engineering Contradiction:
Improveobject segmentation model performanceVSAvoidtraining data acquisition time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The system performs self-service by automatically generating training data through robotic manipulation actions. The robotic manipulator autonomously performs pick-and-place operations on objects, while the imaging system captures depth maps and visual images before and after manipulation. The processing system automatically computes difference masks and generates annotated segmentation masks without human intervention, enabling the system to create its own training data independently.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system creates copies of physical manipulation scenarios in the form of digital training data. By capturing depth maps and visual images during robotic manipulation and processing them into annotated segmentation masks, the system generates synthetic training examples that replicate real-world object manipulation scenarios, enabling the machine learning model to learn from these copied experiences.

Inventive Principle:
Principle #26Copying

2Measurement precision

If traditional object annotation techniques are used, then segmentation accuracy is improved, but automation is reduced due to human operator intervention

Engineering Contradiction:
Improvesegmentation accuracyVSAvoidannotation automation level
Core Design Contradiction:
Measurement precisionVSExtent of automation

Solution Approach 1:

The system replaces the mechanical process of manual annotation by human operators with an automated computational process. Instead of human operators visually inspecting and annotating images, the system uses image processing algorithms to compute depth difference masks and visual difference masks, which are then processed into annotated segmentation masks. This substitution of mechanical human labor with automated computational methods maintains segmentation accuracy while achieving full automation.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The system introduces intermediary computational elements (depth difference masks and visual difference masks) that mediate between the raw image data and the final segmentation annotations. These intermediate representations serve as bridges that automatically translate physical manipulation scenarios into structured training data, eliminating the need for direct human annotation while preserving segmentation quality.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Ease of operation

If objects are scattered on a flat surface for easy manipulation, then ease of operation is improved, but the ability to handle densely stacked objects is reduced

Engineering Contradiction:
Improveobject manipulation easeVSAvoidhandling of densely stacked objects
Core Design Contradiction:
Ease of operationVSAdaptability or versatility

Solution Approach 1:

The system transitions from handling objects on a two-dimensional flat surface to handling objects in three-dimensional stacked configurations. By using depth cameras to capture depth maps and computing depth difference masks, the system gains the ability to perceive and manipulate objects in the vertical dimension, enabling it to handle densely stacked objects while maintaining ease of operation through automated robotic manipulation.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Solution Approach 2:

The system changes the spatial parameters of object arrangement from scattered two-dimensional placement to densely packed three-dimensional stacking. The robotic manipulator is designed to handle objects in various configurations, and the imaging system captures depth information that enables the system to adapt to different stacking densities and arrangements, thereby increasing versatility while maintaining operational ease through automation.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS12592061B2Computer-implemented method for generating (training) images
Publication Date: 2026.03.31 ROBERT BOSCH GMBH
  • US12592061B2 patent drawing
  • US12592061B2 patent drawing
  • US12592061B2 patent drawing

AI summary

A computer-implemented method for generating labelled training images characterizing manipulation of a plurality of stackable objects in a workspace. The method includes: obtaining a first training image subset obtained at a first time index comprising a depth map and a visual image of a plurality of stackable objects in a stacking region of a workspace; obtaining a second training image subset obtained at a second time index comprising a depth map and a visual image of the stacking region in the workspace, wherein the second training image subset characterizes a changed spatial state of the stacking region; computing a depth difference mask based on the depth maps; computing a visual difference mask based on the visual images of the first and second training image subsets; generating an annotated segmentation mask using the depth difference mask and/or the visual difference mask.