Synthetic Image Data for Body Detection from 2D Pose and Text

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The acquisition of 3D body pose annotations is a difficult and slow process, limiting the diversity of scene locations and human appearances in existing datasets, which questions the real-world performance and robustness of current 3D human pose estimators.

Innovation Solution

A method involving the generation of synthetic image data using a conditional image synthesis model, incorporating 2D skeleton representations, dense semantic encodings, and 2D depth maps, conditioned by textual prompts, to enhance the training and validation of downstream neural networks for body detection tasks.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If specialized capture studios and manual annotation methods are used to acquire images with accurate 3D body pose annotations, then measurement precision is improved, but productivity deteriorates

Engineering Contradiction:
Improve3D body pose annotation accuracyVSAvoiddata acquisition speed
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The patent uses pre-trained 3D human pose estimation models to generate synthetic 3D pose annotations by copying and adapting knowledge from large-scale datasets. This allows automated generation of accurate 3D pose data without manual annotation, resolving the contradiction between annotation precision and acquisition productivity

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent replaces the mechanical manual annotation process with automated computer vision algorithms and neural networks. The system uses deep learning models to automatically estimate 3D body poses from images, substituting human labor with computational processes that maintain high precision while dramatically improving productivity

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Measurement precision

If specialized capture studios are used to acquire annotated images, then measurement precision is improved, but device complexity deteriorates

Engineering Contradiction:
Improve3D body pose annotation accuracyVSAvoidcapture studio requirements
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent copies 3D pose annotation capabilities from specialized capture environments into software-based synthetic annotation systems. By using pre-trained models and synthetic data generation, the system replicates the annotation quality of capture studios without requiring physical studio infrastructure

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent substitutes the complex physical capture studio infrastructure with automated computational pipelines. The system uses standard images combined with deep learning algorithms to generate 3D pose annotations, eliminating the need for specialized cameras, motion capture equipment, and controlled studio environments

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Productivity

If existing annotated datasets are used for training, then training efficiency is improved, but adaptability deteriorates

Engineering Contradiction:
Improvetraining efficiencyVSAvoiddiversity of scene locations and human appearances
Core Design Contradiction:
ProductivityVSAdaptability or versatility

Solution Approach 1:

The patent performs preliminary action by generating diverse synthetic training data with accurate 3D pose annotations before downstream model training. By pre-generating augmented datasets with varied appearances, poses, and environments using automated pipelines, the system enables efficient training while improving adaptability to diverse real-world scenarios

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent applies parameter changes through data augmentation techniques that modify synthetic images (lighting, pose, appearance, background) while preserving accurate 3D pose annotations. This creates diverse training scenarios from limited source data, improving model adaptability without sacrificing training efficiency

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentEP4604077A1Technique for generating synthetic image data for body detection-related tasks
Publication Date: 2025.08.20 ROBERT BOSCH GMBH
  • EP4604077A1 patent drawingFigure 1~3
  • EP4604077A1 patent drawingFigure 4~6
  • EP4604077A1 patent drawingFigure 7A~7D

AI summary

A technique for generating synthetic image data, which are usable for training, validating, and/or testing a downstream AI, in particular a downstream neural network, NN, for a body detection-related task based on sensor data is provided. A method comprises a step of receiving visual information in relation to a body, wherein the visual information comprises a two-dimensional, 2D, skeleton representation of the body, a 2D projected (in particular dense) semantic encoding of the body, and a 2D depth map of the body. The method further comprises a step of receiving a textual prompt relating to at least one of an appearance of the body and/or environmental information relative to the body. The method further comprises a step of generating synthetic image data of the body based on the received textual prompt conditioned by the received visual information. The generating is performed by a conditional image synthesis model.