Text-Conditioned Image Synthesis for 3D Body Detection Data

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The acquisition of 3D body pose annotations is a difficult and slow process, limiting the diversity of scene locations and human appearances in existing datasets, which questions the real-world performance and robustness of current 3D human pose estimators.

Innovation Solution

A method for generating synthetic image data using a conditional image synthesis model, incorporating 2D skeleton representations, dense semantic encodings, and 2D depth maps, conditioned by textual prompts, to create diverse and photorealistic images for training and validating downstream neural networks for body detection tasks.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If specialized capture studios are used to acquire images with accurate 3D body pose annotations, then measurement precision is improved, but productivity deteriorates due to slow acquisition process

Engineering Contradiction:
Improve3D body pose annotation accuracyVSAvoidData acquisition speed
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The patent generates synthetic images that copy the essential characteristics of real images with accurate 3D body pose annotations. Instead of capturing real images in specialized studios, the system creates virtual images using neural networks that replicate the appearance, pose, and environmental details of real scenes, thereby obtaining large quantities of high-quality training data without the constraints of physical capture facilities

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The system changes the fundamental parameters of image generation from physical capture to computational synthesis. By using neural networks to generate images from textual descriptions and pose data, the system transforms the data acquisition process from a physical limitation-bound process to a computationally controlled process, enabling unlimited diversity in scene locations, human appearances, and poses

Inventive Principle:
Principle #35Parameter changes

2Adaptability or versatility

If real-world images are used for training downstream AI models, then adaptability to real scenarios is improved, but measurement precision of 3D body pose estimation deteriorates due to lack of accurate annotations

Engineering Contradiction:
ImproveReal-world scenario coverageVSAvoid3D body pose estimation accuracy
Core Design Contradiction:
Adaptability or versatilityVSMeasurement precision

Solution Approach 1:

The patent creates synthetic images that copy the visual characteristics of real-world scenes while incorporating precise 3D body pose annotations. The neural network generates images that replicate lighting, textures, and environmental details of real scenes, but with ground-truth pose information that enables accurate training of pose estimation models without requiring actual captured images

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The system separates the image generation process from the pose annotation process. Instead of requiring simultaneous capture of both image and pose data, the system generates images from textual prompts and pose representations independently, then combines them to create training data with known ground-truth poses, enabling precise pose estimation training

Inventive Principle:
Principle #1Segmentation

3Adaptability or versatility

If diverse datasets with various scene locations and appearances are collected, then adaptability is improved, but device complexity increases due to requirement for specialized capture studios

Engineering Contradiction:
ImproveDiversity of scene locations and appearancesVSAvoidSpecialized capture studio requirements
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent uses neural networks to copy and regenerate diverse scenes without requiring physical capture facilities. The system takes textual descriptions of desired scenes and generates corresponding images with accurate pose annotations, eliminating the need for specialized capture studios while maintaining unlimited diversity in scene locations, human appearances, and poses

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The system replaces the mechanical physical capture process with a computational generation process. Instead of using cameras and lighting equipment in controlled studios to capture diverse scenes, the system uses neural networks to synthetically generate images from textual prompts, substituting physical hardware with computational algorithms that can create any desired scene

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Data Source

PatentUS20250265825A1Technique for generating synthetic image data for body detection-related tasks
Publication Date: 2025.08.21 ROBERT BOSCH GMBH
  • US20250265825A1 patent drawing
  • US20250265825A1 patent drawing
  • US20250265825A1 patent drawing

AI summary

A technique for generating synthetic image data, which are usable for training, validating, and/or testing a downstream AI, in particular a downstream neural network, NN, for a body detection-related task based on sensor data. A method includes receiving visual information in relation to a body, wherein the visual information comprises a two-dimensional, 2D, skeleton representation of the body, a 2D projected (in particular dense) semantic encoding of the body, and a 2D depth map of the body. The method further includes receiving a textual prompt relating to at least one of an appearance of the body and/or environmental information relative to the body. The method further includes generating synthetic image data of the body based on the received textual prompt conditioned by the received visual information. The generating is performed by a conditional image synthesis model.