Metadata-Driven Prompt Generation for Diverse Synthetic Driving Images

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional image synthesis methods generate images that are too generic and lack diversity due to inadequate textual prompts, which limits their effectiveness in training advanced AI perception models, particularly for vehicle control systems.

Innovation Solution

A method that utilizes metadata from vehicle environments, such as weather and location information, to generate detailed and diverse text prompts, which are then used to create synthetic images for training and testing machine learning systems, specifically for autonomous driving applications.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If conventional text-image models are used for image synthesis, then images can be generated, but the generated images are too generic and lack diversity

Engineering Contradiction:
Improvediversity of generated imagesVSAvoiddetail in textual prompts
Core Design Contradiction:
Adaptability or versatilityVSLoss of information

Solution Approach 1:

The system performs preliminary actions by automatically generating comprehensive textual prompts from metadata before image synthesis. This includes creating detailed scene descriptions, object lists, and attribute specifications that are then used as input for the text-image model, ensuring diverse and detailed generated images

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent introduces an intermediary component that translates metadata into enriched textual prompts. This intermediary layer adds descriptive details, contextual information, and varied phrasing to the prompts, which then feed into the image synthesis model to produce more diverse and less generic images

Inventive Principle:
Principle #24Intermediary (Mediator)

2Reliability

If manual creation of training data is used, then data quality can be maintained, but data acquisition efforts and costs increase

Engineering Contradiction:
Improvequality of training dataVSAvoiddata acquisition efficiency
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The system enables self-service by automatically generating training data through a pipeline that combines metadata with text-image synthesis. The automated prompt generation and image creation processes eliminate the need for manual data annotation and curation, significantly improving productivity while maintaining quality through controlled generation parameters

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent uses copying by generating synthetic images that replicate real-world scenarios described in metadata. Instead of manually capturing and annotating real images, the system creates copies of desired training scenarios through text-to-image synthesis,大幅 reducing data acquisition costs and effort while maintaining training data quality

Inventive Principle:
Principle #26Copying

Data Source

PatentEP4697236A1Method for generating a data set for training and/or testing a machine learning system
Publication Date: 2026.02.18 ROBERT BOSCH GMBH
  • EP4697236A1 patent drawingFigure 1~2
  • EP4697236A1 patent drawingFigure 3
  • EP4697236A1 patent drawing

AI summary

The invention relates to a method (100) for generating a data set (60) for training and/or testing a machine learning system (50), comprising: - providing (101) image data specific to images depicting different environmental scenarios, - providing (102) metadata (65) specific to describing the different environmental scenarios, - generating (103) text prompts (70) based on the provided metadata (65), using information contained in the metadata (65) for the text prompts (70), - generating (104) the data set (60) based on the generated text prompts (70) and preferably the provided image data.