Vision-Language Object Detection Through Rare-Scenario Data Augmentation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing autonomous vehicle (AV) training systems face challenges in creating diverse data for rare scenarios, such as encountering objects like caution tape, fire hoses, and plastic bags, which are difficult to gather and label, limiting the effectiveness of machine learning models.

Innovation Solution

Applying a vision-language model to open datasets to segment and augment training data by inserting cropped images of rare objects into training sets, enhancing the diversity and efficiency of data collection without the resource-intensive process of real-world recreation.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If real-world data collection is used to gather rare scenario objects, then data diversity is improved, but time and resource consumption increase significantly

Engineering Contradiction:
Improvedata diversityVSAvoidtime consumption
Core Design Contradiction:
Adaptability or versatilityVSLoss of time

Solution Approach 1:

The patent uses vision-language models to generate synthetic images of rare objects (caution tape, fire hoses, plastic bags) that copy the visual characteristics of real-world objects without requiring physical collection. This allows diverse training data to be created through digital replication rather than time-consuming real-world data gathering.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The vision-language model acts as an intermediary between the desired diverse training data and the practical limitations of data collection. The model processes text descriptions and generates corresponding images, serving as a mediator that transforms abstract data requirements into concrete training samples without direct real-world interaction.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Adaptability or versatility

If real-world data collection is used to gather rare scenario objects, then data diversity is improved, but labor and resource requirements increase

Engineering Contradiction:
Improvedata diversityVSAvoidresource requirements
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

Instead of physically collecting rare objects and managing complex data collection infrastructure, the system uses vision-language models to generate synthetic images. This replaces complex real-world data collection operations with computational processes, reducing physical resource requirements while maintaining data diversity.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent replaces mechanical data collection processes (physically gathering objects, setting up cameras, manual labeling) with computational methods. The vision-language model uses text-to-image generation capabilities to create training data, substituting physical operations with digital processing and eliminating the need for complex data collection infrastructure.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Ease of manufacture

If existing training data is used without augmentation, then data collection simplicity is maintained, but model performance on rare scenarios deteriorates

Engineering Contradiction:
Improvedata collection simplicityVSAvoidmodel performance
Core Design Contradiction:
Ease of manufactureVSReliability

Solution Approach 1:

The system performs preliminary data augmentation by generating diverse synthetic images before they are needed for training. The vision-language model pre-processes data by creating varied representations of rare objects, ensuring that the training data is enriched and ready for effective model training without requiring complex real-world data collection.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent creates synthetic copies of rare objects through vision-language model generation. These copied images replicate the visual characteristics of real rare objects while providing the diversity needed for robust model training, thereby improving model performance without complicating the data collection process.

Inventive Principle:
Principle #26Copying

Data Source

PatentUS20250225760A1Object detection by learning from vision-language model and data
Publication Date: 2025.07.10 GM CRUISE HOLDINGS LLC
  • US20250225760A1 patent drawing
  • US20250225760A1 patent drawing
  • US20250225760A1 patent drawing

AI summary

Aspects of the subject technology relate to systems, methods, and computer-readable media for diversifying training data through application of a vision-language model. A subset of images can be separated from a plurality of images in a dataset based on a presence of a specific object associated with autonomous driving in the subset of images. The specific object can be segmented in a portion of the image in the subset of images through application of a vision-language model. Training data for training a model associated with AV operation can be augmented by inserting the portion of the image into the training data to generate augmented training data. The model can be trained with the augmented data.