Labeled-Object Copying for Theme-Aligned Data Augmentation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing machine learning models face challenges in accuracy due to insufficient training data that aligns with the training theme, and manually generating high-quality data is time-consuming.
Innovation Solution
A method of data augmentation that involves selecting content from labeled areas of original images, generating sample images with border patterns different from the original content, and incorporating these into a sample dataset for training.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If open datasets are used for training machine learning models, then product development speed is improved, but the number of samples meeting the training theme is insufficient
Solution Approach 1:
The patent uses copy-paste operations to replicate target objects from source images and insert them into template images, creating synthetic training data that preserves the characteristics of original samples while generating additional variations to expand the training dataset
Solution Approach 2:
The patent introduces template images as intermediary elements that receive copied target objects and combine them with background patterns, serving as a mediator between source images and final training samples to generate diversified training data
2Manufacturing precision
If manual generation of high-quality training data is performed, then data quality is improved, but product development time and time cost increase significantly
Solution Approach 1:
The patent automatically copies target objects from source images and pastes them into template images through programmatic operations, eliminating manual data generation while maintaining consistent quality standards through automated processing
Solution Approach 2:
The patent varies parameters such as template images, copy positions, scaling factors, and rotation angles to generate diverse training samples automatically, achieving high data quality through systematic parameter variation rather than manual curation
3Quantity of substance
If data augmentation is performed using existing technologies, then training data volume is increased, but the content may not align with the training theme
Solution Approach 1:
The patent extracts target objects from source images based on labeled areas, isolating the relevant content that aligns with the training theme, and then inserts these extracted objects into template images to generate thematically consistent augmented data
Solution Approach 2:
The patent applies different processing treatments to different parts of the image, with target objects being copied and pasted while template images provide consistent backgrounds, ensuring that the augmented data maintains thematic alignment through localized quality control
Data Source
AI summary
A method of data augmentation is provided. The method includes the following operations: selecting an original image from an original dataset including label data configured to indicate a labeled area of the original image; selecting at least part of content, located in the labeled area, of the original image as a first target image; generating a first sample image according to the first target image, in which the first sample image includes the first target image and a first border pattern different from the first target image, and the content, located in the labeled area, of the original image is free from including at least part of the first border pattern; and incorporating the first sample image into a sample dataset, in which the sample dataset is configured to be inputted to a machine learning model.


