Pre-training Apparatus for Image Target Tasks
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing methods for pre-training neural networks for image-target tasks face challenges when a large amount of data cannot be collected due to measurement costs or when unlabeled data cannot be used, and there is a need for a general-purpose pre-training method that can be applied regardless of the task type.
Innovation Solution
A pre-training apparatus that converts input images into conversion images while maintaining local structure and deforming global structure, generates first and second extended images using different methods, and updates feature extractor parameters based on feature amounts calculated from these images.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If pre-training is performed using a large amount of unlabeled data, then the accuracy of the neural network for image-target tasks is improved, but the measurement cost and data collection cost increase
Solution Approach 1:
The patent creates multiple copies and variations of limited labeled images through affine transformations (rotation, scaling, translation) and other image processing techniques. This generates synthetic training data without requiring additional physical measurements, thereby maintaining high training accuracy while avoiding increased measurement costs.
Solution Approach 2:
The patent performs pre-training using synthetically generated images before fine-tuning with actual labeled data. This preliminary training phase prepares the neural network with basic image recognition capabilities using low-cost synthetic data, reducing the need for large amounts of expensive labeled measurement data in subsequent training stages.
2Measurement precision
If pre-training is performed using unlabeled data, then the neural network's performance is improved, but unlabeled data cannot be used in some cases due to task-specific requirements
Solution Approach 1:
The patent employs universal image transformation techniques (affine transformations, color space conversions, geometric modifications) that can be applied to any type of image data regardless of the specific task domain. This creates a general-purpose pre-training methodology that adapts to different task types without requiring task-specific labeled data, enhancing both performance and versatility.
Solution Approach 2:
The patent systematically varies image parameters (rotation angles, scaling factors, translation distances, color channel combinations) to generate diverse training samples from limited original images. This parameter-based approach creates task-agnostic training data that maintains effectiveness across different application domains while requiring minimal original data.
3Adaptability or versatility
If a general-purpose pre-training method is used, then the method can be applied regardless of task type, but existing methods rely on task-specific artificial images generated from random numbers
Solution Approach 1:
The patent introduces an intermediary image transformation process that bridges between limited labeled data and the needs of general-purpose pre-training. By applying controlled affine transformations and other processing as an intermediary step, the system generates diverse training samples that maintain the essential characteristics needed for effective training across different task types without relying on task-specific artificial image generation.
Data Source
AI summary
A pre-training apparatus includes processing circuitry. The processing circuitry is configured to: convert an input image to generate a conversion image; generate a first extended image and a second extended image from the conversion image based on a method different from a method for generating the conversion image; input the first extended image to a first feature extractor to calculate a first feature amount; input the second extended image to a second feature extractor to calculate a second feature amount; and update a parameter of at least one of the first feature extractor and the second feature extractor based on the first feature amount and the second feature amount.


