Student Network Training With Multi-Modal Synthetic Image Selection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing techniques for generating a student network without using the training data of a teacher network face challenges in maintaining estimation accuracy due to discrepancies between training and real environments, particularly when using synthetic images.
Innovation Solution
Generate a student network using synthetic images created from real environment images captured by multiple modalities, such as RGB, ToF, and polarized cameras, and select or fuse these images based on similarity thresholds to improve training data quality.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If synthetic images generated by teacher network are used for training student network, then training can be performed without original training data, but the estimation accuracy deteriorates due to domain discrepancy between training and real environments
Solution Approach 1:
The patent introduces domain adaptation as an intermediary technique that bridges the gap between synthetic training images and real environment images. By applying domain adaptation methods, the system reduces the domain discrepancy and enables the student network to generalize better from synthetic to real images, thus maintaining estimation accuracy while training without original training data
Solution Approach 2:
The patent employs parameter changes by adjusting style parameters of synthetic images to match real environment characteristics. By modifying parameters such as color distribution, texture patterns, and illumination conditions in synthetic images, the system reduces the domain gap and improves estimation accuracy while maintaining the ability to train without original training data
2Manufacturing precision
If multiple modalities are used to acquire real environment images, then the quality of training data improves, but the device complexity increases
Solution Approach 1:
The patent applies partial action by selectively using only the necessary modalities required for the specific application rather than implementing all possible sensor types. This approach maintains training data quality by using the most relevant modalities while reducing device complexity by excluding unnecessary sensors and processing pipelines
Solution Approach 2:
The patent segments the multi-modality acquisition system into modular components, where each modality can be independently configured, processed, and integrated. This segmentation allows flexible combination of modalities based on application requirements, maintaining data quality while managing complexity through modular architecture
Data Source
AI summary
There is provided an information processing device to improve the accuracy of estimation using a student network, the information processing device including an estimation unit that estimates an object class of an object included in an input image using a student network generated based on a teacher network generated by machine learning using images stored in a large-scale image database as training data. The student network is generated by machine learning using, as training data, synthetic images obtained using the teacher network and real environment images acquired by a plurality of different modalities in a real environment in which estimation by the estimation unit is expected to be executed.


