Student Network Training With Multi-Modal Synthetic Image Selection

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing techniques for generating a student network without using the training data of a teacher network face challenges in maintaining estimation accuracy due to discrepancies between training and real environments, particularly when using synthetic images.

Innovation Solution

Generate a student network using synthetic images created from real environment images captured by multiple modalities, such as RGB, ToF, and polarized cameras, and select or fuse these images based on similarity thresholds to improve training data quality.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If synthetic images generated by teacher network are used for training student network, then training can be performed without original training data, but the estimation accuracy deteriorates due to domain discrepancy between training and real environments

Engineering Contradiction:
ImproveAbility to train without original training dataVSAvoidEstimation accuracy
Core Design Contradiction:
Adaptability or versatilityVSMeasurement precision

Solution Approach 1:

The patent introduces domain adaptation as an intermediary technique that bridges the gap between synthetic training images and real environment images. By applying domain adaptation methods, the system reduces the domain discrepancy and enables the student network to generalize better from synthetic to real images, thus maintaining estimation accuracy while training without original training data

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent employs parameter changes by adjusting style parameters of synthetic images to match real environment characteristics. By modifying parameters such as color distribution, texture patterns, and illumination conditions in synthetic images, the system reduces the domain gap and improves estimation accuracy while maintaining the ability to train without original training data

Inventive Principle:
Principle #35Parameter changes

2Manufacturing precision

If multiple modalities are used to acquire real environment images, then the quality of training data improves, but the device complexity increases

Engineering Contradiction:
ImproveQuality of training dataVSAvoidComplexity of multi-modality acquisition system
Core Design Contradiction:
Manufacturing precisionVSDevice complexity

Solution Approach 1:

The patent applies partial action by selectively using only the necessary modalities required for the specific application rather than implementing all possible sensor types. This approach maintains training data quality by using the most relevant modalities while reducing device complexity by excluding unnecessary sensors and processing pipelines

Inventive Principle:
Principle #16Partial or excessive action

Solution Approach 2:

The patent segments the multi-modality acquisition system into modular components, where each modality can be independently configured, processed, and integrated. This segmentation allows flexible combination of modalities based on application requirements, maintaining data quality while managing complexity through modular architecture

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS12530865B2Information processing device and program
Publication Date: 2026.01.20 SONY GROUP CORP
  • US12530865B2 patent drawing
  • US12530865B2 patent drawing
  • US12530865B2 patent drawing

AI summary

There is provided an information processing device to improve the accuracy of estimation using a student network, the information processing device including an estimation unit that estimates an object class of an object included in an input image using a student network generated based on a teacher network generated by machine learning using images stored in a large-scale image database as training data. The student network is generated by machine learning using, as training data, synthetic images obtained using the teacher network and real environment images acquired by a plurality of different modalities in a real environment in which estimation by the estimation unit is expected to be executed.