Image Conversion Model Training with Class Preservation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing image conversion models struggle to convert images while maintaining the class of each image region, leading to loss of objects during environmental conversion.
Innovation Solution
A model training apparatus and method that acquires a training dataset with class information for image regions, trains an image conversion model to output images in a different environment, and uses a discrimination model to calculate losses and update the image conversion model parameters, ensuring that the class of each image region is preserved.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If image conversion is performed without preserving class information, then environmental conversion is achieved, but object class information is lost
Solution Approach 1:
The patent divides the image into multiple image regions and associates each region with class information. The loss function is calculated separately for each region, allowing the model to preserve class information while performing environmental conversion. This segmentation approach enables the model to maintain object identity across different environments.
Solution Approach 2:
The patent introduces a discrimination model that provides feedback about the authenticity and class accuracy of generated image regions. The loss function incorporates discrimination data to guide the image conversion model, ensuring that converted regions maintain their original class characteristics while adapting to the target environment.
2Reliability
If class information is preserved for each image region, then object identity is maintained, but training complexity increases
Solution Approach 1:
The patent combines multiple objectives into a unified loss function that simultaneously evaluates environmental conversion quality and class information preservation. By merging the authenticity loss and class consistency loss into a single optimization framework, the model achieves reliable object identity preservation without proportionally increasing training complexity.
Solution Approach 2:
The discrimination model serves multiple functions: it evaluates image authenticity, preserves class information, and guides the image conversion process. This multi-functionality reduces the need for separate specialized models, thereby limiting the increase in system complexity while maintaining high reliability in object identity preservation.
3Measurement precision
If discrimination data is used for training, then image conversion accuracy improves, but computational resources increase
Solution Approach 1:
The patent applies discrimination data selectively to specific image regions rather than uniformly across the entire image. By focusing computational resources on regions where class preservation is most critical, the model achieves high conversion accuracy while reducing overall computational consumption compared to applying discrimination to all regions equally.
Data Source
AI summary
A model training apparatus acquires a first training data set including a first training image representing a scene in a first environment and first class information indicating a class of each of a plurality of image regions included in the first training image. The model training apparatus inputs the first training image to an image conversion model to acquire an output image representing a scene in a second environment, inputs the output image to a discrimination model to acquire discrimination data, and trains the image conversion model using the discrimination data and the first class information. The discrimination data indicates, for each of a plurality of partial regions included in an image input to the discrimination model, whether or not the partial region is a fake image region, and indicates a class of the partial region when the partial region is not a fake image.


