Neural Network Learning Using Dual Evaluation Functions for CG to Real Image Adaptation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing object recognition systems face challenges in achieving robust recognition results when using computer-generated (CG) images as learning data, especially in scenarios where actual images are scarce, due to differences in image characteristics such as edge enhancements and noise levels.
Innovation Solution
An information processing apparatus and method that utilizes a neural network to perform learning by adjusting weighting coefficients based on evaluation functions, reducing differences between recognition results and training data, as well as intermediate outputs from actual and CG images, to enhance the robustness of inference results.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If CG images are used as learning data, then the coverage of learning scenarios is improved, but the recognition accuracy on actual images deteriorates due to image characteristic differences
Solution Approach 1:
The patent applies parameter changes by modifying the neural network's weighting coefficients through dual evaluation functions. The first evaluation function adjusts weights based on recognition accuracy, while the second evaluation function adjusts weights based on intermediate output differences between CG and actual images. This dynamic parameter adjustment allows the system to adapt to both CG and actual image characteristics, resolving the contradiction between scenario coverage and recognition accuracy.
Solution Approach 2:
The patent introduces intermediate outputs from the hidden layer as a mediator to bridge CG images and actual images. By comparing and minimizing differences in intermediate outputs between CG and actual images through the second evaluation function, the system creates a common representation space that allows the neural network to generalize better from CG images to actual images, thereby improving recognition accuracy while maintaining scenario coverage.
2Measurement precision
If weighting coefficients are adjusted to reduce recognition result differences, then recognition accuracy is improved, but the difference between CG and actual image processing deteriorates
Solution Approach 1:
The patent segments the learning objective into two separate evaluation functions: the first evaluation function focuses on recognition accuracy by comparing recognition results with training data, while the second evaluation function focuses on generalization capability by comparing intermediate outputs between CG and actual images. This segmentation allows independent optimization of both objectives through separate weighting coefficient adjustments, resolving the contradiction between recognition accuracy and generalization capability.
3Reliability
If monochromatic conversion and contrast adjustment are applied, then robustness is improved, but the problem of CG image characteristics is not addressed
Solution Approach 1:
The patent applies dynamics by making the weighting coefficients dynamic rather than fixed. The weighting coefficients are continuously adjusted during learning based on feedback from both evaluation functions. This dynamic adjustment allows the system to adapt to different image types (CG or actual) and their respective characteristics, making the robustness improvement applicable to both CG and actual images rather than being limited to specific preprocessing techniques.
Data Source
AI summary
An information processing apparatus recognizes a target within an actual image by executing processing of a neural network. The information processing apparatus obtains intermediate outputs which correspond to the actual image and a computer graphics (CG) image and which are from a hidden layer when each of the actual image and the CG image has been separately input to the neural network, and causes the neural network to perform learning with use of an evaluation values based on a first evaluation function and a second evaluation function, the first evaluation function causing the evaluation value to decrease as a difference between a recognition result and training data decreases, the second evaluation function causing the evaluation value to decrease as a difference between the intermediate outputs corresponding to the actual image and the CG image decreases.


