Image Identification with Aspect Ratio Distortion and Letterbox Removal
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing image identification technologies struggle to effectively identify attribute information of entire images due to cropping processes that eliminate outside regions, leading to incorrect training and recognition issues.
Innovation Solution
An image identifying apparatus that generates test image data through resizing with predetermined aspect ratio distortion, using a machine learning model trained with diverse aspect ratio distortions, and performs mask processing or erases letterbox information to improve identification accuracy.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If cropping processes are applied to image data for identification, then processing load is reduced, but identification accuracy of entire images deteriorates due to elimination of outside regions
Solution Approach 1:
The patent extracts and removes letterbox information (black border regions) from image data before processing. This extraction principle eliminates unnecessary regions that would interfere with identification while preserving the meaningful content areas, thereby reducing processing load without sacrificing identification accuracy of the entire image.
Solution Approach 2:
The patent applies aspect ratio distortion transformation to convert images with letterbox information into standardized aspect ratios. This dimensional transformation allows the system to process images of varying original dimensions uniformly, reducing processing complexity while maintaining the ability to identify entire image content through comprehensive training on distorted variants.
2Loss of information
If letterbox information is retained in image data, then original image composition is preserved, but identification accuracy deteriorates due to interference with attribute recognition
Solution Approach 1:
The patent specifically extracts and removes letterbox information from image data before feeding it to the identification system. This extraction eliminates the interfering black border regions that would otherwise distort attribute recognition, while the removal process itself is designed to preserve the compositional integrity of the actual image content.
Solution Approach 2:
The patent performs preliminary processing to detect and remove letterbox information before the main identification process. This preliminary action prepares the image data in advance, ensuring that the identification system receives clean input without interfering elements, thereby improving attribute recognition accuracy without losing meaningful image composition.
3Adaptability or versatility
If machine learning model is trained with diverse aspect ratio distortions, then adaptability to different image formats improves, but training complexity increases
Solution Approach 1:
The patent applies parameter changes by systematically transforming training images with various aspect ratio distortions. This allows the machine learning model to learn invariant features across different image formats and aspect ratios, enhancing format adaptability. The parameter transformation approach is more efficient than training with completely different image sets, balancing versatility gains with manageable training complexity.
4Measurement precision
If mask processing is applied to character regions before resizing, then character recognition accuracy improves, but processing steps increase
Solution Approach 1:
The patent segments the image processing into distinct steps: first applying mask processing to character regions to preserve their integrity, then performing resizing on the masked image. This segmentation allows character regions to be handled specially with appropriate masking, improving character recognition accuracy while keeping the overall processing pipeline organized and manageable through clear step separation.
Data Source
AI summary
An image identifying apparatus includes: an obtainer that obtains image data; an image processor that generates test image data by performing resizing to reduce the image data with predetermined aspect ratio distortion; a storage unit that stores a machine learning model used to identify attribute information of the test image data; and an identifier that identifies the attribute information of the test image data, using the machine learning model. The machine learning model includes trained parameters that have been adjusted through machine learning using a training data set including items of second training image data obtained through application of one or more types of aspect ratio distortion including the predetermined aspect ratio distortion to each of items of first training image data, and items of attribute information associated with the items of second training image data.


