Image Identification with Aspect Ratio Distortion and Letterbox Removal

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing image identification technologies struggle to effectively identify attribute information of entire images due to cropping processes that eliminate outside regions, leading to incorrect training and recognition issues.

Innovation Solution

An image identifying apparatus that generates test image data through resizing with predetermined aspect ratio distortion, using a machine learning model trained with diverse aspect ratio distortions, and performs mask processing or erases letterbox information to improve identification accuracy.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If cropping processes are applied to image data for identification, then processing load is reduced, but identification accuracy of entire images deteriorates due to elimination of outside regions

Engineering Contradiction:
Improveprocessing loadVSAvoididentification accuracy
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The patent extracts and removes letterbox information (black border regions) from image data before processing. This extraction principle eliminates unnecessary regions that would interfere with identification while preserving the meaningful content areas, thereby reducing processing load without sacrificing identification accuracy of the entire image.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent applies aspect ratio distortion transformation to convert images with letterbox information into standardized aspect ratios. This dimensional transformation allows the system to process images of varying original dimensions uniformly, reducing processing complexity while maintaining the ability to identify entire image content through comprehensive training on distorted variants.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Loss of information

If letterbox information is retained in image data, then original image composition is preserved, but identification accuracy deteriorates due to interference with attribute recognition

Engineering Contradiction:
Improveimage composition preservationVSAvoidattribute identification accuracy
Core Design Contradiction:
Loss of informationVSMeasurement precision

Solution Approach 1:

The patent specifically extracts and removes letterbox information from image data before feeding it to the identification system. This extraction eliminates the interfering black border regions that would otherwise distort attribute recognition, while the removal process itself is designed to preserve the compositional integrity of the actual image content.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent performs preliminary processing to detect and remove letterbox information before the main identification process. This preliminary action prepares the image data in advance, ensuring that the identification system receives clean input without interfering elements, thereby improving attribute recognition accuracy without losing meaningful image composition.

Inventive Principle:
Principle #10Preliminary action

3Adaptability or versatility

If machine learning model is trained with diverse aspect ratio distortions, then adaptability to different image formats improves, but training complexity increases

Engineering Contradiction:
Improveformat adaptabilityVSAvoidtraining complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent applies parameter changes by systematically transforming training images with various aspect ratio distortions. This allows the machine learning model to learn invariant features across different image formats and aspect ratios, enhancing format adaptability. The parameter transformation approach is more efficient than training with completely different image sets, balancing versatility gains with manageable training complexity.

Inventive Principle:
Principle #35Parameter changes

4Measurement precision

If mask processing is applied to character regions before resizing, then character recognition accuracy improves, but processing steps increase

Engineering Contradiction:
Improvecharacter recognition accuracyVSAvoidprocessing steps
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent segments the image processing into distinct steps: first applying mask processing to character regions to preserve their integrity, then performing resizing on the masked image. This segmentation allows character regions to be handled specially with appropriate masking, improving character recognition accuracy while keeping the overall processing pipeline organized and manageable through clear step separation.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS12205352B2Image identifying apparatus, video reproducing apparatus, image identifying method, and recording medium
Publication Date: 2025.01.21 PANASONIC INTELLECTUAL PROPERTY MANAGEMENT CO LTD
  • US12205352B2 patent drawing
  • US12205352B2 patent drawing
  • US12205352B2 patent drawing

AI summary

An image identifying apparatus includes: an obtainer that obtains image data; an image processor that generates test image data by performing resizing to reduce the image data with predetermined aspect ratio distortion; a storage unit that stores a machine learning model used to identify attribute information of the test image data; and an identifier that identifies the attribute information of the test image data, using the machine learning model. The machine learning model includes trained parameters that have been adjusted through machine learning using a training data set including items of second training image data obtained through application of one or more types of aspect ratio distortion including the predetermined aspect ratio distortion to each of items of first training image data, and items of attribute information associated with the items of second training image data.