Image Normalization via Landmark Warping for Cross-Modal Recognition

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current processor-implemented neural network models for image recognition face challenges in accurately normalizing and comparing images with different fields of view and modalities, such as color and depth images, which affects the reliability of object recognition across varying input formats.

Innovation Solution

An image normalization method and apparatus that extracts object patches from input images, determines landmarks, and normalizes these patches to a standardized format, allowing for effective recognition using a trained object recognition model, even when images have different fields of view or modalities, by warping and aligning landmarks within a predetermined format.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If images with different fields of view and modalities are used for object recognition, then the versatility and adaptability of the recognition system is improved, but the measurement precision and reliability of object recognition deteriorates due to normalization difficulties

Engineering Contradiction:
Improveability to process diverse input formatsVSAvoidobject recognition accuracy
Core Design Contradiction:
Adaptability or versatilityVSMeasurement precision

Solution Approach 1:

The patent applies parameter changes by transforming images with different fields of view and modalities into a standardized format through normalization. Specifically, it converts depth images to RGB format, adjusts image resolutions to match reference images, and transforms 3D spatial coordinates to 2D image coordinates using projection equations. This allows the system to process diverse input formats (different cameras, resolutions, modalities) while maintaining recognition accuracy by ensuring all images are in a comparable standardized form before feature extraction and matching.

Inventive Principle:
Principle #35Parameter changes

2Reliability

If multiple input images with different modalities are processed, then the robustness of the recognition process is improved, but the device complexity increases due to additional normalization steps

Engineering Contradiction:
Improverobustness of recognition processVSAvoidcomplexity of normalization process
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent implements universality by creating a multi-functional normalization system that handles various image modalities (color, depth, infrared) and formats through a single unified process. The normalization apparatus can process reference images and candidate images from different sources (first camera, second camera, depth sensor) using the same normalization workflow: modality conversion, resolution adjustment, and coordinate transformation. This universal approach improves reliability by ensuring consistent processing regardless of input source, while managing complexity through a standardized multi-step process that can be implemented as a single integrated system.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS11475537B2Method and apparatus with image normalization
Publication Date: 2022.10.18 SAMSUNG ELECTRONICS CO LTD
  • US11475537B2 patent drawing
  • US11475537B2 patent drawing
  • US11475537B2 patent drawing

AI summary

A processor-implemented image normalization method includes extracting a first object patch from a first input image and extracting a second object patch from a second input image based on an object area that includes an object detected from any one or any combination of the first input image and the second input image, determining, based on a first landmark detected from the first object patch, a second landmark of the second object patch; and normalizing the first object patch and the second object patch based on the first landmark and the second landmark.