Image Normalization via Landmark Warping for Cross-Modal Recognition
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current processor-implemented neural network models for image recognition face challenges in accurately normalizing and comparing images with different fields of view and modalities, such as color and depth images, which affects the reliability of object recognition across varying input formats.
Innovation Solution
An image normalization method and apparatus that extracts object patches from input images, determines landmarks, and normalizes these patches to a standardized format, allowing for effective recognition using a trained object recognition model, even when images have different fields of view or modalities, by warping and aligning landmarks within a predetermined format.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If images with different fields of view and modalities are used for object recognition, then the versatility and adaptability of the recognition system is improved, but the measurement precision and reliability of object recognition deteriorates due to normalization difficulties
Solution Approach 1:
The patent applies parameter changes by transforming images with different fields of view and modalities into a standardized format through normalization. Specifically, it converts depth images to RGB format, adjusts image resolutions to match reference images, and transforms 3D spatial coordinates to 2D image coordinates using projection equations. This allows the system to process diverse input formats (different cameras, resolutions, modalities) while maintaining recognition accuracy by ensuring all images are in a comparable standardized form before feature extraction and matching.
2Reliability
If multiple input images with different modalities are processed, then the robustness of the recognition process is improved, but the device complexity increases due to additional normalization steps
Solution Approach 1:
The patent implements universality by creating a multi-functional normalization system that handles various image modalities (color, depth, infrared) and formats through a single unified process. The normalization apparatus can process reference images and candidate images from different sources (first camera, second camera, depth sensor) using the same normalization workflow: modality conversion, resolution adjustment, and coordinate transformation. This universal approach improves reliability by ensuring consistent processing regardless of input source, while managing complexity through a standardized multi-step process that can be implemented as a single integrated system.
Data Source
AI summary
A processor-implemented image normalization method includes extracting a first object patch from a first input image and extracting a second object patch from a second input image based on an object area that includes an object detected from any one or any combination of the first input image and the second input image, determining, based on a first landmark detected from the first object patch, a second landmark of the second object patch; and normalizing the first object patch and the second object patch based on the first landmark and the second landmark.


