Machine-Learned Image Transformation Model for Feature Interpolation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing machine-learned models for image transformation struggle with fine-tuned control of specific features, requiring in-depth annotation of training datasets, which is prohibitively difficult and inefficient, especially for tasks like pose interpolation.

Innovation Solution

A computer-implemented method using a machine-learned image transformation model to generate interpolation vectors for selective, fine-grained image transformation and retrieval by processing source and reference images to create channel mappings, allowing for interpolation of specific features without complex dataset annotation.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Manufacturing precision

If in-depth annotation of training datasets is performed to achieve fine-tuned control of specific features, then feature control precision is improved, but annotation complexity and time consumption increase prohibitively

Engineering Contradiction:
Improvefeature control precisionVSAvoidannotation time consumption
Core Design Contradiction:
Manufacturing precisionVSLoss of time

Solution Approach 1:

The patent introduces channel mappings as an intermediary between images and feature controls. Instead of directly annotating features in training data, the system processes images through a machine-learned model to generate channel mappings that automatically encode feature information. This intermediary mechanism eliminates the need for manual feature annotation while preserving fine-grained control capability.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent replaces the mechanical annotation process with an automated computational process. The machine-learned image transformation model automatically extracts and encodes feature information from images through channel mappings, substituting the manual mechanical annotation process with an efficient computational pipeline that generates the same control information without human intervention.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Reliability

If large datasets are processed to average channel contributions, then model robustness is improved, but computational resource consumption increases

Engineering Contradiction:
Improvemodel robustnessVSAvoidcomputational resource consumption
Core Design Contradiction:
ReliabilityVSUse of energy by moving object

Solution Approach 1:

The patent extracts only the necessary feature information from images through channel mappings, rather than processing entire large datasets. By taking out and processing only the relevant channel contributions needed for specific feature control, the system achieves robust results with significantly reduced computational resource consumption compared to processing complete datasets.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent applies partial action by processing only the specific channel mappings required for the desired feature control, rather than exhaustively processing all available training data. This partial processing approach provides sufficient model robustness for the intended application while avoiding the excessive computational costs of complete dataset processing.

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS12008821B2Machine-learned models for unsupervised image transformation and retrieval
Publication Date: 2024.06.11 GOOGLE LLC
  • US12008821B2 patent drawing
  • US12008821B2 patent drawing
  • US12008821B2 patent drawing

AI summary

Systems and methods of the present disclosure are directed to a computer-implemented method. The method can include obtaining a first image depicting a first object and a second image depicting a second object, wherein the first object comprises a first feature set and the second object comprises a second feature set. The method can include processing the first image with a machine-learned image transformation model comprising a plurality of model channels to obtain a first channel mapping indicative of a mapping between the plurality of model channels and the first feature set. The method can include processing the second image with the model to obtain a second channel mapping indicative of a mapping between the plurality of model channels and the second feature set. The method can include generating an interpolation vector for a selected feature.