Local Implicit Normalizing Flow for Arbitrary-Scale Image Super-Resolution

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing arbitrary-scale image super-resolution methods produce blurry outputs due to ignoring the ill-posed problem and using per-pixel absolute loss, limiting their flexibility and quality in real-world applications.

Innovation Solution

The Local Implicit Normalizing Flow (LINF) framework models the distribution of texture details under different scaling factors using a coordinate conditional normalizing flow and local implicit module, enabling the generation of photorealistic high-resolution images by learning the distribution of local texture patches and controlling sampling temperature for improved fidelity and perceptual quality.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If arbitrary-scale SR methods use per-pixel absolute loss to train the model, then the model can perform arbitrary-scale super-resolution, but the output images become blurry

Engineering Contradiction:
Improvearbitrary-scale super-resolution capabilityVSAvoidimage sharpness and detail fidelity
Core Design Contradiction:
Adaptability or versatilityVSManufacturing precision

Solution Approach 1:

The patent changes the loss function parameter from per-pixel absolute loss to perceptual loss (LPIPS) and incorporates style loss, transforming the optimization objective to prioritize perceptual quality and texture fidelity over pixel-level accuracy, thereby producing sharp images with rich details while maintaining arbitrary-scale capability

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent introduces perceptual loss feedback mechanisms that evaluate image quality through learned feature representations rather than direct pixel comparison, enabling the model to iteratively improve texture synthesis and sharpness while maintaining the desired scale flexibility

Inventive Principle:
Principle #23Feedback

2Reliability

If flow-based methods learn the distribution of high-resolution images with normalizing flow, then the method can address the ill-posed nature of super-resolution, but the method is limited to predefined fixed-scale SR

Engineering Contradiction:
Improvehandling of ill-posed super-resolution problemVSAvoidflexibility in output scale
Core Design Contradiction:
ReliabilityVSAdaptability or versatility

Solution Approach 1:

The patent makes the normalizing flow model dynamic by introducing scale-conditioned transformations where the flow parameters are modulated by the target scale factor, allowing the same model architecture to adaptively handle arbitrary scaling operations while maintaining reliable ill-posed problem solving through distribution learning

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The patent creates a universal super-resolution model that can perform multiple scaling operations (2x, 4x, 8x, and arbitrary scales) using a single trained normalizing flow framework, eliminating the need for separate fixed-scale models while maintaining reliable reconstruction quality across all scales

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS20240177269A1Method of local implicit normalizing flow for arbitrary-scale image super-resolution, and associated apparatus
Publication Date: 2024.05.30 MEDIATEK INC
  • US20240177269A1 patent drawing
  • US20240177269A1 patent drawing
  • US20240177269A1 patent drawing

AI summary

A method of local implicit normalizing flow for arbitrary-scale image super-resolution, an associated apparatus and an associated computer-readable medium are provided. The method applicable to a processing circuit may include: utilizing the processing circuit to run a local implicit normalizing flow framework to start performing arbitrary-scale image super-resolution with a trained model of the local implicit normalizing flow framework according to at least one input image, for generating at least one output image, where a selected scale of the output image with respect to the input image is an arbitrary-scale; and during performing the arbitrary-scale image super-resolution with the trained model, performing prediction processing to obtain multiple super-resolution predictions for different locations of a predetermined space in a situation where a same non-super-resolution input image among the at least one input image is given, in order to generate the at least one output image.