Frame Interpolation Quality Metrics for Temporal Artifact Detection

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing video frame interpolation techniques introduce artifacts that degrade video quality, and current metrics like PSNR and SSIM fail to accurately assess these artifacts, particularly those related to temporal changes.

Innovation Solution

A system that uses a feature extraction network trained on human perception to detect video frame interpolation artifacts, incorporating a spatio-temporal network to analyze spatial and temporal features across multiple levels, generating a frame interpolation score that aligns with human perception.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If traditional metrics like PSNR or SSIM are used to assess video quality, then the assessment process is simple and computationally efficient, but the metrics fail to accurately evaluate artifacts introduced by video frame interpolation, particularly temporal artifacts

Engineering Contradiction:
Improvequality assessment accuracyVSAvoidassessment system complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent introduces an intermediary deep learning-based assessment system that bridges the gap between simple traditional metrics and human perception accuracy. This intermediary system uses trained neural networks to detect frame interpolation artifacts, achieving human-perception-aligned quality assessment without requiring direct human evaluation while maintaining computational feasibility through optimized network architectures.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent replaces the mechanical/mathematical computation of traditional metrics (PSNR, SSIM) with a learning-based system that substitutes rigid mathematical formulas with adaptive neural network models. This substitution enables the system to capture complex temporal artifacts and human perception characteristics that cannot be represented by fixed mathematical equations, significantly improving measurement precision for frame interpolation quality assessment.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Measurement precision

If deep learning-based metrics are used to accurately detect frame interpolation artifacts, then measurement precision improves, but computational complexity and processing time increase

Engineering Contradiction:
Improveartifact detection accuracyVSAvoidprocessing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent applies preliminary action by pre-training deep learning models on large datasets of frame interpolation artifacts before deployment. This pre-training phase captures common artifact patterns and temporal characteristics in advance, allowing the deployed system to perform rapid inference on new videos without requiring extensive real-time computation, thus reducing processing time while maintaining high detection accuracy.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent implements dynamics by designing adaptive processing pipelines that adjust computational depth based on video characteristics. The system dynamically selects assessment strategies - using lighter models for real-time applications and heavier models for offline analysis - allowing flexible trade-off between processing speed and detection precision depending on specific application requirements.

Inventive Principle:
Principle #15Dynamics

Data Source

PatentUS20250285253A1Video quality metric for frame interpolated content
Publication Date: 2025.09.11 DISNEY ENTERPRISES INC
  • US20250285253A1 patent drawing
  • US20250285253A1 patent drawing
  • US20250285253A1 patent drawing

AI summary

In some embodiments, a method receives a first video. The first video includes frames that were generated using frame interpolation. A feature extractor extracts first features from frames of the first video. The first features are extracted from a plurality of levels of a network of the feature extractor. A spatio-temporal processing system analyzes the first features spatially and temporally to determine spatial and temporal features for the plurality of levels. The method combines the spatial and temporal features from the plurality of levels to determine a score that measures a quality of the first video.