Deep Learning Video Remastering via Paired Frame Mapping

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional video remastering techniques fail to achieve the same visual quality when enhancing legacy video content without a master copy, as they rely on standard image enhancement tools, resulting in a perceptible difference compared to remastered content from original film scans.

Innovation Solution

A machine learning model is trained to learn a mapping between lower quality and higher quality video representations using paired frames, allowing it to convert lower quality video content into a higher quality version that is visually and stylistically consistent with the original, even in the absence of a master copy.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Manufacturing precision

If conventional image enhancement tools are used to improve legacy video content without a master copy, then some quality improvement is achieved, but the visual quality remains perceptibly inferior compared to content remastered from original film masters

Engineering Contradiction:
Improvevisual qualityVSAvoidconsistency with master copy remaster
Core Design Contradiction:
Manufacturing precisionVSReliability

Solution Approach 1:

The patent creates a digital copy of the film master at high resolution, then uses this digital copy as a reference to guide the enhancement of analog broadcast content. The system learns the mapping between the digital master and analog versions, applying this learned transformation to restore analog content to match the quality and appearance of master copy remasters, thereby achieving both quality improvement and consistency with original intent

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent transforms the enhancement process by changing from conventional fixed-parameter image processing to adaptive parameter adjustment based on learned mappings. The system modifies multiple parameters simultaneously (resolution, color accuracy, noise characteristics, sharpness) based on the specific characteristics of the input analog content and the target digital master style, enabling consistent reproduction across different legacy videos

Inventive Principle:
Principle #35Parameter changes

2Measurement precision

If deep learning models are trained on large datasets, then model accuracy improves, but training time and computational resources increase

Engineering Contradiction:
Improvemodel accuracyVSAvoidtraining time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent performs preliminary actions by pre-processing the training data to ensure proper alignment and pairing between analog and digital master frames. The system prepares synchronized frame pairs with matching temporal and spatial characteristics before training begins, which accelerates convergence and improves training efficiency. This preliminary preparation reduces the overall training time while maintaining high model accuracy

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent segments the training process into distinct stages: first training on frame-level correspondence, then progressing to temporal consistency, and finally optimizing for overall video quality. This segmented approach allows the model to learn complex transformations incrementally, reducing total training time compared to attempting to learn all aspects simultaneously while achieving superior final accuracy

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS20230267706A1Video remastering via deep learning
Publication Date: 2023.08.24 DISNEY ENTERPRISES INC
  • US20230267706A1 patent drawing
  • US20230267706A1 patent drawing
  • US20230267706A1 patent drawing

AI summary

One embodiment of the present invention sets forth a technique for performing remastering of video content. The technique includes determining a first input frame corresponding to a first frame included in a first video and a first target frame corresponding to a second frame included in a second video based on one or more alignments between the first frame and the second frame. The technique also includes executing a machine learning model to convert the first input frame into a first output frame. The technique further includes training the machine learning model based on one or more losses associated with the first output frame and the first target frame.