Neural Network Image Style Translation via Unpaired Keyframes
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing image style transfer techniques face challenges in translating video sequences effectively, particularly as the sequence length increases, leading to degraded quality and requiring manual intervention, and often produce noticeable visual artifacts due to their statistical nature.
Innovation Solution
An image translation network is trained using a pair of keyframes and an unpaired image, employing multiple loss functions to learn how to translate images from a source visual domain to a target visual domain, ensuring temporal stability and preserving style consistency without explicit guidance.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Duration of action of stationary object
If optical flow tracking is used to transfer stylization through video frames, then style transfer can be achieved across multiple frames, but the output quality degrades as video sequence length increases, requiring manual keyframe updates
Solution Approach 1:
The patent introduces an unpaired image as an intermediary reference that mediates between the source and target visual domains. This unpaired image serves as a style guide that the neural network learns from, allowing consistent style transfer across long video sequences without relying on cumulative optical flow tracking. The unpaired image acts as a stable reference point that prevents quality degradation over time.
Solution Approach 2:
The patent replaces the mechanical optical flow tracking system with a neural network-based approach. Instead of tracking individual pixel movements through frames using optical flow algorithms, the system uses a neural network to learn the mapping between source and target visual domains from keyframes and unpaired images, then applies this learned transformation directly to video frames, eliminating the cumulative errors inherent in sequential optical flow tracking.
2Productivity
If neural network techniques are used for style transference, then processing speed is improved, but noticeable visual artifacts are produced due to statistical nature
Solution Approach 1:
The patent applies local quality by using patch-based synthesis within the neural network domain. Instead of processing entire images globally through statistical operations, the system divides images into patches and synthesizes them locally, preserving fine-grained details and reducing the visual artifacts typically produced by global statistical transformations. This allows the neural network to maintain high processing speed while improving output quality.
3Measurement precision
If guidance channels are prepared explicitly by user or generated algorithmically, then style transfer accuracy is improved, but the process becomes cumbersome and requires significant user input
Solution Approach 1:
The patent implements self-service by enabling the neural network to automatically learn the style mapping from unpaired images without requiring users to manually create or prepare guidance channels. The system autonomously extracts style information from the unpaired images and uses this learned representation to perform style transfer, eliminating the cumbersome manual guidance channel preparation process while maintaining high style transfer accuracy.
Data Source
AI summary
Embodiments are disclosed for translating an image from a source visual domain to a target visual domain. In particular, in one or more embodiments, the disclosed systems and methods comprise a training process that includes receiving a training input including a pair of keyframes and an unpaired image. The pair of keyframes represent a visual translation from a first version of an image in a source visual domain to a second version of the image in a target visual domain. The one or more embodiments further include sending the pair of keyframes and the unpaired image to an image translation network to generate a first training image and a second training image. The one or more embodiments further include training the image translation network to translate images from the source visual domain to the target visual domain based on a calculated loss using the first and second training images.


