Comb Neural Network for High-Resolution Image Synthesis

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional image transfer methods produce low-resolution images with heavy artifacting and are time-consuming and costly, requiring manual processes and careful structuring of filmed scenes.

Innovation Solution

The use of a comb neural network architecture for automated image synthesis, which employs a deep learning model that encodes target and source images in a shared latent space and splits into specialist decoders to transfer physical characteristics and behaviors, enabling high-resolution image transfer without human intervention.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Manufacturing precision

If conventional image transfer methods are used, then the process is simpler, but the image resolution is low and artifacting is heavy

Engineering Contradiction:
Improveimage resolutionVSAvoidprocess complexity
Core Design Contradiction:
Manufacturing precisionVSDevice complexity

Solution Approach 1:

The neural network is divided into multiple specialized components including an encoder, multiple decoders for different subjects, and style transfer modules. Each component handles specific aspects of the image transformation, allowing high-resolution output without requiring a single monolithic complex system.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces latent space representations as an intermediary between source and target images. The encoder transforms input images into latent representations, which are then processed by specialized decoders to produce high-resolution output, effectively mediating the complex transformation process.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Manufacturing precision

If manual processes are used for high-resolution image transfer, then image resolution is higher, but the process is time-consuming and costly

Engineering Contradiction:
Improveimage resolutionVSAvoidprocessing speed
Core Design Contradiction:
Manufacturing precisionVSProductivity

Solution Approach 1:

The system performs automated alignment, encoding, and decoding without human intervention. The neural network self-adjusts through training to handle the complex transformations, eliminating the need for manual structuring of filmed scenes and placement of physical landmarks that characterize conventional methods.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent employs progressive training that gradually increases resolution parameters from low to high. Style-matching constraints dynamically adjust during training to balance fidelity and natural appearance, enabling the system to achieve high resolution automatically without manual intervention at each stage.

Inventive Principle:
Principle #35Parameter changes

3Reliability

If manual fitting of computer generated likeness is used, then uncanny aesthetic effect is reduced, but the process requires painstaking manual work

Engineering Contradiction:
Improveaesthetic qualityVSAvoidautomation level
Core Design Contradiction:
ReliabilityVSExtent of automation

Solution Approach 1:

The system uses style-matching constraints that provide feedback during the decoding process to ensure the generated images maintain natural aesthetic qualities. The loss functions monitor and adjust style consistency, enabling automated generation of high-quality results without manual fitting.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The patent employs dynamic style transfer that adapts during the decoding process. The specialized decoders dynamically adjust their output to match the style characteristics of the target subject, automatically achieving natural-looking results without static manual fitting procedures.

Inventive Principle:
Principle #15Dynamics

4Productivity

If conventional image transfer is used, then the process is faster, but image resolution is low and artifacts are present

Engineering Contradiction:
Improveprocessing efficiencyVSAvoidimage quality
Core Design Contradiction:
ProductivityVSManufacturing precision

Solution Approach 1:

The processing pipeline is segmented into specialized stages (encoding, latent representation, specialized decoding) that can be efficiently optimized independently. This segmentation allows the system to maintain high processing efficiency while achieving superior image quality at each stage.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The encoder performs preliminary extraction of essential features and transforms them into compact latent representations before the decoding stage. This preliminary action reduces the computational burden on subsequent stages while preserving the information needed for high-resolution reconstruction.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS10902571B2Automated image synthesis using a comb neural network architecture
Publication Date: 2021.01.26 DISNEY ENTERPRISES INC
  • US10902571B2 patent drawing
  • US10902571B2 patent drawing
  • US10902571B2 patent drawing

AI summary

An image synthesis system includes a computing platform having a hardware processor and a system memory storing a software code including a neural encoder and multiple neural decoders each corresponding to a respective persona. The hardware processor executes the software code to receive target image data, and source data that identifies one of the personas, and to map the target image data to its latent space representation using the neural encoder. The software code further identifies one of the neural decoders for decoding the latent space representation of the target image data based on the persona identified by the source data, uses the identified neural decoder to decode the latent space representation of the target image data as the persona identified by the source data to produce a swapped image data, and blends the swapped image data with the target image data to produce one or more synthesized images.