Image Inpainting via Multi-scale Sketch Tensor Network

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing image inpainting methods struggle to preserve the visual patterns of man-made scenes, particularly edges, lines, and junctions, leading to incomplete structure reconstruction and unreliable pattern transfer, often resulting in blurry or broken line segments and loss of connectivity in building structures.

Innovation Solution

A Multi-scale Sketch Tensor (MST) network is proposed, utilizing a novel encoder-decoder structure that learns a Sketch Tensor space by extracting line and edge maps through an improved wireframe parser and canny detector, and employs Pyramid Decomposing Separable blocks and efficient attention modules to enhance the reconstruction of holistic structures.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Manufacturing precision

If traditional synthesis-based approaches or basic learning-based methods are used for image inpainting, then the processing speed is faster and computation is simpler, but the critical structures (edges, lines, junctions) are missing or broken in the inpainted regions

Engineering Contradiction:
Improvestructure reconstruction accuracyVSAvoidnetwork complexity
Core Design Contradiction:
Manufacturing precisionVSDevice complexity

Solution Approach 1:

The network is segmented into three specialized modules: wireframe parser for holistic structure extraction, canny detector for local edge detection, and MST network for inpainting. Each module focuses on specific structural features, enabling comprehensive structure preservation without requiring a single complex network to handle all aspects simultaneously.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent embeds multiple detection mechanisms within the network architecture. The wireframe parser and canny detector outputs are nested as auxiliary inputs to the MST network, creating a hierarchical structure where simple detectors are contained within the more complex inpainting network, allowing layered feature integration.

Inventive Principle:
Principle #7Nested doll (Nesting)

2Manufacturing precision

If auxiliary information (edges, segmentation) is used to support inpainting, then local visual cues are improved, but the holistic structures of man-made scenes are lost and computation increases

Engineering Contradiction:
Improvelocal visual cue qualityVSAvoidholistic structure information
Core Design Contradiction:
Manufacturing precisionVSLoss of information

Solution Approach 1:

The patent merges two complementary auxiliary information sources: wireframe parsing results (holistic structure) and Canny edge detection (local details). By combining these different types of structural information and feeding them together into the MST network, both holistic and local structural features are preserved simultaneously in the inpainting process.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The MST network is designed to process multiple types of input simultaneously: the original masked image, wireframe structure maps, and Canny edge maps. This multi-functional input processing enables the network to handle both holistic structural reconstruction and local visual detail recovery within a single unified architecture.

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Manufacturing precision

If auxiliary-based inpainting methods are used, then structure guidance is improved, but computation time and resources increase due to additional components

Engineering Contradiction:
Improvestructure guidance qualityVSAvoidinpainting efficiency
Core Design Contradiction:
Manufacturing precisionVSProductivity

Solution Approach 1:

The wireframe parsing and Canny edge detection are performed as preliminary steps before the main inpainting process. By extracting structural features in advance and preparing auxiliary maps beforehand, the MST network receives pre-processed structural guidance, reducing its computational burden during the actual inpainting operation and improving overall efficiency.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The auxiliary structural maps (wireframe and Canny edges) serve as intermediary representations that bridge the gap between the masked input and the desired output. These intermediate structure maps provide guidance to the MST network without requiring direct complex interactions between all input elements, streamlining the computational process.

Inventive Principle:
Principle #24Intermediary (Mediator)

4Manufacturing precision

If learning-based methods with auxiliary detectors are used, then local edge and gradient information is improved, but unreliable image priors are transferred and magnified to masked regions causing degraded results

Engineering Contradiction:
Improveedge and gradient informationVSAvoidinpainted result reliability
Core Design Contradiction:
Manufacturing precisionVSReliability

Solution Approach 1:

The network employs multiple loss functions that provide feedback during training: L1 loss for pixel-level accuracy, perceptual loss for structural fidelity, and adversarial loss for realism. This multi-faceted feedback mechanism guides the network to learn reliable structure priors from unmasked regions while preventing the propagation of unreliable patterns to masked areas, ensuring robust generalization.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS11580622B2System and method for image inpainting
Publication Date: 2023.02.14 FUDAN UNIVERSITY
  • US11580622B2 patent drawing
  • US11580622B2 patent drawing
  • US11580622B2 patent drawing

AI summary

A system for image inpainting is provided, including an encoder, a decoder, and a sketch tensor space of a third-order tensor; wherein the encoder includes an improved wireframe parser and a canny detector, and a pyramid structure sub-encoder; the improved wireframe parser is used to extract line maps from an original image input to the encoder, the canny detector is used to extract edge maps from the original image, and the pyramid structure sub-encoder is used to generate the sketch tensor space based on the original image, the line maps and the edge maps; and the decoder outputs an inpainted image from the sketch tensor space. A method thereof is also provided.