Image Inpainting via Multi-scale Sketch Tensor Network
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing image inpainting methods struggle to preserve the visual patterns of man-made scenes, particularly edges, lines, and junctions, leading to incomplete structure reconstruction and unreliable pattern transfer, often resulting in blurry or broken line segments and loss of connectivity in building structures.
Innovation Solution
A Multi-scale Sketch Tensor (MST) network is proposed, utilizing a novel encoder-decoder structure that learns a Sketch Tensor space by extracting line and edge maps through an improved wireframe parser and canny detector, and employs Pyramid Decomposing Separable blocks and efficient attention modules to enhance the reconstruction of holistic structures.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Manufacturing precision
If traditional synthesis-based approaches or basic learning-based methods are used for image inpainting, then the processing speed is faster and computation is simpler, but the critical structures (edges, lines, junctions) are missing or broken in the inpainted regions
Solution Approach 1:
The network is segmented into three specialized modules: wireframe parser for holistic structure extraction, canny detector for local edge detection, and MST network for inpainting. Each module focuses on specific structural features, enabling comprehensive structure preservation without requiring a single complex network to handle all aspects simultaneously.
Solution Approach 2:
The patent embeds multiple detection mechanisms within the network architecture. The wireframe parser and canny detector outputs are nested as auxiliary inputs to the MST network, creating a hierarchical structure where simple detectors are contained within the more complex inpainting network, allowing layered feature integration.
2Manufacturing precision
If auxiliary information (edges, segmentation) is used to support inpainting, then local visual cues are improved, but the holistic structures of man-made scenes are lost and computation increases
Solution Approach 1:
The patent merges two complementary auxiliary information sources: wireframe parsing results (holistic structure) and Canny edge detection (local details). By combining these different types of structural information and feeding them together into the MST network, both holistic and local structural features are preserved simultaneously in the inpainting process.
Solution Approach 2:
The MST network is designed to process multiple types of input simultaneously: the original masked image, wireframe structure maps, and Canny edge maps. This multi-functional input processing enables the network to handle both holistic structural reconstruction and local visual detail recovery within a single unified architecture.
3Manufacturing precision
If auxiliary-based inpainting methods are used, then structure guidance is improved, but computation time and resources increase due to additional components
Solution Approach 1:
The wireframe parsing and Canny edge detection are performed as preliminary steps before the main inpainting process. By extracting structural features in advance and preparing auxiliary maps beforehand, the MST network receives pre-processed structural guidance, reducing its computational burden during the actual inpainting operation and improving overall efficiency.
Solution Approach 2:
The auxiliary structural maps (wireframe and Canny edges) serve as intermediary representations that bridge the gap between the masked input and the desired output. These intermediate structure maps provide guidance to the MST network without requiring direct complex interactions between all input elements, streamlining the computational process.
4Manufacturing precision
If learning-based methods with auxiliary detectors are used, then local edge and gradient information is improved, but unreliable image priors are transferred and magnified to masked regions causing degraded results
Solution Approach 1:
The network employs multiple loss functions that provide feedback during training: L1 loss for pixel-level accuracy, perceptual loss for structural fidelity, and adversarial loss for realism. This multi-faceted feedback mechanism guides the network to learn reliable structure priors from unmasked regions while preventing the propagation of unreliable patterns to masked areas, ensuring robust generalization.
Data Source
AI summary
A system for image inpainting is provided, including an encoder, a decoder, and a sketch tensor space of a third-order tensor; wherein the encoder includes an improved wireframe parser and a canny detector, and a pyramid structure sub-encoder; the improved wireframe parser is used to extract line maps from an original image input to the encoder, the canny detector is used to extract edge maps from the original image, and the pyramid structure sub-encoder is used to generate the sketch tensor space based on the original image, the line maps and the edge maps; and the decoder outputs an inpainted image from the sketch tensor space. A method thereof is also provided.


