Single-Image 3D Photography With Soft Layering for Thin Object Boundaries

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing methods for generating 3D images from monocular images struggle with accurately modeling appearance effects of thin objects and perform poorly with out-of-distribution scenes due to reliance on hard discontinuities and limited training datasets.

Innovation Solution

A 3D photo system using soft-layering formulations and depth-aware inpainting to decompose scenes into foreground and background layers, generating soft visibility maps and disocclusion masks, and inpainting disoccluded regions to create consistent background layers.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Manufacturing precision

If hard discontinuities are used to separate foreground and background layers, then the layer separation is sharp and well-defined, but the appearance effects of thin objects are not accurately modeled and out-of-distribution scenes are handled poorly

Engineering Contradiction:
Improvelayer separation precisionVSAvoidhandling of out-of-distribution scenes
Core Design Contradiction:
Manufacturing precisionVSAdaptability or versatility

Solution Approach 1:

The patent changes the parameter of layer separation from hard discontinuities (binary classification) to soft-layering formulations (continuous probability distributions). This allows the model to represent uncertain boundaries and thin objects more accurately, improving adaptability to out-of-distribution scenes while maintaining sufficient layer separation precision through the soft masking approach.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent applies local quality by using depth-aware inpainting that adapts to local scene characteristics. The soft-layering formulation allows different regions of the image to have different separation qualities based on depth information, enabling accurate handling of thin objects in specific locations while maintaining overall scene consistency.

Inventive Principle:
Principle #3Local quality

2Measurement precision

If multiple views are used to perform fully automatic alpha matting and refine stereo depths, then the boundary accuracy is improved, but additional optical hardware and multiple views are required

Engineering Contradiction:
Improveboundary accuracyVSAvoidoptical hardware requirements
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent uses a monocular depth estimation model to generate a depth map as a copy of depth information from a single image. This depth map is then used to guide the soft-layering formulation and inpainting process, achieving boundary accuracy similar to multi-view methods without requiring additional cameras or multiple views.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent replaces the mechanical system of multiple cameras and complex multi-view geometry processing with a computational approach using monocular depth estimation and soft-layering formulations. This substitution maintains boundary accuracy while significantly reducing device complexity.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Manufacturing precision

If depth maps are used to improve matting in postproduction, then the matting quality is enhanced, but the process requires separate postproduction steps and increased processing complexity

Engineering Contradiction:
Improvematting qualityVSAvoidprocessing pipeline complexity
Core Design Contradiction:
Manufacturing precisionVSDevice complexity

Solution Approach 1:

The patent merges depth map generation, soft-layering formulation, and inpainting into a single integrated end-to-end system. The depth-aware soft-layering model processes the monocular image through all these steps in one unified pipeline, achieving enhanced matting quality without requiring separate postproduction steps.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent creates a universal model that performs multiple functions simultaneously: depth estimation, layer separation, boundary refinement, and inpainting. This multi-functional approach enhances matting quality while reducing processing pipeline complexity by eliminating the need for separate specialized steps.

Inventive Principle:
Principle #6Universality (Multi-functionality)

4Productivity

If layered texture representations are used for efficient storage and view-dependent rendering, then video-realism is maintained with efficient storage, but the system requires multiple views and complex layered representations

Engineering Contradiction:
Improvestorage efficiencyVSAvoidrepresentation complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent creates a compact 3D representation by copying and processing a single monocular image through depth estimation and soft-layering. This produces a space-efficient representation that enables view-dependent rendering without requiring storage of multiple original views or complex layered texture data structures.

Inventive Principle:
Principle #26Copying

Data Source

PatentEP4283566B1Single image 3D photography with soft-layering and depth-aware inpainting
Publication Date: 2026.03.25 GOOGLE LLC
  • EP4283566B1 patent drawingFigure 1
  • EP4283566B1 patent drawingFigure 2
  • EP4283566B1 patent drawingFigure 3A

AI summary

A computer-implemented method comprising: determining, for each respective pixel of a plurality of pixels of a depth image, a corresponding depth gradient associated with the respective pixel of the depth image, wherein the depth image corresponds to a monocular image having an initial viewpoint, and wherein each respective pixel of the plurality of pixels of the depth image has a corresponding depth value, determining a foreground visibility map comprising, for each respective pixel of the plurality of pixels of the depth image, a visibility value that is inversely proportional to the corresponding depth gradient associated with the respective pixel of the depth image, generating (i) an inpainted image by inpainting portions of the monocular image based on a background disocclusion mask that indicates, for each respective pixel of the plurality of pixels of the depth image, a disocclusion likelihood that a corresponding pixel of the monocular image will be disoccluded by a change in the initial viewpoint and (ii) an inpainted depth image by inpainting portions of the depth image based on the background disocclusion mask, and generating a modified image having an adjusted viewpoint that is different from the initial viewpoint by combining visual information of the monocular image and the inpainted image in accordance with the depth image, the inpainted depth image, and the foreground visibility map.