3D Scene Shadow Generation for Semantic 2D Image Editing

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional image editing systems are inflexible and inefficient, requiring significant user interaction and specialized knowledge to edit digital images, as they operate on a pixel level rather than a semantic level, failing to maintain real-world conditions during edits.

Innovation Solution

A scene-based image editing system that utilizes machine learning to pre-process digital images, identifying objects, relationships, and attributes, allowing intuitive and efficient editing by treating semantic areas as distinct units, and maintaining real-world conditions without user input for preparatory steps.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If conventional image editing systems operate on a pixel level, then they can perform detailed image manipulation, but they require significant user interaction and specialized knowledge, reducing ease of operation

Engineering Contradiction:
Improveease of operationVSAvoiduser interaction time
Core Design Contradiction:
Ease of operationVSLoss of time

Solution Approach 1:

The image is segmented into multiple semantic regions using machine learning models that identify objects, backgrounds, and other meaningful areas. This allows the system to operate on semantic regions rather than individual pixels, automatically understanding image content and enabling users to perform edits at the object level without manual pixel selection, thereby reducing user interaction time while maintaining ease of operation

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

Machine learning models serve as intermediaries between the user and the image editing process. These models automatically analyze image content, identify semantic regions, and prepare edit masks, acting as a mediator that translates high-level user intentions into detailed pixel-level operations without requiring direct user involvement in the tedious segmentation process

Inventive Principle:
Principle #24Intermediary (Mediator)

2Ease of operation

If conventional image editing systems operate on a pixel level, then they can perform detailed image manipulation, but they require specialized knowledge, increasing device complexity

Engineering Contradiction:
Improveease of operationVSAvoidsystem complexity
Core Design Contradiction:
Ease of operationVSDevice complexity

Solution Approach 1:

The system performs self-service by automatically analyzing image content, identifying semantic regions, and generating edit masks without user intervention. The machine learning models autonomously understand the image structure and prepare the necessary data structures for editing, eliminating the need for users to possess specialized knowledge about image processing techniques while the system handles the complex technical operations internally

Inventive Principle:
Principle #25Self-service

3Manufacturing precision

If conventional image editing systems edit pixels independently, then they can provide precise pixel-level control, but they fail to maintain real-world conditions, reducing manufacturing precision

Engineering Contradiction:
Improveediting precisionVSAvoidreal-world condition maintenance
Core Design Contradiction:
Manufacturing precisionVSReliability

Solution Approach 1:

The system performs preliminary actions by pre-processing the image to identify semantic regions and their spatial relationships before executing edits. Machine learning models analyze the image structure, depth information, and lighting conditions in advance, allowing the system to anticipate how edits should affect surrounding areas and maintain real-world consistency, thereby achieving both precise editing and reliable condition maintenance

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system uses feedback from machine learning models that continuously analyze the image content and edit results. The models provide feedback about semantic region boundaries, depth relationships, and lighting consistency, allowing the system to adjust edit parameters to maintain real-world conditions while preserving precise control over the editing outcome

Inventive Principle:
Principle #23Feedback

4Productivity

If conventional image editing systems require manual user input for all operations, then they can provide fine-grained control, but they reduce productivity

Engineering Contradiction:
Improveediting efficiencyVSAvoiduser interaction requirement
Core Design Contradiction:
ProductivityVSEase of operation

Solution Approach 1:

The system performs preliminary actions by automatically segmenting the image into semantic regions and generating edit masks before the user initiates editing operations. This pre-processing eliminates the need for users to manually select regions or specify edit parameters, allowing them to simply choose from pre-computed options and execute edits instantly, thereby dramatically improving productivity while maintaining intuitive ease of operation

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system provides self-service by automatically performing all preparatory editing tasks including region segmentation, mask generation, and parameter optimization without user input. The machine learning models autonomously complete these operations and present ready-to-execute edit commands to the user, who only needs to confirm the desired action, thereby maximizing productivity while keeping the user interface simple and accessible

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS12469194B2Generating shadows for placed objects in depth estimated scenes of two-dimensional images
Publication Date: 2025.11.11 ADOBE INC
  • US12469194B2 patent drawing
  • US12469194B2 patent drawing
  • US12469194B2 patent drawing

AI summary

The present disclosure relates to systems, methods, and non-transitory computer-readable media that modify two-dimensional images via scene-based editing using three-dimensional representations of the two-dimensional images. For instance, in one or more embodiments, the disclosed systems utilize three-dimensional representations of two-dimensional images to generate and modify shadows in the two-dimensional images according to various shadow maps. Additionally, the disclosed systems utilize three-dimensional representations of two-dimensional images to modify humans in the two-dimensional images. The disclosed systems also utilize three-dimensional representations of two-dimensional images to provide scene scale estimation via scale fields of the two-dimensional images. In some embodiments, the disclosed systems utilizes three-dimensional representations of two-dimensional images to generate and visualize 3D planar surfaces for modifying objects in two-dimensional images. The disclosed systems further use three-dimensional representations of two-dimensional images to customize focal points for the two-dimensional images.