3D Human Models for Realistic Semantic Image Editing

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional image editing systems are inflexible and inefficient, requiring significant user interaction and specialized knowledge to edit digital images, as they operate on a pixel level rather than a semantic level, failing to maintain real-world conditions during edits.

Innovation Solution

A scene-based image editing system that utilizes machine learning to pre-process digital images, identifying objects, relationships, and attributes, allowing intuitive and efficient editing by treating semantic areas as distinct units and maintaining real-world conditions without additional user input.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If conventional image editing systems operate on pixel level, then detailed control is achieved, but user interaction complexity and specialized knowledge requirements increase significantly

Engineering Contradiction:
Improveediting operation simplicityVSAvoidsystem operational complexity
Core Design Contradiction:
Ease of operationVSDevice complexity

Solution Approach 1:

The patent segments the image into distinct semantic objects (person, car, building, etc.) rather than treating it as a collection of pixels. Each object is identified and can be independently edited, transforming the complex pixel-level operation into simple object-level operations that users can perform intuitively.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces machine learning models as intermediaries between the user and the image editing process. These models automatically perform complex tasks such as object segmentation, attribute identification, and editing parameter generation, allowing users to interact with simple high-level commands rather than complex pixel-level operations.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Productivity

If machine learning models pre-process images to identify semantic areas, then editing efficiency improves, but computational resources and processing time increase

Engineering Contradiction:
Improveediting efficiencyVSAvoidpre-processing time
Core Design Contradiction:
ProductivityVSLoss of time

Solution Approach 1:

The patent performs preliminary actions by pre-processing the image to identify semantic objects, relationships, and attributes before the actual editing operation. This preparation work is done automatically and efficiently using machine learning models, enabling subsequent editing operations to be executed rapidly with minimal user input.

Inventive Principle:
Principle #10Preliminary action

3Reliability

If editing operations maintain real-world conditions automatically, then realism of edited images improves, but system complexity and computational requirements increase

Engineering Contradiction:
Improverealism maintenanceVSAvoidsystem structural complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent employs feedback mechanisms where the machine learning model continuously analyzes the edited image and adjusts editing parameters to maintain consistency with real-world conditions. The system monitors attributes such as lighting, shadows, and object relationships, automatically correcting discrepancies to preserve realism throughout the editing process.

Inventive Principle:
Principle #23Feedback

4Loss of time

If conventional systems require significant user interaction for editing, then precise control is achieved, but operational time and complexity increase

Engineering Contradiction:
Improveediting timeVSAvoiduser interaction requirement
Core Design Contradiction:
Loss of timeVSEase of operation

Solution Approach 1:

The patent enables the system to perform editing operations autonomously based on user intent. Once the user specifies the desired edit (e.g., move the car, change the building color), the machine learning model automatically executes the entire editing process, including object identification, parameter adjustment, and realism maintenance, without requiring continuous user interaction or specialized knowledge.

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS12499574B2Generating three-dimensional human models representing two-dimensional humans in two-dimensional images
Publication Date: 2025.12.16 ADOBE INC
  • US12499574B2 patent drawing
  • US12499574B2 patent drawing
  • US12499574B2 patent drawing

AI summary

The present disclosure relates to systems, methods, and non-transitory computer-readable media that modify two-dimensional images via scene-based editing using three-dimensional representations of the two-dimensional images. For instance, in one or more embodiments, the disclosed systems utilize three-dimensional representations of two-dimensional images to generate and modify shadows in the two-dimensional images according to various shadow maps. Additionally, the disclosed systems utilize three-dimensional representations of two-dimensional images to modify humans in the two-dimensional images. The disclosed systems also utilize three-dimensional representations of two-dimensional images to provide scene scale estimation via scale fields of the two-dimensional images. In some embodiments, the disclosed systems utilizes three-dimensional representations of two-dimensional images to generate and visualize 3D planar surfaces for modifying objects in two-dimensional images. The disclosed systems further use three-dimensional representations of two-dimensional images to customize focal points for the two-dimensional images.