End-to-End Facial Expression and Pose Transfer for Scene Editing

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional image editing systems are inflexible and inefficient, requiring significant user interaction and specialized knowledge to perform edits, as they operate on a pixel level rather than a semantic level, making it difficult to maintain real-world conditions during modifications.

Innovation Solution

A scene-based image editing system that utilizes machine learning models to pre-process digital images, identifying semantic areas and generating object masks and content fills, allowing for intuitive and efficient editing by treating objects as distinct units and maintaining real-world conditions without user input for preparatory steps.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If conventional image editing systems operate on pixel level, then detailed control is achieved, but user interaction complexity and specialized knowledge requirements increase significantly

Engineering Contradiction:
Improveease of operationVSAvoidediting process complexity
Core Design Contradiction:
Ease of operationVSDevice complexity

Solution Approach 1:

The system segments the image into semantic regions (sky, ground, objects, etc.) and allows users to edit at the region level rather than pixel level. The machine learning model automatically identifies and segments meaningful areas, enabling intuitive selection and manipulation of semantic regions through simple user actions like clicking or hovering over desired areas.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces an intermediary machine learning model that translates between user intent (semantic-level commands) and pixel-level modifications. This intermediary automatically handles the complex pixel manipulation tasks based on high-level user instructions, eliminating the need for users to directly manage pixel-level complexity.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Adaptability or versatility

If pixel-level editing is used, then precise control over image details is possible, but the system cannot maintain real-world conditions and object relationships

Engineering Contradiction:
Improveadaptability to real-world conditionsVSAvoidediting precision
Core Design Contradiction:
Adaptability or versatilityVSManufacturing precision

Solution Approach 1:

The system performs preliminary actions by pre-processing the image to identify semantic regions, object boundaries, and spatial relationships before user editing occurs. The machine learning model prepares the image structure in advance, creating a semantic map that preserves real-world conditions and enables edits that automatically maintain consistency with the original scene geometry and lighting.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent applies local quality by allowing different editing operations on different semantic regions while maintaining their specific properties. Each region can be edited with appropriate constraints and transformations that preserve its real-world characteristics (e.g., sky regions maintain atmospheric perspective, ground regions maintain perspective convergence).

Inventive Principle:
Principle #3Local quality

3Productivity

If traditional image editing workflows are used, then comprehensive control is achieved, but significant user interaction and time are required

Engineering Contradiction:
Improveediting efficiencyVSAvoidtime for user interactions
Core Design Contradiction:
ProductivityVSLoss of time

Solution Approach 1:

The system provides self-service by automatically performing preparatory editing tasks without requiring user input. The machine learning model autonomously identifies edit opportunities, suggests modifications, and executes routine operations based on the overall user intent, significantly reducing the number of user interactions needed while maintaining comprehensive editing control.

Inventive Principle:
Principle #25Self-service

4Ease of operation

If semantic-level editing is implemented, then ease of use improves, but the system complexity increases due to machine learning models

Engineering Contradiction:
Improveintuitiveness of editingVSAvoidsystem architecture complexity
Core Design Contradiction:
Ease of operationVSDevice complexity

Solution Approach 1:

The patent extracts the complex machine learning functionality into a separate, specialized component that operates independently from the user interface layer. This extraction allows the main editing system to remain simple and intuitive, while the embedded AI model handles the complex semantic understanding and translation tasks in the background.

Inventive Principle:
Principle #2Taking out (Extraction)

Data Source

PatentUS12450810B2Animated facial expression and pose transfer utilizing an end-to-end machine learning model
Publication Date: 2025.10.21 ADOBE INC
  • US12450810B2 patent drawing
  • US12450810B2 patent drawing
  • US12450810B2 patent drawing

AI summary

The present disclosure relates to systems, methods, and non-transitory computer-readable media that modify digital images via scene-based editing using image understanding facilitated by artificial intelligence. For example, in one or more embodiments the disclosed systems utilize generative machine learning models to create modified digital images portraying human subjects. In particular, the disclosed systems generate modified digital images by performing infill modifications to complete a digital image or human inpainting for portions of a digital image that portrays a human. Moreover, in some embodiments, the disclosed systems perform reposing of subjects portrayed within a digital image to generate modified digital images. In addition, the disclosed systems in some embodiments perform facial expression transfer and facial expression animations to generate modified digital images or animations.