Scene-Based Facial Expression and Pose Transfer with Semantic Editing

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional image editing systems are inflexible and inefficient, requiring significant user interaction and specialized knowledge to edit digital images, as they operate on a pixel level rather than a semantic level, leading to rigid and cumbersome editing processes.

Innovation Solution

A scene-based image editing system that utilizes machine learning models to pre-process digital images, identifying objects, relationships, and attributes, allowing for intuitive and efficient editing by treating semantic areas as distinct units, reducing the need for user interactions and maintaining real-world conditions.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If conventional image editing systems operate on a pixel level, then detailed control over image elements is achieved, but the editing process becomes rigid and cumbersome requiring significant user interaction

Engineering Contradiction:
Improveediting processVSAvoiduser interaction requirements
Core Design Contradiction:
Ease of operationVSDevice complexity

Solution Approach 1:

The patent segments the image into semantic regions (sky, ground, objects, etc.) rather than treating it as a collection of pixels. This allows users to interact with meaningful image components directly, eliminating the need for pixel-level manipulation and significantly simplifying the editing process while maintaining detailed control over image elements.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces an intermediary layer between the user and the image data - a semantic segmentation model that translates user intentions into precise editing operations. This intermediary automatically identifies and isolates relevant image regions based on semantic meaning, reducing the complexity of user interaction while preserving fine-grained control over image modification.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Productivity

If conventional image editing systems require significant user interaction, then precise control over edits is achieved, but the editing efficiency decreases

Engineering Contradiction:
Improveediting efficiencyVSAvoiduser interaction time
Core Design Contradiction:
ProductivityVSLoss of time

Solution Approach 1:

The patent performs preliminary semantic segmentation and region identification automatically before the user initiates editing. By pre-processing the image to identify semantic regions, objects, and their relationships, the system eliminates the need for users to manually select and define editing areas, significantly improving editing efficiency and reducing time investment.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system performs self-service by automatically analyzing the image content, identifying semantic regions, and preparing editing candidates without user intervention. The machine learning model autonomously understands the image structure and relationships, allowing users to simply specify their editing intent while the system handles the complex task of locating and isolating relevant regions.

Inventive Principle:
Principle #25Self-service

3Adaptability or versatility

If conventional image editing systems treat the image as a two-dimensional array of pixels, then the entire image can be processed uniformly, but the system cannot understand or maintain real-world three-dimensional conditions

Engineering Contradiction:
Improvesemantic understandingVSAvoidreal-world conditions information
Core Design Contradiction:
Adaptability or versatilityVSLoss of information

Solution Approach 1:

The patent transforms the image representation from a simple two-dimensional pixel array to a multi-dimensional semantic structure that includes spatial relationships, object identities, and inferred three-dimensional properties. By changing the parameter space from pixel coordinates to semantic region attributes, the system gains the ability to understand and maintain real-world conditions such as object depth, orientation, and spatial relationships.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent adds semantic dimensions to the traditional two-dimensional image structure by introducing layers of semantic annotation that encode three-dimensional spatial relationships, object hierarchies, and real-world context. This dimensional enrichment allows the system to interpret and preserve real-world conditions while maintaining the underlying pixel data for processing.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

4Reliability

If conventional image editing systems perform edits on selected regions, then localized modifications are achieved, but maintaining consistency with surrounding areas requires additional user guidance

Engineering Contradiction:
Improveedit consistencyVSAvoiduser guidance requirements
Core Design Contradiction:
ReliabilityVSEase of operation

Solution Approach 1:

The patent implements feedback mechanisms where the system automatically analyzes the edited region's relationship with surrounding semantic areas and adjusts the edit to maintain consistency. The machine learning model provides real-time feedback on whether the edit preserves realistic transitions, lighting, and spatial relationships, eliminating the need for users to manually guide the consistency maintenance process.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The patent introduces an intermediary consistency-checking layer that automatically mediates between the user's edit intent and the final output. This intermediary analyzes semantic relationships, spatial continuity, and visual coherence to ensure that localized edits blend seamlessly with surrounding areas, maintaining reliability without requiring additional user guidance.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS12456274B2Facial expression and pose transfer utilizing an end-to-end machine learning model
Publication Date: 2025.10.28 ADOBE INC
  • US12456274B2 patent drawing
  • US12456274B2 patent drawing
  • US12456274B2 patent drawing

AI summary

The present disclosure relates to systems, methods, and non-transitory computer-readable media that modify digital images via scene-based editing using image understanding facilitated by artificial intelligence. For example, in one or more embodiments the disclosed systems utilize generative machine learning models to create modified digital images portraying human subjects. In particular, the disclosed systems generate modified digital images by performing infill modifications to complete a digital image or human inpainting for portions of a digital image that portrays a human. Moreover, in some embodiments, the disclosed systems perform reposing of subjects portrayed within a digital image to generate modified digital images. In addition, the disclosed systems in some embodiments perform facial expression transfer and facial expression animations to generate modified digital images or animations.