Portrait Editing Masks for Preserving Subject Features
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing image editing techniques struggle to achieve desired portrait editing results while preserving subject features, often requiring high-quality training datasets that are difficult to collect, and fail to maintain image quality.
Innovation Solution
A machine learning model is trained using a synthetic dataset generated automatically, leveraging a conditional dataset generation strategy to produce paired data with improved identity and layout alignment, and incorporates editing masks to guide the inference process, ensuring that untargeted features are preserved.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Manufacturing precision
If existing image editing techniques are used, then editing results can be achieved, but subject features are not preserved and image quality deteriorates
Solution Approach 1:
The patent segments the image into editable regions and preserved regions using editing masks. The model learns to distinguish between regions that should be edited and regions that should be preserved, applying edits only to specific segments while maintaining original features in other segments. This segmentation approach directly resolves the contradiction by enabling precise editing control without compromising overall feature preservation.
Solution Approach 2:
The patent applies local quality by differentiating treatment across different regions of the image. Through editing masks and region-aware processing, the model applies edit operations selectively to specific local areas while maintaining original quality and features in other areas. This local differentiation enables precise editing results while preserving subject features that should remain unchanged.
2Manufacturing precision
If high-quality training datasets are used, then editing results improve, but data collection becomes difficult and time-consuming
Solution Approach 1:
The patent creates synthetic training data that copies and adapts existing image patterns to generate realistic training examples. Instead of collecting real portrait images with diverse edits manually, the system generates synthetic versions that replicate the necessary editing variations. This copying approach maintains training quality while dramatically reducing data collection time and effort.
Solution Approach 2:
The patent performs preliminary data generation before actual training. By pre-generating synthetic training datasets through automated processes, the system prepares high-quality training data in advance without requiring manual collection during the development process. This preliminary action eliminates the time-consuming manual data collection step while ensuring adequate training data quality.
3Adaptability or versatility
If editing operations are applied, then desired changes are achieved, but untargeted features are altered and image quality is lost
Solution Approach 1:
The patent extracts and isolates the editing operation from the entire image processing pipeline. By using editing masks to identify only the regions requiring modification, the system separates the edit application from the rest of the image. This extraction ensures that editing flexibility is applied only where needed while preserving image quality in unedited regions, preventing unintended alterations.
Solution Approach 2:
Instead of applying edits and then trying to preserve features, the patent inverts the approach by first identifying what should be preserved through editing masks, then applying edits only to complementary regions. This inversion ensures that feature preservation is built into the editing process itself rather than being an afterthought, maintaining image quality while achieving desired changes.
Data Source
AI summary
The present disclosure describes techniques for implementing portrait editing using a machine learning model. An and a text prompt are input into a first machine learning model. The image comprises a portrait of a subject. The text prompt indicates a target result of editing the image. The first machine learning model is trained to perform portrait editing while preserving untargeted features. An editing mask is generated by the first machine-learning model based on the image. The editing mask indicates a first area for editing and a second area for preserving original content of the image. A mask-guided predicted noise is computed at each timestep and a process of editing the image is guided by the first machine learning model based on the editing mask. An edited image is generated by the first machine learning model. The edited image comprises the target editing result and retains detailed features of the subject.


