Text-to-Image Changer Identity Preservation via Segmentation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current text-to-image models face challenges in maintaining the identity of subjects in reference images while generating realistic outputs and ensuring privacy in modified images.

Innovation Solution

The method employs a trained machine learning model and latent diffusion models to modify reference images based on textual inputs, allowing for changes to the background, areas, or style of the image while preserving the subject's identity and ensuring privacy through a safety check model.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If text-to-image models are used to generate images from textual descriptions, then image generation capability is improved, but subject identity preservation deteriorates

Engineering Contradiction:
Improveimage generation capabilityVSAvoidsubject identity preservation
Core Design Contradiction:
Adaptability or versatilityVSReliability

Solution Approach 1:

The patent segments the image into distinct regions (subject area and background area) and processes them separately. The subject area is identified and masked, then the background is generated independently while preserving the subject's identity features, resolving the contradiction between generative capability and identity preservation

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces an intermediary masking mechanism that separates the subject from the background processing. By creating a mask that isolates the subject area, the system can modify the background while preserving the subject's identity, acting as a mediator between the text prompt and the final image output

Inventive Principle:
Principle #24Intermediary (Mediator)

2Manufacturing precision

If image modification is performed to change background or style, then visual quality is improved, but subject privacy protection deteriorates

Engineering Contradiction:
Improvevisual qualityVSAvoidsubject privacy protection
Core Design Contradiction:
Manufacturing precisionVSObject-affected harmful factors

Solution Approach 1:

The patent extracts the subject area from the image using detection and masking techniques, then processes only the background area for modification. This separation ensures that the subject's privacy is protected while still achieving high visual quality in the modified background regions

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent applies different processing qualities to different regions: high-fidelity preservation for the subject area and creative modification for the background area. This local differentiation maintains visual quality where needed while protecting privacy in the subject region

Inventive Principle:
Principle #3Local quality

3Reliability

If latent diffusion models are used for image generation, then image realism is improved, but computational complexity increases

Engineering Contradiction:
Improveimage realismVSAvoidcomputational complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent segments the computational process into separate stages: subject detection, mask generation, and background generation using latent diffusion models. This segmentation allows the computationally intensive diffusion process to be applied only to the background region, reducing overall computational complexity while maintaining image realism

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent applies the computationally intensive latent diffusion model only partially to the background region rather than the entire image. This partial action reduces computational complexity while still achieving high realism in the modified portions, accepting that the subject area maintains its original quality

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS20250104311A1Text to Image Changer
Publication Date: 2025.03.27 META PLATFORMS INC
  • US20250104311A1 patent drawing
  • US20250104311A1 patent drawing
  • US20250104311A1 patent drawing

AI summary

The application describes method of modifying an image. The method may include a step of receiving, via an user interface of a service, a reference image and an input including text associated with the reference image. The method may also include a step of determining, via a trained machine learning (ML) model, one or more features of the reference image. The method may further include a step of modifying, via one or more trained latent diffusion models (LDMs), the reference image based upon the determined features and the received input. Any one or more of a background of the reference image, an area of the reference image or a style of the reference image may be modified. The method may even further include a step of causing to display, via the user interface of the service, the modified image.