Multi-Person Image Generation via Segmented Diffusion Models

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing systems for generating images of multiple persons in simulated backgrounds are complex, time-consuming, and require significant resources, making it difficult to create high-quality images efficiently.

Innovation Solution

The system uses first and second artificial personalized images generated by generative machine learning models to combine depictions of multiple persons into a foreground image, which is then placed on a background with corresponding visual attributes, allowing for efficient generation of photorealistic images or videos.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Manufacturing precision

If existing systems are used to generate images of multiple persons in simulated backgrounds, then image quality can be achieved, but the process becomes complex and time-consuming

Engineering Contradiction:
Improveimage qualityVSAvoidsystem complexity
Core Design Contradiction:
Manufacturing precisionVSDevice complexity

Solution Approach 1:

The patent divides the image generation process into separate stages: generating individual person images first, then combining them into a composite foreground image, and finally placing it on a background. This segmentation allows each stage to be optimized independently, reducing overall system complexity while maintaining image quality.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent performs preliminary actions by pre-generating individual person images and their corresponding foreground masks before combining them. This preliminary preparation simplifies the subsequent composition process and reduces the complexity of real-time multi-person image generation.

Inventive Principle:
Principle #10Preliminary action

2Manufacturing precision

If existing systems are used to generate images of multiple persons in simulated backgrounds, then image quality can be achieved, but the process becomes time-consuming

Engineering Contradiction:
Improveimage qualityVSAvoidgeneration time
Core Design Contradiction:
Manufacturing precisionVSLoss of time

Solution Approach 1:

By segmenting the generation process into independent stages (individual person generation, foreground combination, background integration), the system can process each stage efficiently and parallelize operations, significantly reducing total generation time while preserving image quality.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent generates individual person images separately and then combines them, effectively creating copies of person images that can be composed together. This copying approach allows for efficient reuse of generated content and reduces redundant processing.

Inventive Principle:
Principle #26Copying

3Manufacturing precision

If existing systems are used to generate images of multiple persons in simulated backgrounds, then image quality can be achieved, but significant resources are required

Engineering Contradiction:
Improveimage qualityVSAvoidcomputational resources
Core Design Contradiction:
Manufacturing precisionVSUse of energy by moving object

Solution Approach 1:

The patent segments the computational workload into separate tasks for generating individual persons and combining them. This allows for more efficient resource allocation and can enable parallel processing, reducing the total computational resources required while maintaining high image quality.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS20250166264A1Diffusion model multi-person image generation
Publication Date: 2025.05.22 SNAP INC
  • US20250166264A1 patent drawing
  • US20250166264A1 patent drawing
  • US20250166264A1 patent drawing

AI summary

Methods and systems are disclosed for generating personalized images using one or more diffusion models. The methods and systems access first and second artificial personalized images generated by first and second generative machine learning models, the first generative machine learning model trained to generate the first artificial personalized image comprising a depiction of a first person, the second generative machine learning model trained to generate the second artificial personalized image comprising a depiction of a second person. The methods and systems generate a foreground image that combines the depiction of the first person in the first artificial personalized image with the depiction of the second person in the second artificial personalized image. The methods and systems access generate a new artificial image comprising the foreground image on a background having visual attributes that correspond to the background information.