Neural Network Image Blending via Driver Image Context

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Machine learning models that use text-based descriptions for image editing often generate output images that do not accurately capture the desired visual attributes, leading to inconsistent results.

Innovation Solution

A technique that combines a source image and a driver image by determining a region of the source image to be blended with the driver image, inputting the source image's exterior region and the driver image into a neural network, and generating an output image that incorporates the driver image's visual attributes while maintaining the source image's context.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If text-based descriptions are used to guide image editing, then the editing process is simple to operate, but the visual attributes of the output image cannot be precisely controlled

Engineering Contradiction:
Improveease of operationVSAvoidvisual attribute precision
Core Design Contradiction:
Ease of operationVSManufacturing precision

Solution Approach 1:

The patent introduces a driver image as an intermediary element that bridges the gap between simple text-based operation and precise visual attribute control. The driver image contains visual examples of the desired attributes, allowing users to guide editing through image selection rather than complex text parameters, thus maintaining ease of operation while achieving precise visual control through the intermediary visual reference.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent changes the parameter representation from text-based descriptions to image-based visual attributes. Instead of using text prompts that are ambiguous and imprecise, the system uses driver images that directly encode visual attributes such as color, texture, and style. This parameter transformation enables precise control of visual attributes while keeping the user interface simple, as users simply select images rather than craft complex text descriptions.

Inventive Principle:
Principle #35Parameter changes

2Manufacturing precision

If visual attributes from driver image are integrated into source image region, then visual attribute precision is improved, but the complexity of the image processing system increases

Engineering Contradiction:
Improvevisual attribute precisionVSAvoidsystem complexity
Core Design Contradiction:
Manufacturing precisionVSDevice complexity

Solution Approach 1:

The patent segments the image editing task into distinct functional components: a region segmentation module that identifies the target area in the source image, a driver image processing module that extracts visual attributes from the driver image, and a synthesis module that combines these elements. This segmentation allows each component to specialize in specific operations, making the overall complex task manageable and the system architecture clearer, thus achieving high visual attribute precision without overwhelming complexity.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent employs a nested structure where the driver image is processed to extract visual attributes, which are then nested into the source image's target region. The driver image contains embedded visual information that is selectively extracted and integrated into the source image, creating a nested relationship between the two images. This nesting approach allows the system to handle complexity by organizing visual attributes in hierarchical layers, making the processing more manageable while achieving precise visual integration.

Inventive Principle:
Principle #7Nested doll (Nesting)

Data Source

PatentUS12333681B1Targeted generative visual editing of images
Publication Date: 2025.06.17 META PLATFORMS TECHNOLOGIES LLC
  • US12333681B1 patent drawing
  • US12333681B1 patent drawing
  • US12333681B1 patent drawing

AI summary

One embodiment of the present invention sets forth a technique for combining a source image and a driver image. The technique includes determining a first region of the source image to be blended with the driver image. The technique also includes inputting a second region of the source image that lies outside of the first region and the driver image into a neural network. The technique further includes generating, via the neural network, an output image that includes a third region corresponding to the first region of the source image and a fourth region corresponding to the second region of the source image, where the third region includes visual attributes of the driver image and a context associated with the source image and the fourth region includes visual attributes of the second region of the source image and the context associated with the source image.