AI Image Generation with Masking and Depth Map Control

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing Generative Artificial Intelligence (Gen AI) systems face challenges in generating high-quality images due to complex prompt engineering, inability to create untrained elements, and poor preservation of reference image integrity and spatial positioning, hindering usability and accuracy in professional applications.

Innovation Solution

A Gen AI-based image generation system that uses advanced prompt engineering, computer vision techniques, and controlled image generation to create high-quality, layout-consistent images by leveraging a Large Language Model (LLM) and structured prompt enhancement, incorporating object segmentation and depth maps to preserve object integrity and position.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Manufacturing precision

If existing Gen AI systems use large-scale diffusion models or transformer-based architectures for image generation, then high-quality images can be created, but complex prompt engineering is required and usability is hindered

Engineering Contradiction:
Improveimage generation qualityVSAvoidusability
Core Design Contradiction:
Manufacturing precisionVSEase of operation

Solution Approach 1:

The patent introduces an intermediary processing layer between user input and the image generation model. This layer automatically translates natural language user inputs into structured prompts with technical details, eliminating the need for users to perform complex prompt engineering while maintaining high image generation quality through the underlying diffusion models.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Reliability

If existing Gen AI systems rely on trained models for image generation, then consistent results are achieved, but inability to generate untrained elements limits creativity

Engineering Contradiction:
Improvegeneration consistencyVSAvoidcreativity
Core Design Contradiction:
ReliabilityVSAdaptability or versatility

Solution Approach 1:

The patent performs preliminary analysis of user inputs to identify both trained and untrained elements. For untrained elements, it generates structured prompts with detailed technical specifications that guide the model to create novel content while maintaining consistency with the overall image composition and style through pre-established generation parameters.

Inventive Principle:
Principle #10Preliminary action

3Adaptability or versatility

If existing Gen AI systems generate images from text prompts, then flexibility is provided, but lack of precise control over object positioning and visual details reduces accuracy

Engineering Contradiction:
ImproveflexibilityVSAvoidaccuracy
Core Design Contradiction:
Adaptability or versatilityVSManufacturing precision

Solution Approach 1:

The patent applies local quality control by generating structured prompts that specify different levels of detail for different regions and elements within the image. Technical details such as object positioning coordinates, scale factors, and visual attribute specifications are embedded in the prompt structure, enabling precise control over specific elements while maintaining overall image flexibility.

Inventive Principle:
Principle #3Local quality

4Manufacturing precision

If existing Gen AI systems use complex prompt engineering, then detailed control is achieved, but device complexity and operational difficulty increase

Engineering Contradiction:
Improvecontrol precisionVSAvoidsystem complexity
Core Design Contradiction:
Manufacturing precisionVSDevice complexity

Solution Approach 1:

The patent implements a self-service mechanism where the system automatically generates structured prompts with technical details based on simple user inputs. The prompt generation process includes automatic extraction of key elements, assignment of technical parameters, and formatting according to the image generation model's requirements, eliminating the need for users to manually construct complex prompts while maintaining detailed control precision.

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS20260051095A1System and method for generative artificial intelligence based image generation
Publication Date: 2026.02.19 ACCENTURE GLOBAL SOLUTIONS LTD
  • US20260051095A1 patent drawing
  • US20260051095A1 patent drawing
  • US20260051095A1 patent drawing

AI summary

System and method for Generative Artificial Intelligence based image generation are disclosed. In an aspect, a user input is received requesting an image of a predetermined object to be generated. An enhanced prompt is then generated from the user input, the enhanced prompt includes one or more of a reference image of the predetermined object and a textual portion describing technical details for generating the image. Further, the enhanced prompt is input to an image generation model that preserves portions of the reference image including the predetermined object unchanged via masking the portions of the reference image, generates a depth map of the reference image, and generates an intermediate image by merging the reference image with the masked portions of the predetermined object and the depth map. Also, an enhanced version of the intermediate image is output as the image generated in response to the user input.