Semantic-Guided Diffusion Image Augmentation for Stable AI Training

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional image augmentation using generative models lacks controllability and accuracy, leading to unstable generation results and affecting the generalization and accuracy of artificial intelligence models.

Innovation Solution

An image augmentation device and method that utilizes a diffusion model to generate a generated image from a noise image and semantic information, incorporating semantic categories and ranges, and composites the generated image with guide images to enhance accuracy.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If a text prompt is input into the generative model for image augmentation, then the generation process is simple, but the controllability and accuracy of the augmentation are poor

Engineering Contradiction:
Improvesimplicity of generation processVSAvoidaccuracy of augmentation
Core Design Contradiction:
Ease of operationVSManufacturing precision

Solution Approach 1:

The patent segments the image into multiple semantic regions using semantic segmentation, allowing different parts of the image to be processed independently with region-specific text prompts. This enables precise control over which objects or regions are augmented and how, resolving the contradiction between simple operation and accurate control by adding semantic structure without significantly complicating the workflow

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent applies different text prompts and augmentation parameters to different semantic regions of the image rather than treating the entire image uniformly. This allows locally optimized control where each region receives appropriate augmentation based on its semantic content, achieving high accuracy while maintaining ease of operation through automated region-specific processing

Inventive Principle:
Principle #3Local quality

2Productivity

If the generative model generates content based on text prompt without semantic guidance, then the process is fast, but the generation result is unstable and inaccurate

Engineering Contradiction:
Improvespeed of generationVSAvoidstability of generation result
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The patent performs semantic segmentation and region identification before the actual image generation process. This preliminary action provides the generative model with structured guidance about which regions to modify and what content to generate, ensuring stable and accurate results while maintaining efficiency through pre-computed semantic information

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent introduces semantic segmentation maps and region masks as intermediary structures between the text prompt and the generative model. These intermediaries translate high-level text descriptions into precise spatial and semantic constraints, enabling the model to generate stable and accurate content in the correct regions without requiring complex direct control mechanisms

Inventive Principle:
Principle #24Intermediary (Mediator)

3Manufacturing precision

If semantic segmentation is performed to guide the diffusion model, then the accuracy of content augmentation is improved, but the device complexity increases

Engineering Contradiction:
Improveaccuracy of content augmentationVSAvoidcomplexity of processing system
Core Design Contradiction:
Manufacturing precisionVSDevice complexity

Solution Approach 1:

The patent employs a diffusion model that can handle both the semantic understanding and the image generation tasks within a single unified framework. This multi-functional approach reduces overall system complexity by eliminating the need for separate specialized modules for each task, while still achieving high accuracy through semantic guidance

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent combines the semantic segmentation module with the diffusion model into an integrated system where the segmentation information is directly fed into the diffusion process. This merging reduces the complexity of coordinating multiple independent components while maintaining the accuracy benefits of semantic guidance through seamless information flow between modules

Inventive Principle:
Principle #5Merging (Combining)

Data Source

PatentUS20260073594A1Image augmentation device and method
Publication Date: 2026.03.12 HON HAI PRECISION INDUSTRY CO LTD
  • US20260073594A1 patent drawing
  • US20260073594A1 patent drawing
  • US20260073594A1 patent drawing

AI summary

An image augmentation device and method are provided. The image augmentation device inputs a noise image corresponding to an original image and semantic information and a text vector corresponding to the original image into a diffusion model to generate a generated image, and the generated image includes a partial contour of the original image. The image augmentation device composites the generated image and a plurality of guide images to generate an augmented image.