Neural Network Image Generation Using Contrastive Loss and Pixel Masking

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing neural networks struggle to accurately generate images from masked pixels due to limitations in training data and poor performance in tasks like optical character recognition (OCR), leading to inaccurate image reconstruction.

Innovation Solution

A system comprising a foundational neural network trained using a combination of pixel-masking encoder, contrastive loss, and caption generation systems, which operates in parallel to enhance image generation and captioning tasks by leveraging both labeled and unlabeled training images.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Quantity of substance

If a neural network is trained using only pixel-masking encoder on unlabeled images, then training data quantity is increased, but image reconstruction accuracy deteriorates due to poor OCR performance and lack of semantic understanding

Engineering Contradiction:
Improvetraining data quantityVSAvoidimage reconstruction accuracy
Core Design Contradiction:
Quantity of substanceVSManufacturing precision

Solution Approach 1:

The patent combines pixel-masking encoder training on unlabeled images with contrastive loss training on labeled image-caption pairs, merging two training approaches to simultaneously leverage large quantities of unlabeled data and the semantic understanding from labeled data, thereby maintaining both training data quantity and reconstruction accuracy

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The neural network is trained to perform multiple functions simultaneously: image reconstruction through pixel-masking and semantic understanding through contrastive loss with captions. This multi-functional training enables the model to benefit from both unlabeled and labeled data for comprehensive image generation capability

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Manufacturing precision

If a neural network is fine-tuned to improve image generation accuracy, then manufacturing precision is improved, but loss of time and computational resources increases due to extensive fine-tuning requirements

Engineering Contradiction:
Improveimage generation accuracyVSAvoidfine-tuning time
Core Design Contradiction:
Manufacturing precisionVSLoss of time

Solution Approach 1:

The patent performs preliminary action by pre-training the neural network on large quantities of unlabeled images using pixel-masking encoder before fine-tuning. This pre-training establishes a strong foundation that reduces the extent and time required for subsequent fine-tuning to achieve high image generation accuracy

Inventive Principle:
Principle #10Preliminary action

3Adaptability or versatility

If a neural network uses multiple training objectives (pixel-masking and contrastive loss), then adaptability is improved, but device complexity increases due to multiple parallel training systems

Engineering Contradiction:
Improvemulti-task capabilityVSAvoidtraining system complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The neural network is designed with universal multi-functionality to handle both pixel-masking reconstruction and contrastive loss-based caption understanding within a single model architecture. This unified approach achieves adaptability across multiple tasks while avoiding the complexity of maintaining separate specialized models for each function

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS20250200821A1Image generation using text
Publication Date: 2025.06.19 NVIDIA CORP
  • US20250200821A1 patent drawing
  • US20250200821A1 patent drawing
  • US20250200821A1 patent drawing

AI summary

Apparatuses, systems, and techniques to perform a neural network to generate an image. In at least one embodiment, for example, one or more neural networks generate one or more portions of one or more images and one or more captions. In at least one embodiment, as another example, a processor uses one or more neural networks to generate one or more images from text based, at least in part, on one or more first images without text indicating content of one or more first images and one or more second images with text indicating content of one or more second images.