Contextually-Persistent Visual Image Generation System

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current computer systems are inadequate in generating and providing context-related digital images for text documents, limiting the enhancement of user understanding and engagement by lacking visual aids.

Innovation Solution

An image generation system that utilizes multiple computer-based models, entity identifiers, and visual entity embeddings to automatically create contextually-persistent synthetic images for text documents, maintaining consistent visual style and entities throughout the image series.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Manufacturing precision

If multiple computer-based models are used to generate contextually-persistent visual images, then the accuracy and consistency of images are improved, but the computing resources required increase

Engineering Contradiction:
Improveimage generation accuracyVSAvoidcomputing resources
Core Design Contradiction:
Manufacturing precisionVSUse of energy by moving object

Solution Approach 1:

The system performs preliminary extraction of visual entity embeddings from reference images before the main image generation process. These embeddings are stored and reused during subsequent image generation, avoiding redundant processing and reducing overall computing resource consumption while maintaining high accuracy through the use of pre-computed visual features

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system extracts visual entity embeddings (copies of visual features) from reference images and uses these embeddings to generate new images. This copying approach allows the system to maintain consistency with reference images while avoiding the need to reprocess the entire reference image data, thus reducing computing resources while preserving accuracy

Inventive Principle:
Principle #26Copying

2Stability of the object's composition

If visual entity embeddings are extracted and stored for multiple entities, then the consistency across image series is improved, but the memory requirements increase

Engineering Contradiction:
Improvevisual consistencyVSAvoidmemory usage
Core Design Contradiction:
Stability of the object's compositionVSQuantity of substance

Solution Approach 1:

The system extracts only the essential visual entity embeddings (key visual features) from reference images and stores them in a separate data structure. This extraction approach allows the system to maintain visual consistency across the image series while using minimal memory, as only critical visual features are stored rather than complete image data

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The system uses visual entity embeddings as intermediary representations between reference images and generated images. These embeddings serve as compact mediators that capture essential visual information for consistency maintenance while occupying minimal memory space, bridging the gap between reference and output images efficiently

Inventive Principle:
Principle #24Intermediary (Mediator)

3Adaptability or versatility

If the system generates additional images in response to user interactions, then the adaptability and user engagement are improved, but the processing time increases

Engineering Contradiction:
Improvesystem flexibilityVSAvoidprocessing time
Core Design Contradiction:
Adaptability or versatilityVSLoss of time

Solution Approach 1:

The system performs preliminary setup by extracting visual entity embeddings and establishing the multi-model framework before user interactions occur. This preliminary preparation enables the system to respond quickly to user requests for additional images, as the foundational processing is already complete and the system is ready to generate new images using pre-configured models and embeddings

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system automatically generates additional images in response to user interactions without requiring manual intervention or reconfiguration. The pre-configured multi-model framework and stored visual entity embeddings enable the system to serve itself by autonomously producing new images based on user feedback, maintaining high adaptability while minimizing processing time through automated operations

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS20250078351A1Generating a series of contextually-persistent visual images for text documents utilizing multiple models
Publication Date: 2025.03.06 MICROSOFT TECHNOLOGY LICENSING LLC
  • US20250078351A1 patent drawing
  • US20250078351A1 patent drawing
  • US20250078351A1 patent drawing

AI summary

This disclosure presents an image generation system designed to generate a series of contextually-persistent visual images for a text document. For instance, the image generation system utilizes multiple computer-based models, entity identifiers, and visual entity embeddings to create multiple synthetic images for a given text document. These synthetic images share a consistent theme and style. Additionally, the synthetic images include the same characters, places, and objects. Indeed, the image generation system implements seamless and consistent visual representations of the entities throughout the text document.