Generative Model Fine-Tuning via Human Visual Appeal Feedback

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Generative machine-learning models (GMLMs) trained to generate images based on textual descriptions often disregard the visual appeal of the images, leading to user dissatisfaction due to the lack of aesthetic pleasure derived from the generated images.

Innovation Solution

A method and system that fine-tune GMLMs by receiving textual descriptions with keywords indicating visual attributes, generating augmented descriptions, and using human assessors to determine the visual appeal of generated image candidates, thereby refining the model to produce more visually appealing images.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If a GMLM is trained on a large training dataset to generate images based on textual descriptions, then the model can generate images according to user queries, but the model disregards quality categories and visual appeal of the generated images

Engineering Contradiction:
Improveimage generation capabilityVSAvoidvisual appeal quality
Core Design Contradiction:
Adaptability or versatilityVSManufacturing precision

Solution Approach 1:

The patent implements a feedback mechanism where human assessors evaluate the visual appeal of generated images and provide ratings. These ratings are then used to fine-tune the GMLM, creating a closed-loop system where the model learns from human feedback to improve image quality while maintaining its image generation capability

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The patent applies preliminary action by fine-tuning the GMLM on a curated training dataset that emphasizes visual appeal qualities before deploying the model for general image generation tasks. This preliminary training ensures the model inherently prioritizes aesthetic qualities when generating images

Inventive Principle:
Principle #10Preliminary action

2Manufacturing precision

If human assessors are used to determine visual appeal of image candidates, then the GMLM can be fine-tuned to generate more visually appealing images, but the process requires significant human involvement and time

Engineering Contradiction:
Improvevisual appeal qualityVSAvoidfine-tuning process time
Core Design Contradiction:
Manufacturing precisionVSLoss of time

Solution Approach 1:

The patent applies partial action by using a limited number of human assessors to evaluate only a subset of generated images for fine-tuning purposes, rather than requiring comprehensive human evaluation of all possible images. This approach achieves sufficient visual appeal improvement without excessive time investment

Inventive Principle:
Principle #16Partial or excessive action

Solution Approach 2:

The patent performs preliminary action by pre-curating a training dataset with images that have known high visual appeal qualities. This pre-prepared dataset enables efficient fine-tuning without requiring extensive real-time human assessment during the training process

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS20240303474A1Method and system for generating training data for a machine-learning algorithm
Publication Date: 2024.09.12 Y E HUB ARMENIA LLC
  • US20240303474A1 patent drawing
  • US20240303474A1 patent drawing
  • US20240303474A1 patent drawing

AI summary

A method and a server for fine-tuning a generative machine-learning model (GMLM) are provided. The method comprises: receiving a given textual description of a testing object a testing image thereof, the given textual description being indicative of what is to be depicted in the testing image in a natural language; receiving keywords associated with the given textual description, a given keyword being indicative of a rendering instruction for rendering the testing object in the testing image; generating, based on the keywords, augmented textual descriptions of the image; feeding to the GMLM, each one of the augmented textual descriptions to generate image candidates of the object; transmitting the image candidates to a plurality of human assessors for pairwise comparison thereof; based on the pairwise comparison, determining for the given image candidate, a respective degree of visual appeal; and using the respective degree of visual appeal for fine-tuning the GMLM.