Generative Model Fine-Tuning via Human Visual Appeal Feedback
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Generative machine-learning models (GMLMs) trained to generate images based on textual descriptions often disregard the visual appeal of the images, leading to user dissatisfaction due to the lack of aesthetic pleasure derived from the generated images.
Innovation Solution
A method and system that fine-tune GMLMs by receiving textual descriptions with keywords indicating visual attributes, generating augmented descriptions, and using human assessors to determine the visual appeal of generated image candidates, thereby refining the model to produce more visually appealing images.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If a GMLM is trained on a large training dataset to generate images based on textual descriptions, then the model can generate images according to user queries, but the model disregards quality categories and visual appeal of the generated images
Solution Approach 1:
The patent implements a feedback mechanism where human assessors evaluate the visual appeal of generated images and provide ratings. These ratings are then used to fine-tune the GMLM, creating a closed-loop system where the model learns from human feedback to improve image quality while maintaining its image generation capability
Solution Approach 2:
The patent applies preliminary action by fine-tuning the GMLM on a curated training dataset that emphasizes visual appeal qualities before deploying the model for general image generation tasks. This preliminary training ensures the model inherently prioritizes aesthetic qualities when generating images
2Manufacturing precision
If human assessors are used to determine visual appeal of image candidates, then the GMLM can be fine-tuned to generate more visually appealing images, but the process requires significant human involvement and time
Solution Approach 1:
The patent applies partial action by using a limited number of human assessors to evaluate only a subset of generated images for fine-tuning purposes, rather than requiring comprehensive human evaluation of all possible images. This approach achieves sufficient visual appeal improvement without excessive time investment
Solution Approach 2:
The patent performs preliminary action by pre-curating a training dataset with images that have known high visual appeal qualities. This pre-prepared dataset enables efficient fine-tuning without requiring extensive real-time human assessment during the training process
Data Source
AI summary
A method and a server for fine-tuning a generative machine-learning model (GMLM) are provided. The method comprises: receiving a given textual description of a testing object a testing image thereof, the given textual description being indicative of what is to be depicted in the testing image in a natural language; receiving keywords associated with the given textual description, a given keyword being indicative of a rendering instruction for rendering the testing object in the testing image; generating, based on the keywords, augmented textual descriptions of the image; feeding to the GMLM, each one of the augmented textual descriptions to generate image candidates of the object; transmitting the image candidates to a plurality of human assessors for pairwise comparison thereof; based on the pairwise comparison, determining for the given image candidate, a respective degree of visual appeal; and using the respective degree of visual appeal for fine-tuning the GMLM.


