Method for acquiring human preference data based on semi-artificial image

The image data is obtained from open source data and company data through semi-artificial methods, combined with multimodal large language model and manual scoring, and the problems of high cost of manual annotation and unstable data quality in the existing technology are solved, and efficient and accurate image human preference data set construction is achieved, which improves the adaptability and evaluation accuracy of the image generation model.

CN120496093APending Publication Date: 2025-08-15GIANT MOBILE TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510573200.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-06
Publication Date
2025-08-15

AI Technical Summary

Technical Problem

The existing human preference data sets mainly rely on manual annotation or large-scale multimodal model generation, which has problems such as high cost, unstable data quality, difficulty in scaling and limited image quality, and are especially difficult to match with new generation models such as Flux, Kolors, MidJourney, etc.

Method used

The semi-artificial method is used to obtain data from open source image generation model and company business data, and the image data and prompts are expanded using DiffusionDB, combined with large language model rewriting and multimodal large language model for image evaluation, and high-quality data sets are constructed through labeler selection and manual scoring to optimize the aesthetic and semantic adaptability of the image generation model.

Benefits of technology

Efficiently and accurately construct large-scale, high-quality preference datasets, optimize the text-image alignment and image authenticity score of image generation models, improve the generalization ability of the model on different styles of images, and provide a robust automated evaluation solution.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120496093A_ABST
    Figure CN120496093A_ABST
Patent Text Reader

Abstract

The invention relates to a method for acquiring human preference data based on a semi-artificial image. The method comprises the following steps: S1, acquiring data; s2, performing optimization and expansion based on the acquired data; s3, rewriting, expanding and enhancing the image data and the prompt words by using a large language model; s4, taking the prompt words processed by the large language model as input, and generating an image by adopting a plurality of different text-to-image generation models; s5, analyzing the image and the corresponding text cue by using a visual language model, and calculating the matching degree of the image and the corresponding text cue; s6, optimizing image evaluation through a multi-modal large language model; and S7, an annotator selects an image which is more consistent with description between the two candidate images to construct comparison preference data. According to the method, a large-scale and high-quality preference data set can be efficiently and accurately constructed, so that the adaptive capacity of an image generation model to human aesthetic appreciation and semantic preference is optimized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data acquisition, and in particular to a method for acquiring image human preference data based on semi-artificial methods. Background Art

[0002] Currently, the optimization of diffusion models increasingly relies on high-quality human preference data to improve the aesthetic quality, semantic consistency, and user satisfaction of generated images. However, existing human preference datasets are mainly generated by manual annotation or large-scale multimodal models (MLLMs), but these methods still have many limitations:

[0003] Manually annotated datasets: The manual annotation process is costly and requires significant human resources, making it difficult to scale to large datasets. Furthermore, manual annotation can easily introduce bias, impacting data quality.

[0004] Datasets generated based on visual language models (VLMs): Although VLM-generated data can reduce costs, its accuracy is usually low, especially in terms of aesthetic scoring and detail recognition, and it is difficult to effectively distinguish between high-quality and low-quality images.

[0005] Limited generative models for datasets: Current open-source datasets do not fully leverage the capabilities of new-generation models, resulting in limited generated image quality. Most existing datasets tend to be generated using Stable Diffusion 1.5 / 2.1. New-generation models such as Flux, Kolors, and MidJourney offer superior image quality and detail. Consequently, the quality of images generated by these traditional datasets struggles to match the latest generative techniques.

[0006] Therefore, it is necessary to provide a method for obtaining semi-artificial image human preference data to efficiently and accurately construct a large-scale, high-quality preference dataset, thereby optimizing the adaptability of image generation models to human aesthetic and semantic preferences. Summary of the Invention

[0007] The purpose of the present invention is to provide a method for obtaining human preference data on images based on semi-artificial methods, so as to efficiently and accurately construct a large-scale, high-quality preference dataset, thereby optimizing the adaptability of image generation models to human aesthetic and semantic preferences.

[0008] In order to solve the problems existing in the prior art, the present invention provides a method for obtaining human preference data of images based on semi-artificial methods, comprising the following steps:

[0009] S1: Obtain data from the generation results of open source image generation models and MJ data accumulated during the company's business process;

[0010] S2: Based on the acquired data, we use DiffusionDB as the basic text prompt dataset and optimize and expand it to generate diverse image data and prompts;

[0011] S3: Use a large language model to rewrite, expand, and enhance the generated diverse image data and prompts to improve the diversity and quality of the prompts;

[0012] S4: Takes the prompt processed by the large language model as input and uses multiple different text-to-image generation models to generate images;

[0013] S5: Analyze the image and the corresponding text prompt using a visual language model, and calculate the degree of matching between the image and the corresponding text prompt to complete the image evaluation;

[0014] S6: Image evaluation is optimized by enhancing the multimodal large language model with chained reasoning and question-answering interaction. The optimization process is as follows:

[0015]

[0016] Among them, M mllm It is a multimodal large language model that performs enhanced chain thinking reasoning and question-answering interaction. CoT stands for enhanced chain thinking reasoning, QA stands for question-answering interaction, Image is the input image, and θ gen are the model parameters that need to be optimized, It is a loss function that measures the gap between the predicted output and the true value, and D represents the distribution of the training data;

[0017] S7: Let the annotator choose the image that better matches the description between two candidate images to build comparative preference data.

[0018] Optionally, in the method for acquiring semi-artificial image human preference data,

[0019] MJ data is the master data in the company's business process. Master data refers to the data shared among various systems within the enterprise.

[0020] Optionally, in the method for obtaining semi-artificial image human preference data, DiffusionDB is a dataset of text-to-image prompts, containing no less than 1.5 million user-written prompts.

[0021] Optionally, in the method for acquiring human preference data for images based on semi-artificial methods, different generation models are randomly selected during the image generation process, and different sampling parameters are set for each prompt, so that multiple images are generated for each prompt.

[0022] Optionally, in the method for acquiring semi-artificial image human preference data, the sampling parameters include a random seed, a sampling step number, and an unclassified guidance factor.

[0023] Optionally, the method for obtaining semi-artificial image human preference data further includes the following steps:

[0024] S8: Have humans rate batches of images to quantify the quality and / or aesthetic value of the images.

[0025] Compared with the prior art, the present invention has the following advantages:

[0026] (1) The present invention can efficiently and accurately construct a large-scale, high-quality preference dataset, thereby optimizing the adaptability of image generation models to human aesthetic and semantic preferences.

[0027] (2) This paper optimizes text-image alignment and image authenticity scoring through multimodal large language model (MLLM) + chain thinking (CoT) reasoning, and improves the model's generalization ability on images of different styles.

[0028] (3) The present invention provides an efficient, robust and more human-friendly solution for automated image evaluation.

[0029] (4) The dataset obtained by the present invention can be effectively trained and improved on a variety of image generation models, demonstrating its wide applicability and high efficiency.

[0030] (5) This paper proposes a semi-artificial method for acquiring image human preference data, which combines CoT-based manual feedback with AI automatic labeling, taking into account data quality, cost and generalization ability. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] Figure 1 This is a flowchart of obtaining preference data provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0032] The following is a more detailed description of the specific embodiments of the present invention with reference to schematic diagrams. The advantages and features of the present invention will become more apparent from the following description. It should be noted that the drawings are greatly simplified and not to exact scale, and are only used for the purpose of conveniently and clearly illustrating the embodiments of the present invention.

[0033] In the description of the present application, it should be understood that the terms "center", "longitudinal", "lateral", "length", "width", "thickness", "up", "down", "front", "back", "left", "right", "vertical", "horizontal", "top", "bottom", "inside", "outside", "clockwise", "counterclockwise" and the like to indicate orientations or positional relationships based on the orientations or positional relationships shown in the accompanying drawings, and are only for the convenience of describing the present application and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be understood as a limitation on the present application.

[0034] Currently, the optimization of diffusion models increasingly relies on high-quality human preference data to improve the aesthetic quality, semantic consistency, and user satisfaction of generated images. However, existing human preference datasets are mainly generated by manual annotation or large-scale multimodal models (MLLMs), but these methods still have many limitations.

[0035] In order to solve the problems existing in the prior art, the present invention provides a method for obtaining human preference data of images based on semi-artificial methods, such as Figure 1 Said method comprises the following steps:

[0036] S1: Obtain data from the generation results of the open source image generation model and the MJ data accumulated in the company's business process. Among them, MJ data is the master data in the company's business process, and master data refers to data shared among various systems within the enterprise.

[0037] S2: Based on the acquired data, we optimized and expanded DiffusionDB, a text prompt dataset, to generate diverse image data and prompts. DiffusionDB is a text-to-image prompt dataset containing no less than 1.5 million user-generated prompts, which is widely used in diffusion model research.

[0038] S3: Use a large language model to rewrite, expand, and enhance the generated diverse image data and prompts to improve the diversity and quality of the prompts;

[0039] S4: Using the prompts processed by the large language model as input, multiple different text-to-image generation models are used to generate images to ensure data diversity and representativeness;

[0040] During the image generation process, different generative models are randomly selected and different sampling parameters are set for each prompt, so that multiple images are generated for each prompt to obtain a wider style distribution. The sampling parameters include the random seed, the number of sampling steps, and the non-classification guidance factor.

[0041] This part of the data covers a wide range of prompt variations and image results generated by multiple models, which helps to improve the diversity and generalization ability of the dataset.

[0042] S5: Analyze the image and the corresponding text prompt using a visual language model (VLM) and calculate the degree of matching between the image and the corresponding text prompt to complete the image evaluation;

[0043] S6: An improved prompt-based scoring method is used to evaluate text-image alignment;

[0044] By optimizing image evaluation with a multimodal large language model that enhances chained reasoning and question-answering interaction, more accurate evaluations can be generated. The optimization process is as follows:

[0045]

[0046] Among them, M mllm It is a multimodal large language model that performs enhanced chain thinking reasoning and question-answering interaction. CoT stands for enhanced chain thinking reasoning, QA stands for question-answering interaction, Image is the input image, and θ gen are the model parameters that need to be optimized, The loss function measures the difference between the predicted output and the true value, where D represents the distribution of the training data. This formula describes how to optimize the Multimodal Large Language Model (MLLM) to achieve more accurate image authenticity scores using CoT+QA, minimizing scoring errors and generating more precise image authenticity scores. First, the model's reasoning capabilities are improved through multiple rounds of CoT reasoning, rather than directly assigning scores. QA further verifies the reasoning results and enhances model stability. By minimizing the loss function (loss = predicted score - true score), the model parameters are continuously optimized to ensure that the scores it generates are more consistent with human annotation standards.

[0047] To optimize image authenticity assessment, the present invention uses multimodal instruction tuning to enhance chain-of-thought (CoT) reasoning capabilities, enabling the model to better perform step-by-step reasoning to improve the accuracy of image authenticity scoring. Through this optimization process, the present invention can ensure that the MLLM maintains stability during the reasoning process of multimodal input (image + text) and improve its performance in image authenticity scoring.

[0048] S7: Let the annotator choose the image that better matches the description between two candidate images to build comparative preference data.

[0049] S8: Have humans rate the batches of images, for example, on a scale of 1-5, to quantify the quality and / or aesthetic value of the images.

[0050] In summary, the present invention has the following advantages compared with the prior art:

[0051] (1) The present invention can efficiently and accurately construct a large-scale, high-quality preference dataset, thereby optimizing the adaptability of image generation models to human aesthetic and semantic preferences.

[0052] (2) This paper optimizes text-image alignment and image authenticity scoring through multimodal large language model (MLLM) + chain thinking (CoT) reasoning, and improves the model's generalization ability on images of different styles.

[0053] (3) The present invention provides an efficient, robust and more human-friendly solution for automated image evaluation.

[0054] (4) The dataset obtained by the present invention can be effectively trained and improved on a variety of image generation models, demonstrating its wide applicability and high efficiency.

[0055] (5) This paper proposes a semi-artificial method for acquiring image human preference data, which combines CoT-based manual feedback with AI automatic labeling, taking into account data quality, cost and generalization ability.

[0056] The above description is merely a preferred embodiment of the present invention and does not limit the present invention in any way. Any person skilled in the art who, without departing from the scope of the present invention, makes any equivalent substitution, modification, or other changes to the technical solution and technical content disclosed in the present invention shall be deemed to be within the scope of the present invention and still fall within the scope of protection of the present invention.

Claims

1. A method for obtaining human preference data of images based on semi-artificial methods, characterized in that: The following steps are involved: S1: Obtain data from the generation results of open source image generation models and MJ data accumulated during the company's business process; S2: Based on the acquired data, we use DiffusionDB as the basic text prompt dataset and optimize and expand it to generate diverse image data and prompts; S3: Use a large language model to rewrite, expand, and enhance the generated diverse image data and prompts to improve the diversity and quality of the prompts; S4: Takes the prompt processed by the large language model as input and uses multiple different text-to-image generation models to generate images; S5: Analyze the image and the corresponding text prompt using a visual language model, and calculate the degree of matching between the image and the corresponding text prompt to complete the image evaluation; S6: Image evaluation is optimized by enhancing the multimodal large language model with chained reasoning and question-answering interaction. The optimization process is as follows: Among them, M mllm It is a multimodal large language model that performs enhanced chain thinking reasoning and question-answering interaction. CoT stands for enhanced chain thinking reasoning, QA stands for question-answering interaction, Image is the input image, and θ gen are the model parameters that need to be optimized, It is a loss function that measures the gap between the predicted output and the true value, and D represents the distribution of the training data; S7: Let the annotator choose the image that better matches the description between two candidate images to build comparative preference data.

2. The method for acquiring semi-artificial image human preference data according to claim 1, characterized in that: MJ data is the master data in the company's business process. Master data refers to the data shared among various systems within the enterprise.

3. The method for acquiring human preference data based on semi-artificial images according to claim 1, characterized in that: DiffusionDB is a dataset of text-to-image prompts containing no less than 1.5 million user-generated prompts.

4. The method for acquiring human preference data based on semi-artificial images according to claim 1, characterized in that: During the image generation process, different generation models are randomly selected, and different sampling parameters are set for each prompt, so that each prompt generates multiple images.

5. The method for acquiring human preference data based on semi-artificial images according to claim 4, characterized in that: Sampling parameters include random seed, number of sampling steps, and non-classified bootstrap factor.

6. The method for acquiring human preference data based on semi-artificial images according to claim 1, characterized in that: The following steps are also included: S8: Have humans rate batches of images to quantify the quality and / or aesthetic value of the images.