A system for optimizing portrait generation quality in Chinese contextualized text-based image scenarios.
By constructing data annotation rules and attention aggregation algorithms, and adjusting parameters using the ChineseCLIP model, the Chinese text-based image model was optimized, solving the problem of insufficient image generation quality in Chinese contexts and achieving higher quality image generation.
Patent Information
- Application Number
- CN202411817909.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-11
- Publication Date
- 2025-10-31
- Estimated Expiration
- 2044-12-11
AI Technical Summary
In the Chinese context, the quality of human portrait generation by the Wenshengtu model is difficult to guarantee, often resulting in image corruption. The generated images do not match the distribution of real human images and fail to achieve the expected results.
We construct data annotation rules and attention aggregation algorithms, adjust parameters through a pre-trained ChineseCLIP model to generate a portrait reward model, and use a partition preference optimization algorithm to optimize the Chinese text-based image model to improve the quality of portrait generation.
It significantly improves the image generation quality of the Chinese context-based image model, enhancing the model's ability to distinguish image quality and improve generation results.
Smart Images

Figure CN119785140B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of artificial intelligence technology, and in particular to a system for optimizing the quality of human portrait generation in Chinese contextual text-based image scenarios. Background Technology
[0002] Text-to-image generation technology is a major research direction in the field of text-image multimodal learning, referring to the generation of images corresponding to given text conditions. Latent diffusion models are an important representative of models for text-to-image generation tasks, capable of generating images based on externally provided text prompts.
[0003] Current research on image quality optimization for text-based image models mainly focuses on English contexts, with limited exploration of models in Chinese contexts. Furthermore, in portrait generation, image corruption often occurs, meaning the generated images do not match the distribution of real human images. This implies that the quality of portrait generation in text-based image scenarios is difficult to guarantee in Chinese contexts, resulting in unsatisfactory image quality. Therefore, we propose a zoning optimization system for portrait generation quality in Chinese-context text-based image scenarios. Summary of the Invention
[0004] The present invention aims to provide a quality partitioning optimization system for portrait generation in Chinese-language text-based image scenarios, which solves the problem that the current quality of portrait generation in Chinese-language text-based image scenarios cannot meet expectations.
[0005] To achieve the above objectives, the technical solution of the present invention is as follows: a system for optimizing the quality of portrait generation in Chinese contextualized text-based image scenarios, comprising the following steps:
[0006] S1. Construct data annotation rules based on the degradation problem of generated human images to build a human image annotation dataset;
[0007] S2. Use the attention aggregation algorithm to transform the human image annotation dataset into a human image annotation preference dataset;
[0008] S3. Using the pre-trained ChineseCLIP model as the base model, the parameters of the ChineseCLIP model are adjusted using the portrait annotation preference dataset and the preset cross-entropy function to obtain the portrait reward model.
[0009] S4. Collect Chinese portrait cues and generate portrait images using a Chinese text-based image model. Use a portrait reward model to assign reward signals to the portrait image samples and construct a portrait partition preference dataset.
[0010] S5. Apply the partition preference optimization algorithm to the Chinese text-to-image model to adjust the parameters of the Chinese text-to-image model until training is complete.
[0011] Furthermore, the data source for step S1 is: collecting commonly used or poorly generated portrait prompts and using a Chinese text-based image model to generate corresponding different images.
[0012] Furthermore, the annotation standards for step S1 are as follows: corresponding annotation standards are constructed based on the issues of the number of limbs, limb shape, hands, interaction, and face in the collapse problem.
[0013] Furthermore, the aggregation rules in step S2 are as follows: an attention aggregation algorithm is adopted, that is, for each annotation, different types of collapse will be assigned a different weight, representing the user's attention to this type of collapse problem, and finally the annotation results are combined; for each image sample, there will be a total attention, which is the sum of the weights of all annotators for the problems they annotated. The lower the attention, the lower the degree of portrait collapse, and vice versa; an image sample and a sample with a higher attention form a preference data pair, thus forming a portrait annotation preference dataset.
[0014] Furthermore, in step S3, the ChineseCLIP model training process minimizes the cross-entropy between the labeled preference distribution and the normalized score of the sample, enabling the portrait reward model to acquire human preference knowledge of portrait quality.
[0015] Furthermore, the method for step S4 is as follows: after collecting a large number of portrait cues and generating corresponding portrait images, the portrait image samples are scored in descending order and the score regions are evenly divided using a portrait reward model. Image samples are then extracted from different score regions to construct the image.
[0016] Furthermore, in step S5, the partition preference optimization algorithm optimizes the Chinese text-based image model to be optimized based on the portrait partition preference dataset and the direct preference loss to improve the portrait generation quality of the Chinese text-based image model.
[0017] Another technical solution provided by the present invention is a computer-readable storage medium storing a computer program, wherein the computer-readable storage medium stores the aforementioned optimization system, and / or controls the device where the computer-readable storage medium is located to execute the aforementioned optimization system when the computer program is running.
[0018] Compared with existing technologies, the beneficial effects of this solution are:
[0019] 1. This solution constructs a portrait quality partitioning preference optimization system based on preference learning, carefully designs annotation criteria and aggregation algorithms, trains a reward model by minimizing the relative entropy between the annotation preference distribution and the sample normalization score, and finally uses the trained reward model to score the portrait quality of portrait image data. A portrait preference dataset is constructed, and the designed partitioning preference optimization algorithm is applied to the raw image model to improve the portrait generation quality of the raw image model. To demonstrate the effectiveness and generalization of this invention, black-box manual evaluation shows that the optimized model significantly improves the portrait generation quality compared to the unoptimized model, proving the generalization of the reward model and the effectiveness of the system.
[0020] 2. This scheme constructs a dataset related to Chinese contextual portraits through the proposed attention aggregation algorithm, and obtains a Chinese contextual portrait reward model through post-training based on the pre-trained Chinese visual-language contrastive learning model (ChineseCLIP). This reward model is suitable for automatic scoring of portrait generation quality in Chinese contextual text-based image scenarios, significantly improving the model's ability to distinguish portrait quality. Based on this model, a partitioning preference optimization algorithm based on preference learning theory is proposed to improve the portrait generation quality of the Chinese text-based image model. Attached Figure Description
[0021] Figure 1 This is a flowchart of the portrait generation quality partitioning optimization system based on Chinese contextual text-based image scenarios of the present invention;
[0022] Figure 2 This is a case study diagram of the image generation quality zoning optimization system based on Chinese contextual text-based image scenarios according to the present invention. Detailed Implementation
[0023] The present invention will be further described in detail below through specific embodiments:
[0024] Example 1
[0025] like Figure 1 and Figure 2 As shown, the image generation quality zoning optimization system based on Chinese contextualized text-based image scenarios includes the following steps:
[0026] S1. Construct a portrait annotation dataset by building data annotation rules based on the five common degradation problems in portraits generated by the Chinese text-based image model. Data sources include: collecting commonly used or poorly generated portrait prompts and generating corresponding images using the Chinese text-based image model. Annotation standards are based on degradation problems in portraits generated by the Chinese text-based image model, namely, issues related to the number of limbs, limb shape, hands, interaction, and faces. Each major problem has different detailed classifications and degrees of degradation to facilitate more accurate annotation by annotators.
[0027] S2. The portrait annotation dataset is transformed into a portrait annotation preference dataset using an attention aggregation algorithm. The aggregation rule is as follows: For each annotation, different degradation types are assigned different weights, representing the user's attention to that type of degradation issue. Finally, the annotation results are combined. Each image sample has a total attention score, which is the sum of the weights of all annotators for the issues they annotated. Lower attention scores indicate lower portrait degradation, and vice versa. An image sample is paired with a sample having a higher attention score to form a preference data pair, thus creating a portrait annotation preference dataset.
[0028] S3. The pre-trained ChineseCLIP model is used as the base model (the ChineseCLIP model in this embodiment is implemented using, but is not limited to, the content described in the literature "Chinese CLIP: Contrastive Vision-Language Pretraining in Chinese", which has already implemented the alignment of Chinese visual information and language information internally), which can meet the needs of Chinese scenarios and obtain the ability of pre-trained Chinese semantic alignment. In order to add the ability to discriminate the quality of human images to the base model, this training combines the base model with the above-mentioned human image annotation preference dataset for reward model training, and completes fine-tuning by minimizing the cross-entropy between the annotation preference distribution and the sample normalization score. That is, the cross-entropy between the annotation preference distribution and the sample normalization score is calculated. Specifically, given two samples, if the first sample is of higher quality, then p = [1, 0], otherwise p = [0, 1], and the scorer s(c, x) receives the prompt c and the human image x.
[0029] make
[0030] The loss function is: Based on the above results, the gradients of all parameters to be optimized in the model are calculated. Then, the gradient descent method is used to update all parameters to be optimized in the model, and the training of the portrait reward model is completed. This step improves the consistency of the existing model with the quality of Chinese portraits in terms of human preferences. Its accuracy in predicting human preferences on the test set relative to the base model increases from 49.02% to 92.46%.
[0031] S4. Collect Chinese image cues and generate image images using a Chinese text-based image model. Use an image reward model to assign reward signals to the image image samples, constructing an image partitioning preference dataset. The method is as follows: After collecting a large number of image cues and generating corresponding image images, use an image reward model to score the image image samples for quality, then sort them in descending order and evenly divide the score regions. Image samples are extracted from different score regions to construct the dataset. That is: collect Chinese image cues and use the initial Chinese text-based image model to be optimized. θ (x, t, c) Generate an image dataset, where x is an image, t is the sampling time step in the diffusion model, and c is a text prompt. The reward model trained using the above process assigns an image quality score to each sample in the image dataset to represent the reward signal.
[0032] S5. A partitioning preference optimization algorithm is used for the Chinese image-generated image model. Preference learning is a machine learning method that constructs a preference prediction model using known, observable preference information. Specifically, it involves collecting human feedback and manually annotating data samples according to predefined standards, then aggregating these annotated data to construct a corresponding dataset for model optimization. In other words, the Chinese image-generated image model is optimized based on the human image partitioning preference dataset and direct preference loss to improve the image generation quality. The parameters of the Chinese image-generated image model are adjusted until training is complete.
[0033] That is: for the N image samples with the same text prompts, first sort them in descending order and divide them into c equal score regions S1, ..., S2. c Image samples located in the score region with smaller indices have higher scores. To construct data pairs with significant differences in sample quality, image samples from different score regions are selected to build preference data pairs (c, x). w x l ), that is, x w ∈S i x l ∈S j If i < j, then construct a human portrait partitioning preference dataset.
[0034] Based on the above portrait zoning preference dataset Optimization based on the direct preference loss model
[0035]
[0036] Where, ∈ ref Is using ∈ θ Initialized reference model, It follows a multidimensional Gaussian distribution, α tand σ t β is a predefined diffusion timeline parameter, and β is a hyperparameter controlling the degree of regularization. Using the constructed portrait partitioning preference dataset and calculating the loss for preference optimization, gradient descent is used to update all parameters to be optimized in the Wensheng image model, completing model training.
[0037] The optimization system of this embodiment is used to optimize the human image generation quality of the Chinese text-based image model, and the optimized Chinese text-based image model and the unoptimized version are subjected to black-box human comparison test, under the conditions of unknown prompts and seed generation.
[0038] The aforementioned black-box testing yielded the following scores, with higher scores indicating higher image quality. The experiment consisted of two groups: one using a normal 25-step sampling strategy, and the other using a reduced 8-step sampling strategy. The testing method was as follows: Under the same sampling strategy and environmental conditions, images were generated using both the pre- and post-optimization models. Subsequently, three evaluators assessed the generated images, selecting either the better image or a "tie" to indicate that the two images were of comparable quality. The following values represent the number of images selected as better. The experimental results show that, regardless of whether the normal or reduced-step sampling strategy is used, the optimization system proposed in this embodiment significantly improves the image generation quality of the Chinese text-based image model, demonstrating the generalization ability of the reward model and the effectiveness of this optimization system.
[0039]
[0040] Example 2
[0041] A computer-readable storage medium storing a computer program, wherein the computer-readable storage medium stores an optimized system of Embodiment 1, and / or controls the device where the computer-readable storage medium is located to execute the optimized system of Embodiment 1 when the computer program is executed.
[0042] The above are merely embodiments of the present invention, and common knowledge such as specific structures and / or characteristics in the solutions are not described in detail here. It should be noted that those skilled in the art can make various modifications and improvements without departing from the structure of the present invention, and these should also be considered within the scope of protection of the present invention. These modifications and improvements will not affect the effectiveness of the implementation of the present invention or the practicality of the patent. The scope of protection claimed in this application should be determined by the content of its claims, and the specific embodiments described in the specification can be used to interpret the content of the claims.
Claims
1. A system for optimizing the quality of portrait generation in Chinese contextualized text-based image scenarios, characterized by: Includes the following steps: S1. Construct data annotation rules based on the degradation problem of generated human images to build a human image annotation dataset; S2. Use the attention aggregation algorithm to transform the human image annotation dataset into a human image annotation preference dataset; S3. Using the pre-trained ChineseCLIP model as the base model, the parameters of the ChineseCLIP model are adjusted using the portrait annotation preference dataset and the preset cross-entropy function to obtain the portrait reward model. S4. Collect Chinese portrait cues and generate portrait images using a Chinese text-based image model. Use a portrait reward model to assign reward signals to the portrait image samples and construct a portrait partition preference dataset. S5. Apply the partition preference optimization algorithm to the Chinese text-to-image model to adjust the parameters of the Chinese text-to-image model until training is complete; The method for step S4 is as follows: After collecting a large number of portrait cues and generating corresponding portrait images, the portrait image samples are scored in descending order and the score regions are evenly divided using a portrait reward model. Image samples are then extracted from different score regions to construct the image. In step S5, the partition preference optimization algorithm optimizes the Chinese text-based image model to be optimized based on the portrait partition preference dataset and the direct preference loss to improve the portrait generation quality of the Chinese text-based image model.
2. The image generation quality zoning optimization system based on Chinese contextualized text-based image scenarios according to claim 1, characterized in that: Step S1 Data Source: Collect commonly used or low-quality human image prompts, and use the Chinese text-based image model to generate corresponding different images.
3. The image generation quality zoning optimization system based on Chinese contextualized text-based image scenarios according to claim 1, characterized in that: The annotation standards for step S1 are constructed by considering the number of limbs, limb shape, hand, interaction, and face issues in the collapse problem.
4. The image generation quality zoning optimization system based on Chinese contextualized text-based image scenarios according to claim 1, characterized in that: The aggregation rule for step S2 is to use a focus aggregation algorithm, which means that for each annotation, different types of breakdowns will be assigned a different weight, representing the user's focus on this type of breakdown problem. Finally, the annotation results are combined. For each image sample, there is a total attention level, which is the sum of the weights of all annotators for the questions they annotated. The lower the attention level, the less the portrait is ruined, and vice versa. An image sample and a sample with a higher attention level form a preference data pair, thus forming a portrait annotation preference dataset.
5. The image generation quality zoning optimization system based on Chinese contextualized text-based image scenarios according to claim 1, characterized in that: In step S3, the training process of the ChineseCLIP model is to minimize the cross-entropy between the labeled preference distribution and the normalized score of the sample, so that the portrait reward model can obtain human preference knowledge of portrait quality.
6. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, wherein the computer-readable storage medium stores an optimization system as described in any one of claims 1-5, and / or controls the device where the computer-readable storage medium is located to execute the optimization system as described in any one of claims 1-5 when the computer program is executed.
Citation Information
Patent Citations
Method, model and device for training text graph model, and electronic equipment
CN116894880A
Method and system for generating image model by optimizing text through artificial feedback reinforcement learning
CN116955972A