Prompt information optimization method and device

By introducing a prompt information optimization model in the literary graphics model, and using semantic alignment and image quality indicators to optimize prompt information, the problem of insufficient flexibility in prompt information design in the prior art is solved, and the generated image is more in line with user intentions and preferences, and the image quality is improved.

CN119962686APending Publication Date: 2025-05-09LENOVO (BEIJING) LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510122758.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-24
Publication Date
2025-05-09

AI Technical Summary

Technical Problem

The prior art is not flexible enough when assisting users in input prompt information, and it is difficult to fully ensure that the generated images meet user intentions and preferences.

Method used

By obtaining the prompt information to be optimized and inputting it into the prompt information to optimize the model, using semantic alignment and image quality as indicators for reasoning, the optimized prompt information is obtained.

Benefits of technology

Ensure that the generated images are not only more in line with user intentions and preferences, but also ensure that the generated images have high quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119962686A_ABST
    Figure CN119962686A_ABST
Patent Text Reader

Abstract

The invention discloses a prompt information optimization method and device, and relates to the technical field of artificial intelligence. The method comprises the following steps: acquiring prompt information to be optimized; inputting the to-be-optimized prompt information into a prompt information optimization model, and obtaining optimized prompt information obtained by reasoning the to-be-optimized prompt information by the prompt information optimization model; wherein in the process that the prompt information optimization model carries out reasoning on the prompt information to be optimized to obtain the optimized prompt information, reasoning is carried out by taking the semantic alignment degree between the prompt information to be optimized and a target image and / or the quality of the target image as indexes; the target image is an image generated by an image generation model based on the prompt information in the inference process of the prompt information optimization model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence technology, and in particular to a prompt information optimization method and device. Background Art

[0002] Wenshengtu is a major branch of artificial intelligence generated content (AIGC) technology. It can transform prompt information into specific and vivid images by inputting a prompt information, and has broad application prospects. In this process, the design of prompt information plays an important role and directly affects the quality of the final generated image.

[0003] In order to improve the quality of images generated by the text-based graph model, related technologies have designed function words to assist users in inputting prompt information for specific text-based graph models. Although this method can help users describe the required images more accurately to a certain extent, it still has the problem of insufficient flexibility and it is difficult to fully ensure that the generated images meet the user's intentions and preferences. Summary of the invention

[0004] In order to solve the above technical problems, the embodiments of the present application provide a prompt information optimization method and device.

[0005] The technical solution of the embodiment of the present application is implemented as follows:

[0006] In a first aspect, an embodiment of the present application provides a prompt information optimization method, comprising:

[0007] Get the prompt information to be optimized;

[0008] Inputting the prompt information to be optimized into a prompt information optimization model to obtain optimized prompt information obtained by the prompt information optimization model through reasoning on the prompt information to be optimized;

[0009] Among them, in the process of the prompt information optimization model inferring the prompt information to be optimized to obtain the optimized prompt information, the reasoning is performed based on the degree of semantic alignment between the prompt information to be optimized and the target image, and / or the quality of the target image; the target image is an image generated by the image generation model based on the prompt information in the reasoning process of the prompt information optimization model.

[0010] In some embodiments, the training process of the prompt information optimization model includes:

[0011] Acquire a first prompt information sample and a first image sample; the first image sample is generated by the image generation model based on the second prompt information sample in the prompt information optimization model training process;

[0012] Scoring the semantic alignment between the first prompt information sample and the first image sample, and / or the quality of the first image sample based on the reward model to obtain a scoring result;

[0013] The parameters of the pre-trained language model are adjusted based on the scoring result to obtain the prompt information optimization model; the prompt information optimization model is used to infer the prompt information to be optimized to obtain the optimized prompt information.

[0014] In some embodiments, the scoring result corresponding to the semantic alignment degree includes at least one of the following: the cosine similarity between the first prompt information sample and the first image sample, and the completion score obtained by answering the question generated based on the first prompt information sample based on the first image sample;

[0015] The scoring result corresponding to the quality of the first image sample includes at least one of the following: a quality score used to characterize whether the first image sample meets human preferences, and a quality score used to characterize the aesthetic degree of the first image sample in human vision.

[0016] In some embodiments, the reward model includes a first reward model, and the first reward model is used to score the semantic alignment between the first prompt information sample and the first image sample;

[0017] The training process of the first reward model includes:

[0018] Determine a semantic alignment between a third prompt information sample and a second image sample; the second image sample is generated by the image generation model based on a fourth prompt information sample, and the fourth prompt information sample is an optimized prompt information sample obtained by inferring the third prompt information sample by the pre-trained language model;

[0019] The third prompt information sample is used as training data, and the semantic alignment between the third prompt information sample and the second image sample is used as a training label, and the parameters of the first language model are adjusted to obtain the first reward model.

[0020] In some embodiments, determining the semantic alignment between the third prompt information sample and the second image sample comprises at least one of the following:

[0021] Determine a cosine similarity between the third prompt information sample and the second image sample as a first semantic alignment;

[0022] Based on the third prompt information sample, generate at least one question related to the content of the third prompt information sample; based on the second image sample, answer the at least one question to obtain results corresponding to each question; based on the results corresponding to each question, determine the second semantic alignment between the third prompt information sample and the second image sample.

[0023] In some embodiments, the method of using the third prompt information sample as training data, and using the semantic alignment between the third prompt information sample and the second image sample as a training label, adjusting the parameters of the first language model, and obtaining the first reward model includes:

[0024] The third prompt information sample is used as training data, and at least one of the first semantic alignment and the second semantic alignment is used as a training label, and the parameters of the first language model are adjusted to obtain the first reward model.

[0025] In some embodiments, the reward model includes a second reward model, the second reward model being used to score the quality of the first image sample;

[0026] The training process of the second reward model includes:

[0027] determining a first quality score of a third image sample, where the first quality score is used to characterize whether the third image sample conforms to human preference; the third image sample is generated by the image generation model based on a fifth prompt information sample, and the fifth prompt information sample is an optimized prompt information sample obtained by inferring the sixth prompt information sample by the pre-trained language model;

[0028] The sixth prompt information sample is used as training data, and the first quality score is used as a training label, and the parameters of the second language model are adjusted to obtain the second reward model.

[0029] In some embodiments, the reward model includes a third reward model, and the third reward model is used to score the quality of the first image sample;

[0030] The training process of the third reward model includes:

[0031] determining a second quality score of a fourth image sample, where the second quality score is used to characterize the aesthetic degree of the fourth image sample in human vision; the fourth image sample is generated by the image generation model based on the seventh prompt information sample, and the seventh prompt information sample is an optimized prompt information sample obtained by inferring the eighth prompt information sample by the pre-trained language model;

[0032] The eighth prompt information sample is used as training data, and the second quality score is used as a training label, and the parameters of the third language model are adjusted to obtain the third reward model.

[0033] In some embodiments, the training process of the pre-trained language model includes:

[0034] Based on at least one prompt information pair, adjusting parameters of an initial pre-trained language model to obtain the pre-trained language model;

[0035] Each of the prompt information pairs includes two prompt information of different qualities, and the semantic similarity of the two prompt information is greater than a specific value.

[0036] In a second aspect, an embodiment of the present application provides a prompt information optimization device, including:

[0037] An acquisition module is used to obtain prompt information to be optimized;

[0038] An optimization module, used for inputting the prompt information to be optimized into a prompt information optimization model, and obtaining optimized prompt information obtained by the prompt information optimization model through reasoning on the prompt information to be optimized;

[0039] Among them, in the process of the prompt information optimization model inferring the prompt information to be optimized to obtain the optimized prompt information, the reasoning is performed based on the degree of semantic alignment between the prompt information to be optimized and the target image, and / or the quality of the target image; the target image is an image generated by the image generation model based on the prompt information in the reasoning process of the prompt information optimization model. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] In order to more clearly illustrate the technical solutions of the embodiments of the present application, a brief introduction will be given below to the drawings required for use in the embodiments. Obviously, the drawings described below are some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0041] Figure 1 A flowchart of a prompt information optimization method provided in an embodiment of the present application;

[0042] Figure 2 One of the flowcharts of a prompt information optimization model training provided in an embodiment of the present application;

[0043] Figure 3 A second flowchart of a prompt information optimization model training method provided in an embodiment of the present application;

[0044] Figure 4 One of the flowcharts of a first reward model training provided in an embodiment of the present application;

[0045] Figure 5 A second flowchart of a first reward model training process provided in an embodiment of the present application;

[0046] Figure 6 A schematic diagram of a process of training a second reward model provided in an embodiment of the present application;

[0047] Figure 7 A schematic diagram of a process of training a third reward model provided in an embodiment of the present application;

[0048] Figure 8 A schematic diagram of the principle of training a pre-trained language model provided in an embodiment of the present application;

[0049] Fig. 9 A schematic diagram of the principle of prompt information optimization model training provided in an embodiment of the present application;

[0050] Fig.10 A schematic diagram of the principle of reward model training provided in an embodiment of the present application;

[0051] Fig.11 A schematic diagram of the structure of a prompt information optimization device provided in an embodiment of the present application;

[0052] Fig.12 A schematic diagram of the physical structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0053] In order to make the purpose, technical solution and advantages of the embodiments of the present application clearer, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the embodiments of the present application.

[0054] It should be noted that in the description of the embodiments of the present application, the terms "first", "second", etc. are used to distinguish similar objects, and are not used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable under appropriate circumstances, so that the embodiments of the present application can be implemented in an order other than those illustrated or described here, and the objects distinguished by "first", "second", etc. are generally of the same type, and the number of objects is not limited. For example, the first object can be one or more.

[0055] In order to facilitate a clearer understanding of the embodiments of the present application, some relevant technical knowledge is first introduced as follows.

[0056] In recent years, the technology of text-based graphs has made significant progress. This technology allows ordinary users without relevant professional knowledge to create diverse images. However, when using the text-based graph model, users need to write text prompts before model inference, and the design of text prompts in this process has a decisive influence on the final image effect. However, ordinary users often need to go through multiple iterations and modify text prompts to achieve satisfactory results required for practical applications. For non-professional users, designing appropriate text prompts is quite challenging, and frequent modifications will also lead to a large consumption of time and computing resources.

[0057] In order to improve the quality of images generated by the text-based image model, the current methods for automatically optimizing text prompts are mainly divided into two categories: completion and rewriting the original text prompts. The completion method is to add some restrictive statements that meet the model's preferences on the basis of the original text prompts, aiming to improve the quality of image generation. For example, the SSP (Self-Support Prototype) method designs an optimal camera matching strategy and implements a classifier that can automatically match the original text prompts with camera descriptions. By adding the corresponding camera description to the original prompt text, the optimized prompt text is generated. However, adding only the camera description does not significantly improve the image quality because there is a difference between the prompt text entered by ordinary users and the prompt text preferred by the model. The rewriting method uses the knowledge learned by the language model to adjust the original text prompt to a text prompt that is more in line with the model's preferences, thereby generating a more aesthetically pleasing image.

[0058] In order to overcome at least some of the above-mentioned defects existing in the related art, the embodiments of the present application provide a prompt information optimization method and device. Through the inference process of the prompt information optimization model, under the constraints of semantic alignment and image quality, it can be ensured that the image generated by the optimized prompt information obtained by reasoning is not only more in line with the user's intentions and preferences, but also can ensure that the generated image has a higher quality.

[0059] The following is an illustrative introduction to the prompt information optimization method and device provided in the embodiments of the present application in conjunction with the drawings in the embodiments of the present application.

[0060] Figure 1 A flowchart of a prompt information optimization method provided in an embodiment of the present application is shown in FIG. Figure 1 As shown, the method includes:

[0061] S101: Obtain prompt information to be optimized.

[0062] It should be noted that the prompt information to be optimized may be text information. Since the image generated by the current image generation model or text graph model directly based on the prompt information may not meet the user's intention and preference, the prompt information may be optimized first, so that the image generated by the image generation model or text graph model based on the optimized prompt information can better meet the user's intention and preference.

[0063] It is understandable that since the quality of the prompt information to be optimized is low, the low-quality prompt information can be optimized to obtain high-quality prompt information, and then the high-quality prompt information can be input into the image generation model or the text image model to obtain an image that better meets the user's intentions and preferences.

[0064] In the embodiment of the present application, the prompt information input by the user may be used as the prompt information to be optimized.

[0065] S102: input the prompt information to be optimized into a prompt information optimization model, and obtain optimized prompt information obtained by the prompt information optimization model through reasoning on the prompt information to be optimized.

[0066] Among them, in the process of the prompt information optimization model inferring the prompt information to be optimized to obtain the optimized prompt information, the reasoning is performed based on the degree of semantic alignment between the prompt information to be optimized and the target image, and / or the quality of the target image; the target image is an image generated by the image generation model based on the prompt information in the reasoning process of the prompt information optimization model.

[0067] It should be noted that the prompt information optimization model can be a model that is pre-built and trained for optimizing prompt information. Moreover, in the process of the prompt information optimization model inferring the prompt information to be optimized to obtain the optimized prompt information, the semantic alignment between the prompt information to be optimized and the target image and / or the quality of the target image can be used as indicators to perform inference to obtain the optimized prompt information, wherein the target image is an image generated by the image generation model based on the prompt information in the inference process of the prompt information optimization model.

[0068] It should be noted that the prompt information optimization model may be a language model. The present embodiment of the present application does not specifically limit the structure of the prompt information optimization model.

[0069] In an embodiment of the present application, the prompt information to be optimized can be input into the prompt information optimization model, so that the prompt information optimization model performs iterative reasoning based on the semantic alignment between the prompt information to be optimized and the target image and / or the quality of the target image as indicators, and the prompt information optimization model continuously adjusts the prompt information it generates based on the feedback of the semantic alignment and image quality until a certain number of iterative reasoning times is reached, or the semantic alignment reaches a specific threshold and / or the image quality reaches a specific threshold to obtain the optimized prompt information. The target image is a dynamically changing image, which is an image generated by the image generation model based on the prompt information continuously inferred during the iterative reasoning process of the prompt information optimization model.

[0070] It should be noted that the image generation model is an artificial intelligence model that can generate a corresponding image according to the input prompt information, and can be a text-generated image model. The embodiment of the present application does not specifically limit the structure of the image generation model.

[0071] It can be understood that the prompt information optimization method provided in the embodiment of the present application inputs the prompt information to be optimized into the prompt information optimization model, and obtains the optimized prompt information obtained by the prompt information optimization model through reasoning the prompt information to be optimized, and in the process of the prompt information optimization model reasoning the prompt information to be optimized to obtain the optimized prompt information, the prompt information optimization model can use the semantic alignment between the prompt information to be optimized and the target image, and / or the quality of the target image as indicators for reasoning, wherein the target image is the image generated by the image generation model based on the prompt information in the reasoning process of the prompt information optimization model. In this way, under the constraints of semantic alignment and image quality during the reasoning process of the prompt information optimization model, it can be ensured that the image generated by the optimized prompt information obtained by reasoning is not only more in line with the user's intentions and preferences, but also can ensure that the generated image has a higher quality.

[0072] In some embodiments, the training process of the prompt information optimization model includes:

[0073] Acquire a first prompt information sample and a first image sample; the first image sample is generated by the image generation model based on the second prompt information sample in the prompt information optimization model training process;

[0074] Scoring the semantic alignment between the first prompt information sample and the first image sample, and / or the quality of the first image sample based on the reward model to obtain a scoring result;

[0075] The parameters of the pre-trained language model are adjusted based on the scoring result to obtain the prompt information optimization model; the prompt information optimization model is used to infer the prompt information to be optimized to obtain the optimized prompt information.

[0076] It should be noted that the reward model may be a pre-built and trained model for scoring the semantic alignment between the prompt information and the image and / or the quality of the image. The embodiment of the present application does not specifically limit the structure of the reward model.

[0077] It should be noted that the pre-trained language model can be a pre-built and trained model that can optimize the prompt information, but compared with the prompt information optimization model in the embodiment of the present application, its optimization performance is lower. The embodiment of the present application does not specifically limit the structure of the pre-trained language model.

[0078] It can be understood that after the first prompt information sample is input into the pre-trained language model as training data, the pre-trained language model will continuously adjust its parameters based on the dynamic scoring results of the reward model on the semantic alignment and / or image quality until a certain number of iterations is reached, or the semantic alignment reaches a specific threshold and / or the image quality reaches a specific threshold, then a prompt information optimization model is obtained.

[0079] It should be noted that, since the training process of the prompt information optimization model is an iterative training process, optimized second prompt information samples will be continuously generated during the training process. Since the second prompt information samples are constantly changing, the first image samples are also constantly changing, and then the semantic alignment between the first prompt information samples and the first image samples and the quality of the first image samples are also constantly changing. Therefore, the pre-trained language model continuously adjusts its parameters based on the constantly changing semantic alignment and / or image quality, and finally obtains a prompt information optimization model with better performance.

[0080] For example, Figure 2 One of the flowcharts of a prompt information optimization model training provided in an embodiment of the present application is as follows: Figure 2 As shown, the process includes:

[0081] S201. Obtain a first prompt information sample and a first image sample; the first image sample is generated by an image generation model based on a second prompt information sample in a prompt information optimization model training process.

[0082] S202: Score the semantic alignment between the first prompt information sample and the first image sample, and / or the quality of the first image sample based on a reward model to obtain a scoring result.

[0083] S203: adjusting the parameters of the pre-trained language model based on the scoring result to obtain the prompt information optimization model; the prompt information optimization model is used to infer the prompt information to be optimized to obtain the optimized prompt information.

[0084] In some embodiments, adjusting the parameters of the pre-trained language model based on the scoring result to obtain the prompt information optimization model includes:

[0085] The scoring result is used as a reward signal fed back by the reward model to the pre-trained language model, and the parameters of the pre-trained language model are adjusted based on the reward signal to obtain the prompt information optimization model.

[0086] In the embodiment of the present application, the reward model is used to score the semantic alignment between the first prompt information sample and the first image sample and / or the quality of the first image sample, and the scoring result is used as the reward signal fed back by the reward model to the pre-trained language model, and then the parameters of the pre-trained language model are adjusted based on the reward signal to obtain the prompt information optimization model. In this way, the training of the prompt information optimization model based on the reinforcement learning strategy is realized, aiming to optimize the pre-trained language model so that the prompt information optimization model finally trained can generate higher quality prompt information.

[0087] For example, Figure 3 The second flowchart of a prompt information optimization model training provided in an embodiment of the present application is as follows: Figure 3 As shown, the process includes:

[0088] S301, obtaining a first prompt information sample and a first image sample; the first image sample is generated by an image generation model based on a second prompt information sample in a prompt information optimization model training process.

[0089] S302: Score the semantic alignment between the first prompt information sample and the first image sample, and / or the quality of the first image sample based on a reward model to obtain a scoring result.

[0090] S303: Using the scoring result as a reward signal fed back by the reward model to the pre-trained language model, and adjusting the parameters of the pre-trained language model based on the reward signal to obtain the prompt information optimization model.

[0091] It should be noted that, for the description of the same steps and the same contents in this embodiment as those in other embodiments, reference can be made to the description in other embodiments and will not be repeated here.

[0092] It can be understood that the embodiment of the present application adjusts the parameters of the pre-trained language model according to the scoring result obtained by scoring the semantic alignment between the first prompt information sample and the first image sample and / or the quality of the first image sample based on the reward model, so as to efficiently find the optimal model parameters, help reduce the training time of the prompt information optimization model, and improve the training efficiency.

[0093] In some embodiments, the scoring result corresponding to the semantic alignment degree includes at least one of the following: the cosine similarity between the first prompt information sample and the first image sample, and the completion score obtained by answering the question generated based on the first prompt information sample based on the first image sample;

[0094] The scoring result corresponding to the quality of the first image sample includes at least one of the following: a quality score used to characterize whether the first image sample meets human preferences, and a quality score used to characterize the aesthetic degree of the first image sample in human vision.

[0095] In an embodiment of the present application, the cosine similarity between the first prompt information sample and the first image sample can be used as the scoring result of the reward model for scoring the semantic alignment between the first prompt information sample and the first image sample, and the completion score obtained by answering the question generated based on the first prompt information sample based on the first image sample can be used as the scoring result of the reward model for scoring the semantic alignment between the first prompt information sample and the first image sample; the quality score that can characterize whether the first image sample conforms to human preferences can be used as the scoring result of the reward model for scoring the quality of the first image sample, and the quality score that can characterize the degree of beauty of the first image sample in human vision can be used as the scoring result of the reward model for scoring the quality of the first image sample.

[0096] It should be noted that since cosine similarity can measure the similarity between two vectors in a vector space, the first prompt information sample and the first image sample can be converted into vectors respectively and then their cosine similarity can be calculated, and the calculation result can reflect the semantic similarity between the first prompt information sample and the first image sample.

[0097] In some embodiments, a set of questions related to the content of the first prompt information sample can be generated based on the first prompt information sample, and then the set of questions can be answered based on the first image sample. A completion score can be obtained based on the obtained answer results, and the completion score can reflect the semantic similarity between the first prompt information sample and the first image sample.

[0098] For example, five questions related to the content of the first prompt information sample are generated based on the first prompt information sample, and three of the five questions are answered correctly based on the first image sample, then the completion score is 3 / 5 = 0.6. In other words, it can be considered that the semantic similarity between the first prompt information sample and the first image sample is 0.6.

[0099] It should be noted that PickScore is a scoring model based on CLIP (Contrastive Language–Image Pre-training), which can be used to evaluate the performance of text-to-image generation (i.e., text-to-image) models. PickScore is trained based on a large amount of user preference data, so it can more accurately capture human expectations and preferences for image generation. Therefore, in some embodiments, the PickScore score can be calculated for the first image sample, and the PickScore score can be used as the scoring result of the reward model to score the quality of the first image sample.

[0100] In some embodiments, the aesthetic degree of the first image sample in human vision may be scored based on the aesthetic scoring model, and the scoring result may be used as the scoring result of the reward model for scoring the quality of the first image sample.

[0101] In some embodiments, the reward model includes a first reward model, and the first reward model is used to score the semantic alignment between the first prompt information sample and the first image sample;

[0102] The training process of the first reward model includes:

[0103] Determine a semantic alignment between a third prompt information sample and a second image sample; the second image sample is generated by the image generation model based on a fourth prompt information sample, and the fourth prompt information sample is an optimized prompt information sample obtained by inferring the third prompt information sample by the pre-trained language model;

[0104] The third prompt information sample is used as training data, and the semantic alignment between the third prompt information sample and the second image sample is used as a training label, and the parameters of the first language model are adjusted to obtain the first reward model.

[0105] It is understandable that in the embodiment of the present application, the fourth prompt information sample is a high-quality prompt information sample corresponding to the third prompt information sample. The second image sample can be generated based on the fourth prompt information sample by the image generation model, and then the third prompt information sample is used as training data, and the semantic alignment between the third prompt information sample and the second image sample is used as a training label, and the parameters of the first language model are adjusted to obtain a first reward model that can be used to score the semantic alignment between the prompt information and the image. Among them, the semantic alignment between the third prompt information sample and the second image sample is calculated by a related semantic alignment algorithm.

[0106] It should be noted that the first language model may be a model that can process prompt information data and learn the semantic relationship between prompt information and images. The embodiment of the present application does not specifically limit the structure of the first language model.

[0107] For example, Figure 4 One of the flowcharts of a first reward model training provided in an embodiment of the present application is as follows: Figure 4 As shown, the process includes:

[0108] S401, determining the semantic alignment between the third prompt information sample and the second image sample; the second image sample is generated by an image generation model based on a fourth prompt information sample, and the fourth prompt information sample is an optimized prompt information sample obtained by inferring the third prompt information sample with a pre-trained language model.

[0109] S402: Use the third prompt information sample as training data, and use the semantic alignment between the third prompt information sample and the second image sample as a training label, adjust the parameters of the first language model, and obtain a first reward model.

[0110] It can be understood that the embodiment of the present application uses the third prompt information sample as training data and the semantic alignment between the third prompt information sample and the second image sample as training labels to adjust the parameters of the first language model, thereby effectively obtaining a first reward model that can accurately evaluate the semantic alignment between the prompt information and the image.

[0111] In some embodiments, determining the semantic alignment between the third prompt information sample and the second image sample comprises at least one of the following:

[0112] Determine a cosine similarity between the third prompt information sample and the second image sample as a first semantic alignment;

[0113] Based on the third prompt information sample, generate at least one question related to the content of the third prompt information sample; based on the second image sample, answer the at least one question to obtain results corresponding to each question; based on the results corresponding to each question, determine the second semantic alignment between the third prompt information sample and the second image sample.

[0114] In the embodiment of the present application, the semantic alignment between the third prompt information sample and the second image sample can be determined in two ways. One of the ways is: after converting the third prompt information sample and the second image sample into vectors respectively, the cosine similarity between the two vectors is calculated, and the cosine similarity is used as the first semantic alignment between the third prompt information sample and the second image sample; the other way is: firstly based on the third prompt information sample, at least one question related to the content of the third prompt information sample can be generated through the relevant model, and then based on the second image sample, at least one generated question can be answered to obtain the answer results corresponding to each question, and then based on the answer results corresponding to each question, a completion score is obtained, and the completion score is used as the second semantic alignment between the third prompt information sample and the second image sample.

[0115] It can be understood that the embodiment of the present application uses the cosine similarity between the third prompt information sample and the second image sample, and the completion score obtained by using the second image sample to answer the question generated based on the third prompt information sample, as a measurement indicator for evaluating the semantic alignment between the third prompt information sample and the second image sample, so as to accurately and comprehensively evaluate the semantic alignment between the third prompt information sample and the second image sample.

[0116] In some embodiments, the method of using the third prompt information sample as training data, and using the semantic alignment between the third prompt information sample and the second image sample as a training label, adjusting the parameters of the first language model, and obtaining the first reward model includes:

[0117] The third prompt information sample is used as training data, and at least one of the first semantic alignment and the second semantic alignment is used as a training label, and the parameters of the first language model are adjusted to obtain the first reward model.

[0118] In an embodiment of the present application, the third prompt information sample can be used as training data, and at least one of the first semantic alignment calculated by cosine similarity and the second semantic alignment calculated by question answering can be used as training labels to adjust the parameters of the first language model. This can effectively obtain a first reward model that can accurately and comprehensively evaluate the semantic alignment between the prompt information and the image.

[0119] For example, Figure 5 A second flow chart of a first reward model training process provided in an embodiment of the present application is as follows: Figure 5 As shown, the process includes:

[0120] S501: Determine the cosine similarity between the third prompt information sample and the second image sample as a first semantic alignment.

[0121] S502: Based on the third prompt information sample, generate at least one question related to the content of the third prompt information sample.

[0122] S503: Based on the second image sample, answer the at least one question to obtain a result corresponding to each question.

[0123] S504: Determine a second semantic alignment between the third prompt information sample and the second image sample based on the results corresponding to each of the questions.

[0124] S505: Use the third prompt information sample as training data, and use at least one of the first semantic alignment and the second semantic alignment as a training label, adjust the parameters of the first language model, and obtain a first reward model.

[0125] It should be noted that, for the description of the same steps and the same contents in this embodiment as those in other embodiments, reference can be made to the description in other embodiments and will not be repeated here.

[0126] It can be understood that the embodiment of the present application uses the third prompt information sample as training data, and uses at least one of the first semantic alignment calculated by cosine similarity and the second semantic alignment calculated by question answering as training labels, and adjusts the parameters of the first language model. This allows the first reward model obtained after the parameter adjustment to accurately and comprehensively evaluate the semantic alignment between the prompt information and the image.

[0127] In some embodiments, the reward model includes a second reward model, the second reward model being used to score the quality of the first image sample;

[0128] The training process of the second reward model includes:

[0129] determining a first quality score of a third image sample, where the first quality score is used to characterize whether the third image sample conforms to human preference; the third image sample is generated by the image generation model based on a fifth prompt information sample, and the fifth prompt information sample is an optimized prompt information sample obtained by inferring the sixth prompt information sample by the pre-trained language model;

[0130] The sixth prompt information sample is used as training data, and the first quality score is used as a training label, and the parameters of the second language model are adjusted to obtain the second reward model.

[0131] It is understandable that in the embodiment of the present application, the fifth prompt information sample is a high-quality prompt information sample corresponding to the sixth prompt information sample. The third image sample can be generated based on the fifth prompt information sample by the image generation model, and then the first quality score that can characterize whether the third image sample meets human preferences is calculated for the third image sample based on a related algorithm (such as a PickScore model), and then the sixth prompt information sample is used as training data, and the first quality score is used as a training label, and the parameters of the second language model are adjusted to obtain a second reward model that can be used to score image quality.

[0132] It should be noted that the second language model may be a model that can process prompt information data and learn whether the generated image meets human preferences. The embodiment of the present application does not specifically limit the structure of the second language model.

[0133] For example, Figure 6 A schematic diagram of a second reward model training process provided in an embodiment of the present application, such as Figure 6 As shown, the process includes:

[0134] S601. Determine a first quality score of a third image sample, where the first quality score is used to characterize whether the third image sample conforms to human preference; the third image sample is generated by an image generation model based on a fifth prompt information sample, and the fifth prompt information sample is an optimized prompt information sample obtained by inferring a sixth prompt information sample with a pre-trained language model.

[0135] S602: Use the sixth prompt information sample as training data and the first quality score as a training label to adjust parameters of a second language model to obtain a second reward model.

[0136] It can be understood that the embodiment of the present application uses the sixth prompt information sample as training data and the first quality score of the third image sample as a training label to adjust the parameters of the second language model, thereby effectively obtaining a second reward model that can accurately evaluate image quality.

[0137] In some embodiments, the reward model includes a third reward model, and the third reward model is used to score the quality of the first image sample;

[0138] The training process of the third reward model includes:

[0139] determining a second quality score of a fourth image sample, where the second quality score is used to characterize the aesthetic degree of the fourth image sample in human vision; the fourth image sample is generated by the image generation model based on the seventh prompt information sample, and the seventh prompt information sample is an optimized prompt information sample obtained by inferring the eighth prompt information sample by the pre-trained language model;

[0140] The eighth prompt information sample is used as training data, and the second quality score is used as a training label, and the parameters of the third language model are adjusted to obtain the third reward model.

[0141] It is understandable that, in the embodiment of the present application, the seventh prompt information sample is a high-quality prompt information sample corresponding to the eighth prompt information sample. A fourth image sample can be generated based on the seventh prompt information sample by an image generation model, and then a second quality score that can characterize the aesthetic degree of the fourth image sample in human vision is calculated for the fourth image sample based on a related aesthetic scoring algorithm, and then the eighth prompt information sample is used as training data, and the second quality score is used as a training label, and the parameters of the third language model are adjusted to obtain a third reward model that can be used to score image quality.

[0142] It should be noted that the third language model may be a model that can process prompt information data and learn the aesthetics of the generated image in human vision. The embodiment of the present application does not specifically limit the structure of the third language model.

[0143] For example, Figure 7 A schematic diagram of a third reward model training process provided in an embodiment of the present application, such as Figure 7 As shown, the process includes:

[0144] S701. Determine a second quality score of the fourth image sample, where the second quality score is used to characterize the degree of beauty of the fourth image sample in human vision; the fourth image sample is generated by an image generation model based on a seventh prompt information sample, and the seventh prompt information sample is an optimized prompt information sample obtained by inferring an eighth prompt information sample with a pre-trained language model.

[0145] S702: Use the eighth prompt information sample as training data and the second quality score as a training label to adjust parameters of a third language model to obtain a third reward model.

[0146] It can be understood that the embodiment of the present application uses the eighth prompt information sample as training data and the second quality score of the fourth image sample as a training label to adjust the parameters of the third language model, thereby effectively obtaining a third reward model that can accurately evaluate image quality.

[0147] In some embodiments, the training process of the pre-trained language model includes:

[0148] Based on at least one prompt information pair, adjusting parameters of an initial pre-trained language model to obtain the pre-trained language model;

[0149] Each of the prompt information pairs includes two prompt information of different qualities, and the semantic similarity of the two prompt information is greater than a specific value.

[0150] In the embodiment of the present application, at least one prompt information pair can be collected, and each prompt information pair includes two prompt information of different quality (for example, one is low-quality prompt information and the other is corresponding high-quality prompt information), and the semantic similarity of the two prompt information is greater than a specific value. Then, the collected at least one prompt information pair is used as training data to adjust the parameters of the initial pre-trained language model to obtain a pre-trained language model after parameter adjustment.

[0151] It should be noted that the architecture of the initial pre-trained language model can be set based on actual applications, and the embodiments of the present application do not specifically limit its architecture.

[0152] It should be noted that the specific values ​​in the embodiments of the present application can be adaptively set based on actual applications, and the embodiments of the present application do not specifically limit this, for example, the specific values ​​are 0.85, 0.90, 0.95, 0.98, etc.

[0153] For example, Figure 8 A schematic diagram of the principle of pre-training language model training provided in an embodiment of the present application, such as Figure 8 As shown, the low-quality prompt information and its corresponding high-quality prompt information are input into the initial pre-trained language model to adjust the parameters of the initial pre-trained language model to obtain the pre-trained language model. Among them, the low-quality prompt information is the prompt information that has not been optimized, and the high-quality prompt information is the prompt information after the low-quality prompt information is optimized. Moreover, the low-quality prompt information and its corresponding high-quality prompt information have a high semantic similarity or semantic consistency.

[0154] In some embodiments, low-quality and high-quality text prompt pairs with high semantic consistency can be collected. In order to ensure the effectiveness of high-quality text prompts, the aesthetic score of the high-quality text prompts generated by the text-generated graph model can be calculated, and the text prompt pairs with low aesthetic scores can be filtered out. Then, the filtered low-quality and high-quality text prompt pairs are used as training sets to fine-tune the initial pre-trained language model, so that the fine-tuned pre-trained language model can output corresponding high-quality text prompts when low-quality text prompts are input.

[0155] It should be noted that although the fine-tuned pre-trained language model has the ability to optimize text prompts, the optimized text prompts cannot ensure that the images generated by the text graph model meet the semantics of low-quality text prompts and high-quality images perceived by humans. In this regard, the embodiment of the present application adopts a reinforcement learning strategy to further fine-tune the pre-trained language model. In order to ensure that the images generated by the text graph model meet the semantics of low-quality text prompts, the CLIP score between the low-quality text prompt and the image generated based on the corresponding high-quality text prompt can be first calculated. The CLIP score is calculated by the cosine similarity between the CLIP embedding vectors (embedding) corresponding to the low-quality text prompt and the image generated based on the corresponding high-quality text prompt. Although the CLIP score can reflect the semantic alignment between the text prompt and the image, this single score evaluation index is coarse-grained. Based on this, the embodiment of the present application introduces a fine-grained method for evaluating the semantic alignment between the text prompt and the image on the basis of this index. This method first generates some atomic, fine-grained and comprehensive questions based on the low-quality text prompt using a method for converting the text prompt into a set of questions related to the text prompt (DSG, Davidsonian SceneGraph). These questions involve the key points of the low-quality text prompt content. Then, a Visual Question Answering (VQA) model is used to answer these questions based on the images generated by the corresponding high-quality text prompts, and then a completion score is obtained based on the answer results. The completion score is used to evaluate the semantic alignment between the text prompt and the image. In summary, the CLIP score and the completion score can be used to train a reward model for evaluating the semantic alignment of images and texts. The output of the reward model is used as part of the reinforcement learning reward value to constrain the semantic alignment of the generated image.

[0156] In some embodiments, the CLIP score between the low-quality text prompt and the image generated based on the corresponding high-quality text prompt can be calculated as a coarse-grained evaluation indicator of the semantic alignment of the image and text. Then, based on the low-quality text prompt, some atomic, fine-grained and comprehensive questions are generated using DSG. These questions involve the key points of the content of the low-quality text prompt. The VQA model is then used to answer these questions based on the images generated by the corresponding high-quality text prompts, and then a completion score is obtained based on the results of the answers. The completion score is used as a fine-grained evaluation indicator of the semantic alignment of the image and text.

[0157] In some embodiments, a language model can be fine-tuned based on low-quality text prompts and corresponding CLIP scores and completion scores to obtain a reward model for evaluating the semantic alignment of images and texts. Low-quality text prompts can be input into the reward model to obtain corresponding semantic alignment scores.

[0158] In some embodiments, in order to improve the quality of image generation and make it more consistent with human perception, the PickScore and aesthetic score of the image generated based on the high-quality text prompt can be calculated and used as another part of the reinforcement learning reward value.

[0159] It should be noted that PickScore can, to a certain extent, characterize whether an image conforms to human preferences. The higher the PickScore score, the more the image conforms to human preferences. The aesthetic score objectively evaluates the quality of image generation.

[0160] In some embodiments, in order to better evaluate the quality of images generated by high-quality text prompts corresponding to low-quality text prompts, two language models can be fine-tuned for PickScore and aesthetic score respectively as reward models for evaluating the quality of generated images. Inputting low-quality text prompts into the corresponding reward model can obtain the corresponding PickScore or aesthetic score.

[0161] In some embodiments, a proximal policy optimization reinforcement learning algorithm (PPO) and a trained reward model can be used to train a fine-tuned pre-trained language model from the two aspects of semantic alignment and raw image quality to obtain a prompt information optimization model. The trained prompt information optimization model can automatically optimize the text prompt information while effectively taking into account the semantic alignment of the text and the raw image quality.

[0162] For example, Fig. 9 A schematic diagram of the principle of prompt information optimization model training provided in an embodiment of the present application, such as Fig. 9 As shown, low-quality prompt information is input into the pre-trained language model to obtain optimized high-quality prompt information, and then the high-quality prompt information is input into the image generation model to obtain the generated image, and then the CLIP score and completion score between the low-quality prompt information and the generated image, as well as the PickScore and aesthetic score of the generated image are calculated, and then based on the CLIP score, completion score, PickScore and aesthetic score, the pre-trained language model is trained using the proximal policy optimization algorithm PPO through the reward model, and the prompt information optimization model can be obtained after the training is completed.

[0163] For example, Fig.10 A schematic diagram of the principle of reward model training provided in an embodiment of the present application, such as Fig.10As shown in the figure, low-quality prompt information is input into the pre-trained language model to obtain optimized high-quality prompt information, which is then input into the image generation model to obtain the generated image, and then the CLIP score and completion score between the low-quality prompt information and the generated image, as well as the PickScore and aesthetic score of the generated image are calculated. Among them, for the calculation of the completion score, a text question set is generated for the low-quality prompt information using DSG, and then the VQA model is used to answer questions based on the generated image to obtain the completion score based on the answer results. Finally, the language model is fine-tuned based on the low-quality prompt information as well as the CLIP score, completion score, PickScore and aesthetic score to obtain the reward model.

[0164] It is understandable that the embodiment of the present application realizes automatic optimization of text prompt information based on the existing language model through supervised fine-tuning and reinforcement learning, and generates high-quality text prompt information. Compared with the cumbersome and time-consuming manual prompt information optimization method, it saves optimization time and improves the image generation effect. At the same time, by combining the coarse-grained CLIP score with the fine-grained completion score, the image generated by the optimized text prompt information has a higher semantic alignment than similar methods. Under the premise of ensuring the user's original prompt semantics, the quality of the generated image is also improved through the constraints of PickScore and aesthetic scoring, which is more in line with human preferences and human perception.

[0165] The prompt information optimization device provided in an embodiment of the present application is described below. The prompt information optimization device described below and the prompt information optimization method described above can be referenced to each other.

[0166] Fig.11 A schematic diagram of the structure of a prompt information optimization device provided in an embodiment of the present application, such as Fig.11 As shown, the device includes: an acquisition module 1110 and an optimization module 1120; wherein:

[0167] An acquisition module 1110 is used to acquire prompt information to be optimized;

[0168] An optimization module 1120 is used to input the prompt information to be optimized into a prompt information optimization model, and obtain optimized prompt information obtained by the prompt information optimization model through reasoning on the prompt information to be optimized;

[0169] Among them, in the process of the prompt information optimization model inferring the prompt information to be optimized to obtain the optimized prompt information, the reasoning is performed based on the degree of semantic alignment between the prompt information to be optimized and the target image, and / or the quality of the target image; the target image is an image generated by the image generation model based on the prompt information in the reasoning process of the prompt information optimization model.

[0170] The prompt information optimization device provided in the embodiment of the present application inputs the prompt information to be optimized into the prompt information optimization model, and obtains the optimized prompt information obtained by the prompt information optimization model through reasoning the prompt information to be optimized. Moreover, in the process of the prompt information optimization model reasoning the prompt information to be optimized to obtain the optimized prompt information, the prompt information optimization model can use the semantic alignment between the prompt information to be optimized and the target image, and / or the quality of the target image as indicators for reasoning, wherein the target image is an image generated by the image generation model based on the prompt information in the reasoning process of the prompt information optimization model. In this way, under the constraints of semantic alignment and image quality during the reasoning process of the prompt information optimization model, it can be ensured that the image generated by the optimized prompt information obtained by reasoning is not only more in line with the user's intentions and preferences, but also can ensure that the generated image has a higher quality.

[0171] In some embodiments, the apparatus further comprises a first training module; the first training module comprises:

[0172] An acquisition unit, configured to acquire a first prompt information sample and a first image sample; the first image sample is generated by the image generation model based on the second prompt information sample in the prompt information optimization model training process;

[0173] a scoring unit, configured to score the semantic alignment between the first prompt information sample and the first image sample, and / or the quality of the first image sample based on a reward model to obtain a scoring result;

[0174] A first parameter adjustment unit is used to adjust the parameters of the pre-trained language model based on the scoring result to obtain the prompt information optimization model; the prompt information optimization model is used to infer the prompt information to be optimized to obtain the optimized prompt information.

[0175] In some embodiments, the scoring result corresponding to the semantic alignment degree includes at least one of the following: the cosine similarity between the first prompt information sample and the first image sample, and the completion score obtained by answering the question generated based on the first prompt information sample based on the first image sample;

[0176] The scoring result corresponding to the quality of the first image sample includes at least one of the following: a quality score used to characterize whether the first image sample meets human preferences, and a quality score used to characterize the aesthetic degree of the first image sample in human vision.

[0177] In some embodiments, the reward model includes a first reward model, and the first reward model is used to score the semantic alignment between the first prompt information sample and the first image sample;

[0178] The device also includes a second training module; the second training module includes:

[0179] A first determining unit is used to determine the semantic alignment between a third prompt information sample and a second image sample; the second image sample is generated by the image generation model based on a fourth prompt information sample, and the fourth prompt information sample is an optimized prompt information sample obtained by inferring the third prompt information sample by the pre-trained language model;

[0180] The second parameter adjustment unit is used to use the third prompt information sample as training data and the semantic alignment between the third prompt information sample and the second image sample as a training label to adjust the parameters of the first language model to obtain the first reward model.

[0181] In some embodiments, the first determining unit is further used for at least one of the following:

[0182] Determine a cosine similarity between the third prompt information sample and the second image sample as a first semantic alignment;

[0183] Based on the third prompt information sample, generate at least one question related to the content of the third prompt information sample; based on the second image sample, answer the at least one question to obtain results corresponding to each question; based on the results corresponding to each question, determine the second semantic alignment between the third prompt information sample and the second image sample.

[0184] In some embodiments, the second parameter adjustment unit is further configured to:

[0185] The third prompt information sample is used as training data, and at least one of the first semantic alignment and the second semantic alignment is used as a training label, and the parameters of the first language model are adjusted to obtain the first reward model.

[0186] In some embodiments, the reward model includes a second reward model, the second reward model being used to score the quality of the first image sample;

[0187] The device further includes a third training module; the third training module includes:

[0188] a second determining unit, configured to determine a first quality score of a third image sample, wherein the first quality score is used to characterize whether the third image sample conforms to human preference; the third image sample is generated by the image generation model based on a fifth prompt information sample, and the fifth prompt information sample is an optimized prompt information sample obtained by inferring the sixth prompt information sample by the pre-trained language model;

[0189] The third parameter adjustment unit is used to use the sixth prompt information sample as training data and the first quality score as a training label to adjust the parameters of the second language model to obtain the second reward model.

[0190] In some embodiments, the reward model includes a third reward model, and the third reward model is used to score the quality of the first image sample;

[0191] The device further includes a fourth training module; the fourth training module includes:

[0192] a third determining unit, configured to determine a second quality score of a fourth image sample, wherein the second quality score is used to characterize the aesthetic degree of the fourth image sample in human vision; the fourth image sample is generated by the image generation model based on the seventh prompt information sample, and the seventh prompt information sample is an optimized prompt information sample obtained by inferring the eighth prompt information sample by the pre-trained language model;

[0193] The fourth parameter adjustment unit is used to use the eighth prompt information sample as training data and the second quality score as a training label to adjust the parameters of the third language model to obtain the third reward model.

[0194] In some embodiments, the apparatus further includes a fifth training module; the fifth training module includes:

[0195] a fifth parameter adjustment unit, configured to adjust parameters of the initial pre-trained language model based on at least one prompt information pair to obtain the pre-trained language model;

[0196] Each of the prompt information pairs includes two prompt information of different qualities, and the semantic similarity of the two prompt information is greater than a specific value.

[0197] It should be noted here that the above-mentioned prompt information optimization device provided in the embodiment of the present application can implement all the method steps implemented in the above-mentioned prompt information optimization method embodiment, and can achieve the same technical effect. The parts and beneficial effects that are the same as the method embodiment in this embodiment will not be described in detail here.

[0198] Fig.12 A schematic diagram of the physical structure of an electronic device provided in an embodiment of the present application, such as Fig.12 As shown, the electronic device may include: a processor 1210, a communication interface 1220, a memory 1230 and a communication bus 1240, wherein the processor 1210, the communication interface 1220 and the memory 1230 communicate with each other through the communication bus 1240. The processor 1210 may call the executable data instructions stored in the memory 1230 to execute part or all of the steps in the prompt information optimization method provided in the above embodiments.

[0199] In addition, the executable data instructions stored in the above-mentioned memory 1230 can be implemented in the form of a software functional unit and can be stored in a computer-readable storage medium when it is sold or used as an independent product. Based on this understanding, the technical solution of the embodiment of the present application can be essentially or partly reflected in the form of a software product that contributes to the relevant technology. The software product is stored in a storage medium, including several instructions to enable a computer device (which can be a personal computer, a server or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM), random access memory (RAM), disk or optical disk, etc. Various media that can store program codes.

[0200] An embodiment of the present application also provides a computer-readable storage medium, in which a computer program is stored. When the computer program is executed by a processor, some or all of the steps in the prompt information optimization method provided in the above embodiments are implemented.

[0201] An embodiment of the present application also provides a computer program product, which includes a computer program stored in a computer-readable storage medium, and the computer program includes program instructions. When the program instructions are executed by a computer, the computer can execute some or all of the steps in the prompt information optimization method provided in the above embodiments.

[0202] The device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. Those of ordinary skill in the art may understand and implement it without creative work.

[0203] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the embodiments of the present application may adopt the form of hardware embodiments, software embodiments, or embodiments in combination with software and hardware. Moreover, the embodiments of the present application may adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage and optical storage, etc.) that contain computer-usable program code.

[0204] The embodiments of the present application are described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0205] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.

[0206] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.

[0207] The above description is merely an optional embodiment of the present application and is not intended to limit the protection scope of the present application.

Claims

1. A prompt information optimization method, comprising: Get the prompt information to be optimized; Inputting the prompt information to be optimized into a prompt information optimization model to obtain optimized prompt information obtained by the prompt information optimization model through reasoning on the prompt information to be optimized; Among them, in the process of the prompt information optimization model inferring the prompt information to be optimized to obtain the optimized prompt information, the reasoning is performed based on the degree of semantic alignment between the prompt information to be optimized and the target image, and / or the quality of the target image; the target image is an image generated by the image generation model based on the prompt information in the reasoning process of the prompt information optimization model.

2. According to the method of claim 1, the training process of the prompt information optimization model comprises: Obtaining a first prompt information sample and a first image sample; The first image sample is generated by the image generation model based on the second prompt information sample in the prompt information optimization model training process; Scoring the semantic alignment between the first prompt information sample and the first image sample, and / or the quality of the first image sample based on the reward model to obtain a scoring result; The parameters of the pre-trained language model are adjusted based on the scoring result to obtain the prompt information optimization model; the prompt information optimization model is used to infer the prompt information to be optimized to obtain the optimized prompt information.

3. According to the method of claim 2, the scoring result corresponding to the semantic alignment degree comprises at least one of the following: the cosine similarity between the first prompt information sample and the first image sample, and the completion score obtained by answering the question generated based on the first prompt information sample based on the first image sample; The scoring result corresponding to the quality of the first image sample includes at least one of the following: a quality score used to characterize whether the first image sample meets human preferences, and a quality score used to characterize the aesthetic degree of the first image sample in human vision.

4. The method according to claim 2 or 3, wherein the reward model comprises a first reward model, and the first reward model is used to score the semantic alignment between the first prompt information sample and the first image sample; The training process of the first reward model includes: Determine the semantic alignment between the third prompt information sample and the second image sample; the second image sample is generated by the image generation model based on the fourth prompt information sample, and the fourth prompt information sample is an optimized prompt information sample obtained by inferring the third prompt information sample by the pre-trained language model; The third prompt information sample is used as training data, and the semantic alignment between the third prompt information sample and the second image sample is used as a training label, and the parameters of the first language model are adjusted to obtain the first reward model.

5. According to the method of claim 4, the determining of the semantic alignment between the third prompt information sample and the second image sample comprises at least one of the following: Determine a cosine similarity between the third prompt information sample and the second image sample as a first semantic alignment; Based on the third prompt information sample, generating at least one question related to the content of the third prompt information sample; Based on the second image sample, answer the at least one question to obtain a result corresponding to each question; Based on the results corresponding to the questions, a second semantic alignment between the third prompt information sample and the second image sample is determined.

6. The method according to claim 5, wherein the third prompt information sample is used as training data, and the semantic alignment between the third prompt information sample and the second image sample is used as a training label, and the parameters of the first language model are adjusted to obtain the first reward model, comprising: The third prompt information sample is used as training data, and at least one of the first semantic alignment and the second semantic alignment is used as a training label, and the parameters of the first language model are adjusted to obtain the first reward model.

7. The method according to claim 2 or 3, wherein the reward model comprises a second reward model, and the second reward model is used to score the quality of the first image sample; The training process of the second reward model includes: determining a first quality score of a third image sample, where the first quality score is used to characterize whether the third image sample conforms to human preference; the third image sample is generated by the image generation model based on a fifth prompt information sample, and the fifth prompt information sample is an optimized prompt information sample obtained by inferring the sixth prompt information sample by the pre-trained language model; The sixth prompt information sample is used as training data, and the first quality score is used as a training label, and the parameters of the second language model are adjusted to obtain the second reward model.

8. The method according to claim 2 or 3, wherein the reward model comprises a third reward model, and the third reward model is used to score the quality of the first image sample; The training process of the third reward model includes: determining a second quality score of a fourth image sample, where the second quality score is used to characterize the aesthetic degree of the fourth image sample in human vision; the fourth image sample is generated by the image generation model based on the seventh prompt information sample, and the seventh prompt information sample is an optimized prompt information sample obtained by inferring the eighth prompt information sample by the pre-trained language model; The eighth prompt information sample is used as training data, and the second quality score is used as a training label, and the parameters of the third language model are adjusted to obtain the third reward model.

9. According to the method of claim 2 or 3, the training process of the pre-trained language model comprises: Based on at least one prompt information pair, adjusting parameters of an initial pre-trained language model to obtain the pre-trained language model; Each of the prompt information pairs includes two prompt information of different qualities, and the semantic similarity of the two prompt information is greater than a specific value.

10. A prompt information optimization device, comprising: An acquisition module is used to obtain prompt information to be optimized; An optimization module, used for inputting the prompt information to be optimized into a prompt information optimization model, and obtaining optimized prompt information obtained by the prompt information optimization model through reasoning on the prompt information to be optimized; Among them, in the process of the prompt information optimization model inferring the prompt information to be optimized to obtain the optimized prompt information, the reasoning is performed based on the degree of semantic alignment between the prompt information to be optimized and the target image, and / or the quality of the target image; the target image is an image generated by the image generation model based on the prompt information in the reasoning process of the prompt information optimization model.

Citation Information

Cited By

  • Prompt information optimization method of generative model, quality detection method and medium

    CN120996110A