Training of Model, Method, Apparatus, Device and Medium for Determining Text-to-Image API

By conducting two-stage training on the Wensheng Diagram API model, the time-consuming and inaccurate selection of Wensheng Diagram API in the existing technology is solved, and fast and accurate API determination is achieved.

CN118052894BActive Publication Date: 2025-06-27SHANGHAI ARTIFICIAL INTELLIGENCE LABORATORY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311796281.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-12-25
Publication Date
2025-06-27
Estimated Expiration
2043-12-25

AI Technical Summary

Technical Problem

The prior art takes a lot of time to select a suitable Wensheng API, and the method of automatically calling Wensheng API is not accurate enough to meet the needs of users.

Method used

The training set of the Wensheng Diagram API model is determined based on the usability filter, the quality filter and the setting language model, and two-stage training is carried out: first, the model is fine-tuned based on the first target training set, and then the model is aligned based on the second target training set to obtain the target set Wensheng Diagram API model.

Benefits of technology

It realizes the fast and accurate determination of the literary picture API, reduces the time of manual selection, improves the accuracy of automatically calling the API, and meets the needs of users.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118052894B_ABST
    Figure CN118052894B_ABST
Patent Text Reader

Abstract

The present invention discloses a method, apparatus, device, and medium for training a model and determining a text-to-image API. The method includes: determining a training set for a set text-to-image API model based on an availability filter, a quality filter, and a first set language model; wherein the training set includes a first target training set and a second target training set; fine-tuning the set text-to-image API model based on the first target training set; and performing alignment training on the fine-tuned set text-to-image API model based on the second target training set to obtain a target set text-to-image API model. In the embodiments of the present invention, by determining the training set for the set text-to-image API model based on the availability filter, the quality filter, and the first set language model, and fine-tuning and performing alignment training on the set text-to-image API model based on the training set, the obtained target set text-to-image API model can quickly and accurately determine the text-to-image API.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present invention relate to the field of artificial intelligence technology, and in particular, to a method, apparatus, device, and medium for training a model and determining a text-to-image API. Background Art

[0002] The text-to-image (T2I) generation technology has received extensive attention in recent years both inside and outside the research field. Currently, there are a large number of text-to-image models in the open source community, which makes it necessary to make multiple attempts and adjustments when selecting a suitable text-to-image application programming interface (API) (including the text-to-image model and corresponding parameters) according to a specific text prompt. In the prior art, when manually selecting a suitable text-to-image API from a large number of text-to-image APIs, it takes a lot of time. For the method of automatically calling the text-to-image API, only a fixed text-to-image API is called or improved based on the general text-to-image API field, so that the determined text-to-image API is not accurate enough to meet the needs of users. Summary of the Invention

[0003] The embodiments of the present invention provide a method, apparatus, device, and medium for training a model and determining a text-to-image API, which can quickly and accurately determine the text-to-image API.

[0004] In a first aspect, an embodiment of the present invention provides a method for training a model, including: determining a training set of a set text-to-image API model based on an availability filter, a quality filter, and a first set language model; where the training set includes a first target training set and a second target training set; fine-tuning the set text-to-image API model based on the first target training set; and performing alignment training on the fine-tuned set text-to-image API model based on the second target training set to obtain a target set text-to-image API model.

[0005] In a second aspect, an embodiment of the present invention further provides a method for determining a text-to-image API, including: obtaining a target text prompt; inputting the target text prompt into a target set text-to-image API model to output a target text-to-image API; where the target set text-to-image API model is obtained by any one of the model training methods.

[0006] In a third aspect, an embodiment of the present invention further provides a training device for a model, including: a training set determination module, configured to determine a training set of a set text-to-image API model based on an availability filter, a quality filter, and a first set language model; wherein the training set includes a first target training set and a second target training set; a fine-tuning module, configured to fine-tune the set text-to-image API model based on the first target training set; an alignment training module, configured to perform alignment training on the fine-tuned set text-to-image API model based on the second target training set to obtain a target set text-to-image API model.

[0007] In a fourth aspect, an embodiment of the present invention further provides a determination device for a text-to-image API, including: a target text prompt acquisition module, configured to acquire a target text prompt; a target text-to-image API output module, configured to input the target text prompt into a target set text-to-image API model to output a target text-to-image API; wherein the target set text-to-image API model is obtained by the model training method according to any one of the above.

[0008] In a fifth aspect, an embodiment of the present invention further provides an electronic device, the electronic device includes:

[0009] at least one processor; and

[0010] a memory communicatively connected to the at least one processor; wherein,

[0011] the memory stores a computer program executable by the at least one processor, and when the computer program is executed by the at least one processor, the at least one processor is enabled to execute the model training or the text-to-image API determination method according to any one of the embodiments of the present invention.

[0012] In a sixth aspect, an embodiment of the present invention further provides a computer-readable storage medium, the computer-readable storage medium stores computer instructions, and when the computer instructions are executed by a processor, the computer instructions are used to implement the model training or the text-to-image API determination method according to any one of the embodiments of the present invention.

[0013] The technical solution of this embodiment determines the training set of the set text-to-image API model based on an availability filter, a quality filter, and a first set language model; wherein, the training set includes a first target training set and a second target training set; fine-tune the set text-to-image API model based on the first target training set; perform alignment training on the fine-tuned set text-to-image API model based on the second target training set to obtain a target set text-to-image API model. In the embodiment of the present invention, by determining the training set of the set text-to-image API model based on an availability filter, a quality filter, and a first set language model, and performing fine-tuning and alignment training on the set text-to-image API model based on the training set, the obtained target set text-to-image API model can quickly and accurately determine the text-to-image API. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] Figure 1 It is a flowchart of a method for training a model provided by an embodiment of the present invention;

[0015] Figure 2 It is a schematic diagram of two-stage training of the set text-to-image API model provided by an embodiment of the present invention;

[0016] Figure 3 It is a flowchart of a method for determining a text-to-image API provided by an embodiment of the present invention;

[0017] Figure 4 It is a schematic structural diagram of a model training device provided by an embodiment of the present disclosure;

[0018] Figure 5 It is a schematic structural diagram of a device for determining a text-to-image API provided by an embodiment of the present disclosure;

[0019] Figure 6 It is a schematic structural diagram of an electronic device for implementing the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0020] The embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although some embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. On the contrary, these embodiments are provided to more thoroughly and completely understand the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes and are not used to limit the protection scope of the present disclosure.

[0021] It should be understood that the steps described in the method embodiments of the present disclosure can be executed in different orders and / or in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present disclosure is not limited in this regard.

[0022] As used herein, the term "including" and its variations are open-ended, i.e., "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". The relevant definitions of other terms will be given in the following description.

[0023] It should be noted that the concepts such as "first", "second", etc. mentioned in this disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules or units.

[0024] It should be noted that the modifications of "one" and "multiple" mentioned in this disclosure are illustrative rather than restrictive. Those skilled in the art should understand that, unless otherwise clearly specified in the context, it should be understood as "one or more".

[0025] In the technical solution of this application, the acquisition, storage, use, processing, etc. of data all comply with the relevant provisions of national laws and regulations.

[0026] Figure 1 FIG. is a flowchart of a method for training a model provided for an embodiment of the present invention. This embodiment is applicable to the case of generating a text-to-image API given a specific text prompt. This method can be executed by a model training device, and specifically includes the following steps:

[0027] S110. Determine a training set of a set text-to-image API model based on an availability filter, a quality filter, and a first set language model.

[0028] Among them, the training set includes a first target training set and a second target training set.

[0029] Among them, the first set language model can be any large language model. The availability filter can be understood as filtering the corresponding text-to-image API in the training set. For example, it can be filtered from multiple dimensions such as whether it is the target model type, whether it is the target basic model architecture, whether it is a compliant model, and whether it can be actually called. The quality filter can be understood as filtering the quality of the text-to-image API in the training set. For example, the quality of the text-to-image API can be evaluated from dimensions such as the number of likes and downloads. In this embodiment, the training set can be divided into a first target training set, a second target training set, and a test set according to a set ratio, and the set ratio can be 8:1:1. Among them, each pair of data includes a text prompt and a text-to-image API.

[0030] In this embodiment, the text-to-image APIs in the original training set can be filtered through an availability filter and a quality filter, and the description information in the text-to-image APIs can be reconstructed using a first predefined language model to obtain the training set of the final predefined text-to-image API model.

[0031] Optionally, determining the training set of the predefined text-to-image API model based on the availability filter, the quality filter, and the first predefined language model includes: obtaining multiple text-to-image APIs and corresponding text prompts from a predefined data open source to form a first training set; filtering the first training set based on the availability filter to obtain a second training set; filtering the second training set based on the quality filter to obtain a third training set; and reconstructing the third training set based on the first predefined language model to obtain the training set of the predefined text-to-image API model.

[0032] Herein, the predefined data open source can be understood as any open source website or community on the Internet, which can be abbreviated as an open source community. Specifically, multiple text-to-image APIs and corresponding text prompts are obtained from the predefined data open source to form a first training set; the first training set can be understood as the original training set. The training set from different data source distributions is filtered based on the availability filter to obtain a second training set; the data in the second training set that does not meet the quality conditions is filtered based on the quality filter to obtain a third training set; and the description information in the third training set is reconstructed based on the first predefined language model to obtain the training set of the predefined text-to-image API model. In this embodiment, there is no limit on the data volume of the training set of the final predefined text-to-image API model. For example, it can be 50,482 pairs.

[0033] Optionally, filtering the first training set based on the availability filter to obtain a second training set includes: filtering the first training set successively according to the filtering conditions corresponding to the model type, the basic model architecture category, the compliant model category, and the available model category to obtain a second training set.

[0034] Exemplarily, the availability filter can include a total of 4 different dimensions of data sources, namely the model type, the basic model architecture category, the compliant model category, and the available model category. Among them, the model type can include Checkpoint and LoRA models, etc.; the basic model architecture category can include other basic model architectures such as Stable Diffusion XL (SDXL) 1.0 and Stable Diffusion (SD) 1.5; the compliant model category can include compliant models and non-compliant models; and the available model category can include available models and unavailable models. An available model can be understood as a model that actually exists and has callability.

[0035] In this embodiment, after filtering the first training set in sequence according to the filtering conditions corresponding to the model type, the basic model architecture category, the compliant model category, and the available model category, a second training set can be obtained. That is, the second training set simultaneously meets the filtering conditions corresponding to the model type, the basic model architecture category, the compliant model category, and the available model category. Among them, the filtering condition for the model type corresponding to the second training set can be to retain the Checkpoint and LoRA models. The filtering condition for the basic model architecture category can be to retain the Stable Diffusion SDXL1.0 basic model architecture and the Stable Diffusion SD1.5 basic model architecture. The filtering condition for the compliant model category can be to retain the compliant model category. The filtering condition for the available model category can be to retain the available models.

[0036] Optionally, filtering the second training set based on the quality filter to obtain a third training set includes: obtaining the quality information corresponding to each text-to-image API in the second training set; where the quality information includes the number of likes, the number of downloads, and the number of votes; scoring the corresponding text-to-image API based on the quality information to obtain the scoring information corresponding to each text-to-image API; if the scoring information is less than the set scoring threshold, then filter out the corresponding text-to-image API and the corresponding text prompt to obtain the third training set.

[0037] Specifically, since the set data open source includes the quality information corresponding to each text-to-image API in the second training set, the quality information corresponding to each text-to-image API can be directly obtained from the set data open source. For the quality information of any one text-to-image API, the number of likes, the number of downloads, and the number of votes included in the quality information can be scored respectively, and the scoring information of the number of likes, the number of downloads, and the number of votes is weighted and summed to obtain the scoring information of the corresponding text-to-image API. Determine whether the scoring information of each text-to-image API meets the quality filtering condition. If the scoring information of the text-to-image API is less than the set scoring threshold, then filter out the corresponding text-to-image API and the corresponding text prompt. If the scoring information of the text-to-image API is greater than or equal to the set scoring threshold, then retain the corresponding text-to-image API and the corresponding text prompt to obtain the third training set.

[0038] Optionally, reconstruct the third training set based on the first preset language model to obtain the training set of the preset text-to-image API model, including: obtaining the additional information corresponding to each text-to-image API in the third training set; inputting the third training set and the additional information into the first preset language model, and outputting the target description information corresponding to each text-to-image API in the third training set; replacing the original description information included in each text-to-image API in the third training set with the corresponding target description information to obtain the training set of the preset text-to-image API model.

[0039] Specifically, all the information related to each text-to-image API in the third training set can be input into the first preset language model. Specifically, the third training set and the additional information can be input into the first preset language model, and the target description information corresponding to each text-to-image API in the third training set is output; wherein, the additional information includes label information and version information; since the preset data open source includes the additional information corresponding to each text-to-image API in the third training set, the additional information corresponding to each text-to-image API can be directly obtained from the preset data open source. The label information can be determined according to the actual situation, and this embodiment does not limit it. For example, it can be the attribute corresponding to the text-to-image API, such as the images generated by the text-to-image API being animals, landscapes, people, etc. Each text-to-image API in the third training set can include the text-to-image model name, the type of the text-to-image model, the basic model architecture, model parameters, and original description information. After obtaining the target description information corresponding to each text-to-image API, replace the original description information included in each text-to-image API in the third training set with the corresponding target description information to obtain the training set of the preset text-to-image API model.

[0040] S120. Fine-tune the preset text-to-image API model based on the first target training set.

[0041] Among them, fine-tuning can be understood as further training the preset text-to-image API model using the first target training set. Among them, the preset text-to-image API model is a second preset language model. The second preset language model can be any open-source large language model, such as the Llama-2 language model. In this embodiment, by fine-tuning the preset text-to-image API model based on the first target training set, the first-stage training is carried out.

[0042] Optionally, fine-tuning the preset text-to-image API model based on the first target training set includes: inputting the first target training set into the second preset language model to output a first predicted text-to-image API; determining a first loss value based on the first predicted text-to-image API and the true text-to-image API in the first target training set; and iteratively training the second preset language model based on the first loss value until a corresponding iterative training stop condition is met, to obtain a fine-tuned preset text-to-image API model.

[0043] Specifically, in the first stage, the second preset language model can be trained based on the following cross-entropy loss function for fine-tuning:

[0044] L sft = -∑ k logP π (r k |t, r <k ) ;

[0045] where t represents the text prompt, r represents the text-to-image API, π represents the model being trained, which is the second preset language model here, and (r k |t, r <k ) represents the probability of predicting r <k given the text prompt t and r k . L sft represents the first loss value. logP π (r k |t, r <k ) represents the log function of the predicted probability P output by the second preset language model π being trained under the condition of (r k |t, r <k ). Here, k represents the kth.

[0046] Among them, the first predicted text-to-image API can be understood as the probability value of predicting the true text-to-image API. Specifically, inputting the first target training set into the second preset language model and iteratively training the second preset language model through the cross-entropy loss function for fine-tuning until a corresponding iterative training stop condition is met, to obtain a fine-tuned preset text-to-image API model, where the corresponding iterative training stop condition can be that the loss value is less than the set loss value for a continuous set number of times, or the number of training iterations is greater than the set number of training iterations.

[0047] S130. Align and train the fine-tuned preset text-to-image API model based on the second target training set to obtain a target preset text-to-image API model.

[0048] In this embodiment, the second target training set is input into the fine-tuned preset text-to-image API model, and alignment training is performed using the Rank Responses with Human Feedback (RRHF) framework, that is, the second-stage training is carried out to obtain the target preset text-to-image API model.

[0049] Optionally, aligning and training the fine-tuned preset text-to-image API model based on the second target training set to obtain the target preset text-to-image API model includes: inputting the second target training set into the fine-tuned preset text-to-image API model to output multiple different second predicted text-to-image APIs; filtering the multiple different second predicted text-to-image APIs based on the second target training set to obtain multiple different third predicted text-to-image APIs; adjusting the formats of the multiple different third predicted text-to-image APIs according to a set standard format to obtain multiple different fourth predicted text-to-image APIs; wherein each text prompt in the second target training set has corresponding multiple different fourth predicted text-to-image APIs; for any text prompt in the second target training set, inputting the text prompt into the corresponding multiple different fourth predicted text-to-image APIs respectively, and each fourth predicted text-to-image API outputs corresponding multiple different text-to-image results; and performing alignment training on the fine-tuned preset text-to-image API model based on the set scoring function corresponding to the text-to-image results to obtain the target preset text-to-image API model.

[0050] Specifically, input the second target training set into the fine-tuned preset text-to-image API model. During the training process, use polynomial sampling to obtain a comprehensive text-to-image API, that is, perform normal output, and use diverse beam search to obtain multiple different text-to-image APIs from the fine-tuned preset text-to-image API model. That is, through polynomial sampling and diverse beam search techniques, the fine-tuned preset text-to-image API model outputs multiple different second predicted text-to-image APIs. Compare each of the multiple different second predicted text-to-image APIs with each text-to-image API in the second target training set. If it is found that a second predicted text-to-image API does not exist in the second target training set, filter out the corresponding second predicted text-to-image API. At the same time, it also indicates that the corresponding second predicted text-to-image API is an unrealistic text-to-image API and is not available. After filtering out the second predicted text-to-image APIs that do not exist in the second target training set, obtain multiple different third predicted text-to-image APIs, and adjust the formats of the multiple different third predicted text-to-image APIs according to the set standard format to obtain multiple different fourth predicted text-to-image APIs; among them, each text prompt in the second target training set has corresponding multiple different fourth predicted text-to-image APIs; that is, if the input is a text prompt, the output is multiple different text-to-image APIs. The set standard format can be understood as a format that meets the aesthetic requirements of users. For example, "Fourth predicted text-to-image API: Name of the text-to-image model: xx; Type of the text-to-image model: xx; Basic model architecture: xx; Model parameters: xx; Description information: xx".

[0051] In this embodiment, for any text prompt in the second target training set, input the text prompt into the corresponding multiple different fourth predicted text-to-image APIs respectively. Each fourth predicted text-to-image API outputs corresponding multiple different images, that is, actually call the fourth predicted text-to-image API, and obtain multiple different images, that is, obtain a list R containing n fourth predicted text-to-image APIs (r1, r2, …, r n ). Use the set scoring function corresponding to the image to enable each fourth predicted text-to-image API to obtain a final scoring information. Based on the scoring information corresponding to each fourth predicted text-to-image API and the RRHF framework, perform alignment training on the fine-tuned preset text-to-image API model to obtain the target preset text-to-image API model. Among them, the set scoring function can be the scoring function of the reward model, that is, the existing evaluation index in the field of text-to-image.

[0052] Optionally, perform alignment training on the fine-tuned preset text-to-image API model based on the preset scoring function corresponding to the text-to-image to obtain a target preset text-to-image API model, including: determining scoring information of multiple different text-to-image corresponding to each fourth predicted text-to-image API based on the preset scoring function; determining the scoring information of the corresponding fourth predicted text-to-image API based on the scoring information of the multiple different text-to-image; determining a second loss value of the fine-tuned preset text-to-image API model based on the scoring information of each fourth predicted text-to-image API and the corresponding prediction probability; performing iterative training on the fine-tuned preset text-to-image API model based on the second loss value until the corresponding iterative training stop condition is satisfied, to obtain a target preset text-to-image API model.

[0053] Exemplarily, the scoring information of multiple different text-to-image corresponding to each fourth predicted text-to-image API can be determined through a preset scoring function; averaging the scoring information of the multiple different text-to-image to obtain the scoring information of the corresponding fourth predicted text-to-image API; denoted as S(t, r i ) = s i , s i represents the scoring information corresponding to the i-th fourth predicted text-to-image API for a text prompt t; in order to align with the scores {s i}, the fine-tuned preset text-to-image API model is used to give a prediction probability p n for each fourth predicted text-to-image API (i.e., r i ): i :

[0054]

[0055] where p i can be the conditional log probability (after length normalization) of r i under the fine-tuned preset text-to-image API model π being trained.

[0056] Exemplarily, use the RRHF framework to make the fine-tuned preset text-to-image API model π in training give a higher probability to the fourth predicted text-to-image API with a higher score and a lower probability to the fourth predicted text-to-image API with a lower score. That is, optimize the fine-tuned preset text-to-image API model through a ranking loss.

[0057]

[0058] where p i is the prediction probability corresponding to r i , p j is the prediction probability corresponding to r j , s i is for ri The corresponding scoring information, s j is r j The corresponding scoring information.

[0059] And the cross-entropy loss is also added:

[0060] L ce = -∑ k log P π (r i',k |t, r i',<k ) ;

[0061] Where r i' represents the r with the highest scoring information i .

[0062] Finally, the fine-tuned text-to-image API model is iteratively trained through the following loss function until the corresponding iterative training stop condition is met, and the target text-to-image API model is obtained.

[0063] L = L rank + L ce ;

[0064] Where L represents the second loss value. The corresponding iterative training stop condition can be that the loss value is less than the set loss value for a continuously set number of times, or the number of training iterations is greater than the set number of training iterations.

[0065] The technical solution of this embodiment determines the training set of the text-to-image API model based on the availability filter, the quality filter, and the first set language model; wherein, the training set includes the first target training set and the second target training set; the text-to-image API model is fine-tuned based on the first target training set; and the fine-tuned text-to-image API model is aligned and trained based on the second target training set to obtain the target text-to-image API model. In the embodiment of the present invention, by determining the training set of the text-to-image API model based on the availability filter, the quality filter, and the first set language model, and fine-tuning and aligning and training the text-to-image API model based on the training set, the obtained target text-to-image API model can quickly and accurately determine the text-to-image API.

[0066] Exemplarily, Figure 2 is a schematic diagram of the two-stage training of the text-to-image API model provided by the embodiment of the present invention. The fine-tuned text-to-image API model can be called DiffAgent-SFT, and the target text-to-image API model can be DiffAgent-RRHF. As Figure 2 shown, for the training in the first stage, the second set language model ( Figure 2The input of the second pre-set language model (the Llama-2 language model) is the text prompt of the first target training set, and the output is the first predicted text-to-image API. After training, DiffAgent-SFT is obtained. For the training in the second stage, the input of the fine-tuned text-to-image API model ( Figure 2 DiffAgent-SFT in this case) is the text prompt of the second target training set. During the training process, polynomial sampling, diverse beam search, unavailable filtering, and text-to-image API format reconstruction are sequentially adopted. The purpose of polynomial sampling and diverse beam search is to ensure that the fine-tuned text-to-image API model outputs multiple different text-to-image APIs during training. Unavailable filtering is to filter out non-existent and non-callable text-to-image APIs. The purpose of text-to-image API format reconstruction is to output multiple different text-to-image APIs in accordance with the set standard format. The output of the fine-tuned text-to-image API model is multiple different text-to-image APIs (multiple different fourth predicted text-to-image APIs), that is, multiple different responses, and the corresponding text prompts are respectively input into the corresponding multiple different fourth predicted text-to-image APIs, that is, the fourth predicted text-to-image APIs are actually called. Each fourth predicted text-to-image API outputs corresponding multiple different images (multiple pictures); the scoring information of the corresponding fourth predicted text-to-image API is determined through the scoring function of the reward model, and the fine-tuned text-to-image API model is trained based on the scoring information of each fourth predicted text-to-image API and the RRHF training framework to obtain DiffAgent-RRHF.

[0067] Figure 3 The figure is a flowchart of a method for determining a text-to-image API provided by an embodiment of the present invention. The specific steps are as follows:

[0068] S310. Obtain a target text prompt.

[0069] Among them, the target text prompt can be a text prompt in actual applications or in a test set. In this embodiment, no specific restrictions are imposed on the text prompt, and it can be applicable to any scenario, such as "I want a picture of a puppy", "I want a picture of the scenery of a certain place", etc.

[0070] S320. Input the target text prompt into a target pre-set text-to-image API model to output a target text-to-image API.

[0071] Among them, the target pre-set text-to-image API model is obtained through the model training method in the above embodiment.

[0072] In this embodiment, the target setting text-to-image API model processes the target text prompt, and the output target text-to-image API is aligned with human preferences, with higher accuracy, improving the user experience. For the text-to-image task, the target setting text-to-image API model can call the target text-to-image API that is most suitable for the target text prompt.

[0073] In the embodiment of the present invention, an API dataset is collected for the text-to-image sub-field, and a two-stage training framework is designed. Not only is supervised fine-tuning alignment performed in the first stage, but also based on the fine-tuning model in the first stage, the RRHF algorithm is used to further train the model to align it with human preferences. Specifically, first, a high-quality text-to-image API dataset (training set) is collected from the open-source community, which includes 50,482 data points consisting of text prompts and text-to-image APIs. Secondly, the training framework proposed by the present invention includes a two-stage training process. The first stage performs standard supervised fine-tuning, and the second stage further aligns the fine-tuned setting text-to-image API model with human preferences by expanding the RRHF algorithm. For the second stage, first, the fine-tuned setting text-to-image API model in the first stage is used to sample multiple text prompts, and then the available and complete text-to-image API responses are screened and reconstructed through the existing data. After that, by actually calling these APIs, multiple images are sampled for each API. Then, the scoring function of the reward model is used for scoring to evaluate the performance of the API. Finally, the RRHF algorithm is used to further fine-tune the model with the text prompts, multiple API responses, and their scoring information to align it with human preferences.

[0074] Figure 4 It is a schematic structural diagram of a training device for a model provided by an embodiment of the present disclosure. The device includes: a training set determination module 410, a fine-tuning module 420, and an alignment training module 430;

[0075] The training set determination module 410 is used to determine the training set of the setting text-to-image API model based on an availability filter, a quality filter, and a first set language model; wherein, the training set includes a first target training set and a second target training set;

[0076] The fine-tuning module 420 is used to fine-tune the setting text-to-image API model based on the first target training set;

[0077] The alignment training module 430 is used to perform alignment training on the fine-tuned setting text-to-image API model based on the second target training set to obtain a target setting text-to-image API model.

[0078] In the technical solution of this embodiment, a training set determining module determines a training set of a set text-to-image API model based on an availability filter, a quality filter, and a first set language model; wherein, the training set includes a first target training set and a second target training set; a fine-tuning module fine-tunes the set text-to-image API model based on the first target training set; an alignment training module performs alignment training on the fine-tuned set text-to-image API model based on the second target training set to obtain a target set text-to-image API model. In the embodiment of the present invention, by determining the training set of the set text-to-image API model based on the availability filter, the quality filter, and the first set language model, and performing fine-tuning and alignment training on the set text-to-image API model based on the training set, the obtained target set text-to-image API model can quickly and accurately determine the text-to-image API.

[0079] Optionally, the training set determining module is specifically configured to: obtain multiple text-to-image APIs and corresponding text prompts from a set data open source to form a first training set; filter the first training set based on the availability filter to obtain a second training set; filter the second training set based on the quality filter to obtain a third training set; reconstruct the third training set based on the first set language model to obtain the training set of the set text-to-image API model.

[0080] Optionally, the training set determining module is specifically configured to: filter the first training set in sequence according to the filtering conditions corresponding to the model type, the basic model architecture category, the compliant model category, and the available model category to obtain a second training set.

[0081] Optionally, the training set determining module is specifically configured to: obtain the quality information corresponding to each text-to-image API in the second training set; wherein, the quality information includes the number of likes, the number of downloads, and the number of votes; score the corresponding text-to-image API based on the quality information to obtain the score information corresponding to each text-to-image API; if the score information is less than a set score threshold, filter out the corresponding text-to-image API and the corresponding text prompt to obtain a third training set.

[0082] Optionally, the training set determining module is specifically configured to: obtain the additional information corresponding to each text-to-image API in the third training set; wherein, the additional information includes tag information and version information; input the third training set and the additional information into the first set language model, and output the target description information corresponding to each text-to-image API in the third training set; replace the original description information included in each text-to-image API in the third training set with the corresponding target description information to obtain the training set of the set text-to-image API model.

[0083] Among them, the set text-to-image API model is the second set language model.

[0084] Optionally, the fine-tuning module is specifically configured to: input the first target training set into the second set language model, and output a first predicted text-to-image API; determine a first loss value based on the first predicted text-to-image API and the true text-to-image API in the first target training set; perform iterative training on the second set language model based on the first loss value until the corresponding iterative training stop condition is met, and obtain a fine-tuned set text-to-image API model.

[0085] Optionally, the fine-tuning module is specifically configured to: input the second target training set into the fine-tuned set text-to-image API model, and output a plurality of different second predicted text-to-image APIs; filter the plurality of different second predicted text-to-image APIs based on the second target training set to obtain a plurality of different third predicted text-to-image APIs; adjust the formats of the plurality of different third predicted text-to-image APIs according to a set standard format to obtain a plurality of different fourth predicted text-to-image APIs; where each text prompt in the second target training set has a corresponding plurality of different fourth predicted text-to-image APIs; for any text prompt in the second target training set, input the text prompt into the corresponding plurality of different fourth predicted text-to-image APIs respectively, and each fourth predicted text-to-image API outputs a corresponding plurality of different text-to-images; perform alignment training on the fine-tuned set text-to-image API model based on the set scoring function corresponding to the text-to-image to obtain a target set text-to-image API model.

[0086] Optionally, the fine-tuning module is specifically configured to: determine scoring information of a plurality of different text-to-images corresponding to each fourth predicted text-to-image API based on the set scoring function; determine scoring information of the corresponding fourth predicted text-to-image API based on the scoring information of the plurality of different text-to-images; determine a second loss value of the fine-tuned set text-to-image API model based on the scoring information of each fourth predicted text-to-image API and the corresponding prediction probability; perform iterative training on the fine-tuned set text-to-image API model based on the second loss value until the corresponding iterative training stop condition is met, and obtain a target set text-to-image API model.

[0087] Figure 5 It is a schematic structural diagram of a device for determining a text-to-image API provided by an embodiment of the present disclosure. The device includes: a target text prompt acquisition module 510 and a target text-to-image API output module 520;

[0088] The target text prompt acquisition module 510 is configured to acquire a target text prompt;

[0089] The target text-to-image API output module 520 is used to input the target text prompt into the target set text-to-image API model and output the target text-to-image API.

[0090] In this embodiment, the target set text-to-image API model processes the target text prompt, and the output target text-to-image API is aligned with human preferences, with higher accuracy, improving the user experience. For the text-to-image task, the target set text-to-image API model can call the target text-to-image API that is most suitable for the target text prompt.

[0091] The above device can execute the methods provided in all the foregoing embodiments of the present invention, and has the corresponding functional modules and beneficial effects for executing the above methods. For the technical details not described in detail in this embodiment, reference may be made to the methods provided in all the foregoing embodiments of the present invention.

[0092] Figure 6 The structural schematic diagram of the electronic device 10 that can be used to implement the embodiments of the present invention is shown. The electronic device is intended to represent various forms of digital computers, such as, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, personal digital processors, cellular phones, smart phones, wearable devices (such as helmets, glasses, watches, etc.) and other similar computing devices. The components shown herein, their connections and relationships, and their functions are only examples and are not intended to limit the implementation of the present invention described and / or claimed herein.

[0093] As Figure 6 shown, the electronic device 10 includes at least one processor 11, and a memory communicatively connected to the at least one processor 11, such as a read-only memory (ROM) 12, a random access memory (RAM) 13, etc. Among them, the memory stores a computer program executable by the at least one processor, and the processor 11 can execute various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 12 or the computer program loaded from the storage unit 18 into the random access memory (RAM) 13. In the RAM 13, various programs and data required for the operation of the electronic device 10 can also be stored. The processor 11, the ROM 12, and the RAM 13 are connected to each other through a bus 14. The input / output (I / O) interface 15 is also connected to the bus 14.

[0094] Multiple components in the electronic device 10 are connected to the I / O interface 15, including: an input unit 16, such as a keyboard, a mouse, etc.; an output unit 17, such as various types of displays, speakers, etc.; a storage unit 18, such as a disk, an optical disc, etc.; and a communication unit 19, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 19 allows the electronic device 10 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0095] The processor 11 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The processor 11 executes the various methods and processes described above, such as the training of the method model and the determination of the text-to-image API.

[0096] In some embodiments, the training of the method model and the determination of the text-to-image API can be implemented as a computer program, which is tangibly contained in a computer-readable storage medium, such as the storage unit 18. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 10 via the ROM 12 and / or the communication unit 19. When the computer program is loaded into the RAM 13 and executed by the processor 11, one or more steps of the training of the method model and the determination of the text-to-image API described above can be executed. Alternatively, in other embodiments, the processor 11 can be configured to execute the training of the method model and the determination of the text-to-image API by any other suitable means (e.g., by means of firmware).

[0097] The various embodiments of the systems and technologies described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: implemented in one or more computer programs, the one or more computer programs can be executed and / or interpreted on a programmable system including at least one programmable processor, the programmable processor can be a special or general programmable processor, can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit the data and instructions to the storage system, the at least one input device, and the at least one output device.

[0098] A computer program for implementing the method of the present invention can be written in any combination of one or more programming languages. These computer programs can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the computer programs are executed by the processor, the functions / operations specified in the flowchart and / or block diagram are implemented. The computer programs can be executed entirely on the machine, partially on the machine, executed partially on the machine and partially on a remote machine as an independent software package, or executed entirely on a remote machine or server.

[0099] In the context of the present invention, a computer-readable storage medium can be a tangible medium that can contain or store a computer program for use by or in connection with an instruction execution system, apparatus, or device. The computer-readable storage medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. Alternatively, the computer-readable storage medium can be a machine-readable signal medium. More specific examples of the machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0100] In order to provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) through which the user can provide input to the electronic device. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0101] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with an implementation of the systems and techniques described herein), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected with each other by digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include: local area network (LAN), wide area network (WAN), blockchain network, and the Internet.

[0102] A computing system can include a client and a server. The client and the server are generally far from each other and typically interact through a communication network. The client-server relationship is created by computer programs running on respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system and solves the defects of difficult management and weak business scalability existing in traditional physical hosts and VPS services.

[0103] It should be understood that various forms of the processes shown above can be used, steps can be reordered, added, or deleted. For example, the steps recited in the present invention can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution of the present invention can be achieved, and no limitation is made herein.

[0104] The above specific embodiments do not constitute a limitation on the protection scope of the present invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A training method for a model, characterized in that, Including: Determining a training set of a set text-to-image API model based on an availability filter, a quality filter, and a first set language model; wherein, the training set includes a first target training set and a second target training set; Fine-tuning the set text-to-image API model based on the first target training set; Performing alignment training on the fine-tuned set text-to-image API model based on the second target training set to obtain a target set text-to-image API model; Wherein, the set text-to-image API model is a second set language model; fine-tuning the set text-to-image API model based on the first target training set includes: Inputting the first target training set into the second set language model to output a first predicted text-to-image API; Determining a first loss value based on the first predicted text-to-image API and the true text-to-image API in the first target training set; Performing iterative training on the second set language model based on the first loss value until a corresponding iterative training stop condition is met to obtain a fine-tuned set text-to-image API model.

2. The method according to claim 1, wherein Determining a training set of a set text-to-image API model based on an availability filter, a quality filter, and a first set language model includes: Obtaining a plurality of text-to-image APIs and corresponding text prompts from a set data open source to form a first training set; Filtering the first training set based on the availability filter to obtain a second training set; Filtering the second training set based on the quality filter to obtain a third training set; Reconstructing the third training set based on the first set language model to obtain a training set of the set text-to-image API model.

3. The method according to claim 2, characterized in that, Filtering the first training set based on the availability filter to obtain a second training set includes: Filtering the first training set successively according to the filtering conditions corresponding to the model type, the basic model architecture category, the compliant model category, and the available model category to obtain a second training set.

4. The method according to claim 2, wherein Filtering the second training set based on the quality filter to obtain a third training set includes: Obtaining quality information corresponding to each text-to-image API in the second training set; wherein, the quality information includes the number of likes, the number of downloads, and the number of votes; Scoring the corresponding text-to-image API based on the quality information to obtain scoring information corresponding to each text-to-image API; If the scoring information is less than a set scoring threshold, filtering out the corresponding text-to-image API and the corresponding text prompt to obtain a third training set.

5. The method according to claim 2, wherein Reconstructing the third training set based on the first set language model to obtain a training set of the set text-to-image API model includes: Obtaining additional information corresponding to each text-to-image API in the third training set; wherein, the additional information includes tag information and version information; Inputting the third training set and the additional information into the first set language model to output target description information corresponding to each text-to-image API in the third training set; Replacing the original description information included in each text-to-image API in the third training set with the corresponding target description information to obtain a training set of the set text-to-image API model.

6. The method according to claim 1, characterized in that Performing alignment training on the fine-tuned set text-to-image API model based on the second target training set to obtain a target set text-to-image API model, including: Inputting the second target training set into the fine-tuned set text-to-image API model to output multiple different second predicted text-to-image APIs; Filtering the multiple different second predicted text-to-image APIs based on the second target training set to obtain multiple different third predicted text-to-image APIs; Adjusting the formats of the multiple different third predicted text-to-image APIs according to a set standard format to obtain multiple different fourth predicted text-to-image APIs; wherein, each text prompt in the second target training set has corresponding multiple different fourth predicted text-to-image APIs; For any text prompt in the second target training set, inputting the text prompt into the corresponding multiple different fourth predicted text-to-image APIs respectively, and each fourth predicted text-to-image API outputs corresponding multiple different text-to-images; Performing alignment training on the fine-tuned set text-to-image API model based on the set scoring function corresponding to the text-to-image to obtain a target set text-to-image API model.

7. The method according to claim 6, wherein Performing alignment training on the fine-tuned set text-to-image API model based on the set scoring function corresponding to the text-to-image to obtain a target set text-to-image API model, including: Determining scoring information of the multiple different text-to-images corresponding to each fourth predicted text-to-image API based on the set scoring function; Determining scoring information of the corresponding fourth predicted text-to-image API based on the scoring information of the multiple different text-to-images; Determining a second loss value of the fine-tuned set text-to-image API model based on the scoring information of each fourth predicted text-to-image API and the corresponding prediction probability; Performing iterative training on the fine-tuned set text-to-image API model based on the second loss value until a corresponding iterative training stop condition is satisfied to obtain a target set text-to-image API model.

8. A method for determining a text-to-image API, characterized in that, Including: Obtaining a target text prompt; Inputting the target text prompt into the target set text-to-image API model to output a target text-to-image API; wherein, the target set text-to-image API model is obtained by the training method of the model described in any one of claims 1-7.

9. A training device for a model, characterized in that Including: A training set determination module, configured to determine a training set of the set text-to-image API model based on an availability filter, a quality filter, and a first set language model; wherein, the training set includes a first target training set and a second target training set; A fine-tuning module, configured to fine-tune the set text-to-image API model based on the first target training set; An alignment training module, configured to perform alignment training on the fine-tuned set text-to-image API model based on the second target training set to obtain a target set text-to-image API model; wherein, the set text-to-image API model is a second set language model; The fine-tuning module is specifically configured to input the first target training set into the second set language model to output a first predicted text-to-image API; Determine a first loss value based on the first predicted text-to-image API and the true text-to-image API in the first target training set; Iteratively train the second preset language model based on the first loss value until the corresponding iterative training stop condition is met, and obtain a fine-tuned preset text-to-image API model.

10. An apparatus for determining a text-to-image API, characterized in that, It includes: A target text prompt acquisition module for acquiring a target text prompt; A target text-to-image API output module for inputting the target text prompt into a target preset text-to-image API model and outputting a target text-to-image API; wherein, the target preset text-to-image API model is obtained by the training method of the model described in any one of claims 1-7.

11. An electronic device, characterized in that, The electronic device includes: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the training method of the model described in any one of claims 1-7 or the determination method of the text-to-image API described in claim 8.

12. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions for causing a processor to implement the training method of the model described in any one of claims 1-7 or the determination method of the text-to-image API described in claim 8 when executed.

Citation Information

Patent Citations

  • Text-based image diffusion model training method and text-based image generation method

    CN116051668A

  • Text-image generation method, system and device and storage medium

    CN117095083A