Prompt modification for creating higher quality images using generative artificial intelligence
Patent Information
- Application Number
- EP2024726818
- Authority / Receiving Office
- EP · EP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-04-19
- Publication Date
- 2025-12-03
AI Technical Summary
Current generative artificial intelligence models often produce inaccurate and low-quality images, and manually designing prompts for image generation is time-consuming and may not lead to optimal results.
A multi-stage training process is employed to develop a prompt enhancement model that uses large language models to automatically generate modified prompts, incorporating mutators to improve image quality, aesthetics, and image-text alignment, and employs reinforcement learning to optimize the model.
The process results in higher quality, aesthetically pleasing images with improved image-text alignment, enhancing the output of text-to-image models without requiring manual prompt design.
Smart Images

Figure US2024025402_23102025_PF_FP_ABST
Abstract
Description
PROMPT MODIFICATION FOR CREATING HIGHER QUALITY IMAGES USING GENERATIVE ARTIFICIAL INTELLIGENCEBACKGROUND
[0001] This specification relates to data processing, artificial intelligence, and generating images using artificial intelligence.
[0002] In a computer networked environment such as the Internet, third-party content providers provide third-party content items for display on end-user computing devices. These third-party7content items, for example, digital images and video, can be displayed on client devices in the environment. Digital images and video can be used, for example, on the Internet, for remote meetings via video conferencing, high-definition video entertainment, and / or sharing of user-generated content.
[0003] Recent developments in artificial intelligence and, in particular, generative artificial intelligence have caused user-produced visual content such as digital images to become ubiquitous. For example, various types of images can be generated by using a text-to-image models based on text prompts. However, current generative artificial intelligence models often produce inaccurate and / or low quality images.SUMMARY
[0004] Effective prompt design significantly improves the text-to-image generation quality. However, manually constructing such prompts can be time-consuming and may not lead to optimal results, considering the inherent complexity of large scale generation models. Instead of manually designing prompts, this document describes techniques for leveraging capabilities of large language models (LLMs) to automatically generate prompt modifications (e g., prompt expansions) that improve the quality of the generated images.
[0005] In some situations, e.g., in digital content distribution, it is important for there to be a high level of image-text alignment (i.e., alignment between the content of the image and the text) in images generated based on the text. For example, it may not be advantageous in this context for an image to include an excess of irrelevant content that may distract the viewer from the important content related to the text used to generate the image.
[0006] The muti-stage model training process described in this document results in a prompt expansion model that is uniquely crafted to meet the specific demands of image digital components, aligning with the nuanced requirements from digital content distribution. Some characteristics of the prompt enhancement models generated using themuti-stage training process described in this document include preserv ing the original content from the input prompt, improving generated image quality’, improving generated image aesthetics, improving image-text alignment, and removing unwanted style / content, e.g., too cartoonish, wrongly spelled text, dark / unhappy theme, etc.
[0007] In general, one innovative aspect of the subject matter described in this specification can be embodied in methods that include the actions obtaining initial training data comprising training samples, wherein each training sample includes a first input prompt and a first modified prompt generated based on the first input prompt; training, using the training samples, a prompt enhancement model to output modified prompts based on input prompts; and updating the prompt enhancement model using reinforcement learning and refinement training data comprising refinement training samples that each include a second input prompt, a second modified prompt generated by the prompt enhancement model using the second input prompt, an image generated using the second input prompt, and an image generated using the second modified prompt, including optimizing a reward that is based on an aesthetic metric for each refinement training sample and an image-text alignment score for each training sample. Other implementations of this aspect include corresponding apparatus, systems, and computer programs, configured to perform the aspects of the methods, encoded on computer storage devices.
[0008] These and other embodiments can each optionally include one or more of the following features. In some aspects, obtaining the initial training data includes generating each training sample using a language model. In some aspects, generating each training sample includes providing the first input prompt of the training sample as an input to the language model and receiving the modified prompt of the training example as an output of the language model.
[0009] In some aspects, generating each training sample includes adding instructions of one or more mutators to a prompt that includes the input prompt; providing the prompt as an input to the language model; and receiving the modified prompt as an output of the language model.
[0010] In some aspects, the one or more mutators include (i) an object mutator, (ii) an adjective and adverb mutator, (iii) a style mutator, (iv) a synonym mutator, (v) an imaging parameter mutator, (vi) a text sanitation mutator, or any combination of (i) to (vi).
[0011] In some aspects, obtaining the initial training data includes, for each training sample, generating a first image using the first input prompt and a second image using the first modified prompt; evaluating the first image and the second image; and determining aquality metric that represents a level of improvement in quality' of the second image relative to the first image based on the evaluation; and filtering one or more training samples from the initial training data based on the quality metric for each training sample.
[0012] In some aspects, the quality7metric for each training sample is based on an inclusive metric that indicates whether the first modified prompt of the training sample includes all information from the first input prompt of the training sample.
[0013] Some aspects include selecting a combination of the mutators for using in generating the training samples based on an aggregate quality metric for each of multiple combinations of mutators.
[0014] Particular embodiments of the subject matter described in this specification can be implemented so as to realize one or more of the following advantages. The techniques described in this document include a muti-stage process for training a prompt enhancement machine learning model to modify input prompts such that the prompts cause a text-to- image generation model (e.g., a text-to-image diffusion model, language model, etc.) to generate higher quality images. The multi-stage process includes a stage for ensuring that the training data used to train the prompt enhancement model results in higher quality images. For example, the prompt enhancement model can be trained on filtered training data that, after filtering, only includes modified prompts that cause the image generation model to generate higher quality images than their corresponding input prompts from which they were modified.
[0015] In some implementations, the training data is generated using one or more mutators, e.g., a combination of multiple mutators. Each mutator can include instructions that cause a machine learning model, e.g., a language model such as a large language model (LLM), to improve an input prompt in one or more ways. By improving various aspects of an input prompt using a combination of mutators, the modified prompts used to train the prompt enhancement model result in higher quality images, w hich improves the qualify of the training data for the prompt enhancement model and therefore the quality of the modified output prompts of the prompt enhancement model and the images generated using the modified prompts.
[0016] The combination of mutators used to generate training data for the prompt enhancement model can be evaluated based on the qualify of the images generated using the modified prompts output by the model and the mutators can be adjusted and / or the combination of mutators can be selected based on the evaluation. This ensures that the training data generated using the mutators is of high quality to ensure that the resultingimages are also of high quality. The evaluation can be based on various aspects of the resulting images such that the mutators that do not result in higher quality images are modified or replaced in the combination of mutators. The combination of mutators and / or the individual mutators can be adjusted to limit the size of the modified prompts output by the prompt enhancement model, which results in higher quality images as excessively long prompts often add artifacts that result in hallucinations and a large number of concepts included in a prompt can cause an image generation model to generate inaccurate and / or otherwise lower quality' images. For example, if a mutator is resulting in the addition of many (e.g., at least a threshold number) of new concepts to an input prompt, that mutator can be removed from the combination of mutators or adjusted to limit the number of concepts, e.g., by changing the instructions of the mutator.
[0017] After generating and / or filtering the training data, another stage of the multi-stage training process includes training the prompt enhancement model using the training data. In some implementations, a sequence-to-sequence text model is trained using training samples that each include an input prompt and a modified prompt which is a modified version of the input prompt. The prompt enhancement model can be trained by fine-tuning an existing language model, e g., an existing LLM using the training samples. By training this prompt enhancement model using high quality training samples obtained as described above and elsewhere herein increases the quality of the modified prompts output by the prompt enhancement model and therefore the images generated by a text-to-image model using the modified prompts.
[0018] Another stage of the multi-stage training process can include improving the prompt enhancement model, e.g.. by improving the image-text alignment and quality of the images output by the text-to-image model using the modified prompts output by the prompt enhancement model, resulting in even higher quality images than using the prompt enhancement model without this fine tuning. In some implementations, reinforcement learning is used to optimize a reward based on (i) aesthetic scores that indicate whether the image generated using the modified prompt is more visually pleasing than the image generated using the input prompt from which the modified prompt was generated and / or (ii) an image to text relevance score that quantifies how much of the original intentions of the input prompt are retained after prompt enhancement.
[0019] Using a multi-stage training process that includes training data generation and / or filtering, initial training based on the high quality training data, fine tuning the initial model using reinforcement learning results in a prompt enhancement model that generates higherquality prompts for use by text-to-image models to generate higher quality' images than using off the shelf language models or other techniques for fine-tuning language models or other types of machine learning models.
[0020] A feedback loop can be used to further improve the prompt enhancement model at one or more stages of the multi-stage training process, resulting in higher quality' images generate using modified prompts output by the prompt enhancement model. For example, images generated using the modified prompts can be evaluated as against images generated using the input prompts from which the modified prompts were generated by the prompt enhancement model. The quality of the images from each pair of prompts (input prompt and its modified prompt) can be measured and used to modify the prompt enhancement model at any time, as a continuous learning process.
[0021] The details of one or more embodiments of the subject matter described in this specification are set forth in the accompanying drawings and the description below. Other features, aspects, and advantages of the subject matter will become apparent from the description, the drawings, and the claims.BRIEF DESCRIPTION OF THE DRAWINGS
[0022] FIG. 1 is a block diagram of an example environment in which a prompt enhancement model is trained and used to generate prompts for generating images.
[0023] FIG. 2 is a block diagram illustrating interactions between an artificial intelligence system, a language model, a prompt enhancement model, a text-to-image model, and a client device.
[0024] FIG. 3 is a flow chart of an example process of training a prompt enhancement model.
[0025] FIG. 4 is a flow chart of an example process of generating images using a prompt enhancement model and a text-to-image model.
[0026] FIG. 5 a block diagram of an example computer.
[0027] Like reference numbers and designations in the various drawings indicate like elements.DETAILED DESCRIPTION
[0028] This specification describes techniques for enabling artificial intelligence to generate new images based on enhanced prompts generated by a trained machine learning model, which can be referred to as a prompt enhancement model. Artificial intelligence (Al) is asegment of computer science that focuses on the creation of models that can perform tasks act autonomously, e.g.. with little to no human intervention. Al systems can utilize, for example, one or more of machine learning, natural language processing, or computer vision. Machine learning, and its subsets, such as deep learning, focus on developing models that can infer outputs from data. The outputs can include, for example, predictions and / or classifications. Natural language processing focuses on analyzing and generating human language. Computer vision focuses on analyzing and interpreting images and videos. Artificial intelligence systems can include generative models that generate new content, such as images, videos, text, audio, and / or other content, in response to input prompts and / or based on other information.
[0029] The techniques described throughout this specification enable artificial intelligence to train a prompt enhancement model using a multi-stage training process that includes a stage for generating and / or filtering training data, a stage for training the prompt enhancement model using the training data, and / or a stage for improving the trained prompt enhancement model, e.g., using reinforcement learning techniques. The muti-stage process can include two or more of these stages. The prompt enhancement model is trained to generate a modified prompt based on an input prompt, e.g., a prompt provided by a digital component provider. The modified prompt can be an expanded prompt or an otherwise enhanced version of the input prompt. The modified prompt can then be used by a text-to- image model to generate one or more images based on the modified prompt. Using the described multi-stage approach, the prompt enhancement model is adapted to generate modified prompts that result in high uality images that are more accurate, more aesthetically pleasing to users, and have better image-text alignment than user generated prompts, e.g.. without adapting the text-to-image model such that the techniques improve generative Al image generation using off the shelf text-to-image models.
[0030] FIG. 1 is a block diagram of an example environment 100 in which customized digital component generation can be performed. The example environment 100 includes a network 102, such as a local area network (LAN), a wide area network (WAN), the Internet, or a combination thereof. The network 102 connects electronic document servers 104, user devices 106, digital component servers 108, and a service apparatus 1 10. The example environment 100 may include many different electronic document servers 104, user devices 106, and digital component servers 108.
[0031] The service apparatus 110 is configured to provide various services to client devices 106, publishers of electronic documents 150, and / or digital component providers thatprovide digital components to client devices 106, e.g., using the service apparatus 110. In some implementations, the service apparatus 110 can provide search services by providing responses to search queries received from client devices 106. For example, the services apparatus 110 can include a search engine and / or an Al agent or other chat agent that enables users to interact with the agent over the course of multiple conversational queries and responses. The service apparatus 110 can also distribute digital components to client devices 106 for presentation with the responses and / or with electronic documents 150. For example, another search service computer system can send component requests 1 12 to the service apparatus 110 and these component requests 112 can include one or more queries. In another example, client devices 106 can send component requests 112 to the service apparatus 110, e.g.. in response to loading an electronic document 150 that includes a digital component slot for presenting a digital component, as described in more detail below.
[0032] The service apparatus 110 is configured to generate images, e.g., digital components that include images (“image digital components”). For example, the service apparatus 110 can receive an input prompt from a device of a digital component provider, enhance the prompt using Al. and generate one or more images based on the modified prompt using Al. The service apparatus 110 and component requests 112 are described in further detail below.
[0033] As used throughout this document, the phrase “digital component” refers to a discrete unit of digital content or digital information (e.g., a video clip, audio clip, multimedia clip, gaming content, image, text, bullet point, artificial intelligence output, language model output, or another unit of content). A digital component can electronically be stored in a physical memory device as a single file or in a collection of files, and digital components can take the form of video files, audio files, multimedia files, image files, or text files and include advertising information, such that an advertisement is a type of digital component.
[0034] A client device 106 is an electronic device capable of requesting and receiving online resources over the network 102. Example client devices 106 include personal computers, gaming devices, mobile communication devices, digital assistant devices, augmented reality devices, virtual reality devices, and other devices that can send and receive data over the network 102. A client device 106 typically includes a user application, such as a web browser, to facilitate the sending and receiving of data over the network 102. but native applications (other than browsers) executed by the client device 106 can also facilitate the sending and receiving of data over the network 102.
[0035] A gaming device is a device that enables a user to engage in gaming applications, for example, in which the user has control over one or more characters, avatars, or other rendered content presented in the gaming application. A gaming device typically includes a computer processor, a memory device, and a controller interface (either physical or visually rendered) that enables user control over content rendered by the gaming application. The gaming device can store and execute the gaming application locally, or execute a gaming application that is at least partly stored and / or served by a cloud server (e.g., online gaming applications). Similarly, the gaming device can interface with a gaming server that executes the gaming application and “streams’’ the gaming application to the gaming device. The gaming device may be a tablet device, mobile telecommunications device, a computer, or another device that performs other functions beyond executing the gaming application.
[0036] Digital assistant devices include devices that include a microphone and a speaker. Digital assistant devices are generally capable of receiving input by way of voice, and respond with content using audible feedback, and can present other audible information. In some situations, digital assistant devices also include a visual display or are in communication with a visual display (e.g.. by way of a wireless or wired connection). Feedback or other information can also be provided visually when a visual display is present. In some situations, digital assistant devices can also control other devices, such as lights, locks, cameras, climate control devices, alarm systems, and other devices that are registered with the digital assistant device.
[0037] As illustrated, the client device 106 is presenting an electronic document 150. An electronic document is data that presents a set of content at a client device 106. Examples of electronic documents include webpages, word processing documents, portable document format (PDF) documents, images, videos, search results pages, and feed sources. Native applications (e.g., “apps” and / or gaming applications), such as applications installed on mobile, tablet, or desktop computing devices are also examples of electronic documents. Electronic documents can be provided to client devices 106 by electronic document servers 104 (“Electronic Doc Servers”).
[0038] For example, the electronic document servers 104 can include servers that host publisher websites. In this example, the client device 106 can initiate a request for a given publisher webpage, and the electronic server 104 that hosts the given publisher webpage can respond to the request by sending machine executable instructions that initiate presentation of the given webpage at the client device 106.
[0039] In another example, the electronic document servers 104 can include app servers from which client devices 106 can download apps. In this example, the client device 106 can download files required to install an app at the client device 106, and then execute the downloaded app locally (i.e., on the client device). Alternatively, or additionally, the client device 106 can initiate a request to execute the app, which is transmitted to a cloud server. In response to receiving the request, the cloud server can execute the application and stream a user interface of the application to the client device 106 so that the client device 106 does not have to execute the app itself. Rather, the client device 106 can present the user interface generated by the cloud server’s execution of the app, and communicate any user interactions with the user interface back to the cloud server for processing.
[0040] Electronic documents can include a variety of content. For example, an electronic document 150 can include native content 152 that is within the electronic document 150 itself and / or does not change over time. Electronic documents can also include dynamic content that may change over time or on a per-request basis. For example, a publisher of a given electronic document (e.g., electronic document 150) can maintain a data source that is used to populate portions of the electronic document. In this example, the given electronic document can include a script, such as the script 154, that causes the client device 106 to request content (e.g., a digital component) from the data source when the given electronic document is processed (e.g., rendered or executed) by a client device 106 (or a cloud server). The client device 106 (or cloud server) integrates the content (e.g.. digital component) obtained from the data source into the given electronic document to create a composite electronic document including the content obtained from the data source.
[0041] In some situations, a given electronic document (e.g., electronic document 150) can include a digital component script (e.g., script 154) that references the service apparatus 110, or a particular service provided by the service apparatus 110. In these situations, the digital component script is executed by the client device 106 when the given electronic document is processed by the client device 106. Execution of the digital component script configures the client device 106 to generate a request for digital components (referred to as a ‘"component request”), which is transmitted over the network 102 to the service apparatus 110. For example, the digital component script can enable the client device 106 to generate a packetized data request including a header and payload data. The component request 112 can include event data specifying features such as a name (or netw ork location) of a server from which the digital component is being requested, a name (or network location) of the requesting device (e.g., the client device 106), and / or information that the service apparatus110 can use to select one or more digital components, or other content, provided in response to the request. The component request 112 is transmitted, by the client device 106, over the network 102 (e.g., a telecommunications network) to a server of the service apparatus 110.
[0042] The component request 112 can include event data specifying other event features, such as the electronic document being requested and characteristics of locations of the electronic document at which the digital component can be presented. For example, event data specifying a reference (e.g., URL) to an electronic document (e.g., webpage) in which the digital component will be presented, available locations of the electronic documents that are available to present digital components, sizes of the available locations, and / or media types that are eligible for presentation in the locations can be provided to the service apparatus 110. Similarly, event data specifying keywords associated with the electronic document (‘'document keywords”) or entities (e.g., people, places, or things) that are referenced by the electronic document can also be included in the component request 112 (e.g., as payload data) and provided to the service apparatus 110 to facilitate identification of digital components that are eligible for presentation with the electronic document.
[0043] The event data can also include a search query that was submitted from the client device 106 to obtain a search results page or a response in a conversational user interface. For example, an Al agent or other form of a chat agent can provide a conversational user interface in which users can provide natural language queries, which can be in the form of prompts for a language model, and receive responses to the queries. The user can refine their expression of their informational needs as the conversation progresses and the Al agent can send component requests 112 with the updated queries. In such examples, the Al agent can include a user session identifier in the component requests so that the service apparatus 110 can correlate queries included in multiple component requests 112 for the same user session with the Al agent and use this information in generating customized digital components.
[0044] Component requests 112 can also include event data related to other information, such as information that a user of the client device has provided, geographic information indicating a state or region from which the component request was submitted, or other information that provides context for the environment in which the digital component will be displayed (e.g., a time of day of the component request, a day of the week of the component request, a type of device at which the digital component will be displayed, such as a mobile device or tablet device). Component requests 112 can be transmitted, for example, over a packetized network, and the component requests 112 themselves can beformatted as packetized data having a header and payload data. The header can specify a destination of the packet and the pay load data can include any of the information discussed above.
[0045] The service apparatus 110 chooses digital components (e.g., third-party content, such as video files, audio files, images, text, gaming content, augmented reality content, and combinations thereof, which can all take the form of advertising content or non-advertising content) that will be presented with the given electronic document (e.g.. at a location specified by the script 154) in response to receiving the component request 112 and / or using information included in the component request 112. In some implementations, choosing a digital component includes choosing a customizable digital component that can be customized based on various data, as described in more detail below.
[0046] In some implementations, a digital component is selected in less than a second to avoid errors that could be caused by delayed selection of the digital component. For example, delays in providing digital components in response to a component request 112 can result in page load errors at the client device 106 or cause portions of the electronic document to remain unpopulated even after other portions of the electronic document are presented at the client device 106.
[0047] Also, as the delay in providing the digital component to the client device 106 increases, it is more likely that the electronic document will no longer be presented at the client device 106 when the digital component is delivered to the client device 106, thereby negatively impacting a user’s experience with the electronic document. Further, delays in providing the digital component can result in a failed delivery7of the digital component, for example, if the electronic document is no longer presented at the client device 106 when the digital component is provided. The described techniques are adapted to generate a customized digital component in a short amount of time such that these errors and user experience impact are reduced or eliminated.
[0048] In some implementations, the service apparatus 110 is implemented in a distributed computing system that includes, for example, a server and a set of multiple computing devices 114 that are interconnected and identify and distribute digital component in response to requests 112. The set of multiple computing devices 114 operate together to identify a set of digital components that are eligible to be presented in the electronic document from among a corpus of millions of available digital components (DCi-x). The millions of available digital components can be indexed, for example, in a digital component database 116. Each digital component index entry can reference thecorresponding digital component and / or include distribution parameters (DPi-DPx) that contribute to (e.g., trigger, condition, or limit) the distribution / transmission of the corresponding digital component. For example, the distribution parameters can contribute to (e.g., trigger) the transmission of a digital component by requiring that a component request include at least one criterion that matches (e.g., either exactly or with some prespecified level of similarity) one of the distribution parameters of the digital component.
[0049] In some implementations, the distribution parameters for a particular digital component can include distribution keywords that must be matched (e.g., by electronic documents, document keywords, and / or terms specified in the component request 112) in order for the digital component to be eligible for presentation. Additionally, or alternatively, the distribution parameters can include embeddings that can use various different dimensions of data, such as website details and / or consumption details (e.g., page viewport, user scrolling speed, or other information about the consumption of data). The distribution parameters can also require that the component request 112 include information specifying a particular geographic region (e.g.. country’ or state) and / or information specifying that the component request 112 originated at a particular type of client device (e.g., mobile device or tablet device) in order for the digital component to be eligible for presentation. The distribution parameters can also specify an eligibility' value (e.g., ranking score, or some other specified value) that is used for evaluating the eligibility of the digital component for distribution / transmission (e.g., among other available digital components).
[0050] The identification of the eligible digital component can be segmented into multiple tasks 117a- 117c that are then assigned among computing devices within the set of multiple computing devices 114. For example, different computing devices in the set 114 can each analyze a different portion of the digital component database 116 to identify various digital components having distribution parameters that match information included in the component request 112. In some implementations, each given computing device in the set 114 can analyze a different data dimension (or set of dimensions) and pass (e.g., transmit) results (Res 1-Res 3) 118a-l 18c of the analysis back to the service apparatus 110. For example, the results 118a- 118c provided by each of the computing devices in the set 114 may identify a subset of digital components that are eligible for distribution in response to the component request and / or a subset of the digital component that have certain distribution parameters. The identification of the subset of digital components can include, for example, comparing the event data to the distribution parameters, and identifying thesubset of digital components having distribution parameters that match at least some features of the event data.
[0051] The service apparatus 110 aggregates the results 118a- 118c received from the set of multiple computing devices 114 and uses information associated with the aggregated results to select one or more digital components that will be provided in response to the request 112. For example, the service apparatus 110 can select a set of winning digital components (one or more digital components) based on the outcome of one or more content evaluation processes, as discussed below. In turn, the service apparatus 110 can generate and transmit, over the network 102, reply data 120 (e.g., digital data representing a reply) that enable the client device 106 to integrate the set of winning digital components into the given electronic document, such that the set of winning digital components (e.g.. winning third-party content) and the content of the electronic document are presented together at a display of the client device 106. In some implementations, the client device 106 executes instructions included in the reply data 120, which configures and enables the client device 106 to obtain the set of winning digital components from one or more digital component servers 108. For example, the instructions in the reply data 120 can include a network location (e.g.. a Uniform Resource Locator (URL)) and a script that causes the client device 106 to transmit a server request (SR) 121 to the digital component server 108 to obtain a given winning digital component from the digital component server 108. In response to the request, the digital component server 108 will identify the given winning digital component specified in the server request 121 (e.g., within a database storing multiple digital components) and transmit, to the client device 106, digital component data (DC Data) 122 that presents the given winning digital component in the electronic document at the client device 106.
[0052] When the client device 106 receives the digital component data 122. the client device will render the digital component (e.g., third-party content), and present the digital component at a location specified by, or assigned to, the script 154. For example, the script 154 can create a walled garden environment, such as a frame, that is presented within, e.g., beside, the native content 152 of the electronic document 150. In some implementations, the digital component is overlay ed over (or adjacent to) a portion of the native content 152 of the electronic document 150, and the service apparatus 110 can specify’ the presentation location within the electronic document 150 in the reply 120. For example, when the native content 152 includes video content, the service apparatus 110 can specify a location or object within the scene depicted in the video content over which the digital component is to be presented.
[0053] The service apparatus 110 can also include an Al system 160 configured to autonomously generate images, e.g., image digital components, either prior to a component request 112 (e.g., offline) and / or in response to a request 1 12 (e.g., online or real-time). The Al system 160 can use one or more Al models, e.g., one or more machine learning models, to generate images based on input prompts, which can be received from digital component providers or other users that request image creation by the Al system 160.
[0054] In some implementations, the Al system 160 is configured to modify, e.g., expand, contract, rewrite, or otherwise enhance, input prompts using a prompt enhancement model 170. The input prompts can include words or phrases that provide instructions for generating an image. For example, the Al system 160 can receive a prompt 172 from a device, e.g., a device of a digital component provider or a client device 106, and provide the prompt 172 as input to the prompt enhancement model 170. The prompt enhancement model 170 is trained to modify the input prompt 170 and generate an output 174 that includes a modified prompt. The Al system 160 can provide the modified prompt as an input 182 to a text-to-image model 180 that is trained to generate, as outputs 184, new images based on prompt inputs 182. The Al system 160 can receive the image outputs 184 and either provide the images to the device that provided the prompt or create digital components using the images and distribute the digital components to client devices 106, as described above.
[0055] One or more of the machine learning models trained and / or used by the Al system 160, e.g., the prompt enhancement model 170 and / or the text-to-image model 180, can include language models, e.g., large language models. A large language model ("LLM") is a model that is trained to generate and understand human language. LLMs are trained on massive datasets of text and code, and they can be used for a variety of tasks. For example, LLMs can be trained to translate text from one language to another; summarize text, such as web site content, search results, news articles, or research papers; answer questions about text, such as "What is the capital of Georgia?”; create chatbots that can have conversations with humans; and generate creative text, such as poems, stories, and code.
[0056] The language model(s) can be any appropriate language model neural network that receives an input sequence made up of text tokens selected from a vocabulary and auto- regressively generates an output sequence made up of text tokens from the vocabulary. For example, the language model can be a Transformer-based language model neural network or a recurrent neural network-based language model.
[0057] In some situations, a language model can be referred to as an auto-regressive neural network when the neural network used to implement the language model auto-regressively generates an output sequence of tokens. More specifically, the auto-regressively generated output is created by generating each particular token in the output sequence conditioned on a current input sequence that includes any tokens that precede the particular text token in the output sequence, i.e., the tokens that have already been generated for any previous positions in the output sequence that precede the particular position of the particular token, and a context input that provides context for the output sequence.
[0058] For example, the current input sequence when generating a token at any given position in the output sequence can include the input sequence and the tokens at any preceding positions that precede the given position in the output sequence. As a particular example, the current input sequence can include the input sequence followed by the tokens at any preceding positions that precede the given position in the output sequence.Optionally, the input and the current output sequence can be separated by one or more predetermined tokens within the current input sequence.
[0059] More specifically, to generate a particular token at a particular position within an output sequence, the neural network of the language model can process the current input sequence to generate a score distribution, e.g., a probability distribution, that assigns a respective score, e.g.. a respective probability, to each token in the vocabulary of tokens. The neural network of the language model can then select, as the particular token, a token from the vocabulary using the score distribution. For example, the neural network of the language model can greedily select the highest-scoring token or can sample, e.g., using nucleus sampling or another sampling technique, a token from the distribution.
[0060] As a particular example, the language model can be an auto-regressive Transformerbased neural network that includes (i) a plurality of attention blocks that each apply a selfattention operation and (ii) an output subnetwork that processes an output of the last attention block to generate the score distribution.
[0061] A language model can have any of a variety of Transformer-based neural network architectures. Examples of such architectures include those described in J. Hoffmann, S. Borgeaud, A. Mensch, E. Buchatskaya, T. Cai, E. Rutherford, D. d. L. Casas, L. A. Hendricks, J. Welbl, A. Clark, et al. Training compute-optimal large language models, arXiv preprint arXiv:2203. 15556, 2022; J.W. Rae, S. Borgeaud, T. Cai, K. Millican, J. Hoffmann, H. F. Song, J. Aslanides, S. Henderson, R. Ring. S. Young, E. Rutherford, T. Hennigan, J. Menick, A. Cassirer, R. Powell, G. van den Driessche, L. A. Hendricks, M.Rauh, P. Huang, A. Glaese, J. Welbl. S. Dathathri, S. Huang, J. Uesato, J. Mellor, I. Higgins, A. Creswell, N. McAleese, A. Wu, E. Eisen, S. M. Jayakumar. E. Buchatskaya, D. Budden, E. Sutherland, K. Simonyan, M. Paganini, L. Sifre, L. Martens, X. L. Li, A. Kuncoro, A. Nematzadeh, E. Gribovskaya, D. Donato, A. Lazaridou, A. Mensch, J. Lespiau, M. Tsimpoukelli, N. Grigorev, D. Fritz, T. Sottiaux, M. Pajarskas, T. Pohlen, Z. Gong, D. Toyama, C. de Masson d’Autume, Y. Li, T. Terzi. V. Mikulik, I. Babuschkin, A. Clark. D. de Las Casas. A. Guy, C. Jones, J. Bradbury, M. Johnson. B. A. Hechtman. L. Weidinger, I. Gabriel, W. S. Isaac, E. Lockhart, S. Osindero, L. Rimell, C. Dyer, O. Vinyals, K. Ayoub, J. Stanway, L. Bennett, D. Hassabis, K. Kavukcuoglu, and G. Irving. Scaling language models: Methods, analysis & insights from training gopher. CoRR, abs / 2112. 11446, 2021 ; Colin Raffel, Noam Shazeer. Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. Exploring the limits of transfer learning with a unified text-to-text transformer. arXiv preprint arXiv: 1910. 10683, 2019; Daniel Adiwardana, Minh-Thang Luong, David R. So, Jamie Hall, Noah Fiedel, Romal Thoppilan, Zi Yang, Apoorv Kulshreshtha, Gaurav Nemade. Yifeng Lu. and Quoc V. Le. Towards a human-like open-domain chatbot. CoRR, abs / 2001.09977, 2020; and Tom B Brown. Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. Language models are few-shot learners. arXiv preprint arXiv:2005. 14165, 2020.
[0062] Generally, however, the Transformer-based neural network includes a sequence of attention blocks, and, during the processing of a given input sequence, each attention block in the sequence receives a respective input hidden state for each input token in the given input sequence. The attention block then updates each of the hidden states at least in part byapplying self-attention to generate a respective output hidden state for each of the input tokens. The input hidden states for the first attention block are embeddings of the input tokens in the input sequence and the input hidden states for each subsequent attention block are the output hidden states generated by the preceding attention block.
[0063] In this example, the output subnetwork processes the output hidden state generated by the last attention block in the sequence for the last input token in the input sequence to generate the score distribution.
[0064] Generally, because the language model is auto-regressive, the service apparatus 110 can use the same language model to generate multiple different candidate output sequences in response to the same request, e.g., by using beam search decoding from score distributions generated by the language model, using a Sample-and-Rank decoding strategy.by using different random seeds for the pseudo-random number generator that’s used in sampling for different runs through the language model or using another decoding strategy that leverages the auto-regressive nature of the language model.
[0065] In some implementations, the language model is pre-trained, e.g., trained on a language modeling task that does not require providing evidence in response to user questions, and the service apparatus 110 (e.g., using Al system 160) causes the language model to generate output sequences according to the pre-determined syntax through natural language prompts in the input sequence.
[0066] For example, the sendee apparatus 110 (e.g., Al system 160), or a separate training system, pre-trains a language model (e.g., the neural network) on a language modeling task, e.g., a task that requires predicting, given a current sequence of text tokens, the next token that follows the current sequence in the training data. As a particular example, the language model can be pre-trained on a maximum-likelihood objective on a large dataset of text, e.g., text that is publicly available from the Internet or another text corpus. In some implementations, as described below, the Al system 160 can train the prompt enhancement model 170 by adapting, e.g., fine tuning, a language model using training data.
[0067] FIG. 2 is a block diagram illustrating interactions between an artificial intelligence system 160, a language model 210, a prompt enhancement model 170, a text-to-image model 180, and a client device 106. The Al system 160 can include a training apparatus 206, a prompt apparatus 208. an image evaluation apparatus 210, and a digital component apparatus 212. The Al system 1 0 can be configured to interact with a memory structure 230 to extract and / or store information and content. In particular, the memory structure 230 can store the digital component database 116, a training database 232 that stores training data for training one or more of the machine learning models, e.g., models 170, 180, and / or 220) described herein, and a prompt database 234 that stores prompts. The memory structure 230 can include one or more databases or other data structures stored on one or more memories and / or data storage devices.
[0068] The training apparatus 206 is configured to obtain the training data stored in the training database 232. The training apparatus 206 can obtain the training data by generating the training data, collecting the training data from another source or sources, and / or filtering the obtained training data, e.g., to remove low quality data from the training data stored in the training database 232.
[0069] In some implementations, the training apparatus 206 collects data for training the prompt enhancement model 170 from the prompt database 234. The prompt database 234can include prompts received by the Al system 160 or another system for the purposes of generating images. For example, the prompt database 234 can include actual prompts received from devices of digital component providers and / or client devices 106. The prompts stored in the prompt database 234 can be used as seed prompts for training the prompt enhancement model 170.
[0070] In this example, the training apparatus 206 can modify each seed prompt using the language model 220. For example, the training apparatus 206 can provide an input prompt 222 to the language model 220 that includes the seed prompt and instructions for improving the seed prompt and the language model 220 can generate a modified prompt output 224 based on the input prompt 222. The instructions can include, for example, one or more customized style modifiers for the language model 220 to add to the seed prompt to generate a modified prompt. In a particular example, the instructions can include aesthetic suffixes, e.g., “professional photography”, “photorealistic”, “realistic”, “studio medium”, “super detailed”, “high detail”, “very detailed,” etc. In general, each aesthetic suffix is selected to cause the language model 220 to expand the prompt to generate a modified prompt that causes the text-to-image model 180 to generate a higher quality image than the corresponding seed prompt from which the modified prompt 224 is generated. Each seed prompt and its corresponding modified prompt 224 can be referred to as a training sample that includes a prompt pair, where the prompt pair is the input prompt (i.e., seed prompt in this example) and the modified prompt 224.
[0071] In another example, the training apparatus 206 can generate training data using mutators 207. Each mutator 207 can include instructions that cause the language model 220 to improve an input prompt 172 in one or more ways. Some example mutators 207 include an object mutator, an adjective and adverb mutator, a style mutator, a synonym mutator, an imaging parameter mutator, and a text sanitation mutator.
[0072] The object mutator can be configured to cause the language model 220 to modify the input prompt such that the modified prompt causes the text-to-image model 180 to add or modify objects in the resulting image that is generated using the modified prompt, e.g., relative to the image that is generated based on the input prompt. For example, the object mutator can include instructions that cause the language model 220 to add or modify objects to the input prompt. The instructions can be part of the input 222 provided to the language model 220.
[0073] For example, the object mutator’s instructions can be “Please try to improve the above PROMPT by adding more objects or concepts to enrich the design.” The input 222can include the input prompt and the instructions of the object mutator, and optionally the instructions of one or more other mutators. The input prompts used with the mutators can be part of a training set of input prompts. This training set can include actual prompts submitted by users and / or input prompts generated for the purpose of generating training data for the prompt enhancement model 170. For example, an LLM or other machine learning model, or a prompt designer can generate the input prompts.
[0074] The adjective and adverb mutator can be configured to cause the language model 220 to modify the input prompt such that the modified prompt causes the text-to-image model 180 to add or change properties of the object or image itself, e.g., by adding adjectives or adverbs to the input prompt to generate the modified prompt. For example, the adjective and adverb mutator's instructions can be ‘"Please try to improve the above PROMPT by adding more relevant adjectives or adverbs to enrich the design, they can be either added to the obj ects in the image, or added to the style / imaging condition of the image.’' The input 222 can include the input prompt and the instructions of the adjective and adverb mutator and optionally the instructions of one or more other mutators.
[0075] The style mutator can be configured to cause the language model 220 to modify the input prompt such that the modified prompt causes the text-to-image model 180 to change image sty les of the resulting image, e.g., by suggesting sty les to the language model 220. For example, the style mutator's instructions can be “Sometimes changing the style of the image can create a stunning experience for the audience. For example, instead of ‘photorealistic’ image, we can consider the following styles, [‘watercolor”, sketch’, ’oil painting’, ‘neo wallpaper’]. Also, it is strongly recommended to think of creative styles out of this list. Please tty to improve the above PROMPT by changing the style of the image, make sure the style you choose is relevant to the context of the content of the digital component.” The list can include any combination of sty les.
[0076] The synonym mutator can be configured to cause the language model 220 to modify the input prompt such that the modified prompt includes synonyms for words or phrases in the input prompt. For example, the synonym mutator's instructions can be “Replace some of the words in the above PROMPT with synonyms to improve the PROMPT.” An input prompt may be “ Children hand-feeding goats, collecting colorful eggs from a chicken coop, riding a pony (with supervision, of course), or learning to plant seeds in a vegetable garden” and example output using this input may be “Children offering food to goats, harvesting brightly-colored eggs from a chicken coop, riding a pony (under careful watch, of course), or studying how to plant seeds in a vegetable garden.”
[0077] The imaging parameter mutator can be configured to cause the language model 220 to modify the input prompt such that the modified prompt causes the text-to-image model 180 to change imaging parameters, e.g., lighting condition, focus, camera angle, etc., of the resulting image, e.g., by suggesting image parameters for the language model 220 to change in the modified prompt. For example, the imaging parameter mutator’s instructions can be “Changing the imaging parameters of the image can create a stunning experience for the audience. For example, you can add to or change the following imaging parameters, [‘lighting conditions’, ‘focus’, ’shutter speed’, ‘resolution’, ‘camera quality’, 'object distance’]. Also, it is strongly recommended to think of creative imaging parameters out of this list. Please try to improve the above PROMPT by adding to or changing the imaging parameters of the image, make sure the imaging parameters you choose or change are relevant to the content of the digital component.” The list can include any combination of imaging parameters.
[0078] The text sanitation mutator can be configured to cause the language model 220 to modify the input prompt to remove text related keywords from the input prompt. For example, the text sanitation mutator’s instructions can be “We have an image generation model that can generate digital component images based on a text prompt, ft has a limitation that it cannot render text correctly. Could you help me correct them so that they can be used with our image generation model? Example: original design idea: a photo of a living room with a custom-made 3D printed coffee table in the center. The text reads: 'Tailored furniture just for you! .’ Corrected design: A photo of a living room with a custom- made 3D printed coffee table in the center. Please help me correct the following design idea. Please output an object with field names ‘original design idea’ and ‘corrected design idea’.”
[0079] The training apparatus 206 can generate the input 222 using the input prompt and the instructions of one or more of the mutators, e.g., any one mutator or any combination of two or more mutators. For each input 222 and modified prompt output by the language model 220, the training apparatus 206 can generate a training example. Similar to the training samples described above, each training sample can include a prompt pair, where the prompt pair is the input prompt and the modified prompt 224.
[0080] In some implementations, the Al system 160 can adjust the one or more mutators that are used to generate the modified prompts of the training samples based on the quality (e.g., accuracy) of the images generated using the modified prompts relative to the qualityof the images generated using the input prompts corresponding to the modified prompts, as described in more detail below.
[0081] The training data used to train the prompt enhancement model 170 can include training samples generated from the seed prompts and / or the training samples generated using the mutators. In some implementations, the training apparatus 206 can filter the training data, e.g., the training samples, to remove low quality training samples. For example, the training apparatus 206 can remove training samples for which the modified prompt does not improve the image relative to its corresponding input prompt. That is, the training apparatus 206 can remove training samples for which an image generated by the text-to-image model 180 using the modified prompt is not an improvement over an image generated by the text-to-image model 180 using the input prompt.
[0082] For each training sample, the training apparatus 206 can provide the input prompt to the text-to-image model 180 and receive a first image generated by the text-to-image model 180 using the input prompt. The training apparatus 206 can also provide the modified prompt of the training sample to the text-to-image model 180 and receive a second image generated by the text-to-image model 180 using the modified prompt.
[0083] The training apparatus 206 can then provide both the first image and the second image to the image evaluation apparatus 210. The image evaluation apparatus 210 is configured to evaluate the quality of images. The quality of an image can be based on aesthetics of the image, the quality of the modified prompt, the style and / or content of the image, the image-text alignment of the image relative to the text of the prompt, and / or the accuracy of the image in representing this text. For example, the image evaluation apparatus 210 can be configured to generate, as a quality metric, an aesthetic metric that indicates whether the second image generated using the modified prompt is more visually pleasing than the first image generated using the input prompt. Example models that can be used to generate the aesthetic metric include VILA AVA and VILA SAC.
[0084] The image evaluation apparatus 210 can be configured to generate an inclusive metric that indicates whether the modified prompt includes all information from the original prompt, which is a reflection of the accuracy of the image generation. The language model 220 can be used to generate the inclusive metric based on the input prompt and the modified prompt. For example, the image evaluation apparatus 210 can provide an input 222 that includes both prompts and instructions for generating the inclusive metric to the language model 220 and the language model 220 can provide an output 224 that includes the inclusive metric.
[0085] The image evaluation apparatus 210 can be configured to perform a style and / or content check on the second image. For example, the image evaluation apparatus 210 can generate a style and content metric that indicates the quality of the style and / or content of the second image. As a particular example, if the second image is too cartoonish or has unwanted content such as text or collage, the style and content metric may be low. A machine learning model can be trained to output the style and content metric using the second image as an input. In another example, one or more persons can evaluate the second image and provide the style and content metric.
[0086] The image evaluation apparatus 210 can determine an overall quality metric for each training sample based one or more of the aesthetic metric, the inclusive metric, and the style and content metric. The overall quality metric can be a combination, e.g., a weighted combination, of the individual metrics. For example, the overall quality metric can be a sum, an average, a weighted average, or another measure of central tendency between the metrics. The training apparatus 206 can filter, from the training data, each training sample for which the overall quality metric fails to satisfy a threshold. The threshold can be a specified value. In this example, the remaining training samples of the training data would be those having an overall quality metric that satisfies, e.g., meets or exceeds, the threshold.
[0087] In another example, the training apparatus 206 can keep a specified number of training samples having the highest overall metrics and remove the rest of the training samples. In this example, the threshold can be considered to be the overall metric for the Nth training sample if the specified number of training samples is N.
[0088] The training apparatus 206 can train the prompt enhancement model 170 using the training samples remaining after the filtering is performed. The prompt enhancement model 170 is trained to modify’ input prompts to generate modified prompts that are enhanced version of the input prompts. In some implementations, the prompt enhancement model 170 is a sequence-to-sequence text model. For example, the training apparatus 206 can train the sequence-to-sequence text model by adjusting, e.g., fine tuning, an existing LLM such as the language model 220. The training apparatus 206 can fine tune the LLM using the training samples that remain after the filtering is performed. For example, the training apparatus 206 can fine tune the LLM by applying generative loss on the modified prompts of the training samples.
[0089] The prompt enhancement model can be trained using <input prompt, modified prompt pairs> pairs. Given the input prompt, the sequence-to-sequence model can be trained with sequence generation loss (e.g., next token prediction) on the modified prompt.
[0090] The training apparatus 206 can also improve the trained prompt enhancement model 170 using reinforcement learning. A goal of the reinforcement learning can be to optimize a reward that is determined by a composite of metrics that compare a first image generated using an input prompt and a second image generated using a modified prompt output by the trained prompt enhancement model 170. For example, a refinement set of training samples can include a set of seed input prompts. These seed input prompts can include actual user prompts or prompts generated for the purposes of training, e.g., by users, an LLM, or another machine learning model. The training apparatus 206 can provide each seed input prompt as input to the trained prompt enhancement model 170 and the trained prompt enhancement model 170 can output a modified prompt.
[0091] For each training sample in the refinement set, the training apparatus 206 can provide the input prompt to the text-to-image model 180 and receive a first image generated by the text-to-image model 180 using the input prompt. The training apparatus 206 can also provide the modified prompt to the text-to-image model 180 and receive a second image generated by the text-to-image model 180 using the input prompt.
[0092] The training apparatus 206 can provide the first and second image for each training sample in the refinement set to the image evaluation apparatus 210. The image evaluation apparatus 210 can determine a composite score for each training sample based on an aesthetic metric and / or an image-text alignment metric. The image evaluation apparatus 210 can evaluate the first and second images to determine the aesthetic metric based on whether the second image generated using the modified prompt is more visually pleasing than the first image generated using the input prompt. Example models that can be used to generate the aesthetic metric include VILA AV A and VILA SAC.
[0093] The image evaluation apparatus 210 can evaluate the input prompt and the second image generated using the modified prompt to determine the image-text alignment metric. In general, the image-text alignment metric quantifies how much the original user intentions are retained after the input prompt is modified to generate the modified prompt. In some implementations, the image evaluation apparatus 210 can evaluate the input prompt and the second image using a machine learning model trained to generate CLIP scores that are text- to-image similarity metrics.
[0094] The image evaluation apparatus 210 can then determine a composite metric for each training sample in the refinement set using the aesthetic metric and the image-text alignment metric for the training sample. In some implementations, the composite metric for each training sample can also be based on a performance metric for digital components thatinclude the first or second image of the training sample and / or an evaluation metric provided by a human based on images and / or prompts of the training sample. The performance metric for the digital components can include a click-through or conversion rate for the second image of the training sample or a difference between the click -through or conversion rate for digital components that include the second image and the same metric for digital components that include the first image.
[0095] The composite metric can be a combination, e.g., a weighted combination, of the metrics. For example, the composite metric can be a sum, an average, a weighted average, or another measure of central tendency between the metrics. The training apparatus 206 can use the composite score for each training sample in the refinement set as a reward in the reinforcement training process.
[0096] The Al system 160 can use the trained prompt enhancement model 170 to generate modified prompts for image generation, e.g., in response to requests received from devices of digital component providers or from client devices 106. The prompt apparatus 208 can be configured to receive an input prompt from a device and provide an input prompt 172. e.g., the received prompt, to the prompt enhancement model 170. The prompt enhancement model 170 can generate a modified prompt and provide the modified prompt as an output 174 to the prompt apparatus 208.
[0097] To generate an image using the modified prompt, the prompt apparatus 208 can provide the modified prompt as an input 182 to the text-to-image model 180. The text-to- image model is trained to generate images based on text prompts, e.g., based on natural language text prompts. The text-to-image model 180 can be implemented as, for example, a text-to-image diffusion model, a multi-modal language model, or another appropriate type of text-to-image model. The text-to-image model can generate an image based on the input 182 and provide the image as an output 184 to the prompt apparatus 208.
[0098] The digital component apparatus 212 is configured to generate digital components using images generated by the text-to-image model 180. For example, the digital component apparatus 212 can add text to the image, e.g., a text headline or description of an item that is the subject of the digital component. The digital component apparatus 212 can also generate a digital component creative file that includes a link to a landing page and data that enables a client device 106 to render the digital component. The service apparatus 110 can then distribute the digital component data 122 that includes and / or presents the created digital component in response to requests 122. as described above.
[0099] As noted above, the Al system 160 can adjust the one or more mutators that are used to generate the modified prompts of the training samples based on the quality of the images generated using the modified prompts relative to the quality of the images generated using the input prompts corresponding to the modified prompts. To do so, the training apparatus 206 can generate training samples using a combination of the mutators and a set of seed prompts. Each training sample can include a seed prompt and a modified prompt generated by the language model 220 using the seed prompt and the combination of mutators. Both prompts of each training sample can then be provided to the text-to-image model 180 to generate images for each prompt. The image evaluation apparatus 210 can evaluate the quality of each image generated using a modified prompt relative to the image generated using the seed prompt from which the modified prompt was generated. For example, the image evaluation apparatus 210 can generate an overall quality metric for each training sample, as described above.
[0100] The training apparatus 206 can select the mutators to use in generating training data for training the prompt enhancement model 170 based on the quality metrics of the training samples for each combination of mutators. For example, the training apparatus 206 can generate a set of training samples, as described above, using each combination of mutators. The training apparatus 206 can then select the combination resulting in the highest aggregate quality metric. The aggregate quality metric for a combination of mutators can be the sum, average, median, or other measure of central tendency of the overall quality scores for the training samples generated using the combination of mutators.
[0101] In some cases, the training apparatus 206 can train a prompt enhancement model 170 using each combination of mutators. In this example, the image evaluation apparatus 210 can evaluate the quality of images generated using the seed prompts and the images generated using the modified prompts output by the prompt enhancement models 170. For example, the aggregate quality metric for a combination of mutators can be based on the overall quality metric of each prompt pair that includes a seed prompt and a modified prompt generated using the trained prompt enhancement model 170 using the seed prompt. The training apparatus 206 can then keep the prompt enhancement model 170 having the best, e.g., highest, aggregate quality metric for use in generating modified prompts in response to requests 112.
[0102] In some implementations, the training apparatus 206 can be configured to adapt the combination of mutators and the resulting prompt enhancement models to different categories of images. For example, one combination of mutators may performbeter for training a prompt enhancement model 170 to generate modified prompts for images of cars while another combination of mutators may perform beter for training a prompt enhancement model 170 to generate modified prompts for images of clothing. The training process 206 can be configured to perform the process of identifying the best combination of mutators for each category of images and train a dedicated prompt enhancement model 170 for each category of image using the training samples generated using its combination or mutators. These dedicated prompt enhancement models 170 can then be used to generate modified prompts in response to requests 112 related to their categories.
[0103] The training apparatus 206 can also adjust the one or more mutators that are used to generate the modified prompts based on the length of the modified prompts. For example, excessively long prompts often add artifacts that result in hallucinations and a large number of concepts included in a prompt can cause an image generation model to generate inaccurate and / or otherwise lower quality images. To prevent such hallucinations, the training apparatus 206 can remove, from the combination of mutators, at least one mutator to reduce the length of the prompts output by the language model 220 and / or the modified prompts output by the prompt enhancement model 170. For example, the training apparatus 206 can be configured to monitor the length in words and / or characters of each prompt output by the language model 170 in generating training samples and / or the length in words and / or characters of each prompt output by the prompt enhancement model 180. If the length of the output prompts is too long, e.g., if the average length of the output prompts satisfies a threshold or if a threshold percentage of the output prompts exceeds a threshold, the training apparatus 206 can either use a different combination of mutators or remove a mutator and generate new training samples and / or retrain the prompt enhancement model 170.
[0104] In some implementations, removing a mutator can include selecting a mutator to remove in order of priority' or at random. In some implementations, the training apparatus 206 can evaluate the prompts output by the language model 220 and / or the prompts output by the prompt enhancement model 170 to determine which mutator is causing the excessive length of the prompts. For example, if there are a lot of added synonyms in output prompts, the training apparatus 206 can determine that the synonym mutator is the culprit and remove the synonym mutator from the combination of mutators used to generate the training samples.
[0105] FIG. 3 is a flow chart of an example process 300 of training a prompt enhancement model. Operations of the process 300 can be performed, for example, by the service apparatus 110 of FIG. 1 (e.g., by the Al system 160), or another data processing apparatus. The operations of the process 300 can also be implemented as instructions stored on a computer readable medium, which can be non-transitory. Execution of the instructions, by one or more data processing apparatus, causes the one or more data processing apparatus to perform operations of the process 300. For brevity, the process 300 is described in terms of being performed by a system.
[0106] The system obtains training data (310). As described above, the system can obtain training data by generating modified prompts based on input prompts, e.g., using a language model. For example, the system can use a language model to generate modified prompts using input prompts and a combination of mutators. Each training sample in the training data can include an input prompt and a modified prompt. As described above, the training samples can be filtered based on the quality of images generated using the input prompts and the modified prompts.
[0107] The system trains the prompt enhancement model using the training data (320). For example, the system can train a sequence-to-sequence model by fine tuning a language model using the training samples.
[0108] The sy stem improves the prompt enhancement model (330). For example, the system can use reinforcement learning to improve the trained prompt enhancement model. The reward that is optimized in the reinforcement learning process can be a composite metric for each of multiple training samples in a refinement set of training samples. As described above, the composite metric for each training sample can include an aesthetic metric for an image generated using the modified prompt of the training sample and an image-text alignment metric for the training sample.
[0109] FIG. 4 is a flow chart of an example process 400 of generating images using a prompt enhancement model and a text-to-image model. Operations of the process 400 can be performed, for example, by the service apparatus 110 of FIG. 1 (e.g.. by the Al system 160), or another data processing apparatus. The operations of the process 400 can also be implemented as instructions stored on a computer readable medium, which can be non- transitory. Execution of the instructions, by one or more data processing apparatus, causes the one or more data processing apparatus to perform operations of the process 400. For brevity, the process 400 is described in terms of being performed by a system.
[0110] The system receives an input prompt (410). The system can receive the input prompt from another device, e.g., the device of a digital component provider. For example, the digital component provider can provide an input prompt for use in generating one or more images for a digital component, e.g., a digital component for a product, service, or other item. The prompt can be in the form of natural language input, e.g., a phrase or sentence that includes instructions for a machine learning model to use to generate the image(s).
[0111] The system provides the input prompt as an input to a prompt enhancement model (420). As described above, a prompt enhancement model can be trained to modify input prompts to generate modified prompts that expand on or otherwise enhance the input prompt.
[0112] The system receives a modified prompt as an output of the prompt enhancement model (430). The prompt enhancement model can generate the modified prompt based on the input prompt. The modified prompt can be a natural language prompt that expands on the input prompt or that otherwise is a modified version of the input prompt.
[0113] The system provides the modified prompt to a text-to-image model (440). As described above, the text-to-image model can be trained to generate images based on prompts, e.g., natural language text prompts.
[0114] The system receives an image as an output of the text-to-image model (450). The system sends the image to one or more recipients (460). For example, the system can send the image to a device of a digital component provider that provided the input prompt. In some implementations, the system can generate a digital component using the image, e.g., by adding text to the image and / or by generating a digital component file or set of files that include a link to a landing page and data that enables a client device to render the digital component. The system can distribute the digital component to client devices, e.g., in response to component requests, as described above.
[0115] FIG. 5 is a block diagram of an example computer system 500 that can be used to perform operations described above. The system 500 includes a processor 510. a memory 520, a storage device 530, and an input / output device 540. Each of the components 510, 520, 530, and 540 can be interconnected, for example, using a system bus 550. The processor 510 is capable of processing instructions for execution within the system 500. In one implementation, the processor 510 is a single-threaded processor. In anotherimplementation, the processor 510 is a multi -threaded processor. The processor 510 is capable of processing instructions stored in the memory 520 or on the storage device 530.
[0116] The memory 520 stores information within the system 500. In one implementation, the memory 520 is a computer-readable medium. In one implementation, the memory 520 is a volatile memory unit. In another implementation, the memory 520 is a non-volatile memory unit.
[0117] The storage device 530 is capable of providing mass storage for the system 500. In one implementation, the storage device 530 is a computer-readable medium. In various different implementations, the storage device 530 can include, for example, a hard disk device, an optical disk device, a storage device that is shared over a network by multiple computing devices (e.g., a cloud storage device), or some other large capacity storage device.
[0118] The input / output device 540 provides input / output operations for the system 500. In one implementation, the input / output device 540 can include one or more of a network interface devices, e.g.. an Ethernet card, a serial communication device, e.g.. and RS-232 port, and / or a wireless interface device, e.g.. and 802. 11 card. In another implementation, the input / output device can include driver devices configured to receive input data and send output data to other devices, e.g., keyboard, printer, display, and other peripheral devices 560. Other implementations, however, can also be used, such as mobile computing devices, mobile communication devices, set-top box television client devices, etc.
[0119] Although an example processing system has been described in FIG. 5, implementations of the subject matter and the functional operations described in this specification can be implemented in other types of digital electronic circuitry, or in computer software, firmware, or hardware, including the structures disclosed in this specification and their structural equivalents, or in combinations of one or more of them.
[0120] An electronic document (which for brevity will simply be referred to as a document) does not necessarily correspond to a file. A document may be stored in a portion of a file that holds other documents, in a single file dedicated to the document in question, or in multiple coordinated files.
[0121] For situations in which the systems discussed here collect and / or use personal information about users, the users may be provided with an opportunity to enable / disable or control programs or features that may collect and / or use personal information (e.g., information about a user’s social network, social actions or activities, auser’s preferences, or a user’s current location). In addition, certain data may be treated in one or more ways before it is stored or used, so that personally identifiable information associated with the user is removed. For example, a user’s identity may be anonymized so that the no personally identifiable information can be determined for the user, or a user’s geographic location may be generalized where location information is obtained (such as to a city, ZIP code, or state level), so that a particular location of a user cannot be determined.
[0122] Embodiments of the subject matter and the operations described in this specification can be implemented in digital electronic circuitry, or in computer software, firmware, or hardware, including the structures disclosed in this specification and their structural equivalents, or in combinations of one or more of them. Embodiments of the subject matter described in this specification can be implemented as one or more computer programs, i.e., one or more modules of computer program instructions, encoded on computer storage medium for execution by, or to control the operation of, data processing apparatus. Alternatively, or in addition, the program instructions can be encoded on an artificially-generated propagated signal, e.g.. a machine-generated electrical, optical, or electromagnetic signal, that is generated to encode information for transmission to suitable receiver apparatus for execution by a data processing apparatus. A computer storage medium can be, or be included in, a computer-readable storage device, a computer-readable storage substrate, a random or serial access memory array or device, or a combination of one or more of them. Moreover, while a computer storage medium is not a propagated signal, a computer storage medium can be a source or destination of computer program instructions encoded in an artificially -generated propagated signal. The computer storage medium can also be, or be included in, one or more separate physical components or media (e.g., multiple CDs. disks, or other storage devices).
[0123] The operations described in this specification can be implemented as operations performed by a data processing apparatus on data stored on one or more computer-readable storage devices or received from other sources.
[0124] The term "‘data processing apparatus” encompasses all kinds of apparatus, devices, and machines for processing data, including by way of example a programmable processor, a computer, a system on a chip, or multiple ones, or combinations, of the foregoing. The apparatus can include special purpose logic circuitry, e.g., an FPGA (field programmable gate array) or an ASIC (application-specific integrated circuit). The apparatus can also include, in addition to hardware, code that creates an execution environment for the computer program in question, e.g., code that constitutes processorfirmware, a protocol stack, a database management system, an operating system, a crossplatform runtime environment, a virtual machine, or a combination of one or more of them. The apparatus and execution environment can realize various different computing model infrastructures, such as web services, distributed computing and grid computing infrastructures.
[0125] This document refers to a service apparatus. As used herein, a service apparatus is one or more data processing apparatus that perform operations to facilitate the distribution of content over a network. The service apparatus is depicted as a single block in block diagrams. However, while the service apparatus could be a single device or single set of devices, this disclosure contemplates that the service apparatus could also be a group of devices, or even multiple different systems that communicate in order to provide various content to client devices. For example, the service apparatus could encompass one or more of a search system, a video streaming service, an audio streaming service, an email service, a navigation service, an advertising service, a gaming service, or any other service.
[0126] A computer program (also known as a program, software, software application, script, or code) can be written in any form of programming language, including compiled or interpreted languages, declarative or procedural languages, and it can be deployed in any form, including as a stand-alone program or as a module, component, subroutine, object, or other unit suitable for use in a computing environment. A computer program may, but need not, correspond to a file in a file system. A program can be stored in a portion of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document), in a single file dedicated to the program in question, or in multiple coordinated files (e.g.. files that store one or more modules, sub-programs, or portions of code). A computer program can be deployed to be executed on one computer or on multiple computers that are located at one site or distributed across multiple sites and interconnected by a communication network.
[0127] The processes and logic flows described in this specification can be performed by one or more programmable processors executing one or more computer programs to perform actions by operating on input data and generating output. The processes and logic flows can also be performed by, and apparatus can also be implemented as, special purpose logic circuitry', e.g., an FPGA (field programmable gate array) or an ASIC (application-specific integrated circuit).
[0128] Processors suitable for the execution of a computer program include, by way of example, both general and special purpose microprocessors, and any one or moreprocessors of any kind of digital computer. Generally, a processor will receive instructions and data from a read-only memory or a random access memory or both. The essential elements of a computer are a processor for performing actions in accordance with instructions and one or more memory devices for storing instructions and data. Generally, a computer will also include, or be operatively coupled to receive data from or transfer data to, or both, one or more mass storage devices for storing data, e.g., magnetic, magnetooptical disks, or optical disks. However, a computer need not have such devices.Moreover, a computer can be embedded in another device, e.g., a mobile telephone, a personal digital assistant (PDA), a mobile audio or video player, a game console, a Global Positioning System (GPS) receiver, or a portable storage device (e.g., a universal serial bus (USB) flash drive), to name just a few. Devices suitable for storing computer program instructions and data include all forms of non-volatile memory, media and memory devices, including by way of example semiconductor memory devices, e.g., EPROM, EEPROM, and flash memory devices; magnetic disks, e.g., internal hard disks or removable disks; magneto-optical disks; and CD-ROM and DVD-ROM disks. The processor and the memory can be supplemented by, or incorporated in, special purpose logic circuitry.
[0129] To provide for interaction with a user, embodiments of the subject matter described in this specification can be implemented on a computer having a display device, e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor, for displaying information to the user and a keyboard and a pointing device, e.g., a mouse or a trackball, by which the user can provide input to the computer. Other kinds of devices can be used to provide for interaction with a user as well; for example, feedback provided to the user can be any form of sensory feedback, e.g., visual feedback, auditory feedback, or tactile feedback; and input from the user can be received in any form, including acoustic, speech, or tactile input. In addition, a computer can interact with a user by sending documents to and receiving documents from a device that is used by the user; for example, by sending web pages to a web browser on a user’s client device in response to requests received from the web browser.
[0130] Embodiments of the subject matter described in this specification can be implemented in a computing system that includes a back-end component, e.g., as a data server, or that includes a middleware component, e.g., an application server, or that includes a front-end component, e g., a client computer having a graphical user interface or a Web browser through which a user can interact with an implementation of the subject matter described in this specification, or any combination of one or more such back-end,middleware, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication, e.g., a communication network. Examples of communication networks include a local area network (“LAN’’) and a wide area network (“WAN”), an inter-network (e.g., the Internet), and peer-to-peer networks (e.g., ad hoc peer-to-peer networks).
[0131] The computing system can include clients and servers. A client and serv er are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other. In some embodiments, a server transmits data (e.g., an HTML page) to a client device (e.g., for purposes of displaying data to and receiving user input from a user interacting with the client device). Data generated at the client device (e.g., a result of the user interaction) can be received from the client device at the server.
[0132] While this specification contains many specific implementation details, these should not be construed as limitations on the scope of any inventions or of what may be claimed, but rather as descriptions of features specific to particular embodiments of particular inventions. Certain features that are described in this specification in the context of separate embodiments can also be implemented in combination in a single embodiment. Conversely, various features that are described in the context of a single embodiment can also be implemented in multiple embodiments separately or in any suitable subcombination. Moreover, although features may be described above as acting in certain combinations and even initially claimed as such, one or more features from a claimed combination can in some cases be excised from the combination, and the claimed combination may be directed to a subcombination or variation of a subcombination.
[0133] Similarly, while operations are depicted in the drawings in a particular order, this should not be understood as requiring that such operations be performed in the particular order shown or in sequential order, or that all illustrated operations be performed, to achieve desirable results. In certain circumstances, multitasking and parallel processing may be advantageous. Moreover, the separation of various system components in the embodiments described above should not be understood as requiring such separation in all embodiments, and it should be understood that the described program components and systems can generally be integrated together in a single software product or packaged into multiple software products.
[0134] Thus, particular embodiments of the subject matter have been described.Other embodiments are within the scope of the following claims. In some cases, the actions recited in the claims can be performed in a different order and still achieve desirable results. In addition, the processes depicted in the accompanying figures do not necessarily require the particular order shown, or sequential order, to achieve desirable results. In certain implementations, multitasking and parallel processing may be advantageous.
[0135] What is claimed is:
Claims
CLAIMS1. A method, comprising: obtaining initial training data comprising training samples, wherein each training sample includes a first input prompt and a first modified prompt generated based on the first input prompt; training, using the training samples, a prompt enhancement model to output modified prompts based on input prompts; and updating the prompt enhancement model using reinforcement learning and refinement training data comprising refinement training samples that each include a second input prompt, a second modified prompt generated by the prompt enhancement model using the second input prompt, an image generated using the second input prompt, and an image generated using the second modified prompt, including optimizing a reward that is based on an aesthetic metric for each refinement training sample and an image-text alignment score for each training sample.
2. The method of claim 1. wherein obtaining the initial training data comprises generating each training sample using a language model.
3. The method of claim 2, wherein generating each training sample comprises providing the first input prompt of the training sample as an input to the language model and receiving the modified prompt of the training example as an output of the language model.
4. The method of claim 2, wherein generating each training sample comprises: adding instructions of one or more mutators to a prompt that includes the input prompt; providing the prompt as an input to the language model; and receiving the modified prompt as an output of the language model.
5. The method of claim 2. wherein the one or more mutators comprise (i) an object mutator, (ii) an adjective and adverb mutator, (iii) a style mutator, (iv) a synonym mutator,(v) an imaging parameter mutator, (vi) a text sanitation mutator, or any combination of (i) to(vi).
6. The method of claim 4 or 5, wherein obtaining the initial training data comprises: for each training sample, generating a first image using the first input prompt and a second image using the first modified prompt; evaluating the first image and the second image; and determining a quality metric that represents a level of improvement in quality of the second image relative to the first image based on the evaluation; and filtering one or more training samples from the initial training data based on the quality metric for each training sample.
7. The method of claim 6, wherein the quality’ metric for each training sample is based on an inclusive metric that indicates whether the first modified prompt of the training sample includes all information from the first input prompt of the training sample.
8. The method of claim 6 or 7, further comprising selecting a combination of the mutators for using in generating the training samples based on an aggregate quality metric for each of multiple combinations of mutators.
9. A system comprising: one or more processors; and one or more storage devices storing instructions that, when executed by the one or more processors, cause the one or more processors to carry' out the method of any preceding claim.
10. A computer readable storage medium carry ing instructions that, when executed by one or more processors, cause the one or more processors to cany out the method of any one of claims 1 to 8.
11. A computer program product comprising instructions which, when executed by one or more computers, cause the one or more computers to carry out the steps of the method of any of claims 1 to 8.