Image generation using enhanced prompts for artificial intelligence models

EP4713753A2Pending Publication Date: 2026-03-25GOOGLE LLC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
EP · EP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-06-17
Publication Date
2026-03-25

AI Technical Summary

Technical Problem

Current generative artificial intelligence models often produce inaccurate and low-quality digital images due to the direct use of unstructured initial data, leading to computational inefficiencies and hallucinations.

Method used

A chain of operations is employed to generate enhanced prompts using structured descriptions and content outputs, distilling initial data into structured form before generating digital components, thereby improving accuracy and reducing computational resources.

Benefits of technology

This approach results in more accurate and higher-quality digital components by optimizing the prompt generation process, reducing computational burden, and preventing hallucinations, while enhancing the relevance of the generated content to the digital component's context.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2024034315_26122025_PF_FP_ABST
    Figure US2024034315_26122025_PF_FP_ABST
Patent Text Reader

Abstract

Methods, systems, and apparatus, including computer programs encoded on a computer storage medium, for generating enhanced prompts and digital components using the enhanced prompts. In one aspect, a method includes generating, by an AI system, a first prompt that includes a first request for one or more models to generate a structured description for a subject of a new digital component to be generated and initial data for use by the one or more models to generate the structured description. The system generates a second prompt that includes a second request for the one or more models to generate a content output. The system generates a third prompt that includes a third request for the one or more models to generate the new digital component based on the content output, and obtains the digital component output by the one or more models.
Need to check novelty before this filing date? Find Prior Art

Description

IMAGE GENERATION USING ENHANCED PROMPTS FOR ARTIFICIALINTELLIGENCE MODELSBACKGROUND

[0001] This specification relates to data processing, artificial intelligence, and generating digital content.

[0002] In a computer networked environment such as the Internet, third-party content providers provide third-party content items for display on end-user computing devices. These third-party content items, for example, digital images and video, can be displayed on client devices in the environment. Digital images and video can be used, for example, on the Internet, for remote meetings via video conferencing, high-definition video entertainment, and / or sharing of user-generated content.

[0003] Recent developments in artificial intelligence and, in particular, generative artificial intelligence have caused user-produced visual content such as digital images to become ubiquitous. For example, various types of images can be generated by using text-to-image models based on text prompts. However, current generative artificial intelligence models often produce inaccurate and / or low quality images.SUMMARY

[0004] In general, a first innovative aspect of the subject matter described in this specification can be embodied in methods that include the actions of generating, by an artificial intelligence system, a first prompt that includes (i) a first request for one or more models to generate a structured description for a subject of a new digital component to be generated by the one or more models and (ii) initial data for use by the one or more models to generate the structured description; obtaining, by the artificial intelligence system and from the one or more models, the structured description as generated based on the first prompt; generating, by the artificial intelligence system, a second prompt that includes a second request for the one or more models to generate a content output that includes at least one of (i) a content item, (ii) a description of a content item, or (iii) one or more characteristics of the new digital component, wherein the second prompt includes the structured description, the initial data, or both; obtaining, by the artificial intelligence system and from the one or more models, the content output as generated based on the second prompt; generating, by the artificial intelligence system, a third promptthat includes a third request for the one or more models to generate the new digital component based on the content output, wherein the third prompt includes the content output, the structured description, and the initial data; and obtaining, by the artificial intelligence system and from the one or more models, the new digital component generated based on the third prompt. Other implementations of this aspect include corresponding apparatus, systems, and computer programs, configured to perform the aspects of the methods, encoded on computer storage devices.

[0005] These and other embodiments can each optionally include one or more of the following features. Some aspects can include distributing the new digital component to one or more client devices.

[0006] In some aspects, the initial data includes at least one of: a name of the subject of the new digital component, a text description of the subject of the new digital component, or other digital components for the subject of the new digital component.

[0007] In some aspects, the initial data includes a specified source of online content, wherein the structured description generated based on the first prompt is generated by the one or more models based on at least one piece of content obtained from the specified source of online content.

[0008] In some aspects, the second request of the second prompt is indicative of an image as the content item, characteristics of text content to be used when generating the new digital component, or both.

[0009] In some aspects, obtaining from the one or more models includes obtaining the image, and wherein the second request in the second prompt is indicative of a type of the image to be generated.

[0010] In some aspects, the content item is a text content item, wherein the second request of the second prompt is indicative of a visual characteristic of the text content item, wherein the visual characteristic includes at least one of a dimensional or text effect characteristic.

[0011] In some aspects, the one or more characteristics of the new digital component comprises visual characteristics of content depicted by the new digital component.

[0012] In some aspects, obtaining the new digital component includes providing a request for at least one digital component by invoking the one or more models to provide clauses and images based on the initial data, the structured description, and the content output andgenerating, by the one or more models, a new clause that indicates visual characteristics generated based on the content output.

[0013] In some aspects, the one or more models include a single large language model.

[0014] In some aspects, the one or more models include multiple machine learning models.

[0015] In some aspects, the third prompt includes at least a portion of the initial data.

[0016] Similar operations and processes may be performed in a system comprising at least one process and a memory communicatively coupled to the at least one processor where the memory stores instructions, that when executed cause the at least one processor to perform the operations. Further, a computer-readable medium, which can be non-transitory, storing instructions which, when executed, cause at least one processor to perform the operations may also be contemplated. In other words, while generally described as computer implemented software embodied on tangible, non-transitory media that processes and transforms the respective data, some or all of the aspects may be computer implemented methods or further included in respective systems or other devices for performing this described functionality. The details of these and other aspects and embodiments of the present disclosure are set forth in the accompanying drawings and the description below. Other features, objects, and advantages of the disclosure will be apparent from the description and drawings, and from the claims.

[0017] Particular embodiments of the subject matter described in this specification can be implemented so as to realize one or more of the following advantages. The techniques described in this specification enable artificial intelligence (Al) to be used to generate customized content, e.g., customized digital components, based on data related to the digital components and / or one or more queries received from a client device. Digital components can be created by Al models based on enhanced prompts (e.g., new or regenerated prompts) that support improved relevance of the content of the generated digital component to the context for the digital component. The techniques described in this specification support generating enhanced prompts through a chain of operations to provide an optimized, or at least improved, prompt that can lead to more relevant and higher quality (e.g., more accurate) digital components than digital components that are directly generated from initial input without enhancing the prompt. For example, by generating an understanding of initial data for the digital component (e.g., in the form of a structured description), generating various ideas forthe digital component (e.g., content for the digital component and / or visual characteristics of the digital component), and then generating the prompt for creating the digital component based on the understanding and the ideas ensures that the digital component accurately conveys a message and / or accurately reflects a context defined by the initial data. This also prevents hallucinations of Al models that would otherwise occur if a large amount of initial data was provided as input to the model without distilling the initial data into an understanding and generating the ideas and digital component based on the understanding.

[0018] Further, the digital component generation can be performed faster and with fewer resources, e.g., computational resources, for the generation of the digital components and / or for the interaction with users during the generation. The enhancement of the prompt can streamline the digital component generation and reduce post-processing of generated output or regeneration of digital components to provide final digital components as a result. The chain of operations that are defined to obtain data, e.g., structured data generated based on initial data and / or a description of content items or images, that can be used to enhance the prompt to for a generative Al model to improve the relevance of the generated digital component to the objective of the generation process.

[0019] In particular, the chain of operations can include a stage that uses Al to distill a large amount of information (e.g., multi-modal information that includes images, text, and / or other forms of content) and, based on this information, output an understanding of this large amount of information, e g., in the form of a structured description. This understanding can then be used in later stages to formulate one or more prompts for generating further content or digital components that convey that understanding using one or more types of content, e.g., text, images, audio, video, etc. By obtaining the understanding before assembling prompts for content generation, the accuracy of the generated content is improved and the computational burden placed on the Al system that generates the content is reduced relative to techniques that would provide all of the information (e.g., without modifying the initial data used for the prompt) directly to this Al model. For example, providing too much information to an Al model often results in hallucinations and other errors in content generation. By first generating an understanding of a large amount of information (e.g., initial data that can be unstructured data), and using the understanding to generate assets as options or alternatives (e.g., design ideation) that can be used to further enhance a prompt so that the enhanced prompt can bedefined as an enhanced prompt that when used can result in improved content generation prompts that result in more accurate and higher quality content. The operations to process initial data (e.g., creative data) to generate further content (e.g., text or images) to improve the understanding and to generate content to be used for defining a prompt for a particular objective of a digital component result in the consumption of fewer resources compared to the resources that may be needed if the initial content is directly used by an Al model for generating a digital component that is relative to the particular objective. In some cases, if the initial data is used directly, the processing load on the Al model may be higher than if the initial data is distilled into a structured description. For example, if the initial content is used directly for generating digital components, multiple digital components may need to be generated, e.g., in an iterative process, until a final result is selected that is of substantially similar relevancy to the objective of the digital component generation.

[0020] Using the described techniques in this document supports efficient utilization of computer resources based on the ability to obtain input, and create prompts in iterations using one or more trained models, e.g., one or more trained language models, so that prompts are redefined or otherwise created and used to generate digital components that can be output as a response. Such processes can reduce the network calls between the system and a user device if the processes involve interaction with a user to arrive at the generation of the content. Further, the Al system can generate content based on an understanding an item that is the subject of a digital component in a context that is relevant to the user and / or a digital component’s target context.

[0021] The details of one or more embodiments of the subject matter described in this specification are set forth in the accompanying drawings and the description below. Other features, aspects, and advantages of the subject matter will become apparent from the description, the drawings, and the claims.BRIEF DESCRIPTION OF THE DRAWINGS

[0022] FIG. 1 is a block diagram of an example environment in which generative artificial intelligence can be implemented to generate digital components.

[0023] FIG. 2 is a block diagram illustrating interactions between an artificial intelligence system, Al models, and a client device.

[0024] FIG. 3 A is a block diagram of an example process for generating enhanced prompts for use in generating digital components using a trained Al model.

[0025] FIG. 3B is a block diagram of an example process for generating digital components based on an enhanced prompt including generatively designed images.

[0026] FIG. 4A is a block diagram of an example process for generating enhanced prompts and using the enhanced prompts to generate digital components.

[0027] FIG. 4B is a block diagram of an example process for generating a prompt with obtained descriptions of images from a trained Al model to generate an enhanced prompt to be used for generating digital components.

[0028] FIG. 5 a block diagram of an example computer.

[0029] Like reference numbers and designations in the various drawings indicate like elements.DETAILED DESCRIPTION

[0030] This specification describes techniques for enabling artificial intelligence to generate new digital components in faster and more efficient ways using enhanced prompts so that they are adapted to, e.g., optimized to, a target context. Artificial intelligence (Al) is a segment of computer science that focuses on the creation of models that can perform tasks act autonomously, e.g., with little to no human intervention. Al systems can utilize, for example, one or more of machine learning, natural language processing, or computer vision. Machine learning, and its subsets, such as deep learning, focus on developing models that can infer outputs from data. The outputs can include, for example, predictions and / or classifications. Natural language processing focuses on analyzing and generating human language. Computer vision focuses on analyzing and interpreting images and videos. Artificial intelligence systems can include generative models that generate new content, such as images, videos, text, audio, and / or other content, in response to input prompts and / or based on other information.

[0031] The techniques described throughout this specification enable Al to generate digital components using various combinations of text, images, and / or other types of content (e.g., emojis, graphics, etc.) by generating and providing enhanced prompts as inputs to one or more Al models, e.g., one or more language models, diffusion models, and / or other types of Al models. The techniques can include the generation or regeneration of prompts over multiple stages to include relevant data to be used for the digital component generation by a trained Almodel to create more accurate and relevant digital components. For example, an Al system can obtain information from various sources, such as online sources (e.g., web pages or other trusted online sources) and / or other data sources (e.g., at a local storage), and use this information in different ways to generate an enhanced prompt to be used to create different digital components. The digital components can be distributed to client devices of users, e.g., for presentation with other content to enhance the users’ online experience.

[0032] Generally speaking, the Al system can generate an initial prompt (which can be referred to as a first prompt) based on initial data. The initial data can be unstructured data, e g., data that is not structured to include data attribute values associated with a defined structured model for generating the digital component. The initial data can include data related to a subject of the digital component that is to be generated using the Al model. For example, the initial data can include the name of an entity (e g., the name of a business) that is the subject of the digital component or that is related to the subject (e.g., the subject can be a product and the entity can be the business that offers the product), other digital components for the subject of the digital component to be created, headlines or other text of these other digital components, text descriptions of the subject or entity, and / or other data related to the subject or entity.

[0033] The initial prompt can be provided as input to a language model, e.g., a large language model (LLM), or other type of Al model that is trained to output a structured description of the subject based on the initial prompt. The structured description can be used to generate a second prompt that can include, for example, a request to generate a content output based on the structured description. The content output can include content or data for use in generating the digital component. For example, the content output can include a content item (e.g., an image, emoji, text, graphics, etc.), a description of a content item, or one or more characteristics of the new digital component. The one or more characteristics can indicate visual characteristics of content of the digital component, e.g., colors, style features, text effects, fonts, etc. In some implementations, multiple second prompts can be generated and used by the Al model(s) to generate different content outputs for use in generating a new digital component.

[0034] A third prompt can be generated based on the content output(s) and can include a request for an Al model to generate a new digital component based on the content output. The third prompt can include the structured description of the subject and / or the initial data or a portion of the initial data. This enhanced prompt for generating the digital component isspecialized (e.g., created or augmented) to improve the overall quality of the content generated as part of the digital component that can include customized images generated by an image model (e.g., a language model or diffusion model) and the structured description and / or initial data. Post-processing operations can then be used to detect errors associated with generating the new digital component, and the new digital component can be output to a client device (e.g., user computer, mobile device, tablet device, audio device, gaming device, etc.).

[0035] Using the enhanced prompt that includes structured description and content output(s) for the digital component can reduce wasted computing resources that would otherwise generate more low quality digital components if a more general prompt was used or if multiple prompts related to different concepts (e.g., some for query data and others for digital component data) were used.

[0036] Similarly, using the enhanced prompt for creating a new digital component based on query data (a request), as well as a structured one or more content outputs for the digital component can save computing resources and result in generating an output faster by using the enhanced prompt to constrain the parameters used by the language model to generate the new digital component. This prompt for generating a digital component can be enhanced by regenerating an initial prompt that includes initial data, which is used by a language model (or other Al model) to understand the goal and content to be used for the digital component and to generate the new prompt that is more accurate and corresponds to the digital components and the goal for the digital component generation. For example, the data included in the new prompt can be expanded to include not only initial data related to the digital component but also additional structured data that reflects underlying concepts of the digital component that is learned from the initial data, as well as the design of the digital component (e.g., included image data, visualization of content, etc ). This reduces the amount of processing performed using the language model and results in higher quality prompts that result in higher quality content (including text and / or images) for the customized digital components. Using a single prompt that is based on both initial and structured data for use when generating the digital component enables the use of a single high quality prompt that reduces computational resources that would be required for multiple prompts to regenerate an already generated digital component and enables the language model to better understand the concepts of boththe request and the digital component data (e.g., structured and initial data, image content, characteristics of the digital component, etc.).

[0037] As used throughout this document, the phrase “digital component” refers to a discrete unit of digital content or digital information (e.g., a video clip, audio clip, multimedia clip, gaming content, image, text, bullet point, artificial intelligence output, language model output, or another unit of content). A digital component can electronically be stored in a physical memory device as a single file or in a collection of files, and digital components can take the form of video files, audio files, multimedia files, image files, or text files and include advertising information, such that an advertisement is a type of digital component.

[0038] FIG. 1 is a block diagram of an example environment 100 in which generative Al can be implemented to generate digital components. The example environment 100 includes a network 102, such as a local area network (LAN), a wide area network (WAN), the Internet, or a combination thereof. The network 102 connects electronic document servers 104, user devices 106, digital component servers 108, and a service apparatus 110. The example environment 100 may include many different electronic document servers 104, user devices 106, and digital component servers 108.

[0039] A client device 106 is an electronic device capable of requesting and receiving online resources over the network 102. Example client devices 106 include personal computers, gaming devices, mobile communication devices, digital assistant devices, augmented reality devices, virtual reality devices, and other devices that can send and receive data over the network 102. A client device 106 typically includes a user application, such as a web browser, to facilitate the sending and receiving of data over the network 102, but native applications (other than browsers) executed by the client device 106 can also facilitate the sending and receiving of data over the network 102.

[0040] A gaming device is a device that enables a user to engage in gaming applications, for example, in which the user has control over one or more characters, avatars, or other rendered content presented in the gaming application. A gaming device typically includes a computer processor, a memory device, and a controller interface (either physical or visually rendered) that enables user control over content rendered by the gaming application. The gaming device can store and execute the gaming application locally or execute a gaming application that is at least partly stored and / or provided by a cloud server (e.g., online gaming applications).Similarly, the gaming device can interface with a gaming server that executes the gaming application and “streams” the gaming application to the gaming device. The gaming device may be a tablet device, mobile telecommunications device, a computer, or another device that performs other functions beyond executing the gaming application.

[0041] Digital assistant devices include devices that include a microphone and a speaker. Digital assistant devices are generally capable of receiving input by way of voice, and respond with content using audible feedback, and can present other audible information. In some situations, digital assistant devices also include a visual display or are in communication with a visual display (e.g., by way of a wireless or wired connection). Feedback or other information can also be provided visually when a visual display is present. In some situations, digital assistant devices can also control other devices, such as lights, locks, cameras, climate control devices, alarm systems, and other devices that are registered with the digital assistant device.

[0042] As illustrated, the client device 106 is presenting an electronic document 150. An electronic document is data that presents a set of content on a client device 106. Examples of electronic documents include online resources, webpages, word processing documents, portable document format (PDF) documents, images, videos, search results pages, and feed sources. Native applications (e.g., “apps” and / or gaming applications), such as applications installed on mobile, tablet, or desktop computing devices are also examples of electronic documents. Electronic documents can be provided to client devices 106 by electronic document servers 104 (“Electronic Doc Servers”).

[0043] For example, the electronic document servers 104 can include servers that host publisher websites. In this example, the client device 106 can initiate a request for a given publisher webpage, and the electronic server 104 that hosts the given publisher webpage can respond to the request by sending machine executable instructions that initiate the presentation of the given webpage at the client device 106.

[0044] In another example, the electronic document servers 104 can include app servers from which client devices 106 can download apps. In this example, the client device 106 can download files required to install an app at the client device 106, and then execute the downloaded app locally (i.e., on the client device). Alternatively, or additionally, the client device 106 can initiate a request to execute the app, which is transmitted to a cloud server. In response to receiving the request, the cloud server can execute the application and stream auser interface of the application to the client device 106 so that the client device 106 does not have to execute the app itself. Rather, the client device 106 can present the user interface generated by the cloud server’s execution of the app and communicate any user interactions with the user interface back to the cloud server for processing.

[0045] Electronic documents can include a variety of content. For example, an electronic document 150 can include native content 152 that is within the electronic document 150 itself and / or does not change over time. Electronic documents can also include dynamic content that may change over time or on a per-request basis. For example, a publisher of a given electronic document (e.g., electronic document 150) can maintain a data source that is used to populate portions of the electronic document. In this example, the given electronic document can include a script, such as the script 154, that causes the client device 106 to request content (e.g., a digital component) from the data source when the given electronic document is processed (e.g., rendered or executed) by a client device 106 (or a cloud server). The client device 106 (or cloud server) integrates the content (e.g., digital component) obtained from the data source into the given electronic document to create a composite electronic document including the content obtained from the data source.

[0046] In some situations, a given electronic document (e.g., electronic document 150) can include a digital component script (e.g., script 154) that references the service apparatus 110, or a particular service provided by the service apparatus 110. In these situations, the digital component script is executed by the client device 106 when the given electronic document is processed by the client device 106. Execution of the digital component script configures the client device 106 to generate a request for digital components 112 (referred to as a “component request”), which is transmitted over the network 102 to the service apparatus 110. For example, the digital component script can enable the client device 106 to generate a packetized data request including a header and payload data. The component request 112 can include event data specifying features such as the name (or network location) of a server from which the digital component is being requested, a name (or network location) of the requesting device (e.g., the client device 106), and / or information that the service apparatus 110 can use to select one or more digital components, or other content, provided in response to the request. The component request 112 is transmitted, by the client device 106, over the network 102 (e.g., a telecommunications network) to a server of the service apparatus 110.

[0047] The component request 112 can include event data specifying other event features, such as the electronic document being requested and characteristics of locations of the electronic document at which digital components can be presented. For example, event data specifying a reference (e.g., URL) to an electronic document (e.g., webpage) in which the digital component will be presented, available locations of the electronic documents that are available to present digital components, sizes of the available locations, and / or media types that are eligible for presentation in the locations can be provided to the service apparatus 110. Similarly, event data specifying keywords associated with the electronic document (“document keywords”) or entities (e.g., people, places, or things) that are referenced by the electronic document can also be included in the component request 112 (e.g., as payload data) and provide to the service apparatus 110 to facilitate identification of digital components that are eligible for presentation with the electronic document. The event data can also include a query that was submitted from the client device 106 to obtain a search results page.

[0048] Component requests 112 can also include event data related to other information, such as information that a user of the client device has provided, geographic information indicating a state or region from which the component request was submitted, or other information that provides context for the environment in which the digital component will be displayed (e.g., a time of day of the component request, a day of the week of the component request, a type of device at which the digital component will be displayed, such as a mobile device or tablet device). Component requests 112 can be transmitted, for example, over a packetized network, and the component requests 112 themselves can be formatted as packetized data having a header and payload data. The header can specify a destination of the packet and the payload data can include any of the information described above.

[0049] The service apparatus 110 chooses digital components (e.g., third-party content, such as video files, audio files, images, text, gaming content, augmented reality content, and combinations thereof, which can all take the form of advertising content or non-advertising content) that will be presented with the given electronic document (e.g., at a location specified by the script 154) in response to receiving the component request 112 and / or using information included in the component request 112.

[0050] In some implementations, a digital component is selected in less than a second to avoid errors that could be caused by delayed selection of the digital component. For example, delaysin providing digital components in response to a component request 112 can result in page load errors at the client device 106 or cause portions of the electronic document to remain unpopulated even after other portions of the electronic document are presented at the client device 106.

[0051] Also, as the delay in providing the digital component to the client device 106 increases, it is more likely that the electronic document will no longer be presented at the client device 106 when the digital component is delivered to the client device 106, thereby negatively impacting a user's experience with the electronic document. Further, delays in providing the digital component can result in a failed delivery of the digital component, such as, if the electronic document is no longer presented at the client device 106 when the digital component is provided.

[0052] In some implementations, the service apparatus 110 is implemented in a distributed computing system that includes, for example, a server and a set of multiple computing devices 114 that are interconnected and identify and distribute digital components in response to requests 112. The set of multiple computing devices 114 operate together to identify a set of digital components that are eligible to be presented in the electronic document from among a corpus of millions of available digital components (DCi.x). The millions of available digital components can be indexed, for example, in a digital component database 116. Each digital component index entry can reference the corresponding digital component and / or include distribution parameters (DPi-DPx) that contribute to (e.g., trigger, condition, or limit) the distribution / transmission of the corresponding digital component. For example, the distribution parameters can contribute to (e.g., trigger) the transmission of a digital component by requiring that a component request include at least one criterion that matches (e.g., either exactly or with some pre-specified level of similarity) one of the distribution parameters of the digital component.

[0053] In some implementations, the distribution parameters for a particular digital component can include distribution keywords that must be matched (e.g., by electronic documents, document keywords, or terms specified in the component request 112) in order for the digital component to be eligible for presentation. Additionally, or alternatively, the distribution parameters can include embeddings that can use various different dimensions of data, such as website details and / or consumption details (e.g., page viewport, user scrolling speed, or otherinformation about the consumption of data). The distribution parameters can also require that the component request 112 include information specifying a particular geographic region (e.g., country or state) and / or information specifying that the component request 112 originated at a particular type of client device (e.g., mobile device or tablet device) in order for the digital component to be eligible for presentation. The distribution parameters can also specify an eligibility value (e.g., ranking measure, or some other specified value) that is used for evaluating the eligibility of the digital component for distribution / transmission (e.g., among other available digital components).

[0054] The identification of the eligible digital component can be segmented into multiple tasks 117a- 117c that are then assigned among computing devices within the set of multiple computing devices 114. For example, different computing devices in the set 114 can each analyze a different portion of the digital component database 116 to identify various digital components having distribution parameters that match the information included in the component request 112. In some implementations, each given computing device in the set 114 can analyze a different data dimension (or set of dimensions) and pass (e.g., transmit) results (Res 1-Res 3) 118a-l 18c of the analysis back to the service apparatus 110. For example, the results 118a-l 18c provided by each of the computing devices in the set 114, may identify a subset of digital components that are eligible for distribution in response to the component request and / or a subset of the digital components that have certain distribution parameters. The identification of the subset of digital components can include, for example, comparing the event data to the distribution parameters, and identifying the subset of digital components having distribution parameters that match at least some features of the event data.

[0055] The service apparatus 110 aggregates the results 118a-118c received from the set of multiple computing devices 114, and uses information associated with the aggregated results to select one or more digital components that will be provided in response to the request 112. For example, the service apparatus 110 can select a set of winning digital components (one or more digital components) based on the outcome of one or more content evaluation processes, as described below. In turn, the service apparatus 110 can generate and transmit, over the network 102, reply data 120 (e.g., digital data representing a reply) that enable the client device 106 to integrate the set of winning digital components into the given electronic document, suchthat the set of winning digital components (e.g., winning third-party content) and the content of the electronic document are presented together at a display of the client device 106.

[0056] In some implementations, the client device 106 executes instructions included in the reply data 120, which configures and enables the client device 106 to obtain the set of winning digital components from one or more digital component servers 108. For example, the instructions in the reply data 120 can include a network location (e.g., a URL) and a script that causes the client device 106 to transmit a server request (SR) 121 to the digital component server 108 to obtain a given winning digital component from the digital component server 108. In response to the request, the digital component server 108 will identify the given winning digital component specified in the server request 121 (e.g., within a database storing multiple digital components) and transmit, to the client device 106, digital component data (DC Data) 122 that presents the given winning digital component in the electronic document at the client device 106.

[0057] When the client device 106 receives the digital component data 122, the client device will render the digital component (e.g., third-party content), and present the digital component at a location specified by, or assigned to, the script 154. For example, the script 154 can create a walled garden environment, such as a frame, that is presented within, e.g., beside, the native content 152 of the electronic document 150. In some implementations, the digital component is overlayed over (or adj acent to) a portion of the native content 152 of the electronic document 150, and the service apparatus 110 can specify the presentation location within the electronic document 150 in the reply 120. For example, when the native content 152 includes video content, the service apparatus 110 can specify a location or object within the scene depicted in the video content over which the digital component is to be presented.

[0058] The service apparatus 110 can also include an artificial intelligence (“Al”) system 160 configured to autonomously generate digital components, either prior to a request 112 (e.g., offline) and / or in response to a request 112 (e.g., online or real-time). As described in more detail throughout this specification, the Al system 160 can collect online content about a specific entity (e.g., digital component provider or another entity) and summarize the collected online content (e.g., to generate a structured description) using one or more language models 170, which can include large language models.

[0059] A large language model (“LLM”) is a model that is trained to generate and understand human language. LLMs are trained on massive datasets of text and code, and they can be used for a variety of tasks. For example, LLMs can be trained to translate text from one language to another; summarize text, such as website content, search results, news articles, or research papers; answer questions about text, such as “What is the capital of Georgia?”; create chatbots that can have conversations with humans; and generate creative text, such as poems, stories, and code.

[0060] The language model 170 can be any appropriate language model neural network that receives an input sequence made up of text tokens selected from a vocabulary and auto- regressively generates an output sequence made up of text tokens from the vocabulary. For example, the language model 170 can be a Transformer-based language model neural network or a recurrent neural network-based language model.

[0061] In some situations, the language model 170 can be referred to as an auto-regressive neural network when the neural network used to implement the language model 170 auto- regressively generates an output sequence of tokens. More specifically, the auto-regressively generated output is created by generating each particular token in the output sequence conditioned on a current input sequence that includes any tokens that precede the particular text token in the output sequence, i.e., the tokens that have already been generated for any previous positions in the output sequence that precedes the particular position of the particular token, and a context input that provides context for the output sequence.

[0062] For example, the current input sequence when generating a token at any given position in the output sequence can include the input sequence and the tokens at any preceding positions that precede the given position in the output sequence. As a particular example, the current input sequence can include the input sequence followed by the tokens at any preceding positions that precede the given position in the output sequence. Optionally, the input and the current output sequence can be separated by one or more predetermined tokens within the current input sequence.

[0063] More specifically, to generate a particular token at a particular position within an output sequence, the neural network of the language model 170 can process the current input sequence to generate a score distribution, e.g., a probability distribution, that assigns a respective score, e g., a respective probability, to each token in the vocabulary of tokens. The neural networkof the language model 170 can then select, as the particular token, a token from the vocabulary using the score distribution. For example, the neural network of the language model 170 can greedily select the highest- scoring token or can sample, e.g., using nucleus sampling or another sampling technique, a token from the distribution.

[0064] As a particular example, the language model 170 can be an auto-regressive Transformer-based neural network that includes (i) a plurality of attention blocks that each apply a self-attention operation and (ii) an output subnetwork that processes an output of the last attention block to generate the score distribution.

[0065] The language model 170 can have any of a variety of Transformer-based neural network architectures. Examples of such architectures include those described in J. Hoffmann, S. Borgeaud, A. Mensch, E. Buchatskaya, T. Cai, E. Rutherford, D. d. L. Casas, L. A. Hendricks, J. Welbl, A. Clark, et al. Training compute-optimal large language models, arXiv preprint arXiv:2203.15556, 2022; J.W. Rae, S. Borgeaud, T. Cai, K. Millican, J. Hoffmann, H. F. Song, J. Aslanides, S. Henderson, R. Ring, S. Young, E. Rutherford, T. Hennigan, J. Menick, A. Cassirer, R. Powell, G. van den Driessche, L. A. Hendricks, M. Rauh, P. Huang, A. Glaese, J. Welbl, S. Dathathri, S. Huang, J. Uesato, J. Mellor, I. Higgins, A. Creswell, N. McAleese, A.Wu, E. Eisen, S. M. Jayakumar, E. Buchatskaya, D. Budden, E. Sutherland, K. Simonyan, M. Paganini, L. Sifre, L. Martens, X. L. Li, A. Kuncoro, A. Nematzadeh, E. Gribovskaya, D. Donato, A. Lazaridou, A. Mensch, J. Lespiau, M. Tsimpoukelli, N. Grigorev, D. Fritz, T. Sottiaux, M. Pajarskas, T. Pohlen, Z. Gong, D. Toyama, C. de Masson d’Autume, Y. Li, T. Terzi, V. Mikulik, I. Babuschkin, A. Clark, D. de Las Casas, A. Guy, C. Jones, J. Bradbury, M. Johnson, B. A. Hechtman, L. Weidinger, I. Gabriel, W. S. Isaac, E. Lockhart, S. Osindero, L. Rimell, C. Dyer, O. Vinyals, K. Ayoub, J. Stanway, L. Bennett, D. Hassabis, K. Kavukcuoglu, and G. Irving. Scaling language models: Methods, analysis & insights from training gopher. CoRR, abs / 2112.11446, 2021; Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. Exploring the limits of transfer learning with a unified text-to-text transformer. arXiv preprint arXiv: 1910.10683, 2019; Daniel Adiwardana, Minh-Thang Luong, David R. So, Jamie Hall, Noah Fiedel, Romal Thoppilan, Zi Yang, Apoorv Kulshreshtha, Gaurav Nemade, Yifeng Lu, and Quoc V. Le. Towards a human-like open-domain chatbot. CoRR, abs / 2001.09977, 2020; and Tom B Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, PrafullaDhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. Language models are few-shot learners. arXiv preprint arXiv:2005.14165, 2020.

[0066] Generally, however, the Transformer-based neural network includes a sequence of attention blocks, and, during the processing of a given input sequence, each attention block in the sequence receives a respective input hidden state for each input token in the given input sequence. The attention block then updates each of the hidden states at least in part by applying self-attention to generate a respective output hidden state for each of the input tokens. The input hidden states for the first attention block are embeddings of the input tokens in the input sequence and the input hidden states for each subsequent attention block are the output hidden states generated by the preceding attention block.

[0067] In this example, the output subnetwork processes the output hidden state generated by the last attention block in the sequence for the last input token in the input sequence to generate the score distribution.

[0068] Generally, because the language model is auto-regressive, the service apparatus 110 can use the same language model 170 to generate multiple different candidate output sequences in response to the same request, e.g., by using beam search decoding from score distributions generated by the language model 170, using a Sample-and-Rank decoding strategy, by using different random seeds for the pseudo-random number generator that is used in sampling for different runs through the language model 170, or using another decoding strategy that leverages the auto-regressive nature of the language model.

[0069] In some implementations, the language model 170 is pre-trained, i.e., trained on a language modeling task that does not require providing evidence in response to user questions, and the service apparatus 110 (e.g., using Al system 160) causes the language model 170 to generate output sequences according to the pre-determined syntax through natural language prompts in the input sequence.

[0070] For example, the service apparatus 110 (e.g., Al system 160), or a separate training system, pre-trains the language model 170 (e.g., the neural network) on a language modeling task, e.g., a task that requires predicting, given a current sequence of text tokens, the next token that follows the current sequence in the training data. As a particular example, the language model 170 can be pre-trained on a maximum-likelihood objective on a large dataset of text, e g., text that is publicly available from the Internet or another text corpus.

[0071] In some implementations, the Al system 160 can generate a prompt 172 that is submitted to the language model 170, and causes the language model 170 to generate the output sequences 174, also referred to simply as “output”. The Al system 160 can generate prompts 172 in a manner (e.g., having a structure) that is predefined for performing a prompt regeneration process as described in this document, and for example, in relation to FIG. 3A. In some implementations, a prompt can specify a set of constraints for the language model 160 that are to be used to generate content and other outputs, which can be used for the generation of an enhanced prompt for generating digital components and / or content items for the digital components, e.g., as described in relation to FIGS. 3A and 3B.

[0072] In a particular example, assume that the Al system 160 is generating a digital component to provide in response to a component request 112, which includes a keyword or query. In this example, the Al system 160 can generate the prompt 172 to include the query and additional content received in the prior output 174. The additional prompt 172 can also include instructions regarding how clauses generated by the language model 170 using the additional prompt 172 are to be formatted, styled, or semantically styled, among other example configurations for the output clauses (e.g., specifying content that should be excluded from the clauses, e.g., granular details such as numbers).

[0073] In some implementations, submission of this prompt 172 to the language model 170 can cause the language model 170 to generate an output 174, which includes multiple sets of clauses generated according to the query and constraints. The output with the clauses is communicated electronically to the Al system 160. The Al system 160 receives the clauses of the output 174 and generates multiple candidate digital components that are candidates for distribution to the client device 106 in response to the request 112. In some implementations, each different candidate digital component includes a different combination of the clauses received from the language model 170 in the output 174. For example, assume that the output 174 includes 12 different clauses, and that the formatting of the digital components being generated by the Al system 160 each includes space for three different clauses, the Al system 160 could make 220 different candidate digital components using 3 different clauses in each of the candidate digital components (e.g., 12! / (3 !(12-3)!)=220).

[0074] In some instances, the number of candidate digital components that can be generated can include all, or some of all possible combination of available clauses. In some instances,the available clauses can include model generated clauses and / or user-defined clauses provided as input. In some implementations, a prediction can be made for the possible combinations to determine the relevance of the respective digital components from the respective combinations to the digital component page. In some cases, based on the determined relevance, a set of candidate digital components can be generated, where the set is associated with measured relevance (e.g., according to a predefined relevance scale) above a threshold level of relevance. For example, the relevance can be measured based on relevance criteria defining the scoring of clauses to be included in a candidate digital component based on a received query for generating a digital component. In some cases, the relevance can be determined based on the evaluation of relevance associated with a quality criterion (e.g., user defined), a query element, or a usefulness criterion. In some instances, the set of candidate digital components include clauses that are evaluated with high relevance as meeting a relevance threshold for one or more relevance criteria as explained before. In some implementations, the determination of how many digital components to be generated (as candidate digital components) can be based on considerations for optimization of the resources used for the generations. For example, the number of digital components to be generated can be determined according to rules to perform a comparison between resources needed to generate all possible digital components as the available combinations from the generated clauses and resources needed to generate digital components of certain relevance that meets a defined digital component objective criterion.

[0075] In some implementations, the Al system 160 can also create the candidate digital components using a set of different links to online content (e.g., second level domain links to web pages discussing a topic of the candidate digital components, phone numbers, etc.), which can continue to exponentially increase the number of different candidate digital components that the Al system 160 can create using the clauses of the output 174 of the language model 170.

[0076] Furthermore, although a single language model 170 is shown in FIG. 1, different language models can be specially trained to process different prompts at different stages of the processing pipeline. For example, a more general (e.g., larger) language model can be used to generate prompts that are input to a more specialized and faster language model in an online process, e.g., real-time in response to receiving the request 112. Additionally, the Al system 160 can generate a set of candidate digital components as an offline process (e.g., prior toreceiving the request 112), and store the set of candidate digital components in a database. In this scenario, when the Al system 160 receives the request 112, the Al system 160 can further evaluate the stored candidate digital components, e.g., for selection, based on additional information included in the request 112 and / or other contextual data (e.g., time of day, day of week, weather conditions, etc.).

[0077] The Al system 160 can use a chain of prompts to generate a digital components. For example, the Al system 160 can use multiple prompts to arrive at an enhanced prompt that includes a request (e.g., in the form of instructions) for the language model 170 or another Al model (e.g., a diffusion model or another type of text-to-image model) to generate a digital component. This enhanced prompt can include the request, a structured description of a subject of the digital component to be generated, one or more content items (e.g., images, graphics, emojis, etc.) to depict in the digital component, characteristics of the digital component (e.g., style features, colors, text font, etc.), and / or initial data (or a portion thereof) related to the subject of the digital component. Any combination of this data and content can be included with the request in the enhanced prompt.

[0078] The chain of prompts can include a first prompt that instructs the language model 170 to generate a structured description of initial data related to the digital component. The chain of prompts can also include one or more second prompt that instructs the language model 170 to generate a content output that includes the content items and / or characteristics. For example, a single second prompt can be used for all content items and / or characteristics or an individual prompt can be used for each content item and each characteristic. The chain of prompts can include a third prompt that instructs the language model 170 (or another Al model) to generate a digital component using the content output and optionally the structured description and / or initial data (or a portion of the initial data). The chain of prompts is described in more detail below.

[0079] In some implementations, the same language model 170 can be used to process each prompt and generate an output for each prompt. In some implementations, a different Al model, e.g., a different language model 170, is used to generate an output for one or more of the prompts in the chain of prompts. For example, the language model 170 can be used to process the first and second prompts, while a different multi-model model or diffusion model is used to generate the digital component using the third prompt.

[0080] FIG. 2 is a block diagram illustrating interactions between an artificial intelligence system 160, models 251, 252, and 273, and a client device 106. Although this example includes three models 251, 252, and 273, different numbers of models can be used in other examples. For example, a single model can be used to perform the functions of the three models 251, 252, and 273, or two or more models can be used. Each model 251, 252, and 273 is an Al model, e.g., a machine learning model, that is trained to generate an output based on an input. For example, each model 251, 252, and 273 can be substantially similar to the language model 170 of FIG. 1 and trained to generate digital components (offline or in realtime) as described herein. For ease of subsequent description, the models 251, 252, and 273 are separate models.

[0081] The Al system 160 can include separate apparatus to enhance prompts that are to be used in a language model or other type of machine learning model for generating digital components, e.g., the second model 272. In some instances, the second model 252 can be a machine learning model trained to generate text content and, based on the text content, to generate images. In some instances, the image generation can be performed through invoking another model that can be a diffusion model for image generation based on prompts including text description provided as output from the second model 252 based on the respective prompt received as input. In some instances, the third model 272 can be a machine learning model trained for generating images based on prompt including text description and / or images. The third model 272 can be trained to generate digital components in response to input including at least one of a request, descriptive text, images, data, page content (e.g., digital component page content of a web page provided as a reference link or as direct content), or other input such as a product, service, location, time period (e.g., a holiday period).

[0082] In some implementations, a digital component is requested to be generated in relation to an object or an item that will be the subject of the digital component. The Al system 160 can process an initial input that can be used to sequentially generate an enhanced prompt for creating one or more candidate digital components. The enhanced prompt can be generated by adding additional data and / or content items to an initial or first prompt that includes the initial data or a portion of the initial data. The enhancement of the initial prompt can be performed in a chain of processing stages where the initial prompt is used for querying language models to improve the understanding of the initial data (and therefore the subject ofthe digital component, the context of the subject, and / or an entity related to the subject) and to support the generation of multiple options for the digital component that are accurate and relevant to the initial input. In that example, one or more of the candidate digital components can be provided in response to the prompt 245 provided to the second model 272.

[0083] FIG. 2 is an illustration of a utilization of language models that are trained to generate content that can include data generation, image generation, instructions generation, digital components, or another type of content as described throughout this document. For example, content can be generated based on prompts provided to language models as described in relation to FIGS. 3A, 3B, 4A, and 4B.

[0084] The Al system 160 is configured to autonomously generate digital components, either prior to a request (e.g., offline) and / or in response to a request (e.g., real-time or on-the-fly) as described in relation to FIG. 1. For example, the request can be received from a client device 106 that can be the same as or substantially similar to the client device 106 of FIG. 1. In accordance with the present disclosure, the Al system 160 can collect content, e.g., content that can be accessed on the Internet and / or from other networked or non-networked sources, about a specific entity (e.g., digital component provider or another entity such as an organization that publishes a digital component) and / or about a subject of the digital component to be provided. The Al system 160 can collect the content in response to a received request or as part of a process for creating digital components.

[0085] The Al system 160 includes a request apparatus 220 that is configured to obtain initial content such as input data (or other forms of data) relevant for generating a digital component. The input data can be initial data for use in generating a digital component. For example, the initial data, or a portion thereof, can be included in one or more prompts of a chain of prompts for generating a digital component. One example prompt can include the initial data to generate a structured description for the digital component.

[0086] In some instances, the request apparatus 220 can include logic to select at least a portion of the initial data and provide the selected data to be used as part of a prompt 243 that is provided to a first model 251. The prompt 243 can be generated by a prompt apparatus 230 that generates the prompts directed to the models 251, 252, and 272. Thus, the request apparatus 220 can provide the selected data to the prompt apparatus 230 for use in generating one or more prompts that are provided to the model(s).

[0087] The initial data can include data and / or content related to a subject of a digital component to be generated and / or an entity related to the subject. For example, the subject can be an item and the entity can be an entity that produces or distributes the item. In a particular example, the subject can be a product and the entity can be an organization that produces and / or offers the product. In another example, the subject can be an event and the entity can be an organizer of the event.

[0088] The initial data can include, for example, text related to the subject and / or entity, images of the subject and / or entity, other digital components for the subject and / or entity, and / or other appropriate types of content and / or data. The text related to the subject and / or entity can include a name of the subject and / or entity, a description of the subject and / or entity, text extracted from the other digital components (e.g., headlines or descriptions depicted by the other digital components, text depicted by online resources (e.g., web pages or app content) related to the subject or entity, and / or other types of text related to the subject and / or entity. The other digital components can include digital components generated by the Al system 160 for the subject and / or entity, digital components previously distributed by the service apparatus 110 for the subject and / or entity, digital components provided to the service apparatus 110 by the entity, and / or other digital components. The request apparatus 220 can extract text and / or images from the other digital components and include this extracted text and / or images in the initial data.

[0089] In some implementations, the request apparatus 220 receives the initial data from the entity. In some implementations, the request apparatus 220 is configured to identify online resources related to the subject and / or entity, extract data and / or images from the identified resources, and generate the initial data using the extracted data and / or images.

[0090] The request apparatus 220 can select at least a portion of this initial data and provide the initial data to the prompt apparatus 230. The prompt apparatus 230 can generate the prompt 243 (which can be referred to as a first prompt) using the initial data. The prompt 243 can include the initial data and a request for the first model 251 to generate a structured description of the initial data. The structured description can indicate an objective for the digital component, a context for the digital component, and / or another understanding of the subject, entity, and / or purpose of the digital component. The request can include instructions for thefirst model 251 . For example, the instructions can include text that instructs the first model 251 to output a description of what the entity is about and / or its audience.

[0091] In some implementations, the prompt 243 can be configured to instruct the first model 251 to determine whether the initial data violates a condition, e.g., a policy condition. If so, the Al system 160 can halt the process of generating a digital component based on the initial data.

[0092] The prompt apparatus 230 can generate the prompt 243 using a prompt template. The prompt template can include the request (or a portion thereof) and populatable fields for the initial data. For example, the prompt template can include a populatable field for a description of the entity and / or subject, the request, and populatable fields for other initial data. An example prompt 243 for generating a structured description is provided below.

[0093] The prompt apparatus 230 can provide the prompt 243 to the first model 251 to request the generating of a structured description. The first model 251 can process the prompt 243, generate the structured description based on the prompt 243, and provide the structured description to the Al system 160 as a response 253.

[0094] The prompt apparatus 230 can be configured to include logic to generate one or more prompts 244, which can be referred to as second prompts. Each prompt 244 can be configured to generate a content output based on the structured description and optionally the initial data or a portion of the initial data, e.g., the initial data included in the prompt 244. Each content output can include one or more content items to be included in the digital component that is being generated, a description of a content item to be included in the digital component, or one or more characteristics of the digital component.

[0095] The content item can be text, an image, graphics, an emoji, audio, or other types of content. For example, a prompt 244 can be configured to instruct the second model 252 to generate an image of an object or a description of the object that can be used by another model to generate the image of the object. In another example, a prompt can be configured to instruct the second model to recommend, from a set of emojis, one or more emojis to include in the digital component that is being generated.

[0096] The characteristics of the digital component can include visual characteristics of the digital component or portions thereof. For example, the characteristics can include characteristics of the font of text to be depicted by the digital component (e.g., text color, fonttype, font size, font combinations for multiple text portions etc ), visual style features, background color and / or patterns, identification of words or phrases to highlight in the digital component, etc.

[0097] The prompt 244 can include a request for the second model 252 to generate a content output, structured description obtained from the first model 251, and optionally the initial data (e.g., the initial data of the prompt 251). The request can be indicative of character! stic(s) or content item(s) that are requested to be generated as the content output. For example, the request can include instructions for generating a content output of a particular type. In a particular example, the request can be “please recommend the most relevant emojis to be used as decorations in this digital component” and can follow the structured description and / or initial data in the prompt 244. In another example, the request can be “I am a UX designer and am going to design a visual digital component for this campaign. Please recommend some fonts based on the overview above” and this can follow the structure description and / or initial data in the prompt 244.

[0098] The prompt apparatus 230 can generate each prompt 244 using a prompt template. The prompt apparatus 230 can include one or more prompt templates for each type of content output. The prompt template can include the request (or a portion thereof) and populatable fields for the structured description and optionally the initial data. For example, the prompt template can include a populatable field for the structured description, the request, and populatable fields for other initial data. Example prompts 244 for generating content output are provided below.

[0099] The prompt apparatus 230 can provide the prompt 244 to the second model 252 to obtain a content output as a response 254. If multiple content outputs are desired, the prompt apparatus 230 can generate and provide the prompt 244 for each content output to the second model 252. The second model 252 can process each prompt 244, generate a content output based on the prompt 252, and provide the content output to the Al system 160 as a response 254. In some implementations, different second models 252 can be trained for each type of content output. For example, a different second model can be trained for generating images than a second model 252 trained for generating font characteristics.

[0100] The prompt apparatus 230 can include logic for generating a prompt 254 (which can be referred to as a third prompt or enhanced prompt) for generating digital components.The prompt 254 can include a request, the content output(s) generated by the second model(s) 252, and optionally the structured description generated by the first model 251 and / or the initial data (e.g., the initial data of the prompt 243).

[0101] The request can include instructions for generating a digital component using the content output(s) and based on the structured description and / or initial data, if included in the prompt. Similar to the other prompts, the prompt apparatus 230 can generate the prompt 254 using a prompt template.

[0102] The third model 272 can process the prompt 245, generate a digital component based on the prompt, and provide the digital component as a response 255 to the Al system 160. If multiple digital components are desired, the third model 272 can process the prompt 245 multiple times and generate a digital component each time. Or, the request can include instructions to generate multiple digital components using the prompt 245.

[0103] In some implementations, the third model can be configured to generate one or more clauses for use in generating digital components, e.g., rather than generating the digital components themselves. For example, the models 215, 252, and 272 can be trained to generate clauses, content, or image or text, among other example content items that are to be used for generating new prompts and / or regenerate the prompts to arrive at an enhanced prompt for use in generating a new digital component generation that meets a grounding threshold so that output clauses are classified as factual and grounded. In some implementations, the models can be machine-language models trained based on relevant training data including data sets of text items, images, digital components, and other data obtained for the training. In some instances, at least some of the models 215, 252, and 272, can be large language models.

[0104] The digital components used for training the model 272 can include, for example, previously generated digital components. In some examples, the digital components used for the training of the language model can be existing digital components generated by the language model in response to previous requests or can be digital components generated through other means (e.g., other language models or other techniques). The response 255 can include generated clauses that can be used to generate a new digital component 260 at the Al system 160 that can be provided to the client device 106. In some implementations, the Al system 160 can use the clauses to create multiple different digital components and then performpost-processing to select, from among the different candidate digital components, a set of output digital components.

[0105] The post training apparatus 240 can include logic to perform post training analysis to evaluate the obtained clauses from the model 272 and to filter those of the clauses that are grounded in the content of the respective online source. In some implementations, generated clauses can be evaluated based on quality error rating criteria that can measure the performance of generated clauses with respect to different quality characteristics. For example, the quality characteristics can be related to linguistic aspects of the generated causes (e.g., grammar and spelling errors) and / or semantics and context (e.g., not supported by the digital component page provided with a request for digital component generation, awkward wording, or else). In some implementations, automatically created digital components can be determined immediately as eligible to be provided upon generation in response to requests. In some cases, if it is determined that a digital component does not meet quality criteria for generation and / or is determined to not be used after generations, can be automatically removed from the set of digital assets stored for providing in response to requests.

[0106] In some implementations, the digital component apparatus 241 can generate digital components by using one or more of the created clauses and can provide at least one digital component 260 to the user device 106 in response to the query 270. In some implementations, the generation of the digital components can be performed after or before receipt of a query 270. In the latter case, the Al system 160 can pre-store digital components that are generated for different online sources (e.g., digital component pages) and stored at a digital components 285 storage at the memory structure 275. In some implementations, generated clauses from the model 250 as obtained through the response 255, can be stored by the Al system 160 at the memory structure 275, for example, at a clause data 280 storage.

[0107] The digital component apparatus241 can also include logic to evaluate the relevance of generated clauses to query constraints received with requests for generating digital components. The relevance can be determined as a level of completeness of one or more clauses to content located in a candidate digital component, and / or an evaluation of the tone (e.g., positive or negative) of the clause.

[0108] FIG. 3A is a block diagram of an example process 300 for generating enhanced prompts for use in generating digital components using a trained Al model. Operations of theprocess 300 can be performed, for example, by the service apparatus 110 and / or the Al apparatus 160 of FIG. 1, or another data processing apparatus. The operations of the process 300 can also be implemented as instructions stored on a computer readable medium, which can be non-transitory. Execution of the instructions, by one or more data processing apparatus, causes the one or more data processing apparatus to perform operations of the process 300.

[0109] In some implementations, when a digital component is generated using generative Al techniques, a prompt is created that includes text and / or images to define the task and / or sub-tasks of the digital component generation. However, if the prompt is generated based on data that is initial, a misunderstanding during the interpretation of the data can occur that can influence the quality of the digital component that is generated.

[0110] In some implementations, when input for generating a prompt is received, the input can be used as part of a prompt to regenerate the input so that structured data is provided that is descriptive of the digital component. For example, an initial prompt can be created to include initial data 301 that can include a name 302 an object associated with the digital component (e.g., a subject of the digital component or an entity associated with the subject), images 303, and / or text 303. To illustrate, the initial data 301 can include data as shown in Table 1 below.Table 1

[0111] The input, as presented in Table 1 , includes initial data that can be provided for generating a digital component. The initial data can be data of any format that may not be structured according to a data model defined for understanding the request for the generation of a digital component. In the example of Table 1, even if the data is provided in a table structure, the data may not be structured according to an expected (predefined) model that has a structure that is defined to include attributes associated with a domain of the request and / or objective of the request for the digital component. For example, the predefined structure that can be used for generating the structured data can be as shown in Table 2 and can include at least for example attributes such as relevant product / services, audience, properties of the audience, objectives of the digital component. The initial data can be first processed to enhance a prompt to be used for the generation of the digital component rather than being used directly to request the digital component generation. The initial data is in this way pre-processed to support a regeneration of a prompt to obtain an enhanced prompt that when used at a language model can result in a digital component that is more relevant to the initial data 301 compared to a digital component that can be generated if the initial data 301 was used as part of a prompt directly used by a language model to generate the digital component. Further, by using a regenerated prompt, the process of the digital component generation can be enhanced so that fewer computational resources and / or interactions are performed to obtain a result that meets a digital component criterion (e.g., quality, content, size, evaluation metric threshold, etc.).

[0112] In some instances, at 310, the initial data 301 can be used to be provided to a first model, for example, the first model 251 of FIG. 2 to generate, at 311, a creative understanding of the initial data 301. The first model can generate as part of the creative understanding 311 structured data that is a structured description generated based on the creative data 301. The structured data can be considered as an understanding of the digital component’s content that is to be based on the provided initial data 301. The generated structured description from the creative understanding 311 can include text indicative of an overview of the object 305 that is associated with the digital asset (e g., a subject or entity associated with the digital component) and structured data 306 that is descriptive of the digital component. For example, the creative understanding 311 can result in generating structured data such as the data presented in Table 2 below.LLM Understanding of the initial dataTable 2

[0113] The data in Table 2 can be considered a structured description. In some implementations, based on the generated structured description from the creative understanding 311, design ideas 320 can be generated using a language model (e.g., such as the second model of FIG. 2) to request ideation 321 for design ideas 320. The design ideas 320 that can be generated by querying a language model according to a prompt including the generated structured description, as well as the creative data, and a request associated with the digital component generation. The design ideas 320 that can be generated include content items, such as image items or description of images (e.g., background images 307 and marketing image 308) and visual characteristics 309 to be applied for content visualization (e g., recommendation of a font for text). Table 3 below presents an example of a generated content item that includes descriptive content for marketing images to be used by a language model when generating the digital component that is to include marketing images.Table 3

[0114] The creative understanding 311 and design ideas 320 generation can be implemented as processes executed at a computer system that is for generation of enhanced prompts based on initial input that is the initial data 301. The generated output, which is also referred to herein at content output, from the creative understanding 311 and the design ideas 320 generation, such as text, images, (e.g., background images 307 and / or marketing images 308), and visual characteristics description 308 can be provided for generating enhancedprompts at the prompt enhancement 330 so that one or more enhanced prompts can be generated as part of the enhanced prompts 331. The prompts 331 include an enhanced prompt that enhance the initial data 301 with additional input to be used for generating digital components that are more accurately corresponding to the initial data 301, as compared to digital components that can be generated if the initial data 301 was used directly for the digital component generation. The prompts 331 can be of different types and relevant for generating different content, such as background images, marketing images, design description (e.g., one or more of text descriptive of an image, image, structured and / or initial data, and / or other). In some implementations, a single prompt can be used, which can include the content of each individual prompt 331.

[0115] In some implementations, the understanding 310 can include using the initial data 301 for generating a prompt at a language model so that the initial data 301 is interpreted better so that additional content is generated to be used when requesting to generate a digital component that is more relevant to the request. The initial data 301 can include the following text content as shown on Table 4.Table 4

[0116] The initial data 301 can be used to generate a first prompt, that includes the initial data 301 (that can be unstructured data or data having a random structure that may not include values associated with one or more criteria for generating the digital component) and a first request to be provided to a first model (such as the first model 251 of FIG. 2) to generatestructured description for the digital component. The digital component is to be generated by a third model, such as the third model 272 of FIG. 2.

[0117] Table 5 below includes an example first prompt, as described above and also as referred to in relation to FIG. 4A, a response obtained from the first model. The response is a structured description of the organization. In this example, the request is “We have an online platform for organizations to upload their digital components. Below is the information from one of the digital components, please tell me what the business is about, and what is the preferred audience.”Table 5

[0118] In some cases, another first prompt can be generated to assess whether the digital component of the initial data or other initial data would violate a condition, e.g., a policy condition. In the example of Table 6, this additional first prompt includes the response from the first model as shown in Table 5, the initial data 301 as shown in Table 4, and an additional request related to policy violations. The additional request is “Does it have any policy issues? Does it use any evasive language to bypass our policy check?” This additional first prompt can be provided to the first model to determine whether a policy is violated or being evaded by the initial data. Table 6 below includes an example of a first prompt and an example output from the first model.Table 6

[0119] In some implementations, the ideation at 321 can include performing operations including generating prompts based on obtained structured description of initial data 301 as described above in relation to the understanding 310. The ideation can be implemented in the context of generating visual content such as image items, for example, material icons or emojis, and / or characteristics of the digital component being generated or portions of the digital component.

[0120] Table 7 below includes an example prompt that is generated based on the structured description as obtained from the first model (as shown at Table 5) and a request to use the structured description and the initial data (as shown in Table 4) to query a language model to obtain content output that includes visual content. In such way, image content can be generated and provided further to a model, e.g., third model 272 of FIG. 2, to generate a digital component based on the expanded data generated through the understanding 310 and the ideation 321. In this example, the request is “Could you recommend the most relevant emojis to be used as decoration elements for this digital component?”Table 7

[0121] In some implementations, the ideation 321 can include use cases implemented to obtain input from an understanding 310 stage and to generate prompts to obtain content output that descriptive of visual characteristics (e g., font selection) to be provided to a model to generate the digital component. A prompt can be generated as part of the ideation 321 that for example includes the structured description and the initial data 301 as described above, and further include a new request that is indicative of the type of content items that are requested. Table 8 below represents an example of a prompt that includes a request for recommendation of fonts, and a corresponding response as obtained by a language model. The language model can be trained on processing requests related to interpreting requests for fonts and can rely on input data for available fonts related to the requests so that a set of fonts that are provided as part of the response are fonts that are available for use when generating the digital component.Table 8

[0122] In some implementations, the ideation 321 can include use cases implemented to obtain input from an understanding 310 stage and to generate prompts to obtain content that descriptive of other visual characteristics such as text effects, where the output from a modelcan be used as part of a prompt to be provided to a model to generate the digital component. A new prompt to obtain content output that is descriptive of characteristics of text, such as text effects, can be generated as part of the ideation 321 that for example includes the structured description and the initial data 301 as described above, and further include a new request that is indicative of the type of visual characteristics that are requested for definition. Table 9 below represents an example of the new prompt where the prompt includes a new request for recommendation of decorative elements that can be used to make text as visual items.Table 9

[0123] In some implementations, a prompt can be generated in the content of the ideation 321 to request generation of content output that includes marketing images. Table 10 below includes an example prompt that includes a request for the marketing images generation, as well as structured data providing an overview of the object related to the digital component, and example of initial data 301. Further, the prompt includes as part a request for an output format for the provided result by the model. Table 10 also includes an example output provided by the model that includes data that is indicative of characteristic(s) and / or content items to be used when generating the digital component. As such the output provided as a response can be used to generate the enhanced prompt for querying a model to generate the digital component, as described in relation to FIG. 1, 2, 4 A, and 4B.Table 10[00124J The response example as shown at Table 10 includes different options as generated design ideas for images that are provided in the form of descriptive content and data (e g., in a specified format in the request such as JSON).

[0125] In some implementations, the generated responses as shown in Tables 5, 6, 7, 8, 9, and 10 can be used as part of the prompt enhancement 330, for example, to form a chain of prompt regenerations so that an enhanced prompt is generated. The chain of prompts includes as part of an initial prompt the provided initial data 301, where subsequent prompts as generated are expanded with data obtained from executing an understanding 310 and ideation 321 operations as described in relation to FIGS. 2, 3A, 3B, 4A and 4B. The expansion of the prompts can be performed by generating new prompts according to defined execution logic and including specific requests to obtain targeted content that can support the generation of improved digital components in a fast and computationally efficient manner.

[0126] FIG. 3B is a block diagram of an example process 340 for generating digital components 385 based on an enhanced prompt including generatively designed images. Operations of the process 340 can be performed, for example, by the service apparatus 110 of FIG. 1, and for example, by the Al system 160 as described in relation to FIGS. 1 and 2, or another data processing apparatus. The operations of the process 340 can be implemented as instructions stored on a computer readable medium, which can be non-transitory.

[0127] The process 340 can be performed as part of a request to generate digital components according to generative Al techniques, where the request for the generation includes text 346 and images 347 to be used for the digital component generation. In some implementations, the input text and images can be generated as part of a flow of executing understanding and ideation as described in relation to FIG. 3A. The input text and images can be created based on initial data substantially similar to the initial data 301. At initial prompt 345 can be generated to include:Text 346 descriptive of images, for example, text generated based on structured data used as part of a prompt provided to a language model. For example, the text can include the examples presented in Table 3 above of generated content item that includes descriptive content for marketing images to be used by a model when generating the digital component that is to include marketing images.Generated image 347 by a trained image model that is queried based on a prompt including the text 346.

[0128] A first language model 350 can be queried based on the initial prompt. A response can be obtained that includes descriptions of the images 347. The descriptions of the images can be considered as expanded text compared to the text 346 as part of the initial prompt 345. An expanded prompt 355 can be generated and provided for querying a second language model 365 that can provide images 370. The generated expanded prompt with the obtained descriptions can be substantially similar to prompt obtained at 460 of FIG. 4B where the generated images 370 are the images obtained at 470 of FIG. 4B. The images 370 are generatively designed images created by the second language model that is a trained image editing model to generate candidate images based on interpreting the descriptions of the images 360. In some instances, the second language model 365 can receive other information as part of the expanded prompt 365 to generate the images 370. For example, the expanded prompt 365 can include additional structured and initial data and specific request for the image generation (e.g., type of images, including format of image, visual characteristics of images (such as defined based on querying a language model for ideation of visual characteristics, e.g., font or text effects). The images 370 can be used as part of an enhanced prompt 375, for example, as described in relation to operation 480 of FIG. 4B to obtain digital components at 490 of FIG. 4B. The enhanced prompt 375 can be provided for querying a third model 380 such as the third model 272 of FIG. 2 and as described in relation to FIG. 4A and 4B for a model that is provided with an enhanced prompt to generate digital components 385 that are improved compared to digital components 385 that can be generated by the same language model if that was queries with the initial prompt 345.

[0129] FIG. 4A is a block diagram of an example process 400 for generating enhanced prompts and using the enhanced prompts to generate digital components. Operations of the process 400 can be performed, for example, by the service apparatus 110 of FIG. 1 (e.g., bythe Al system 160), or another data processing apparatus. The operations of process 400 can be performed as described in relation to FIG. 2 and the interaction between the Al system 160 and one or more models, e.g., one or more Al models which can include one or more language models. The operations of the process 400 can also be implemented as instructions stored on a computer readable medium, which can be non-transitory. Execution of the instructions, by one or more data processing apparatus, causes the one or more data processing apparatus to perform operations of the process 400. For brevity, the process 400 is described in terms of being performed by a system.

[0130] The system generates a first prompt that includes a first request and initial data for use by one or more models to generate a structured description for a digital component to be generated (405). The first request can include instructions for generating the structure description using the initial data. The initial data can include data related to a subject of the digital component that is to be generated. For example, the initial data can include the name of an entity (e.g., the name of a business) that is the subject of the digital component or that is related to the subject (e.g., the subject can be a product and the entity can be the business that offers the product), other digital components for the subject of the digital component to be created, headlines or other text of these other digital components, text descriptions of the subject or entity, and / or other data related to the subject or entity.

[0131] The system obtains the structured description as generated based on the first prompt from the one or more models (410). As described above, the structured description can represent an understanding of the subject of the digital component to be generated and / or an entity associated with the subject based on the initial data.

[0132] The system generates a second prompt that includes a second request for the one or more models to generate a content output (415). The content output can be, for example, a content item, a description of a content item, or one or more characteristics of the digital component to be generated. The second prompt can include, in addition to the second request, the structured description and / or the initial data. As described above, the system can use multiple second prompts to obtain multiple content outputs. The system obtains content output for use in generating the digital component (420).

[0133] The system generates a third prompt that includes a third request for the one or more models to generate the new digital component based on the content output (425). Inaddition to the content output, the third prompt can also include the structured description and / or the initial data. The third prompt is an enhanced prompt for input into a model to generate the digital component.

[0134] The system obtains the digital component from the model based on the third prompt (430). The model can generate the digital component by combining the content output(s) based on the other content of the third prompt. For example, the model can generate a digital component that includes an image of a content output, e.g., as a background image, and any additional content output(s) of the third prompt. The model can use the other content, e.g., structured description and / or initial data, of the third prompt to arrange the content output(s) within an image-based digital component. The model can also generate a new image as an image-based digital component based on the content output(s) and other content of the third prompt.

[0135] The same model can be used to generate the structured description, the content output(s), and the digital component. In some implementations, a different model can be used to generate each of these outputs. In some implementations, a language model is used to generate the structured description and the content outputs, and a text-to-image model is used to generate the digital component.

[0136] FIG. 4B is a block diagram of an example process 450 for enhancing a prompt with obtained descriptions of images from a trained language model to generate a new prompt to be used for generating digital components. Operations of the process 450 can be performed, for example, by the service apparatus 110 of FIG. 1 (e.g., by the Al system 160), or another data processing apparatus. The operations of the process 400 can also be implemented as instructions stored on a computer readable medium, which can be non-transitory. Execution of the instructions, by one or more data processing apparatus, causes the one or more data processing apparatus to perform operations of the process 400. For brevity, the process 400 is described in terms of being performed by a system.

[0137] In some implementations, generative Al techniques can be used based on an enhanced prompt that is generated in a chain of iterations that include an initial querying of a language model based on a first prompt that include structured description, such as the structured description as described in relation to FIG. 2, 3A, and 3B, for example provided from execution of a creative understanding process. The system generates a prompt thatincludes a first request and the structured description for use by a one or more models (460). The first request is indicative of an image to be generated for use in generating a digital component by the one or more models. The system obtains descriptions of images to be used for generating the digital component from the one or more models (470). The system generates a second prompt that includes the description of images as generated by the one or more models and an image generated based on the descriptions (480). The second prompt is for generating new image content to be included in an enhanced prompt for input into the one or more models to generate the digital component. The system obtains, from the one or more models, the digital component based on the enhanced prompt (490).

[0138] For example, the process 450 can be executed in the context of a request for generation of image content to be used for multiple digital components that are generated according to generative Al techniques as candidate digital components that can be provided for display at display devices, for example, after post-processing of the candidate digital components. The post-processing can be performed to filter a subset of the generated output that meets a defined criterion (e.g., quality threshold, visual criteria with regard to image quality, colors, resolution, etc.).

[0139] FIG. 5 is a block diagram of an example computer system 500 that can be used to perform operations described above. The system 500 includes a processor 510, a memory 520, a storage device 530, and an input / output device 540. Each of the components 510, 520, 530, and 540 can be interconnected, for example, using a system bus 550. The processor 510 is capable of processing instructions for execution within the system500. In one implementation, the processor 510 is a single-threaded processor. In another implementation, the processor 510 is a multi -threaded processor. The processor 510 is capable of processing instructions stored in the memory 520 or on the storage device 530.

[0140] The memory 520 stores information within the system 500. In one implementation, the memory520 is a computer-readable medium. In one implementation, the memory520 is a volatile memory unit. In another implementation, the memory 520 is a nonvolatile memory unit.

[0141] The storage device 530 is capable of providing mass storage for the system 500. In one implementation, the storage device 530 is a computer-readable medium. In various different implementations, the storage device 530 can include, for example, a hard disk device,an optical disk device, a storage device that is shared over a network by multiple computing devices (e.g., a cloud storage device), or some other large capacity storage device.

[0142] The input / output device 540 provides input / output operations for the system 500. In one implementation, the input / output device 540 can include one or more of a network interface devices, e.g., an Ethernet card, a serial communication device, e g., and RS-232 port, and / or a wireless interface device, e.g., and 802.11 card. In another implementation, the input / output device can include driver devices configured to receive input data and send output data to other devices, e.g., keyboard, printer, display, and other peripheral devices 550. Other implementations, however, can also be used, such as mobile computing devices, mobile communication devices, set-top box television client devices, etc.

[0143] Although an example processing system has been described in FIG. 5, implementations of the subject matter and the functional operations described in this specification can be implemented in other types of digital electronic circuitry, or in computer software, firmware, or hardware, including the structures disclosed in this specification and their structural equivalents, or in combinations of one or more of them.

[0144] An electronic document (which for brevity will simply be referred to as a document) does not necessarily correspond to a file. A document may be stored in a portion of a file that holds other documents, in a single file dedicated to the document in question, or in multiple coordinated files.

[0145] For situations in which the systems described here collect and / or use personal information about users, the users may be provided with an opportunity to enable / disable or control programs or features that may collect and / or use personal information (e.g., information about a user’s social network, social actions or activities, a user’s preferences, or a user’s current location). In addition, certain data may be treated in one or more ways before it is stored or used, so that personally identifiable information associated with the user is removed. For example, a user’s identity may be anonymized so that the no personally identifiable information can be determined for the user, or a user’s geographic location may be generalized where location information is obtained (such as to a city, ZIP code, or state level), so that a particular location of a user cannot be determined.

[0146] Embodiments of the subject matter and the operations described in this specification can be implemented in digital electronic circuitry, or in computer software,firmware, or hardware, including the structures disclosed in this specification and their structural equivalents, or in combinations of one or more of them. Embodiments of the subject matter described in this specification can be implemented as one or more computer programs, i.e., one or more modules of computer program instructions, encoded on a computer storage medium for execution by, or to control the operation of, a data processing apparatus. Alternatively, or in addition, the program instructions can be encoded on an artificially- generated propagated signal, e.g., a machine-generated electrical, optical, or electromagnetic signal, that is generated to encode information for transmission to a suitable receiver apparatus for execution by a data processing apparatus. A computer storage medium can be, or be included in, a computer-readable storage device, a computer-readable storage substrate, a random or serial access memory array or device, or a combination of one or more of them. Moreover, while a computer storage medium is not a propagated signal, a computer storage medium can be a source or destination of computer program instructions encoded in an artificially-generated propagated signal. The computer storage medium can also be, or be included in, one or more separate physical components or media (e.g., multiple CDs, disks, or other storage devices).

[0147] The operations described in this specification can be implemented as operations performed by a data processing apparatus on data stored on one or more computer-readable storage devices or received from other sources.

[0148] The term “data processing apparatus” encompasses all kinds of apparatuses, devices, and machines for processing data, including by way of example a programmable processor, a computer, a system on a chip, or multiple ones, or combinations, of the aforementioned. The apparatus can include special purpose logic circuitry, e.g., an FPGA (field programmable gate array) or an ASIC (application-specific integrated circuit). The apparatus can also include, in addition to hardware, code that creates an execution environment for the computer program in question, e.g., code that constitutes processor firmware, a protocol stack, a database management system, an operating system, a cross-platform runtime environment, a virtual machine, or a combination of one or more of them. The apparatus and execution environment can realize various different computing model infrastructures, such as web services, distributed computing and grid computing infrastructures.

[0149] This document refers to a service apparatus. As used herein, a service apparatus is one or more data processing apparatuses that perform operations to facilitate the distribution of content over a network. The service apparatus is depicted as a single block in block diagrams. However, while the service apparatus could be a single device or single set of devices, this disclosure contemplates that the service apparatus could also be a group of devices, or even multiple different systems that communicate in order to provide various content to client devices. For example, the service apparatus could encompass one or more of a search system, a video streaming service, an audio streaming service, an email service, a navigation service, an advertising service, a gaming service, or any other service.

[0150] A computer program (also known as a program, software, software application, script, or code) can be written in any form of programming language, including compiled or interpreted languages, declarative or procedural languages, and it can be deployed in any form, including as a stand-alone program or as a module, component, subroutine, object, or other unit suitable for use in a computing environment. A computer program may, but need not, correspond to a file in a file system. A program can be stored in a portion of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document), in a single file dedicated to the program in question, or in multiple coordinated files (e.g., files that store one or more modules, sub-programs, or portions of code). A computer program can be deployed to be executed on one computer or on multiple computers that are located at one site or distributed across multiple sites and interconnected by a communication network.

[0151] The processes and logic flows described in this specification can be performed by one or more programmable processors executing one or more computer programs to perform actions by operating on input data and generating output. The processes and logic flows can also be performed by, and apparatus can also be implemented as, special purpose logic circuitry, e.g., an FPGA (field programmable gate array) or an ASIC (application-specific integrated circuit).

[0152] Processors suitable for the execution of a computer program include, by way of example, both general and special purpose microprocessors, and any one or more processors of any kind of digital computer. Generally, a processor will receive instructions and data from a read-only memory or a random-access memory or both. The essential elements of a computer are a processor for performing actions in accordance with instructions and one or more memorydevices for storing instructions and data. Generally, a computer will also include, or be operatively coupled to receive data from or transfer data to, or both, one or more mass storage devices for storing data, e.g., magnetic, magneto-optical disks, or optical disks. However, a computer need not have such devices. Moreover, a computer can be embedded in another device, e.g., a mobile telephone, a personal digital assistant (PDA), a mobile audio or video player, a game console, a Global Positioning System (GPS) receiver, or a portable storage device (e.g., a universal serial bus (USB) flash drive), to name just a few. Devices suitable for storing computer program instructions and data include all forms of non-volatile memory, media and memory devices, including by way of example semiconductor memory devices, e.g., EPROM, EEPROM, and flash memory devices; magnetic disks, e.g., internal hard disks or removable disks; magneto-optical disks; and CD-ROM and DVD-ROM disks. The processor and the memory can be supplemented by, or incorporated in, special purpose logic circuitry.

[0153] To provide for interaction with a user, embodiments of the subject matter described in this specification can be implemented on a computer having a display device, e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor, for displaying information to the user and a keyboard and a pointing device, e.g., a mouse or a trackball, by which the user can provide input to the computer. Other kinds of devices can be used to provide for interaction with a user as well; for example, feedback provided to the user can be any form of sensory feedback, e.g., visual feedback, auditory feedback, or tactile feedback; and input from the user can be received in any form, including acoustic, speech, or tactile input. In addition, a computer can interact with a user by sending documents to and receiving documents from a device that is used by the user; for example, by sending web pages to a web browser on a user’ s client device in response to requests received from the web browser.

[0154] Embodiments of the subject matter described in this specification can be implemented in a computing system that includes a back-end component, e.g., as a data server, or that includes a middleware component, e.g., an application server, or that includes a frontend component, e.g., a client computer having a graphical user interface or a Web browser through which a user can interact with an implementation of the subject matter described in this specification, or any combination of one or more such back-end, middleware, or front-end components. The components of the system can be interconnected by any form or medium ofdigital data communication, e.g., a communication network. Examples of communication networks include a local area network (“LAN”) and a wide area network (“WAN”), an internetwork (e.g., the Internet), and peer-to-peer networks (e.g., ad hoc peer-to-peer networks).

[0155] The computing system can include clients and servers. A client and server are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other. In some embodiments, a server transmits data (e.g., an HTML page) to a client device (e.g., for purposes of displaying data to and receiving user input from a user interacting with the client device). Data generated at the client device (e.g., a result of the user interaction) can be received from the client device at the server.

[0156] While this specification contains many specific implementation details, these should not be construed as limitations on the scope of any inventions or of what may be claimed, but rather as descriptions of features specific to particular embodiments of particular inventions. Certain features that are described in this specification in the context of separate embodiments can also be implemented in combination in a single embodiment. Conversely, various features that are described in the context of a single embodiment can also be implemented in multiple embodiments separately or in any suitable sub-combination. Moreover, although features may be described above as acting in certain combinations and even initially claimed as such, one or more features from a claimed combination can in some cases be excised from the combination, and the claimed combination may be directed to a subcombination or variation of a sub-combination.

[0157] Similarly, while operations are depicted in the drawings in a particular order, this should not be understood as requiring that such operations be performed in the particular order shown or in sequential order, or that all illustrated operations be performed, to achieve desirable results. In certain circumstances, multitasking and parallel processing may be advantageous. Moreover, the separation of various system components in the embodiments described above should not be understood as requiring such separation in all embodiments, and it should be understood that the described program components and systems can generally be integrated together in a single software product or packaged into multiple software products.

[0158] Thus, particular embodiments of the subject matter have been described. Other embodiments are within the scope of the following claims. In some cases, the actions recited in the claims can be performed in a different order and still achieve desirable results. In addition, the processes depicted in the accompanying figures do not necessarily require the particular order shown, or sequential order, to achieve desirable results. In certain implementations, multitasking and parallel processing may be advantageous.

[0159] What is claimed is:

Claims

CLAIMS1. A computer-implemented method comprising: generating, by an artificial intelligence system, a first prompt that includes (i) a first request for one or more models to generate a structured description for a subject of a new digital component to be generated by the one or more models and (ii) initial data for use by the one or more models to generate the structured description; obtaining, by the artificial intelligence system and from the one or more models, the structured description as generated based on the first prompt; generating, by the artificial intelligence system, a second prompt that includes a second request for the one or more models to generate a content output that includes at least one of (i) a content item, (ii) a description of a content item, or (iii) one or more characteristics of the new digital component, wherein the second prompt includes the structured description, the initial data, or both; obtaining, by the artificial intelligence system and from the one or more models, the content output as generated based on the second prompt; generating, by the artificial intelligence system, a third prompt that includes a third request for the one or more models to generate the new digital component based on the content output, wherein the third prompt includes the content output, the structured description, and the initial data; and obtaining, by the artificial intelligence system and from the one or more models, the new digital component generated based on the third prompt.

2. The method of claim 1, further comprising distributing the new digital component to one or more client devices.

3. The method of claim 1 or 2, wherein the initial data includes at least one of: a name of the subject of the new digital component, a text description of the subject of the new digital component, or other digital components for the subject of the new digital component.

4. The method of any preceding claim, wherein the initial data includes a specified source of online content, wherein the structured description generated based on the first prompt isgenerated by the one or more models based on at least one piece of content obtained from the specified source of online content.

5. The method of any preceding claim, wherein the second request of the second prompt is indicative of an image as the content item, characteristics of text content to be used when generating the new digital component, or both.

6. The method of claim 5, wherein the obtaining from the one or more models includes obtaining the image, and wherein the second request in the second prompt is indicative of a type of the image to be generated.

7. The method of any preceding claim, wherein the content item is a text content item, wherein the second request of the second prompt is indicative of a visual characteristic of the text content item, wherein the visual characteristic includes at least one of a dimensional or text effect characteristic.

8. The method of any preceding claim, wherein the one or more characteristics of the new digital component comprises visual characteristics of content depicted by the new digital component.

9. The method of any preceding claim, wherein obtaining the new digital component comprises: providing a request for at least one digital component by invoking the one or more models to provide clauses and images based on the initial data, the structured description, and the content output; and generating, by the one or more models, a new clause that indicates visual characteristics generated based on the content output.

10. The method of any preceding claim, wherein the one or more models include a single large language model.

11. The method of any one of claims 1 to 9, wherein the one or more models include multiple machine learning models.

12. The method of any preceding claim, wherein the third prompt includes at least a portion of the initial data.

13. A computer implemented method comprising: generating, by an artificial intelligence system, a prompt that includes a first request and structured description for use by one or more models, wherein the first request is indicative of an image to be generated for use in generating a digital component by a second language model; obtaining, by the artificial intelligence system and from the one or more models, descriptions of images to be used for generating the digital component; and generating, by the artificial intelligence system, a second prompt that includes the description of images as generated by the one or more models and an image generated based on the descriptions, wherein the second prompt is for generating new image content to be included in an expanded prompt for input into the second language model to generate the digital component; and obtaining, by the artificial intelligence system and from the one or more models, the digital component based on the expanded prompt.

14. The method of claim 13, wherein obtaining the digital component comprises: requesting generation of one or more new images generated based on the second prompt; and providing the one or more new images together with the second prompt for obtaining the digital component.

15. The method of claim 13 or 14, wherein the structured description is generated based on a second request and initial data descriptive of a subject of the digital component.

16. The method of any one of claims 13 to 15, wherein generating the second prompt comprises generating the image based on a third language model comprising: generating, by the artificial intelligence system, a third prompt that includes the second request and initial data for use by the one or more models to generate the structured description for the digital component to be generated by the one or more models; obtaining, by the artificial intelligence system and from the fourth language model, the structured description as generated based on the third prompt; generating, by the artificial intelligence system, a fourth prompt that includes a third request, the structured description, and the initial data for use by the one or more models, wherein the third request is indicative of an image as a content item to be used for generating the digital component by the one or more models; and obtaining, by the artificial intelligence system and from the one or more models, an image as the image content item to be used for generating the digital component.

17. A system comprising: one or more processors; and one or more storage devices storing instructions that, when executed by the one or more processors, cause the one or more processors to carry out the method of any preceding claim.

18. A computer readable storage medium carrying instructions that, when executed by one or more processors, cause the one or more processors to carry out the method of any one of claims 1 to 16.

19. A computer program product comprising instructions which, when executed by one or more computers, cause the one or more computers to carry out the operations of the method of any of claims 1 to 16.