Artificial intelligence model training for image generation

By parsing static images into content items and structured representations, and training an image design model with high-quality data and noise techniques, the method enhances the quality and diversity of generated images, addressing inaccuracies in current generative AI models.

WO2026049733A1PCT designated stage Publication Date: 2026-03-05GOOGLE LLC
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
PCT/US2024/044413
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-08-29
Publication Date
2026-03-05

AI Technical Summary

Technical Problem

Current generative artificial intelligence models produce inaccurate and low-quality images, lacking robust training data and often incorporating hallucinations and artifacts, which affects the quality and diversity of generated images.

Method used

A method involving an image parser model to parse static images into content items and structured representations, combined with an image design model trained using high-quality static images and noise techniques to generate diverse and accurate training data, resulting in higher quality image designs.

Benefits of technology

The approach generates more diverse and higher quality images by expanding training data through noise techniques and using structured representations, reducing model complexity and computational resources while improving accuracy and relevance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2024044413_05032026_PF_FP_ABST
    Figure US2024044413_05032026_PF_FP_ABST
Patent Text Reader

Abstract

Methods, systems, and apparatus, including computer programs encoded on a computer storage medium, for training AI models to generate images. In one aspect, a method includes obtaining first training data including a set of first training samples. Each first training sample includes a first image. An image parser model is trained using the first training data to output structured representations of input images. Second images are processed using the image parser model to generate second training data including a set of second training samples. Each second training sample includes a second image and a corresponding structured representation of content items depicted by the second image output by the image parser model based on the second image. An image design model is trained using the second training data to output a structured representation for an input including a set of content items.
Need to check novelty before this filing date? Find Prior Art

Description

Attorney Docket No. t>6113-0768WO1ARTIFICIAL INTELLIGENCE MODEL TRAINING FOR IMAGE GENERATIONTECHNICAL FIELD

[0001] This specification relates to data processing, artificial intelligence, and generating images using artificial intelligence.BACKGROUND

[0002] In a computer networked environment such as the Internet, third-party content providers provide third-party content items for display on end-user computing devices. These third-party content items, for example, digital images and video, can be displayed on client devices in the environment. Digital images and video can be used, for example, on the Internet, for remote meetings via video conferencing, high-definition video entertainment, and / or sharing of usergenerated content.

[0003] Recent developments in artificial intelligence and, in particular, generative artificial intelligence have caused user-produced visual content such as digital images to become ubiquitous. For example, various types of images can be generated by using a text-to-image models based on text prompts. However, current generative artificial intelligence models often produce inaccurate and / or low quality images.SUMMARY

[0004] In general, one innovative aspect of the subject matter described in this specification can be embodied in methods that include the actions of obtaining first training data including a set of first training samples, wherein each first training sample includes a first image; training, using the first training data, an image parser model to output structured representations of input images; obtaining a set of second images; processing the second images using the image parser model to generate second training data including a set of second training samples, wherein each second training sample includes a second image and a corresponding structured representation of content items depicted by the second image output by the image parser model based on the second image; training, using the second training data, an image design model to output a structured representation for an input including a set of content items, the set of content items including at least one of text or one or more images; obtaining a new set of content items for a new image; processing the set of content items for the new image using the image design model to generate a corresponding new structured representation for the new image; and generatingAttorney Docket No. t>6113-0768WO1 the new image using the set of content items and the new structured representation. Other implementations of this aspect include corresponding apparatus, systems, and computer programs, configured to perform the aspects of the methods, encoded on computer storage devices.

[0005] These and other embodiments can each optionally include one or more of the following features. Some aspects include distributing the new image to one or more client devices.

[0006] In some aspects, each structured representation indicates locations at which content is depicted by the image corresponding to the structured representation.

[0007] In some aspects, obtaining the first training data includes generating additional first training samples in the first training data by adding noise to the first images, the corresponding structured representations of the content items depicted by the first images, or both. Generating additional first training samples in the first training data by adding noise to the first images, the corresponding structured representations of the content items depicted by the first images, or both can include generating an additional first training sample by adjusting one or more of a color, size, or location of a content item of a given first image.

[0008] In some aspects, the image parser comprises a neural network includes a first sub-neural network trained to split a rendering into multiple layers and a second sub-neural network trained to convert the multiple layers into content items and a corresponding structured representation.

[0009] In some aspects, each first training sample includes a corresponding structured representation of content items depicted by the first image of the first training sample. Training, using the first training data, the image parser model to output structured representations of input images can include training the image parser model using supervised learning.

[0010] In some aspects, training, using the first training data, the image parser model to output structured representations of input images includes training the image parser model to optimize a loss function based on a difference between each first image and a rendered image that is rendered based on a structured representation of the first image output by the image parser model.

[0011] Particular embodiments of the subject matter described in this specification can be implemented so as to realize one or more of the following advantages. The techniquesAttorney Docket No. t>6113-0768WO1 described in this document create robust sets of training data for training artificial intelligence image design models to generate data that is used to create new images that can be used to create digital components that are distributed to users. Using the robust training data described in this document enables the models to create more diverse and higher quality (e.g., improved relevance, aesthetics, and / or performance) image designs than those resulting from training data that is currently available. The described techniques can generate robust training data by training an image parser model to parse static images into its content items (e.g., text, images, logos, buttons, etc.) and a structured representation that indicates characteristics of the content items, the location of the content items in the images. This enables the image design models to be trained using both available training data that includes this rich data about images as well as static images for which such data was previously unavailable. As the number of available static images can greatly exceed the number of images for which rich data is available, this greatly increases the amount of training data and resulting quality of the images generated using models trained using the robust training data. For example, training models using a larger number of training samples typically results in higher quality models, especially when the larger number of training samples are of high quality as described herein.

[0012] Training data can also be generated by using artificial intelligence to generate new images and corresponding structured representations. However, images generated using artificial intelligence are often of lower quality than actual static images obtained from various sources, e.g., from the Internet or digital component providers that publish image digital components. Additionally, images generated using generative artificial intelligence models often include hallucinations and / or other artifacts that would reduce the quality of models trained using these images and the images created based on the outputs of these models. The techniques described in this document include training an image parser model to generate accurate training data that includes high quality images provided as input to the image parser model and the corresponding content items and structured representations output by the image parser model. Using real static images to create a large, robust, and accurate set of training data and then training an image design model using this training data results in a higher quality image design model that can generate higher quality structured representations for a set of new content items than using images and corresponding structured representations created using generative artificial intelligence.Attorney Docket No. t>6113-0768WO1

[0013] Additionally, the static images that are selected for use in generating the training data can be selected such that they are all of high quality and / or obtained from high quality sources, e.g., by obtaining actual image digital components published by digital component providers. Using images that have been published for a particular purpose, e.g., for use in creating digital components, also ensures that the trained image design model is adapted to that particular purpose. Thus, the techniques described in this document solve various problems in training an image design model, including a lack of training data and a lack of techniques for generating high quality training data.

[0014] Training the models using content items and structured representations also increases the quality of the images generated using the models while also reducing the complexity of the models and corresponding computational resources required to process the models relative to models trained using pixel data for images. Thus, the described techniques enable more and better training data than existing solutions, resulting in higher quality and more diverse designs than before. The complexity of the models can be reduced as compared to large language models (LLMs) that include billions of parameters whereas the models described in this document can be implemented as neural networks with substantially fewer parameters (e.g., millions, tens of millions, or even fewer).

[0015] In addition, the training data can be further expanded to include even more diverse designs by adding noise to the training data before training the image parser model. For example, characteristics of content items depicted by an image can be adjusted to create new training samples. In this way, the models can be trained based on many different training samples created from each training image. This increased number of training samples and more diverse training designs also increases the accuracy of the image parser model in parsing static images into content items and structured representations.

[0016] The details of one or more embodiments of the subject matter described in this specification are set forth in the accompanying drawings and the description below. Other features, aspects, and advantages of the subject matter will become apparent from the description, the drawings, and the claims.BRIEF DESCRIPTION OF THE DRAWINGS

[0017] FIG. l is a block diagram of an example environment in which generative artificial intelligence can be implemented to generate images and digital components.Attorney Docket No. t>6113-0768WO1

[0018] FIG. 2 is a block diagram illustrating interactions between an artificial intelligence system, an image parser model, an image design model, and a client device.

[0019] FIG. 3 is a block diagram of an example process for training an image parser model and using the image parser model to generate structured representations of images.

[0020] FIG. 4 is a block diagram of an example process for training an image parser model.

[0021] FIG. 5 is a block diagram of an example process for generating images using an image design model.

[0022] FIG. 6 is a flow chart of an example process of generating images using an image parser model and an image design model.

[0023] FIG. 7 a block diagram of an example computer.

[0024] Like reference numbers and designations in the various drawings indicate like elements.DETAILED DESCRIPTION

[0025] This specification describes techniques for training and using Al models to generate new images. Artificial intelligence (Al) is a segment of computer science that focuses on the creation of models that can perform tasks and actions autonomously, e.g., with little to no human intervention. Al systems can utilize, for example, one or more of machine learning, natural language processing, or computer vision. Machine learning, and its subsets, such as deep learning, focus on developing models that can infer outputs from data. The outputs can include, for example, predictions and / or classifications. Natural language processing focuses on analyzing and generating human language. Computer vision focuses on analyzing and interpreting images and videos. Artificial intelligence systems can include generative models that generate new content, such as images, videos, text, audio, and / or other content, in response to input prompts and / or based on other information.

[0026] The techniques described throughout this specification enable Al to generate training data from static images and use the training data to train image design models to generate image data that can be used to generate new images. In general, a new image can be created using content items that are arranged in a particular manner. For example, the content items can include text, one or more images, logos, graphics, interactive elements (e.g., buttons, icons, etc.), and / or other types of visual content. The image design model can be trained to accept as input a set of content items and output image data that includes a structured representation of the content items. The structured representation can specify the arrangement of each contentAttorney Docket No. t>6113-0768WO1 item to form a new image, the color(s) depicted by each content item, the size of each content item (e.g., in number of pixels or in pixel dimensions), data indicating shape cropping (e.g., irregular shape cropping), decorations (e.g., gradients), geometric elements, font characteristics of text content items (e.g., font type and / or size), transparencies of content items, and / or other characteristics of the image or individual content items in the image. For example, the structured representation can specify the vertical and horizontal coordinates of each content item or the relative position of the content items with respect to each other (e.g., one content item at the top, another content item below that content item, etc ). The vertical coordinates can indicate the boundary of the content item. A new image can be generated, e.g., by a rendering engine, using the set of content items and the structured representation.

[0027] To generate the training data for the image design model, an image parser model can be trained to accept static images (e.g., images without a known structured representation) as input and output the content items depicted by the image and a structured representation of the content items depicted by the image. The content items, corresponding structured representations, and / or the rendering of each image can then be used as training samples to train the image design model. Using the image parser model to generate the training data in this way results in a substantially larger number of training samples relative to using only those images for which a structure representation is already known. In addition, noise techniques can be used to generate additional training samples. For example, noise techniques can be used to generate multiple training samples from a single training sample by adjusting characteristics of a training sample. In a particular example, the noise techniques can include adjusting the location, color, and / or size of one or more content items in the structured representation of a training sample. Expanding the training data in this way results in a more diverse training data that enables the image design model to generate more diverse and higher quality (e.g., more relevant, accurate, and / or aesthetically pleasing) image designs than using the few images for which structured representations are known. For example, the structured representation ensures that there is no misspelled text, gibberish text, imperfect and / or distorted shapes (e.g., circles and squares), etc., similar to vectorized graphics and rasterized pixels.

[0028] FIG. 1 is a block diagram of an example environment 100 in which generative Al can be implemented to generate images and digital components. The example environment 100 includes a network 102, such as a local area network (LAN), a wide area network (WAN), theAttorney Docket No. t>6113-0768WO1Internet, or a combination thereof. The network 102 connects electronic document servers 104, client devices 106, digital component servers 108, and a service apparatus 110. The example environment 100 may include many different electronic document servers 104, client devices 106, and digital component servers 108.

[0029] The service apparatus 110 is configured to provide various services to client devices 106, publishers of electronic documents 150, and / or digital component providers that provide digital components to client devices 106, e.g., using the service apparatus 110. In some implementations, the service apparatus 110 can provide search services by providing responses to search queries received from client devices 106. For example, the services apparatus 110 can include a search engine and / or an Al agent or other chat agent that enables users to interact with the agent over the course of multiple conversational queries and responses. The service apparatus 110 can also distribute digital components to client devices 106 for presentation with the responses and / or with electronic documents 150. For example, another search service computer system can send component requests 112 to the service apparatus 110 and these component requests 112 can include one or more queries. In another example, client devices 106 can send component requests 112 to the service apparatus 110, e.g., in response to loading an electronic document 150 that includes a digital component slot for presenting a digital component, as described in more detail below.

[0030] The service apparatus 110 is configured to generate images, e.g., digital components that include images (“image digital components”). The service apparatus 110 and component requests 112 are described in further detail below.

[0031] As used throughout this document, the phrase “digital component” refers to a discrete unit of digital content or digital information (e.g., a video clip, audio clip, multimedia clip, gaming content, image, text, bullet point, artificial intelligence output, language model output, or another unit of content). A digital component can electronically be stored in a physical memory device as a single file or in a collection of files, and digital components can take the form of video files, audio files, multimedia files, image files, or text files and include advertising information, such that an advertisement is a type of digital component.

[0032] A client device 106 is an electronic device capable of requesting and receiving online resources over the network 102. Example client devices 106 include personal computers, gaming devices, mobile communication devices, digital assistant devices, augmented realityAttorney Docket No. t>6113-0768WO1 devices, virtual reality devices, and other devices that can send and receive data over the network 102. A client device 106 typically includes a user application, such as a web browser, to facilitate the sending and receiving of data over the network 102, but native applications (other than browsers) executed by the client device 106 can also facilitate the sending and receiving of data over the network 102.

[0033] A gaming device is a device that enables a user to engage in gaming applications, for example, in which the user has control over one or more characters, avatars, or other rendered content presented in the gaming application. A gaming device typically includes a computer processor, a memory device, and a controller interface (either physical or visually rendered) that enables user control over content rendered by the gaming application. The gaming device can store and execute the gaming application locally, or execute a gaming application that is at least partly stored and / or served by a cloud server (e.g., online gaming applications). Similarly, the gaming device can interface with a gaming server that executes the gaming application and “streams” the gaming application to the gaming device. The gaming device may be a tablet device, mobile telecommunications device, a computer, or another device that performs other functions beyond executing the gaming application.

[0034] Digital assistant devices include devices that include a microphone and a speaker. Digital assistant devices are generally capable of receiving input by way of voice, and respond with content using audible feedback, and can present other audible information. In some situations, digital assistant devices also include a visual display or are in communication with a visual display (e.g., by way of a wireless or wired connection). Feedback or other information can also be provided visually when a visual display is present. In some situations, digital assistant devices can also control other devices, such as lights, locks, cameras, climate control devices, alarm systems, and other devices that are registered with the digital assistant device.

[0035] As illustrated, the client device 106 is presenting an electronic document 150. An electronic document is data that presents a set of content at a client device 106. Examples of electronic documents include webpages, word processing documents, portable document format (PDF) documents, images, videos, search results pages, and feed sources. Native applications (e.g., “apps” and / or gaming applications), such as applications installed on mobile, tablet, or desktop computing devices are also examples of electronic documents. ElectronicAttorney Docket No. t>6113-0768WO1 documents can be provided to client devices 106 by electronic document servers 104 (“Electronic Doc Servers”).

[0036] For example, the electronic document servers 104 can include servers that host publisher websites. In this example, the client device 106 can initiate a request for a given publisher webpage, and the electronic server 104 that hosts the given publisher webpage can respond to the request by sending machine executable instructions that initiate presentation of the given webpage at the client device 106.

[0037] In another example, the electronic document servers 104 can include app servers from which client devices 106 can download apps. In this example, the client device 106 can download fdes required to install an app at the client device 106, and then execute the downloaded app locally (i.e., on the client device). Alternatively, or additionally, the client device 106 can initiate a request to execute the app, which is transmitted to a cloud server. In response to receiving the request, the cloud server can execute the application and stream a user interface of the application to the client device 106 so that the client device 106 does not have to execute the app itself. Rather, the client device 106 can present the user interface generated by the cloud server’s execution of the app, and communicate any user interactions with the user interface back to the cloud server for processing.

[0038] Electronic documents can include a variety of content. For example, an electronic document 150 can include native content 152 that is within the electronic document 150 itself and / or does not change over time. Electronic documents can also include dynamic content that may change over time or on a per-request basis. For example, a publisher of a given electronic document (e.g., electronic document 150) can maintain a data source that is used to populate portions of the electronic document. In this example, the given electronic document can include a script, such as the script 154, that causes the client device 106 to request content (e.g., a digital component) from the data source when the given electronic document is processed (e.g., rendered or executed) by a client device 106 (or a cloud server). The client device 106 (or cloud server) integrates the content (e.g., digital component) obtained from the data source into the given electronic document to create a composite electronic document including the content obtained from the data source.

[0039] In some situations, a given electronic document (e.g., electronic document 150) can include a digital component script (e.g., script 154) that references the service apparatus 110, orAttorney Docket No. t>6113-0768WO1 a particular service provided by the service apparatus 110. In these situations, the digital component script is executed by the client device 106 when the given electronic document is processed by the client device 106. Execution of the digital component script configures the client device 106 to generate a request for digital components (referred to as a “component request”), which is transmitted over the network 102 to the service apparatus 110. For example, the digital component script can enable the client device 106 to generate a packetized data request including a header and payload data. The component request 112 can include event data specifying features such as a name (or network location) of a server from which the digital component is being requested, a name (or network location) of the requesting device (e.g., the client device 106), and / or information that the service apparatus 110 can use to select one or more digital components, or other content, provided in response to the request. The component request 112 is transmitted, by the client device 106, over the network 102 (e.g., a telecommunications network) to a server of the service apparatus 110.

[0040] The component request 112 can include event data specifying other event features, such as the electronic document being requested and characteristics of locations of the electronic document at which the digital component can be presented. For example, event data specifying a reference (e.g., URL) to an electronic document (e.g., webpage) in which the digital component will be presented, available locations of the electronic documents that are available to present digital components, sizes of the available locations, and / or media types that are eligible for presentation in the locations can be provided to the service apparatus 110.Similarly, event data specifying keywords associated with the electronic document (“document keywords”) or entities (e.g., people, places, or things) that are referenced by the electronic document can also be included in the component request 112 (e.g., as payload data) and provided to the service apparatus 110 to facilitate identification of digital components that are eligible for presentation with the electronic document.

[0041] The event data can also include a search query that was submitted from the client device 106 to obtain a search results page or a response in a conversational user interface. For example, an Al agent or other form of a chat agent can provide a conversational user interface in which users can provide natural language queries, which can be in the form of prompts for a language model, and receive responses to the queries. The user can refine their expression of their informational needs as the conversation progresses and the Al agent can send componentAttorney Docket No. t>6113-0768WO1 requests 1 12 with the updated queries. In such examples, the Al agent can include a user session identifier in the component requests so that the service apparatus 110 can correlate queries included in multiple component requests 112 for the same user session with the Al agent and use this information in generating customized digital components.

[0042] Component requests 112 can also include event data related to other information, such as information that a user of the client device has provided, geographic information indicating a state or region from which the component request was submitted, or other information that provides context for the environment in which the digital component will be displayed (e.g., a time of day of the component request, a day of the week of the component request, a type of device at which the digital component will be displayed, such as a mobile device or tablet device). Component requests 112 can be transmitted, for example, over a packetized network, and the component requests 112 themselves can be formatted as packetized data having a header and payload data. The header can specify a destination of the packet and the payload data can include any of the information discussed above.

[0043] The service apparatus 110 chooses digital components (e.g., third-party content, such as video fdes, audio fdes, images, text, gaming content, augmented reality content, and combinations thereof, which can all take the form of advertising content or non-advertising content) that will be presented with the given electronic document (e.g., at a location specified by the script 154) in response to receiving the component request 112 and / or using information included in the component request 112. In some implementations, choosing a digital component includes choosing a customizable digital component that can be customized based on various data, as described in more detail below.

[0044] In some implementations, a digital component is selected in less than a second to avoid errors that could be caused by delayed selection of the digital component. For example, delays in providing digital components in response to a component request 112 can result in page load errors at the client device 106 or cause portions of the electronic document to remain unpopulated even after other portions of the electronic document are presented at the client device 106.

[0045] Also, as the delay in providing the digital component to the client device 106 increases, it is more likely that the electronic document will no longer be presented at the client device 106 when the digital component is delivered to the client device 106, thereby negativelyAttorney Docket No. t>6113-0768WO1 impacting a user’s experience with the electronic document. Further, delays in providing the digital component can result in a failed delivery of the digital component, for example, if the electronic document is no longer presented at the client device 106 when the digital component is provided. The described techniques are adapted to generate a customized digital component in a short amount of time such that these errors and user experience impact are reduced or eliminated.

[0046] In some implementations, the service apparatus 110 is implemented in a distributed computing system that includes, for example, a server and a set of multiple computing devices 114 that are interconnected and identify and distribute digital component in response to requests 112. The set of multiple computing devices 114 operate together to identify a set of digital components that are eligible to be presented in the electronic document from among a corpus of millions of available digital components (DCi.x). The millions of available digital components can be indexed, for example, in a digital component database 116. Each digital component index entry can reference the corresponding digital component and / or include distribution parameters (DPi-DPx) that contribute to (e.g., trigger, condition, or limit) the distribution / transmission of the corresponding digital component. For example, the distribution parameters can contribute to (e.g., trigger) the transmission of a digital component by requiring that a component request include at least one criterion that matches (e.g., either exactly or with some pre-specified level of similarity) one of the distribution parameters of the digital component.

[0047] In some implementations, the distribution parameters for a particular digital component can include distribution keywords that must be matched (e.g., by electronic documents, document keywords, and / or terms specified in the component request 112) in order for the digital component to be eligible for presentation. Additionally, or alternatively, the distribution parameters can include embeddings that can use various different dimensions of data, such as website details and / or consumption details (e.g., page viewport, user scrolling speed, or other information about the consumption of data). The distribution parameters can also require that the component request 112 include information specifying a particular geographic region (e.g., country or state) and / or information specifying that the component request 112 originated at a particular type of client device (e.g., mobile device or tablet device) in order for the digital component to be eligible for presentation. The distribution parameters can also specify anAttorney Docket No. t>6113-0768WO1 eligibility value (e.g., ranking score, or some other specified value) that is used for evaluating the eligibility of the digital component for distribution / transmission (e.g., among other available digital components).

[0048] The identification of the eligible digital component can be segmented into multiple tasks 117a- 117c that are then assigned among computing devices within the set of multiple computing devices 114. For example, different computing devices in the set 114 can each analyze a different portion of the digital component database 116 to identify various digital components having distribution parameters that match information included in the component request 112. In some implementations, each given computing device in the set 114 can analyze a different data dimension (or set of dimensions) and pass (e.g., transmit) results (Res 1-Res 3) 118a-l 18c of the analysis back to the service apparatus 110. For example, the results 118a- 118c provided by each of the computing devices in the set 114 may identify a subset of digital components that are eligible for distribution in response to the component request and / or a subset of the digital component that have certain distribution parameters. The identification of the subset of digital components can include, for example, comparing the event data to the distribution parameters, and identifying the subset of digital components having distribution parameters that match at least some features of the event data.

[0049] The service apparatus 110 aggregates the results 118a- 118c received from the set of multiple computing devices 114 and uses information associated with the aggregated results to select one or more digital components that will be provided in response to the request 112. For example, the service apparatus 110 can select a set of winning digital components (one or more digital components) based on the outcome of one or more content evaluation processes, as discussed below. In turn, the service apparatus 110 can generate and transmit, over the network 102, reply data 120 (e.g., digital data representing a reply) that enable the client device 106 to integrate the set of winning digital components into the given electronic document, such that the set of winning digital components (e.g., winning third-party content) and the content of the electronic document are presented together at a display of the client device 106. In some implementations, the client device 106 executes instructions included in the reply data 120, which configures and enables the client device 106 to obtain the set of winning digital components from one or more digital component servers 108. For example, the instructions in the reply data 120 can include a network location (e.g., a Uniform Resource Locator (URL))Attorney Docket No. t>6113-0768WO1 and a script that causes the client device 106 to transmit a server request (SR) 121 to the digital component server 108 to obtain a given winning digital component from the digital component server 108. In response to the request, the digital component server 108 will identify the given winning digital component specified in the server request 121 (e.g., within a database storing multiple digital components) and transmit, to the client device 106, digital component data (DC Data) 122 that presents the given winning digital component in the electronic document at the client device 106.

[0050] When the client device 106 receives the digital component data 122, the client device will render the digital component (e.g., third-party content), and present the digital component at a location specified by, or assigned to, the script 154. For example, the script 154 can create a walled garden environment, such as a frame, that is presented within, e.g., beside, the native content 152 of the electronic document 150. In some implementations, the digital component is overlayed over (or adjacent to) a portion of the native content 152 of the electronic document 150, and the service apparatus 110 can specify the presentation location within the electronic document 150 in the reply 120. For example, when the native content 152 includes video content, the service apparatus 110 can specify a location or object within the scene depicted in the video content over which the digital component is to be presented.

[0051] The service apparatus 110 can also include an Al system 160 configured to autonomously generate images, e.g., image digital components, either prior to a component request 112 (e.g., offline) and / or in response to a request 112 (e.g., online or real-time). For example, the Al system 160 can generate images in an offline or online process using content items received from a digital component provider. As described above the content items can include text, one or more images, logos, graphics, interactive elements (e.g., buttons, icons, etc.), and / or other types of visual content. For example, a digital component provider can provide, to the service apparatus 110, a set of content items for use in generating digital components for an item (e.g., a product or service) or for a digital component distribution effort for one or more items.

[0052] The Al system 160 can use one or more Al models, e.g., one or more machine learning models, to generate images based on input prompts and / or content items, which can be received from digital component providers or other users that request image creation by the Al system 160. One or more of the machine learning models trained and / or used by the Al system 160Attorney Docket No. t>6113-0768WO1 can include language models 170, e.g., large language models. A large language model (“LLM”) is a model that is trained to generate and understand human language. LLMs are trained on massive datasets of text and code, and they can be used for a variety of tasks. For example, LLMs can be trained to translate text from one language to another; summarize text, such as web site content, search results, news articles, or research papers; answer questions about text, such as “What is the capital of Georgia?”; create chatbots that can have conversations with humans; and generate creative text, such as poems, stories, and code.

[0053] The language model(s) 170 can be any appropriate language model neural network that receives an input sequence made up of text tokens selected from a vocabulary and auto- regressively generates an output sequence made up of text tokens from the vocabulary. For example, the language model 170 can be a Transformer-based language model neural network or a recurrent neural network-based language model.

[0054] In some situations, a language model 170 can be referred to as an auto-regressive neural network when the neural network used to implement the language model 170 auto-regressively generates an output sequence of tokens. More specifically, the auto-regressively generated output is created by generating each particular token in the output sequence conditioned on a current input sequence that includes any tokens that precede the particular text token in the output sequence, i.e., the tokens that have already been generated for any previous positions in the output sequence that precede the particular position of the particular token, and a context input that provides context for the output sequence.

[0055] For example, the current input sequence when generating a token at any given position in the output sequence can include the input sequence and the tokens at any preceding positions that precede the given position in the output sequence. As a particular example, the current input sequence can include the input sequence followed by the tokens at any preceding positions that precede the given position in the output sequence. Optionally, the input and the current output sequence can be separated by one or more predetermined tokens within the current input sequence.

[0056] More specifically, to generate a particular token at a particular position within an output sequence, the neural network of the language model 170 can process the current input sequence to generate a score distribution, e.g., a probability distribution, that assigns a respective score, e.g., a respective probability, to each token in the vocabulary of tokens. The neural network ofAttorney Docket No. t>6113-0768WO1 the language model 170 can then select, as the particular token, a token from the vocabulary using the score distribution. For example, the neural network of the language model 170 can greedily select the highest-scoring token or can sample, e.g., using nucleus sampling or another sampling technique, a token from the distribution.

[0057] As a particular example, the language model 170 can be an auto-regressive Transformer-based neural network that includes (i) a plurality of attention blocks that each apply a self-attention operation and (ii) an output subnetwork that processes an output of the last attention block to generate the score distribution.

[0058] A language model can have any of a variety of Transformer-based neural network architectures. Examples of such architectures include those described in J. Hoffmann, S. Borgeaud, A. Mensch, E. Buchatskaya, T. Cai, E. Rutherford, D. d. L. Casas, L. A. Hendricks, J. Welbl, A. Clark, et al. Training compute-optimal large language models, arXiv preprint arXiv:2203.15556, 2022; J.W. Rae, S. Borgeaud, T. Cai, K. Millican, J. Hoffmann, H. F. Song, J. Aslanides, S. Henderson, R. Ring, S. Young, E. Rutherford, T. Hennigan, J. Menick, A. Cassirer, R. Powell, G. van den Driessche, L. A. Hendricks, M. Rauh, P. Huang, A. Glaese, J. Welbl, S. Dathathri, S. Huang, J. Uesato, J. Mellor, I. Higgins, A. Creswell, N. McAleese, A. Wu, E. Eisen, S. M. Jayakumar, E. Buchatskaya, D. Budden, E. Sutherland, K. Simonyan, M. Paganini, L. Sifre, L. Martens, X. L. Li, A. Kuncoro, A. Nematzadeh, E. Gribovskaya, D. Donato, A. Lazaridou, A. Mensch, J. Lespiau, M. Tsimpoukelli, N. Grigorev, D. Fritz, T. Sottiaux, M. Pajarskas, T. Pohlen, Z. Gong, D. Toyama, C. de Masson d’Autume, Y. Li, T.Terzi, V. Mikulik, I. Babuschkin, A. Clark, D. de Las Casas, A. Guy, C. Jones, J. Bradbury, M. Johnson, B. A. Hechtman, L. Weidinger, I. Gabriel, W. S. Isaac, E. Lockhart, S. Osindero, L. Rimell, C. Dyer, O. Vinyals, K. Ayoub, J. Stanway, L. Bennett, D. Hassabis, K. Kavukcuoglu, and G. Irving. Scaling language models: Methods, analysis & insights from training gopher. CoRR, abs / 2112.11446, 2021; Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. Exploring the limits of transfer learning with a unified text-to-text transformer. arXiv preprint arXiv: 1910.10683, 2019; Daniel Adiwardana, Minh-Thang Luong, David R. So, Jamie Hall, Noah Fiedel, Romal Thoppilan, Zi Yang, Apoorv Kulshreshtha, Gaurav Nemade, Yifeng Lu, and Quoc V. Le. Towards a human-like open-domain chatbot. CoRR, abs / 2001.09977, 2020; and Tom B Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, ArvindAttorney Docket No. t>6113-0768WO1Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. Language models are fewshot learners. arXiv preprint arXiv:2005.14165, 2020.

[0059] Generally, however, the Transformer-based neural network includes a sequence of attention blocks, and, during the processing of a given input sequence, each attention block in the sequence receives a respective input hidden state for each input token in the given input sequence. The attention block then updates each of the hidden states at least in part by applying self-attention to generate a respective output hidden state for each of the input tokens. The input hidden states for the first attention block are embeddings of the input tokens in the input sequence and the input hidden states for each subsequent attention block are the output hidden states generated by the preceding attention block.

[0060] In this example, the output subnetwork processes the output hidden state generated by the last attention block in the sequence for the last input token in the input sequence to generate the score distribution.

[0061] Generally, because the language model 170 is auto-regressive, the service apparatus 110 can use the same language model 170 to generate multiple different candidate output sequences in response to the same request, e.g., by using beam search decoding from score distributions generated by the language model, using a Sample-and-Rank decoding strategy, by using different random seeds for the pseudo-random number generator that’s used in sampling for different runs through the language model 170 or using another decoding strategy that leverages the auto-regressive nature of the language model 170.

[0062] In some implementations, the language model 170 is pre-trained, e.g., trained on a language modeling task that does not require providing evidence in response to user questions, and the service apparatus 110 (e.g., using Al system 160) causes the language model to generate output sequences according to the pre-determined syntax through natural language prompts in the input sequence.

[0063] For example, the service apparatus 110 (e.g., Al system 160), or a separate training system, pre-trains a language model (e.g., the neural network) on a language modeling task, e.g., a task that requires predicting, given a current sequence of text tokens, the next token that follows the current sequence in the training data. As a particular example, the language model 170 can be pre-trained on a maximum-likelihood objective on a large dataset of text, e.g., text that is publicly available from the Internet or another text corpus.Attorney Docket No. t>6113-0768WO1

[0064] The Al system 160 is configured to train and use one or more Al models, e.g., language models 170, neural networks, etc. to generate image data for new images, e.g., for new image digital components. In some implementations, the Al system 160 trains an image parser model to generate parsed image data and an image design model to generate image data for new image digital components using content items for the images. The parsed image data generated by the image parser model can include, for example, a set of content items depicted in the image and a structured representation of the content items. The image data generated by the image design model can include, for example, a structured representation for an input set of content items provided as input to the image design model. The service apparatus 110 can distribute these image digital components to the client devices 106, e.g., using distribution parameters for the image digital components as described above.

[0065] For example, a digital component provider can provide to the service apparatus 110 a set of content items and distribution parameters for a new digital component distribution effort (e.g., a digital component campaign). The Al system 160 can use the image design model to generate digital components using the set of content items as input. The Al system 160 can then distribute these digital components to the client devices 160 using the distribution parameters.

[0066] FIG. 2 is a block diagram illustrating interactions between an artificial intelligence system 160, an image parser model 222, an image design model 224, and a client device 106. In this example, the Al system 160 includes a training apparatus 202, a digital component apparatus 204, and an image evaluation apparatus 206. Each apparatus 202, 204, and 206 can be implemented as one or more computers in one or more locations. In some implementations, two or more of the apparatus 202, 204, and 206 can be implemented using the same apparatus (e.g., the same one or more computers).

[0067] The Al system 160 can be configured to interact with a memory structure 240 to extract and / or store information and content. In particular, the memory structure 240 can store the digital component database 116, images 242, training data 244, and content items 246.

[0068] The digital component database 116 can store rich data about the digital components indexed in the digital component database 116. For example, the digital component database 116 can store digital component data 122 and / or other data that can be used by a client device 106 to render the digital components. In a particular example, the digital component databaseAttorney Docket No. t>6113-0768WO1116 can store, for at least some of the digital components, a set of content items used to render the digital component and a corresponding structured representation that can indicate to the client device 106 how to arrange the set of content items to render the digital component.

[0069] The images 242 can include static images for which rich data, e.g., a structured representation of content items of the static image, is not available. For example, the Al system 160 can obtain these images 242 in their rendered form, e.g., from online resources or from digital component providers.

[0070] As described below, the Al system 160 can generate at least some of the training data 244 using the static images. The training data 244 can also include the rich data stored in the digital component database 116.

[0071] The content items 246 can include content items received from digital component providers. As described above, a digital component provider can provide a set of content items for use in generating digital components for an item (e.g., a product or service) or for a digital component distribution effort for one or more items. The content items 246 can include text, one or more images, logos, graphics, interactive elements (e.g., buttons, icons, etc.), and / or other types of visual content.

[0072] The training apparatus 202 is configured to train the image parser model 222 and the image design model 224. The image parser model 222 is trained to output parsed image data based on an input image. The parsed image data can include, for example, a set of content items depicted by the input image and a structured representation of the content items in the input image. The structured representation can specify the arrangement of each content item to form a new image, the color(s) depicted by each content item, the size of each content item (e.g., in number of pixels or in pixel dimensions), and / or other characteristics of the image.

[0073] In some implementations, the image parser model 222 is a neural network or language model. The training apparatus 202 can train the image parser model 222 using supervised learning techniques or using reinforcement learning (e.g., RLNN) techniques. An example process of training the image parser model 222 using supervised learning is shown in FIG. 3 and described below. An example process for training the image parser model 222 using RLNN is shown in FIG. 4 and described below.

[0074] In RLNN techniques, the image parser model 222 can be trained to minimize the difference between an input static image and a rendered image that is rendered based on theAttorney Docket No. t>6113-0768WO1 output of the image parser model 222. The image evaluation apparatus 206 can evaluate each input image and rendered image pair to determine a measure of the difference between the two images during this training process. For example, the image evaluation apparatus 206 can evaluate the difference between the pixels of the two images to determine the measure of difference between the two images. The measure of difference can be an aggregation of (e.g., sum, mean, median, or other measure of central tendency of) differences in characteristics (e.g., color, intensity, etc.) between each corresponding pixel of the two images.

[0075] The training apparatus 202 can use the trained image parser model 222 to generate training data 244 for use in training the image design model 224. For example, the training apparatus 222 can provide, to the image parser model 222, inputs 232 in the form of static images. The image parser model 222 is trained to provide, for each input 232, an output 234 that includes parsed image data for the input static image. The parsed image data, which can include a set of content items depicted by the input static image and the structured representation of the set of content items, can be stored along with the static image as training data 244 for training the image design model 224. Generating such rich training data using static images allows for a substantial amount of training data 244 for the image design model 224.

[0076] The image design model 224 is trained to output image data that includes a structured representation of a set of input content items. In some implementations, the image design model 224 is a neural network or language model. The training apparatus 202 can train the image design model 224 using supervised learning techniques. For example, the training apparatus 202 can train the image design model 224 by optimizing a loss function using labeled training samples. Each training sample can include features that include a static image and a set of content items depicted by the image. Each training sample can also include a label that indicates the structured representation of the content items in the rendered static image.

[0077] The digital component apparatus 204 is configured to use the image design model 114 to generate new images, e.g., new image digital components. For example, the digital component apparatus 204 can provide, as an input 236 to the image design model 224, a set of content items for an item or digital component distribution effort. The image design model 224 can provide an output 224 that includes a structured representation (e.g., arrangement and / orAttorney Docket No. t>6113-0768WO1 characteristics) of the set of content items that can be used to generate an image digital component.

[0078] The digital component apparatus 204 includes a rendering engine 205 that is configured to render images based on a set of content items and a structured representation of the content items. For example, the rendering engine 205 can use the structured representation to arrange the set of content items within a bounded area to form an image. If the structured representation indicates other characteristics of the content items, e.g., the color, size, or shape of the content items, the rendering engine 205 can adjust the content items according to their characteristics when rendering the content items within the image.

[0079] In some implementations, the structured representation is in a defined structural format, e.g., a JSON format. The rendering image 205 can be configured to read the appropriate data from the defined structural format and use that data for render the content items within the image.

[0080] In some implementations, the content items are considered layers within an image. For example, when parsing a static image into content items and structured representations, the imager parser model 222 can output a set of layers and a structured representation of the layers. Each layer can be for a particular content item detected in the static image. The structured representation can indicate the location of each layer in the image (e.g., using coordinates or relative locations) and / or the location of each content item within its layer. For example, if the image parser model 222 identifies nine content items and thus nine layers in the static image, the image parser model 222 can output the content items in a layered representation, e.g., a 3x3 layered representation with three rows and three columns. Similarly, the rendering image 205 can be configured to render images based on a layered representation of the content items. Example rendering engines include Lottie, HTML, and Latex, to name a few. The digital component apparatus 204 can render the image in various formats, e.g., HTML, PNG, JPG, GIF, etc.

[0081] The digital component apparatus 204 can store the new image digital components in the digital component database 116 with a reference to their distribution parameters. In some implementations, the digital component apparatus 204 generates digital component data 122 for each new image to enable client devices 106 to render the image digital component. In addition to the rendered image, the digital component data 122 can include a link to a landing page forAttorney Docket No. t>6113-0768WO1 the item corresponding to the digital component, a script that causes the client devices 106 to navigate to the landing page in response to user interaction with the image digital component, and / or other data that can be used by a client device 106 to render the image digital component.

[0082] The services apparatus 110 can distribute the image digital components to client device 106, e.g., in response to component requests 112 received from the client devices 106. If an image digital component is selected for distribution to a client device 106, the service apparatus 110 can provide the digital component data 122 for the image digital component to the client device 106.

[0083] FIG. 3 is a block diagram of an example process 300 for training an image parser model 222 and using the image parser model 122 to generate structured representations of images. Operations of the process 300 can be performed, for example, by the service apparatus 110 of FIG. 1 (e.g., by the Al system 160), or another data processing apparatus. The operations of the process 300 can also be implemented as instructions stored on a computer readable medium, which can be non-transitory. Execution of the instructions, by one or more data processing apparatus, causes the one or more data processing apparatus to perform operations of the process 300. For brevity, the process 300 is described in terms of being performed by a system.

[0084] The system obtains training data 302 for use in training the image parser model 222. This training data 302 can include training samples. Each training sample can include a rendered image, a set of content items depicted in the image, and a structured representation of the content items depicted in the image. The system can obtain the training data 302 from a digital component database 116 or other source of rich data for images.

[0085] The system can optionally expand the training data 302 to generate expanded training data 305 by generating additional training samples using noise techniques. For example, the system can apply noise to images and associated data in the training data 302 by adjusting one or more of a color, size, location, geometric decorations, transparencies, special effects (e.g., water splash parameters), or other characteristic of the image and / or a content item depicted by the image, e.g., by adjusting the structured representation of the content items depicted by the image. In a particular example, the image data 302 can include a training sample that includes an image of a red truck in front of a mountain, the set of content items depicted in this image (e.g., an image of the truck, an image of the mountain, etc.) and a structured representation of the content items. The system can generate additional training samples using this one image byAttorney Docket No. t>6113-0768WO1 adjusting the color of the truck and / or moving the truck to either side of the mountain or on the mountain. By applying noise in this way, the system can generate a substantial amount of additional training samples that can be used to train the image parser model 222.

[0086] For each newly created training sample, the system can render a new image using the adjusted structured representation and / or adjusted content item(s). Each created training sample can include the rendered image and corresponding content items and structured representation.

[0087] The system trains the image parser model 222 using the training data 302 or the expanded training data 205. For example, the system can train the image parser model 222 using supervised learning techniques where the label is the structured representation and set of content items and the features are the rendered image. Thus, the image parser model 222 is trained to output a set of content items and structured representation of the content items for input static images.

[0088] Once trained, the system can use the image parser model 222 to parse static images 312 into content items and corresponding structured representations 320. The system can obtain the input images from one or more of various sources, e.g., from digital component providers, online resources (e.g., web pages), etc. The system can provide each static image as input to the image parser model 222 and receive, for each static image, and output of the image parser model 222. The output for each static image can include the content items depicted by the image and the corresponding structured representations 320 of the content items.

[0089] FIG. 4 is a block diagram of an example process 400 for training an image parser model 222. In this example, the system trains the image parser model 222 using reinforcement learning (e.g., RLNN) techniques. Operations of the process 400 can be performed, for example, by the service apparatus 110 of FIG. 1 (e.g., by the Al system 160), or another data processing apparatus. The operations of the process 400 can also be implemented as instructions stored on a computer readable medium, which can be non-transitory. Execution of the instructions, by one or more data processing apparatus, causes the one or more data processing apparatus to perform operations of the process 400. For brevity, the process 400 is described in terms of being performed by a system.

[0090] The system can train the image parser model 122 based on characteristics of pixels of input images 410 to output the content items and structured representations of the contentAttorney Docket No. t>6113-0768WO1 items. The characteristics of the pixels can include, for example, the color and / or intensity of each pixel in an input image 410. The input images 410 can be static images.[00911 The system can train the image parser model 222 to minimize the difference between an input image 410 and a rendered image 460 that is rendered based on the output of the image parser model 222. For example, the system can train the image parser model 222 by optimizing a loss function based on the difference between pixels of input images 410 and pixels of rendered output images 460.

[0092] During training, the system can process pixels of an input image 410 using the image parser model 222 to output content items 430 and a structured representation 440 of the content items identified by the image parser model 222 as being depicted by the input image 410. The system can then use the rendering engine 205 to render an output image 460 based on the output content items 430 and a structured representation 440 of the content items identified by the image parser model 222.

[0093] The system can then use the image evaluation apparatus 206 to evaluate the input image 410 and the rendered output image 460 to determine a measure of the difference between the two images. For example, the image evaluation apparatus 206 can evaluate the difference between the pixels of the two images to determine the measure of difference between the two images. As described above, the measure of difference can be an aggregation of (e.g., sum, mean, median, or other measure of central tendency of) differences in characteristics (e.g., color, intensity, etc.) between each corresponding pixel of the two images.

[0094] The system can then adjust the parameters (e.g., weights and biases) of the image parser model 222 based on the measure of difference. The system can train the image parser model 222 in this way using many static images until the measure of difference reaches an acceptable stopping point. In some implementations, the stopping point for training the model can be based on straining steps / epochs, loss function values, and / or evaluation metrics.

[0095] FIG. 5 is a block diagram of an example process 500 for generating images using an image design model 224. Operations of the process 500 can be performed, for example, by the service apparatus 110 of FIG. 1 (e.g., by the Al system 160), or another data processing apparatus. The operations of the process 500 can also be implemented as instructions stored on a computer readable medium, which can be non-transitory. Execution of the instructions, by one or more data processing apparatus, causes the one or more data processing apparatus toAttorney Docket No. t>6113-0768WO1 perform operations of the process 500. For brevity, the process 500 is described in terms of being performed by a system.

[0096] The system obtains content items 502 for a new image. In some implementations, the content items are received in the form of vectors or bytes of data representing text, images, graphics, logos, and / or other types of content items. The content items can be received from a digital component provider.

[0097] The system provides the content items as input to an image design model 224 that is trained to output a structured representation for a new image design based on input content items. The image design model 224 can process the input content items 502 and output a structured representation 504 for the content items.

[0098] The rendering engine 205 can then generate a new image 506 using the content items 502 and the structured representation 504. For example, the rendering engine 205 can arrange the content items 502 within an image based on the structured representation 504 output by the image design model 224.

[0099] FIG. 6 is a flow chart of an example process 600 of generating images using an image parser model and an image design model. Operations of the process 600 can be performed, for example, by the service apparatus 110 of FIG. 1 (e.g., by the Al system 160), or another data processing apparatus. The operations of the process 600 can also be implemented as instructions stored on a computer readable medium, which can be non -transitory. Execution of the instructions, by one or more data processing apparatus, causes the one or more data processing apparatus to perform operations of the process 600. For brevity, the process 600 is described in terms of being performed by a system.

[0100] The system obtains first training data that includes a set of first training samples (602). Each first training sample can include a first image. Each first training sample can also include a set of content items depicted by the first image and a structured representation of the content items depicted by the first image.

[0101] The system trains, using the first training data, an image parser model to output structured representations of input images (604). As described above, the system can train the image parser model using supervised learning or reinforcement learning. In some implementations, the system uses expanded training data to train the image parser model. ForAttorney Docket No. t>6113-0768WO1 example, the system can use noise techniques to generate new training samples from the first training samples, as described above.[001021 The system obtains a set of second images (606). The second images can include static images for which rich data (e.g., data indicating content items and structured representation of the content items in the image) is not available.

[0103] The system processes the second images using the image parser model to generate second training data that includes a set of second training samples (608). For example, the system can provide each static image as an input to the image parser model and the image parsed model can output the content items depicted in the image and the structured representation of the content items in the second image. Each second training sample can include a second image and a corresponding structured representation of content items depicted by the second image output by the image parser model based on the second image. Each training sample can also include the set of content items depicted by the second image as output by the image parser model.

[0104] The system trains, using the second training data, an image design model to output a structured representation for an input that includes a set of content items (610). As described above, the image design model can be trained using supervised learning.

[0105] The system obtains a new set of content items for a new image (612). For example, a digital component provider and provide the set of content items for the new image to the system.

[0106] The system processes the set of content items for the new image using the image design model to generate a corresponding new structured representation for the new image (614). For example, the system can provide the set of content items as input to the image design model and receive as an output of the image design model a new structured representation for the content items to generate a new image.

[0107] The system generates the new image using the set of content items and the new structured representation (616). For example, a rendering image can render the content items within the image according to the structured representation output by the image design model.

[0108] FIG. 7 is a block diagram of an example computer system 700 that can be used to perform operations described above. The system 700 includes a processor 710, a memory 720, a storage device 730, and an input / output device 740. Each of the components 710, 720,Attorney Docket No. t>6113-0768WO1730, and 740 can be interconnected, for example, using a system bus 750. The processor 710 is capable of processing instructions for execution within the system 700. In one implementation, the processor 710 is a single-threaded processor. In another implementation, the processor 710 is a multi-threaded processor. The processor 710 is capable of processing instructions stored in the memory 720 or on the storage device 730.

[0109] The memory 720 stores information within the system 700. In one implementation, the memory 720 is a computer-readable medium. In one implementation, the memory 520 is a volatile memory unit. In another implementation, the memory 720 is a nonvolatile memory unit.

[0110] The storage device 730 is capable of providing mass storage for the system 700. In one implementation, the storage device 730 is a computer-readable medium. In various different implementations, the storage device 730 can include, for example, a hard disk device, an optical disk device, a storage device that is shared over a network by multiple computing devices (e.g., a cloud storage device), or some other large capacity storage device.

[0111] The input / output device 740 provides input / output operations for the system 700. In one implementation, the input / output device 740 can include one or more of a network interface devices, e.g., an Ethernet card, a serial communication device, e.g., and RS-232 port, and / or a wireless interface device, e.g., and 802.11 card. In another implementation, the input / output device can include driver devices configured to receive input data and send output data to other devices, e.g., keyboard, printer, display, and other peripheral devices 760. Other implementations, however, can also be used, such as mobile computing devices, mobile communication devices, set-top box television client devices, etc.

[0112] Although an example processing system has been described in FIG. 7, implementations of the subject matter and the functional operations described in this specification can be implemented in other types of digital electronic circuitry, or in computer software, firmware, or hardware, including the structures disclosed in this specification and their structural equivalents, or in combinations of one or more of them.

[0113] An electronic document (which for brevity will simply be referred to as a document) does not necessarily correspond to a file. A document may be stored in a portion of a file that holds other documents, in a single file dedicated to the document in question, or in multiple coordinated files.Attorney Docket No. t>6113-0768WO1

[0114] For situations in which the systems discussed here collect and / or use personal information about users, the users may be provided with an opportunity to enable / disable or control programs or features that may collect and / or use personal information (e.g., information about a user’s social network, social actions or activities, a user’s preferences, or a user’s current location). In addition, certain data may be treated in one or more ways before it is stored or used, so that personally identifiable information associated with the user is removed. For example, a user’s identity may be anonymized so that the no personally identifiable information can be determined for the user, or a user’s geographic location may be generalized where location information is obtained (such as to a city, ZIP code, or state level), so that a particular location of a user cannot be determined.

[0115] Embodiments of the subject matter and the operations described in this specification can be implemented in digital electronic circuitry, or in computer software, firmware, or hardware, including the structures disclosed in this specification and their structural equivalents, or in combinations of one or more of them. Embodiments of the subject matter described in this specification can be implemented as one or more computer programs, i.e., one or more modules of computer program instructions, encoded on computer storage medium for execution by, or to control the operation of, data processing apparatus.Alternatively, or in addition, the program instructions can be encoded on an artificially- generated propagated signal, e.g., a machine-generated electrical, optical, or electromagnetic signal, that is generated to encode information for transmission to suitable receiver apparatus for execution by a data processing apparatus. A computer storage medium can be, or be included in, a computer-readable storage device, a computer-readable storage substrate, a random or serial access memory array or device, or a combination of one or more of them. Moreover, while a computer storage medium is not a propagated signal, a computer storage medium can be a source or destination of computer program instructions encoded in an artificially-generated propagated signal. The computer storage medium can also be, or be included in, one or more separate physical components or media (e.g., multiple CDs, disks, or other storage devices).

[0116] The operations described in this specification can be implemented as operations performed by a data processing apparatus on data stored on one or more computer-readable storage devices or received from other sources.Attorney Docket No. t>6113-0768WO1

[0117] The term “data processing apparatus” encompasses all kinds of apparatus, devices, and machines for processing data, including by way of example a programmable processor, a computer, a system on a chip, or multiple ones, or combinations, of the foregoing. The apparatus can include special purpose logic circuitry, e.g., an FPGA (field programmable gate array) or an ASIC (application-specific integrated circuit). The apparatus can also include, in addition to hardware, code that creates an execution environment for the computer program in question, e.g., code that constitutes processor firmware, a protocol stack, a database management system, an operating system, a cross-platform runtime environment, a virtual machine, or a combination of one or more of them. The apparatus and execution environment can realize various different computing model infrastructures, such as web services, distributed computing and grid computing infrastructures.

[0118] This document refers to a service apparatus. As used herein, a service apparatus is one or more data processing apparatus that perform operations to facilitate the distribution of content over a network. The service apparatus is depicted as a single block in block diagrams. However, while the service apparatus could be a single device or single set of devices, this disclosure contemplates that the service apparatus could also be a group of devices, or even multiple different systems that communicate in order to provide various content to client devices. For example, the service apparatus could encompass one or more of a search system, a video streaming service, an audio streaming service, an email service, a navigation service, an advertising service, a gaming service, or any other service.

[0119] A computer program (also known as a program, software, software application, script, or code) can be written in any form of programming language, including compiled or interpreted languages, declarative or procedural languages, and it can be deployed in any form, including as a stand-alone program or as a module, component, subroutine, object, or other unit suitable for use in a computing environment. A computer program may, but need not, correspond to a fde in a file system. A program can be stored in a portion of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document), in a single file dedicated to the program in question, or in multiple coordinated files (e.g., files that store one or more modules, sub-programs, or portions of code). A computer program can be deployed to be executed on one computer or on multiple computers that are located at one site or distributed across multiple sites and interconnected by a communication network.Attorney Docket No. t>6113-0768WO1

[0120] The processes and logic flows described in this specification can be performed by one or more programmable processors executing one or more computer programs to perform actions by operating on input data and generating output. The processes and logic flows can also be performed by, and apparatus can also be implemented as, special purpose logic circuitry, e.g., an FPGA (field programmable gate array) or an ASIC (application-specific integrated circuit).

[0121] Processors suitable for the execution of a computer program include, by way of example, both general and special purpose microprocessors, and any one or more processors of any kind of digital computer. Generally, a processor will receive instructions and data from a read-only memory or a random access memory or both. The essential elements of a computer are a processor for performing actions in accordance with instructions and one or more memory devices for storing instructions and data. Generally, a computer will also include, or be operatively coupled to receive data from or transfer data to, or both, one or more mass storage devices for storing data, e.g., magnetic, magneto-optical disks, or optical disks. However, a computer need not have such devices. Moreover, a computer can be embedded in another device, e.g., a mobile telephone, a personal digital assistant (PDA), a mobile audio or video player, a game console, a Global Positioning System (GPS) receiver, or a portable storage device (e.g., a universal serial bus (USB) flash drive), to name just a few. Devices suitable for storing computer program instructions and data include all forms of non-volatile memory, media and memory devices, including by way of example semiconductor memory devices, e.g., EPROM, EEPROM, and flash memory devices; magnetic disks, e.g., internal hard disks or removable disks; magneto-optical disks; and CD-ROM and DVD-ROM disks. The processor and the memory can be supplemented by, or incorporated in, special purpose logic circuitry.

[0122] To provide for interaction with a user, embodiments of the subject matter described in this specification can be implemented on a computer having a display device, e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor, for displaying information to the user and a keyboard and a pointing device, e.g., a mouse or a trackball, by which the user can provide input to the computer. Other kinds of devices can be used to provide for interaction with a user as well; for example, feedback provided to the user can be any form of sensory feedback, e.g., visual feedback, auditory feedback, or tactile feedback; and input from the user can be received in any form, including acoustic, speech, or tactile input. In addition, aAttorney Docket No. t>6113-0768WO1 computer can interact with a user by sending documents to and receiving documents from a device that is used by the user; for example, by sending web pages to a web browser on a user’s client device in response to requests received from the web browser.

[0123] Embodiments of the subject matter described in this specification can be implemented in a computing system that includes a back-end component, e.g., as a data server, or that includes a middleware component, e.g., an application server, or that includes a frontend component, e.g., a client computer having a graphical user interface or a Web browser through which a user can interact with an implementation of the subject matter described in this specification, or any combination of one or more such back-end, middleware, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication, e.g., a communication network. Examples of communication networks include a local area network (“LAN”) and a wide area network (“WAN”), an internetwork (e.g., the Internet), and peer-to-peer networks (e.g., ad hoc peer-to-peer networks).

[0124] The computing system can include clients and servers. A client and server are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other. In some embodiments, a server transmits data (e.g., an HTML page) to a client device (e.g., for purposes of displaying data to and receiving user input from a user interacting with the client device). Data generated at the client device (e.g., a result of the user interaction) can be received from the client device at the server.

[0125] While this specification contains many specific implementation details, these should not be construed as limitations on the scope of any inventions or of what may be claimed, but rather as descriptions of features specific to particular embodiments of particular inventions. Certain features that are described in this specification in the context of separate embodiments can also be implemented in combination in a single embodiment. Conversely, various features that are described in the context of a single embodiment can also be implemented in multiple embodiments separately or in any suitable subcombination.Moreover, although features may be described above as acting in certain combinations and even initially claimed as such, one or more features from a claimed combination can in someAttorney Docket No. t>6113-0768WO1 cases be excised from the combination, and the claimed combination may be directed to a subcombination or variation of a subcombination.[001261 Similarly, while operations are depicted in the drawings in a particular order, this should not be understood as requiring that such operations be performed in the particular order shown or in sequential order, or that all illustrated operations be performed, to achieve desirable results. In certain circumstances, multitasking and parallel processing may be advantageous. Moreover, the separation of various system components in the embodiments described above should not be understood as requiring such separation in all embodiments, and it should be understood that the described program components and systems can generally be integrated together in a single software product or packaged into multiple software products.

[0127] Thus, particular embodiments of the subject matter have been described. Other embodiments are within the scope of the following claims. In some cases, the actions recited in the claims can be performed in a different order and still achieve desirable results. In addition, the processes depicted in the accompanying figures do not necessarily require the particular order shown, or sequential order, to achieve desirable results. In certain implementations, multitasking and parallel processing may be advantageous.

[0128] What is claimed is:

Claims

Attorney Docket No. t>6113-0768WO1CLAIMS1. A method comprising: obtaining first training data comprising a set of first training samples, wherein each first training sample comprises a first image; training, using the first training data, an image parser model to output structured representations of input images; obtaining a set of second images; processing the second images using the image parser model to generate second training data comprising a set of second training samples, wherein each second training sample comprises a second image and a corresponding structured representation of content items depicted by the second image output by the image parser model based on the second image; training, using the second training data, an image design model to output a structured representation for an input comprising a set of content items, the set of content items comprising at least one of text or one or more images; obtaining a new set of content items for a new image; processing the set of content items for the new image using the image design model to generate a corresponding new structured representation for the new image; and generating the new image using the set of content items and the new structured representation.

2. The method of claim 1, further comprising distributing the new image to one or more client devices.

3. The method of claim 1 or 2, wherein each structured representation indicates locations at which content is depicted by the image corresponding to the structured representation.

4. The method of any preceding claim, wherein obtaining the first training data comprises generating additional first training samples in the first training data by adding noise to the first images, the corresponding structured representations of the content items depicted by the first images, or both.Attorney Docket No. t>6113-0768WO15. The method of claim 4, wherein generating additional first training samples in the first training data by adding noise to the first images, the corresponding structured representations of the content items depicted by the first images, or both comprises generating an additional first training sample by adjusting one or more of a color, size, or location of a content item of a given first image.

6. The method of any preceding claim, wherein the image parser comprises a neural network comprises a first sub-neural network trained to split a rendering into multiple layers and a second sub-neural network trained to convert the multiple layers into content items and a corresponding structured representation.

7. The method of any preceding claim, wherein: each first training sample comprises a corresponding structured representation of content items depicted by the first image of the first training sample; and training, using the first training data, the image parser model to output structured representations of input images comprises training the image parser model using supervised learning.

8. The method of any one of claims 1 to 6, wherein training, using the first training data, the image parser model to output structured representations of input images comprises training the image parser model to optimize a loss function based on a difference between each first image and a rendered image that is rendered based on a structured representation of the first image output by the image parser model.

9. A system comprising: one or more processors; and one or more storage devices storing instructions that, when executed by the one or more processors, cause the one or more processors to carry out the method of any preceding claim.Attorney Docket No. t>6113-0768WO110. A computer readable storage medium carrying instructions that, when executed by one or more processors, cause the one or more processors to carry out the method of any one of claims 1 to 8.

11. A computer program product comprising instructions which, when executed by one or more computers, cause the one or more computers to carry out the steps of the method of any of claims 1 to 8.

Citation Information

Patent Citations

  • Utilizing a transformer-based generative language model to generate digital design document variations

    US20230305690A1