Artificial intelligence for efficient image editing

Through the combination of multimodal models and language models, images are evaluated and edited to meet specific conditions, and the problems of inefficient image editing and privacy risks in the prior art are solved, and efficient and accurate image editing is achieved.

CN120303690APending Publication Date: 2025-07-11GOOGLE LLC
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202480004917.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-11-09
Filing Date
2024-10-07
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

In the prior art, it is difficult for users to efficiently edit images to satisfy specific conditions, resulting in waste of computing resources and network bandwidth and privacy risks, and different types of images or projects may require different conditional settings.

Method used

Using a combination of multimodal models and language models, information redundancy and errors of a single model are avoided by evaluating why images fail to meet conditions and generating image editing prompts.

Benefits of technology

It realizes efficient editing of images to meet conditions, reduces the consumption of computing resources and network bandwidth, improves the accuracy and quality of image generation, and reduces privacy risks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120303690A_ABST
    Figure CN120303690A_ABST
Patent Text Reader

Abstract

Methods, systems, and apparatus, including computer programs encoded on a computer storage medium, are described for enabling artificial intelligence to generate updated images based on one or more conditions. In one aspect, a method includes receiving data indicating that a first image violates one or more conditions. In response to receiving the data indicating that the first image violates the one or more conditions, an image editing prompt is generated that instructs the image editing model to edit the first image to meet the one or more conditions. The image editing prompt and the first image are provided as inputs to the image editing model. A second image is received as an output of the image editing model. The second image is provided to one or more devices.
Need to check novelty before this filing date? Find Prior Art

Description

BACKGROUND OF THE INVENTION

[0001] This specification relates to data processing, artificial intelligence, and using artificial intelligence to generate images.

[0002] Advances in machine learning have enabled artificial intelligence to be implemented in more applications. For example, large language models have been implemented to allow for image editing. This allows for more efficient image editing using the provided information associated with the image. SUMMARY OF THE INVENTION

[0003] In general, one innovative aspect of the subject matter described in this specification can be embodied in a method that includes the following actions: receiving data indicating that a first image violates one or more conditions; in response to receiving data indicating that the first image violates one or more conditions, generating an image editing prompt that instructs an image editing model to edit the first image to meet the one or more conditions, the generating including: generating explanation data indicating the location of content within the image that violates the one or more conditions; and using a language model to generate the image editing prompt based on the explanation data and the one or more conditions; providing the image editing prompt and the first image as inputs to the image editing model; receiving a second image as an output of the image editing model; and providing the second image to one or more devices. Other implementations of this aspect include corresponding devices, systems, and computer programs configured to perform the aspects of the method encoded on a computer storage device.

[0004] These and other embodiments may each optionally include one or more of the following features. In some aspects, generating the explanation data includes providing the image and the one or more conditions to a multimodal model that is trained to identify the locations within the image that violate the input conditions.

[0005] In some aspects, generating the explanation data includes providing the image and the one or more conditions to a multimodal model that is trained to predict whether the image violates the input conditions and output data indicating the locations of content within the image that are likely to violate at least one of the input conditions.

[0006] In some aspects, generating the explanation data includes: providing the first image to a first machine learning model that is trained to generate an image caption for the image; receiving the image caption for the image from the first machine learning model; providing the image caption and the one or more conditions to a second machine learning model that is trained to output explanation data for the image based on the input image caption and the input conditions; and receiving the explanation data for the first image from the second machine learning model.

[0007] In some aspects, the interpreted data includes a location indicator. The location indicator indicates the location of the content in the first image that is determined to violate at least one of one or more conditions. The location indicator may include a bounding box depicted around the content in the first image that is determined to violate at least one of one or more conditions. The location indicator may include coordinates defining the bounding box around the content in the first image that is determined to violate at least one of one or more conditions. Providing the image editing hint and the first image as an input to the image editing model may include providing the location indicator to the image editing model.

[0008] In some aspects, the interpreted data includes an explanation of why the first image violates one or more conditions. In some aspects, the image editing hint includes at least a portion of the interpreted data.

[0009] In some aspects, using a language model to generate an image editing hint based on the interpreted data and one or more conditions includes using the interpreted data to generate a hint for the language model and providing the hint to the language model. The hint may include an instruction to instruct the language model to generate an image editing hint based on the interpreted data and data defining each condition violated by the first image. The interpreted data includes the name of each condition violated by the first image.

[0010] In some aspects, generating an image editing hint includes: obtaining a hint template adapted to the image editing model; and filling the hint with at least a portion of the interpreted data, the at least a portion of the interpreted data including the name of the condition that the first image is determined to violate. The image editing hint output by the language model is adapted to the image editing model.

[0011] Certain embodiments of the subject matter described in this specification can be implemented to achieve one or more of the following advantages. The techniques described in this document enable artificial intelligence (AI) to be used to generate updated images based on data related to one or more conditions. An AI system can use one or more machine learning models to evaluate an image that fails to meet the conditions and determine why the image fails to meet the conditions and / or which part of the image causes the image to not meet the conditions. Without the described techniques, it may be difficult for a user to determine why an image is rejected by a system that evaluates images based on conditions and attempt to modify the image in various ways to try to meet the conditions. This can lead to unnecessary modifications by the user, which result in lower-quality images. This can also lead to the user uploading multiple versions of the image to the system in an attempt to find a version that meets the conditions, thereby imposing an unnecessary burden on the system and the network of the device connecting the system to the user and thus wasting computing resources and network bandwidth. Additionally, uploading multiple versions of the image can lead to data privacy / confidentiality issues because continuous uploading allows adversarial attacks on the machine learning model to learn how the model works. Even worse, some systems employ many different conditions for different types of images or images for different types of items. Thus, an image may meet the conditions for one type of item (e.g., one type of product) but may not meet the conditions for another type of item.

[0012] The techniques described herein can address these problems by using AI to determine a set of reasons why an image fails to meet the conditions and editing the image so that the image meets the conditions. In this way, the user may be able to obtain a compliant image by uploading the image only once, thereby saving resources that would otherwise be wasted on uploading and evaluating multiple images. For example, a user can send a first image to the system, and the system can determine whether the first image violates the conditions. If so, the system can use a machine learning model (e.g., an image editing model) to update the image so that the updated image does not violate the conditions, thus allowing the user to efficiently submit the image to the system without having to upload additional images that the user hopes will be compliant. In this way, a compliant image can be generated without having to upload multiple images, which reduces the computational burden on the system that evaluates the images (e.g., the processing cycles used to evaluate the images, the data storage used to store the images, etc.), reduces the amount of network bandwidth consumed and the associated burden on network resources for sending multiple images, and reduces the computational burden on the user's device when modifying the image and sending the image to the system that evaluates the images.

[0013] A chain of prompts to one or more machine learning models can be used to evaluate an image to determine a set of reasons why the image fails to meet the conditions and to generate an updated image based on the evaluation. In this way, the tasks are separated, and models trained for specific tasks can be used to generate higher-quality outputs that take into account the outputs of previous models. This enables the system to generate higher-quality explanations of why the image fails to meet the conditions and to generate high-quality specialized image editing prompts that accurately instruct an image editing model to correct the image. Separating the process into multiple discrete AI tasks can reduce the amount of information provided to the AI model for each task, which prevents hallucinations and other AI model errors that often occur when large amounts of information are provided as input to the AI model.

[0014] Multiple machine learning models can be used to evaluate and edit an image based on conditions. For example, a multimodal model with question-and-answer and image generation capabilities can evaluate the image and a set of conditions and output explanation data indicating why the image does not meet the conditions and / or which parts of the image cause the image to not meet the conditions. A language model can then generate an image editing prompt based on the explanation data, and an image editing model (e.g., a text-to-image model) can generate an updated version of the image based on the prompt and the image. Using multiple models in this way achieves more accurate outputs and thus higher-quality images compared to using a single model to perform all these different tasks. Additionally, since multiple models rather than a single model are used to handle all the tasks in the process, the system can scale each corresponding model for deployment to generate higher-quality images while efficiently managing resources. In particular, compared to other models (e.g., a model for generating an image editing prompt for an image editing model), the image editing model can be relatively large and more complex, which can result in higher-quality images, higher throughput, reduced memory usage, and reduced latency.

[0015] Using a multimodal model to generate explanatory data and using a language model to generate prompts for an image editing model enables the use of off-the-shelf, general-purpose, or other pre-trained image editing models without customizing or retraining the image editing model. Instead, the techniques described herein can include generating specialized prompts using an instruction language model that are specifically adapted to the image editing model. This thus eliminates the need to adapt or retrain the image editing model to accept inputs generated by another model while ensuring that the image editing model generates high-quality edited images. Using a combination of models in this way for the first time enables an entire AI system to generate a large set of qualified edited images, thereby reducing the number of images uploaded for evaluation and the number of evaluations performed on the images and providing the various computational savings described above and elsewhere herein. Accordingly, the techniques described herein provide a particular application of AI models and a hint for AI models to solve problems that arise in applying conditions to images and generating images that meet the conditions and more generally in the field of automated image generation.

[0016] Details of one or more embodiments of the subject matter described in this specification are set forth in the accompanying drawings and the following description. Other features, aspects, and advantages of the subject matter will become apparent from the description, the drawings, and the claims. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Figure 1 is a block diagram of an example environment for evaluating and using artificial intelligence to edit images based on conditions.

[0018] Figure 2 is a block diagram showing the interaction between an artificial intelligence system, a multimodal model, a language model, and an image editing model.

[0019] Figure 3 is a flowchart of an example process for generating an image based on one or more conditions.

[0020] Figure 4 is a block diagram of an example computer.

[0021] Like reference numerals and names in the various figures indicate like elements. DETAILED DESCRIPTION

[0022] This specification describes techniques for enabling artificial intelligence to generate updated images based on one or more conditions. For example, the techniques can be used to evaluate an image that does not meet one or more conditions and edit the image such that the edited image meets the conditions. Artificial intelligence (AI) is a branch of computer science focused on creating agents that can learn and act autonomously (e.g., without human intervention). Artificial intelligence can leverage machine learning, which focuses on developing algorithms that can learn from data, natural language processing, which focuses on understanding and generating human language, and / or computer vision, which is a field focused on understanding and interpreting images and videos.

[0023] The techniques described throughout this specification enable an AI model to edit an image that fails to meet one or more conditions such that the edited image meets the conditions. The conditions can be the policy conditions of an entity (e.g., an entity that distributes digital components of an image to users) that makes the image viewable by others. Generally speaking, an AI system can receive data indicating that an image violates a condition, and the AI system can use an image editing model to generate an updated version of the image, e.g., by editing the image, that does not violate the condition. The AI system can use one or more machine learning models (e.g., a language model and / or a multimodal model, which can be a language model or other type of multimodal model) to evaluate the image to determine why the image fails to meet the condition and generate an image editing prompt that instructs an image editing model (which can also be a language model) to edit the image in a particular way such that the edited image meets the condition. The AI system can provide the image editing prompt and the image to the image editing model, and the image editing model can edit the image and output an updated version of the image.

[0024] Using AI to evaluate an image based on conditions and generate an image editing prompt as described herein enables the creation of a specialized image editing prompt that instructs an image editing model to appropriately edit the image, which results in an image that meets the conditions without unnecessary editing, thereby producing a high-quality image that is as close as possible to the original image while meeting the conditions. Additionally, providing the specialized prompt to a language model ensures that the image editing prompt provided to the image editing model is adapted to the image editing model, such that the image editing model does not have to be adapted or retrained for use with the language model. Editing an image using AI in this manner reduces the computational resources wasted on generating and uploading multiple versions of an image and evaluating the multiple versions until one that meets the conditions is found. This also increases the likelihood of creating an image that conforms to the image as opposed to a user editing the image without fully understanding the conditions that the image does not meet. This all contributes to a system that can more quickly create an updated image that meets the conditions such that the updated image can be created and supplied in a real-time interactive environment, such as in response to a user search query or a component request for a digital component to be displayed at the user's device.

[0025] As used throughout this document, the phrase "digital component" refers to a discrete unit of digital content or digital information (e.g., a video clip, an audio clip, a multimedia clip, game content, an image, text, a bullet point, an artificial intelligence output, a language model output, or another unit of content). Digital components can be electronically stored in a physical memory device as a single file or as a collection of files, and digital components can take the form of a video file, an audio file, a multimedia file, an image file, or a text file, and include advertising information such that an advertisement is a type of digital component.

[0026] Figure 1 is a block diagram of an example environment 100 that evaluates and uses artificial intelligence to edit an image based on conditions. The example environment 100 includes a network 102, such as a local area network (LAN), a wide area network (WAN), the Internet, or a combination thereof. The network 102 connects an electronic document server 104, a user device 106, a digital component server 108, and a service device 110. The example environment 100 can include many different electronic document servers 104, user devices 106, and digital component servers 108.

[0027] The service device 110 is configured to provide various services to the client device 106 and / or the publisher of the electronic document 150. In some implementations, the service device 110 may provide a search service by providing a response to a search query received from the client device 106. For example, the service device 110 may include a search engine and / or an AI agent or other chat agent that enables a user to interact with the agent during a plurality of conversational queries and responses. The service device 110 may also distribute digital components to the client device 106 for presentation with the response and / or with the electronic document 150. For example, another search service computer system may send a component request 112 to the service device 110, and these component requests 112 may include one or more queries. The service device 110 and the component request 112 will be described in further detail below.

[0028] The client device 106 is an electronic device capable of requesting and receiving online resources over the network 102. Example client devices 106 include personal computers, gaming devices, mobile communication devices, digital assistant devices, augmented reality devices, virtual reality devices, and other devices that can send and receive data over the network 102. The client device 106 typically includes a user application (such as a web browser) to facilitate sending and receiving data over the network 102, but native applications executed by the client device 106 may also facilitate sending and receiving data over the network 102.

[0029] A gaming device is a device that enables a user to participate in a gaming application. For example, in the device, the user can control one or more characters, avatars, or other rendered content presented in the gaming application. A gaming device typically includes a computer processor, a memory device, and a controller interface (physical or visually rendered) that enables the user to control the content rendered by the gaming application. The gaming device may store and execute the gaming application locally, or execute a gaming application that is at least partially stored and / or served by a cloud server (e.g., an online gaming application). Similarly, the gaming device may interface with a gaming server that executes the gaming application and "streams" the gaming application to the gaming device. The gaming device may be a tablet device, a mobile telecommunications device, a computer, or another device that performs other functions in addition to executing the gaming application.

[0030] The digital assistant device includes a device having a microphone and a speaker. The digital assistant device is generally capable of receiving input via voice, and responding to content using audible feedback, and may present other audible information. In some cases, the digital assistant device also includes a visual display or communicates with a visual display (e.g., via a wireless or wired connection). When there is a visual display, feedback or other information may also be provided visually. In some cases, the digital assistant device may also control other devices, such as lights, locks, cameras, climate control devices, alarm systems, and other devices registered with the digital assistant device.

[0031] As shown, client device 106 is presenting electronic document 150. An electronic document is a collection of data that presents content at client device 106. Examples of electronic documents include web pages, word processing documents, Portable Document Format (PDF) documents, images, videos, search result pages, and feeds. Native applications (e.g., "apps" and / or game applications) such as those installed on mobile, tablet, or desktop computing devices are also examples of electronic documents. The electronic document may be provided to client device 106 by an electronic document server 104 ("Electronic Doc Servers").

[0032] For example, electronic document server 104 may include a server hosting a publisher website. In this example, client device 106 may initiate a request for a given publisher web page, and the electronic server 104 hosting the given publisher web page may respond to the request by sending machine-executable instructions that initiate the presentation of the given web page at client device 106.

[0033] In another example, electronic document server 104 may include an app server from which client device 106 can download an app. In this example, client device 106 may download the files required to install the app at client device 106 and then execute the downloaded app locally (i.e., on the client device). Alternatively or additionally, client device 106 may initiate a request to execute the app, which is sent to a cloud server. In response to receiving the request, the cloud server may execute the application and stream the user interface of the application to client device 106 such that client device 106 does not have to execute the app itself. Instead, client device 106 may present the user interface generated by the cloud server's execution of the app and communicate any user interactions with the user interface back to the cloud server for processing.

[0034] An electronic document can include a variety of content. For example, the electronic document 150 can include native content 152 that is within the electronic document 150 itself and / or does not change over time. The electronic document can also include dynamic content that can change over time or based on each request. For example, a publisher of a given electronic document (e.g., the electronic document 150) can maintain a data source for populating portions of the electronic document. In this example, a given electronic document can include a script, such as the script 154, that causes the client device 106 (or a cloud server) to request content (e.g., digital components) from the data source when the given electronic document is processed (e.g., rendered or executed) by the client device 106. The client device 106 (or a cloud server) integrates the content (e.g., digital components) obtained from the data source into the given electronic document to create a composite electronic document that includes the content obtained from the data source.

[0035] In some cases, a given electronic document (e.g., the electronic document 150) can include a digital component script (e.g., the script 154) that references the service device 110 or a particular service provided by the service device 110. In these cases, the digital component script is executed by the client device 106 when the given electronic document is processed by the client device 106. The execution of the digital component script configures the client device 106 to generate a request for a digital component 112 (referred to as a "component request") that is sent over the network 102 to the service device 110. For example, the digital component script can enable the client device 106 to generate a packetized data request that includes a header and payload data. The component request 112 can include event data specifying characteristics such as the name (or network location) of the server from which the digital component is being requested, the name (or network location) of the requesting device (e.g., the client device 106), and / or information that the service device 110 can use to select one or more digital components or other content to provide in response to the request. The component request 112 is sent by the client device 106 over the network 102 (e.g., a telecommunications network) to the server of the service device 110.

[0036] The component request 112 may include event data specifying other event characteristics, such as the electronic document being requested and characteristics of the location where the digital component may be presented. For example, event data specifying a reference (e.g., URL) to an electronic document (e.g., a web page) in which the digital component will be presented, an available location in the electronic document for presenting the digital component, the size of the available location, and / or the media type eligible to be presented at the location may be provided to the service device 110. Similarly, event data specifying keywords associated with the electronic document ("document keywords") or entities referenced by the electronic document (e.g., a person, place, or thing) may also be included in the component request 112 (e.g., as payload data) and provided to the service device 110 to facilitate identification of digital components eligible to be presented with the electronic document. The event data may also include a search query submitted from the client device 106 to obtain a search result page.

[0037] The component request 112 may also include event data related to other information, such as information provided by a user of the client device, geographic information indicating the state or region from which the component request is submitted, or other information providing the context of the environment in which the digital component will be displayed (e.g., the time of day of the component request, the date of the week of the component request, the type of device on which the digital component will be displayed, such as a mobile device or a tablet device). The component request 112 may be sent, for example, over a packetized network, and the component request 112 itself may be formatted as packetized data having a header and payload data. The header may specify the destination of the packet, and the payload data may include any of the information discussed above.

[0038] In response to receiving the component request 112 and / or using the information included in the component request 112, the service device 110 selects digital components (e.g., third-party content, such as video files, audio files, images, text, game content, augmented reality content, and combinations thereof, all of which may take the form of advertising content or non-advertising content) to be presented with a given electronic document (e.g., at the location specified by the script 154). In some implementations, selecting the digital components includes selecting the digital components based on text characteristics.

[0039] In some implementations, the digital components are selected in less than one second to avoid errors that may result from a delayed selection of the digital components. For example, a delay in providing the digital components in response to the component request 112 may cause a page load error at the client device 106 or result in parts of the electronic document remaining unfilled even after other parts of the electronic document are presented at the client device 106. The techniques described are adapted to generate digital components in a short amount of time such that these errors and user experience impacts are reduced or eliminated.

[0040] Moreover, as the latency in providing digital components to the client device 106 increases, it becomes more likely that the electronic document will no longer be presented at the client device 106 when the digital components are delivered to the client device 106, thereby negatively affecting the user's experience of the electronic document. Additionally, the latency in providing digital components may result in the failure of delivery of the digital components, for example, if the electronic document is no longer presented at the client device 106 when the digital components are provided.

[0041] In some implementations, the service device 110 is implemented in a distributed computing system that includes a set of, for example, servers and multiple computing devices 114, which are interconnected and identify and distribute digital components in response to requests 112. The set of multiple computing devices 114 operate together to identify a set of digital components eligible for presentation in an electronic document from a corpus of millions of available digital components (DC 1-x ). The millions of available digital components may be indexed, for example, in a digital component database 116. Each digital component index entry may reference the corresponding digital component and / or include distribution parameters (DP1 to DP x ) that facilitate (e.g., trigger, condition, or limit) the distribution / transmission of the corresponding digital component. For example, the distribution parameters may facilitate (e.g., trigger) the transmission of a digital component by requiring that the component request include at least one criterion that matches (e.g., exactly or at some pre-specified level of similarity) one of the distribution parameters of the digital component.

[0042] In some implementations, the distribution parameters of a particular digital component may include distribution keywords that must be matched (e.g., match the electronic document, document keywords, or terms specified in the component request 112) for the digital component to be eligible for presentation. Additionally or alternatively, the distribution parameters may include embeddings that can use various different data dimensions, such as website details and / or consumption details (e.g., page viewport, user scroll speed, or other information regarding data consumption). The distribution parameters may also require that the component request 112 include information specifying a particular geographic region (e.g., country or state) and / or information specifying that the component request 112 originates from a particular type of client device (e.g., a mobile device or a tablet device) for the digital component to be eligible for presentation. The distribution parameters may also specify an eligibility value (e.g., a ranking score, or some other specified value) that is used to evaluate the eligibility of the digital component for distribution / transmission (e.g., as compared to other available digital components).

[0043] The identification of eligible digital components can be segmented into multiple tasks 117a through 117c, which are then assigned among the computing devices within a set of multiple computing devices 114. For example, different computing devices 114 in the set can each analyze a different portion of the digital component database 116 to identify various digital components having distribution parameters that match the information included in the component request 112. In some implementations, each given computing device 114 in the set can analyze different data dimensions (or sets of dimensions) and pass (e.g., send) the results (Res 1 through Res 3) 118a through 118c of the analysis back to the service device 110. For example, the results 118a through 118c provided by each of the computing devices 114 in the set can identify subsets of digital components eligible for distribution in response to the component request and / or subsets of digital components having certain distribution parameters. The identification of subsets of digital components can include, for example, comparing event data with distribution parameters and identifying subsets of digital components having distribution parameters that match at least some of the characteristics of the event data.

[0044] The service device 110 aggregates the results 118a through 118c received from the set of multiple computing devices 114 and uses the information associated with the aggregated results to select one or more digital components to be provided in response to the request 112. For example, the service device 110 can select a set of winning digital components (one or more digital components) based on the results of one or more content evaluation processes, as described below. Further, the service device 110 can generate and send, over the network 102, reply data 120 (e.g., digital data representing a reply) that enables the client device 106 to integrate the set of winning digital components into a given electronic document such that the set of winning digital components (e.g., winning third-party content) and the content of the electronic document are presented together at the display of the client device 106.

[0045] In some implementations, the client device 106 executes the instructions included in the reply data 120, which configure the client device 106 and enable it to obtain the set of winning digital components from one or more digital component servers 108. For example, the instructions in the reply data 120 can include a network location (e.g., a Uniform Resource Locator (URL)) and a script that causes the client device 106 to send a server request (SR) 121 to the digital component server 108 to obtain a given winning digital component from the digital component server 108. In response to the request, the digital component server 108 will identify the given winning digital component specified in the server request 121 (e.g., within a database storing multiple digital components) and send digital component data (DC data) 122 to the client device 106, which presents the given winning digital component in the electronic document at the client device 106.

[0046] When client device 106 receives digital component data 122, the client device will render the digital component (e.g., third-party content) and present the digital component at the location specified by or assigned to script 154. For example, script 154 can create a walled garden environment (such as a frame) that is presented within the native content 152 of electronic document 150, e.g., presented beside the native content. In some implementations, the digital component overlays (or is adjacent to) a portion of the native content 152 of electronic document 150, and service device 110 can specify the presentation location within electronic document 150 in response 120. For example, when native content 152 includes video content, service device 110 can specify a location or object within the scene depicted in the video content, and the digital component will be presented above the location or object.

[0047] Service device 110 may also include an artificial intelligence system 160 that is configured to autonomously generate digital components before request 112 (e.g., offline) and / or in response to request 112 (e.g., online or in real time). As described in more detail throughout this specification, the artificial intelligence (“AI”) system 160 can collect online content about a particular entity (e.g., a digital component provider or another entity) and use one or more language models 170 to summarize the collected online content, and the one or more language models may include large language models.

[0048] A large language model (“LLM”) is a model that is trained to generate and understand human language. LLMs are trained on large datasets of text and code, and they can be used for a variety of tasks. For example, an LLM can be trained to translate text from one language to another; summarize text, such as website content, search results, news articles, or research papers; answer questions about text, such as “What is the capital of Georgia?”; create chatbots that can converse with humans; and generate creative text, such as poems, stories, and code.

[0049] Language model 170 can be any suitable language model neural network that receives an input sequence composed of text tokens selected from a vocabulary and autoregressively generates an output sequence composed of text tokens from the vocabulary. For example, language model 170 can be a Transformer-based language model neural network or a recurrent neural network-based language model.

[0050] In some situations, when the neural network used to implement the language model 170 autoregressively generates an output sequence of tokens, the language model 170 can be referred to as an autoregressive neural network. More specifically, the autoregressively generated output is created by generating each particular token in the output sequence conditioned on the current input sequence, which includes any tokens in the output sequence that are before the particular token (i.e., tokens that have been generated for any previous positions before the particular position of the particular token in the output sequence) and a context input that provides context for the output sequence.

[0051] For example, when generating a token at any given position in the output sequence, the current input sequence can include the input sequence and any tokens at previous positions in the output sequence that are before the given position. As a specific example, the current input sequence can include the input sequence, followed by any tokens at previous positions in the output sequence that are before the given position. Optionally, the input and the current output sequence can be separated by one or more predetermined tokens within the current input sequence.

[0052] More specifically, to generate a particular token at a particular position within the output sequence, the neural network of the language model 170 can process the current input sequence to generate a score distribution, e.g., a probability distribution, that assigns a corresponding score (e.g., a corresponding probability) to each token in the vocabulary of tokens. The neural network of the language model 170 can then use the score distribution to select a token from the vocabulary as the particular token. For example, the neural network of the language model 170 can greedily select the token with the highest score or can sample tokens according to the distribution, e.g., using nucleus sampling or another sampling technique.

[0053] As a specific example, the language model 170 can be an autoregressive Transformer-based neural network that includes (i) a plurality of attention blocks, each attention block applying self-attention operations; and (ii) an output subnetwork that processes the output of the last attention block to generate a score distribution.

[0054] The language model 170 can have any of a variety of transformer-based neural network architectures. Examples of such architectures include those described in the following: J. Hoffmann, S. Borgeaud, A. Mensch, E. Buchatskaya, T. Cai, E. Rutherford, D. d. L. Casas, L. A. Hendricks, J. Welbl, A. Clark et al. Training compute-optimal large language models, arXiv preprint arXiv:2203.15556, 2022; J.W. Rae, S. Borgeaud, T. Cai, K. Millican, J. Hoffmann, H. F. Song, J. Aslanides, S. Henderson, R. Ring, S. Young, E. Rutherford, T. Hennigan, J. Menick, A. Cassirer, R. Powell, G. van den Driessche, L. A. Hendricks, M. Rauh, P. Huang, A. Glaese, J. Welbl, S. Dathathri, S. Huang, J. Uesato, J. Mellor, I. Higgins, A. Creswell, N. McAleese, A. Wu, E. Elsen, S. M. Jayakumar, E. Buchatskaya, D. Budden, E. Sutherland, K. Simonyan, M. Paganini, L. Sifre, L. Martens, X. L. Li, A. Kuncoro, A. Nematzadeh, E. Gribovskaya, D. Donato, A. Lazaridou, A. Mensch, J. Lespiau, M. Tsimpoukelli, N. Grigorev, D. Fritz, T. Sottiaux, M. Pajarskas, T. Pohlen, Z. Gong, D. Toyama, C. de Masson d’Autume, Y. Li, T. Terzi, V. Mikulik, I.Babuschkin, A. Clark, D. de Las Casas, A. Guy, C. Jones, J. Bradbury, M. Johnson, B. A. Hechtman, L. Weidinger, I. Gabriel, W. S. Isaac, E. Lockhart, S. Osindero, L. Rimell, C. Dyer, O. Vinyals, K. Ayoub, J. Stanway, L. Bennett, D. Hassabis, K. Kavukcuoglu and G. Irving. Scaling language models: Methods, analysis & insights from training gopher (Scaling Language Models: Methods, Analysis & Insights from Training Gopher). CoRR, abs / 2112.11446, 2021; Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li and Peter J Liu. Exploring the limits of transfer learning with a unified text-to-text transformer (Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer). arXiv preprint arXiv:1910.10683, 2019; Daniel Adiwardana, Minh-Thang Luong, David R. So, Jamie Hall, Noah Fiedel, Romal Thoppilan, Zi Yang, Apoorv Kulshreshtha, Gaurav Nemade, Yifeng Lu and Quoc V. Le. Towards a human-like open-domain chatbot (Towards a Human-Like Open-Domain Chatbot). CoRR, abs / 2001.09977, 2020; and Tom B Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell et al. Language models are few-shot learners. arXiv preprint arXiv:2005.14165, 2020.

[0055] However, generally, a transformer-based neural network includes a sequence of attention blocks, and during the processing of a given input sequence, each attention block in the sequence receives the corresponding input hidden states of each input token in the given input sequence. The attention blocks then update each of the hidden states at least in part by applying self-attention to generate the corresponding output hidden states of each of the input tokens. The input hidden states for the first attention block are the embeddings of the input tokens in the input sequence, and the input hidden states for each subsequent attention block are the output hidden states generated by the previous attention block.

[0056] In this example, the output subnetwork processes the output hidden state generated by the last attention block in the sequence for the last input token in the input sequence to generate a score distribution.

[0057] Generally, since the language model is autoregressive, the service device 110 can use the same language model 170 to generate multiple different candidate output sequences in response to the same request, e.g., by using beam search decoding based on the score distribution generated by the language model 170, using sampling and ranking decoding strategies, by using different random seeds for the pseudo-random number generator used in sampling during different runs of the language model 170, or by using another decoding strategy that exploits the autoregressive nature of the language model.

[0058] In some implementations, the language model 170 is pre-trained, i.e., trained on a language modeling task that does not require providing evidence in response to a user question, and the service device 110 (e.g., using the AI system 160) causes the language model 170 to generate an output sequence according to a predetermined grammar via a natural language prompt in the input sequence.

[0059] For example, a service device 110 (e.g., an AI system 160) or a separate training system pre-trains a language model 170 (e.g., a neural network) on a language modeling task (e.g., a task that requires predicting the next token in the training data following the current sequence of text tokens given the current sequence). As a specific example, the language model 170 can be pre-trained on a large dataset of text (e.g., text publicly available from the Internet or another text corpus) according to a maximum likelihood objective.

[0060] The AI system 160 can use the language model 170 to generate image editing prompts for editing images that fail to meet one or more conditions. The service device 110 can maintain conditions for digital components and / or other images sent to the client device 106. For example, the conditions can be used to ensure that digital components do not include explicit content or language. The conditions can vary based on the type of item corresponding to the image and / or based on the type of the image. For example, for a first type of product, there can be a first set of conditions, and for a second type of product, there can be a second set of conditions. Additionally, the service device 110 can maintain conditions for each publisher of the electronic document 150. For example, a publisher who publishes web pages for children may have enhanced conditions for the digital components presented on their web pages.

[0061] In some implementations, the language model 170 can be adapted to use zero-shot learning or few-shot learning to generate image editing prompts. In some examples, the system can provide manually created examples to the language model 170 for few-shot learning. In few-shot learning, the AI system 160 can provide a small number (e.g., three to seven, ten, or another number) of training examples to the language model 170. These training examples also include the original images, a set of conditions, and image editing prompts that include instructions for editing the images to meet the conditions.

[0062] The service device 110 can evaluate image digital components and / or other images to determine whether these images meet the set of conditions for each image. For images that do not meet the conditions, the service device 110 can use the AI system 160 to evaluate why the images do not meet the conditions and / or edit the images so that the images meet the conditions.

[0063] The AI system 160 can use the language model 170 to generate an image editing prompt that instructs the image editing model to generate an updated image that meets the condition. For example, the AI system 160 can generate a prompt 172 that instructs the language model 170 to generate an image editing prompt based on the explanation data related to the explanation of why the image does not meet the condition. In some implementations, the prompt 172 includes the explanation data. The explanation data can include data indicating that the image violates the condition, an explanation of the violation (e.g., the name of the policy violated), the part of the image that causes the image to violate the condition, the condition itself, and / or other data. The data indicating the part of the image that causes the image to violate the condition can include a location indicator depicted in the image.

[0064] For example, the definition of the condition can state: "An image violates [the policy name] policy if [policy definition].(If [policy definition], then the image violates [the policy name] policy.)" In this example, the explanation data can state: "This particular image violates [policy name] policy because [explanation of why the image violates the policy].(Due to [the explanation of why the image violates the policy], this particular image violates [the policy name] policy.)" Hint 172 could be "An image violates [the policy name] policy if [policy definition]. This particular image violates [policy name] policy because [explanation of why the image violates the policy]. My task is to edit the image so that the image doesn't violate [the policy name] policy, and I am going to use a diffusion model by providing it with the original image and the image editing prompt. Please generate the image editing prompt to make the image policy compliant. Make sure the image editing prompt doesn't remove or edit the important things that the image is trying to convey.(If [policy definition], then the image violates [the policy name] policy. Due to [the explanation of why the image violates the policy], this particular image violates [the policy name] policy. My task is to edit the image so that the image does not violate [the policy name] policy, and I will use a diffusion model by providing the original image and the image editing prompt. Please generate the image editing prompt to make the image compliant with the policy. Make sure the image editing prompt does not remove or edit the important things that the image is trying to convey.)Ensure that the image editing prompt does not remove or edit the important things that the image is trying to convey.) The text in parentheses can be filled by the AI system 160 with appropriate data. For example, the AI system 160 can fill [policy name] with the name of the policy violated by the original image.

[0065] Here, the prompt 172 is specifically adapted to instruct the language model 170 to generate an image editing prompt for the diffusion model by informing the language model 170 of the following subsequent task: "I am going to use a diffusion model by providing it with the original image and the image editing prompt. (I will use a diffusion model by providing the original image and the image editing prompt to the diffusion model.)" In this way, the language model 170 can generate an image editing prompt specifically adapted to the diffusion model, which accepts the image editing prompt and the original image as inputs. This eliminates the need to adapt or retrain the image editing model.

[0066] In some implementations, the AI system 160 can maintain a prompt template for each of one or more image editing models. Each prompt template can be in the form of an example prompt 172 with fields in parentheses, which can be filled by the AI system 160 using policy information and explanatory information. When evaluating and editing an image that violates a condition (e.g., a policy condition), the AI system 160 can receive explanatory data from a multimodal model with question-answering and image editing capabilities and use the explanatory data to fill the template.

[0067] If a multimodal language model is used as the language model 170, the prompt 172 can include an image. If the language model 170 only accepts text input, the prompt 172 can include caption text that explains the content of the image to the language model 170.

[0068] In some examples, the explanatory data can include an image with a bounding box around the part of the image that is considered the reason for the image to violate the condition. In this example, the bounding box is a position indicator, and the language model 170 is a multimodal model that accepts text and the image prompt 172 as inputs.

[0069] In some implementations, the AI system 160 uses the language model 170 to generate explanatory data. For example, the AI system 160 can use the language model 170 or another model to generate caption text for the image. The caption text can explain the content of the image. For example, the AI system 160 can generate instructions for the language model 170 or another model (e.g., as referenced Figure 2Another multimodal model) generates a caption for the input image as prompt 172. The AI system 160 can receive the caption and generate another prompt 172 for the instruction language model 170 to evaluate the set of the caption and conditions and output an explanation of why the image violates the conditions or which part of the image as described by the caption violates the conditions. In some examples, the prompt 172 includes a list of objects identified in the image based on the bounding boxes. In this example, the prompt 172 can include the caption and the set of conditions.

[0070] For example, the prompt 172 can state: "An image violates [policy name] policy if [policy definition]. This image includes [description of image]. The following objects and their bounding boxes will give you more spatial awareness context of the image: [object 1: <x1, y1, x2, y2>, object 2: <x1, y1, x2, y2>. Based on the context I provided you about the image, predict if the image violates [policy name] policy. Provide a very detailed explanation for your decision including the regions that violate the policy." In this example, x1 and x2 are the coordinates along one dimension (e.g., the x-axis or the horizontal direction) within the image, and y1 and y2 are the coordinates along another dimension (e.g., the y-axis or the vertical direction) within the image. These coordinates inform the language model 170 of the locations where objects 1 and 2 can be found in the image.

[0071] In some implementations, the AI system 160 uses a multi-modal model to generate explanation data based on a set of images and conditions. In this example, the AI system 160 provides the multi-modal model with the set of images and conditions and requests the multi-modal model to output explanation data, e.g., an explanation of why the image does not meet the conditions and / or an image with a bounding box or other location indicator indicating the part of the image that violates the conditions.

[0072] In any example, the language model 170 can use the explanation data to evaluate the prompt 172 and generate an output 174 including an image editing prompt based on the input data. The language model 170 can generate an image editing prompt (e.g., having a structure that instructs the image editing model to generate an updated image based on the input image and the prompt) in a way that instructs the image editing model to generate an updated image based on the input image and the prompt.

[0073] The AI system 160 can use the image editing prompt to generate an updated image using the image editing model, as described in further detail with reference to Figure 2 The updated image is an edited version of the initial image that does not violate the conditions. For example, the updated image may not contain or depict a region (e.g., the location of the content) in the image that violates the conditions. In a specific example, the image editing model can replace the content in that part of the image with content that meets the conditions.

[0074] For example, an image editing prompt for hiding a specific area of a person in an image or removing an item can state: "Cover the person’s [body part] and remove the [item] from the image." Another example of an image editing prompt for editing a person in an image can state: "Make the facial expressions of the person(s) in the image neutral." Another example of an image editing prompt for removing an item from an image can state: "Crop the image to remove the item from the image." Another example of an image editing prompt for removing an item from an image can state: "Erase the adult beverage bottles from the image."

[0075] The AI system 160 can then send (e.g., provide) the updated image to one or more devices (e.g., one or more client devices 106) as a response 120. For example, the AI system 160 can generate a digital component to provide in response to a request 112 from a user. The digital component can include the updated image. The digital component can include a link to an electronic document related to the subject of the digital component (e.g., the item depicted by the image), metadata, and / or other data and / or files that enable the client device 106 to render the updated image.

[0076] Although Figure 1 a single language model 170 is shown, different language models can be specifically trained to process different prompts at different stages of the processing pipeline. For example, one language model can be trained to generate interpretation data for an image, while another language model can be trained to generate image editing prompts based on the interpretation data.

[0077] Figure 2 FIG. 200 is a block diagram showing the interaction between the AI system 160, the multimodal model 202, the language model 170, and the image editing model 204. The AI system 160 can include an image evaluation device 206, a prompt device 208, and a digital component device 210.

[0078] The language model 170 can be trained to perform various tasks as described above. The AI system 160 can use the language model 170 to generate interpretation data and / or generate image editing prompts for the image editing model 204. Although Figure 2 a single language model 170 is shown, the AI system 160 can interact with any number of language models 170 to generate image editing prompts that instruct the image editing model 204 to generate an updated image that meets one or more conditions (e.g., one or more policy conditions).

[0079] The multimodal model 202 can be implemented as a machine learning model trained to generate interpretation data 212. For example, the training process can use a set of training images and ground truth interpretation data corresponding to the training data. For example, for each image that violates a condition, the ground truth training data can include a label indicating the condition that was violated and the reason the image violated the condition. The label can also indicate the part of the image that violated the condition. Based on this set of training images, the multimodal model 202 can be trained to generate interpretation data 212.

[0080] The multimodal model 202 can be trained to generate text based on text and image inputs. For example, the multimodal model 202 can be trained to output, as explanation data, an explanation that interprets why an image violates one or more conditions based on an input that includes an image and text indicating one or more conditions. In some implementations, the multimodal model 202 can be trained to output, as explanation data, an image with a location indicator (e.g., a bounding box) indicating the portion of the image that violates the condition. Similar to the examples provided above, the x and y coordinates of the image can be used to outline the bounding box. The multimodal model 202 can take an image and text as inputs, and the multimodal model 202 can generate text as an output. During training, the image and a question (e.g., text asking for explanation data) are used as inputs to the model, and the multimodal model 202 is trained to generate an answer (e.g., text answering the question about the explanation data). In a supervised learning example, the training samples can include an image and a question that includes a condition, where the label has an answer that includes an explanation of why the image does not meet the condition. In some implementations, the multimodal model 202 can be a neural network or other type of machine learning model that is trained to provide an answer and edit an image in response to a question.

[0081] The image editing model 204 can be a machine learning model, such as a text-to-image neural network, that is trained to generate an image based on an input image and an image editing prompt 215 that instructs the image editing model 204 how to edit the image. In some implementations, the image editing model 204 is a language model that is trained to edit an image. In some implementations, the image editing model 204 is a diffusion model.

[0082] During training, the image editing model 204 can take the original image caption of the image and the image editing prompt as inputs, and the image editing model 204 can be trained to generate a target text prompt by applying the image editing prompt to the image caption. Given the image and the target text prompt, the image editing model 204 can encode the target text prompt to generate an initial text embedding. The image editing model 204 then processes (e.g., optimizes) the initial text embedding to reconstruct the input image. The system then fine-tunes the image editing model 204 (e.g., the diffusion model of the image editing model 204) by interpolating the target text prompt with the input image to generate the output of the image editing model (e.g., the edited image) to improve the overall accuracy.

[0083] The AI system 160 may also include a memory structure 218 or be configured to interact with the memory structure 218 to extract and / or store information and content. The memory structure 218 may include one or more databases or other data structures stored on one or more memories and / or data storage devices. In particular, the memory structure 218 may store a digital component database 116, digital components 220, images 222, and condition data 224.

[0084] As described above, the digital component database 116 may include distribution parameters for the digital components 220. The distribution parameters for the digital components 220 may include, for example, keywords and / or geographical locations to which the digital components 220 are eligible to be distributed to the client devices 106. For each digital component 220, the digital component database 116 may also include metadata of the digital component, descriptive text for each image 222 corresponding to the digital component, data related to the digital component provider providing the digital component, and / or other data related to the digital component. The digital components 220 may include candidate digital components that may be provided in response to a component request 112 and / or a query received by the service device 110. The images 222 may include one or more images for each digital component 220. The AI system 160 may obtain the images for the digital components 220 from a digital component provider or other sources.

[0085] The condition database 224 may store conditions for the images. The condition database 224 may store a set of one or more conditions for each type of image, for each type of item depicted by the image, for each publisher, and / or for other entities.

[0086] The AI system 160 may interact with the memory structure 218 and the models 170, 202, 204 to evaluate the images and generate updated images of those images that do not meet one or more conditions for the images. In some examples, the AI system 160 may receive an image (e.g., a first image 211) from the client device 106 (e.g., the user's user device or the device of a digital component provider). The image evaluation device 206 may obtain a set of conditions for the first image 211 from the condition database 224. For example, the image evaluation device 206 may obtain a set of conditions 214 for the image based on the type of item (e.g., the type of product) that is the subject of the digital component including the first image 211. The image evaluation device 206 may evaluate the first image 211 based on the conditions 214 and output data indicating whether the first image 211 meets the conditions 214. In another example, a human may review the first image 211 and provide data indicating whether the first image 211 meets the conditions 214 to the AI system 160.

[0087] If the first image 211 does not meet the condition 214, the AI system 160 may generate an image editing prompt 215 that instructs the image editing model 204 to edit the first image 211 to create a second image 216 that meets the condition 214. The AI system 160 may use the multimodal model 202 and / or the language model 170 to generate the image editing prompt 215.

[0088] In some implementations, the AI system 160 sends the first image 211 and the condition 214 along with a request or query to the multimodal model 202 that requests the multimodal model 202 to output explanation data 212 that indicates (e.g., in text) why the first image 211 violates the condition 214 and / or indicates the location of the portion of the first image 211 that violates the condition 214 (e.g., using a bounding box or other location indicator such as another type of visual indicator superimposed over the portion of the first image 211, a textual description of the location, or the coordinates of the portion of the first image 211 (e.g., pixel coordinates)).

[0089] In this example, the AI system 160 may generate a prompt 172 based on the explanation data 212. The prompt 172 may include instructions that direct the language model 170 to generate an image editing prompt 215 that instructs the image editing model 204 to generate a second image 216 that meets the condition 214. The prompt 172 may include the first image 211, the explanation data 212, and / or the condition 214.

[0090] In some implementations, the AI system 160 sends the first image 211 and a request for a caption that the multimodal model 202 generated for the first image 211 to the multimodal model 202. The caption may explain the content of the image. In this example, the AI system 160 may generate a prompt 172 based on the caption and the condition 214. For example, the prompt 172 may instruct the language model 170 instead of the multimodal model 202 to output the explanation data 212 based on the caption and the condition 214. In this example, the explanation data 212 may indicate which portion of the first image 211 violates the condition 214 in terms of the caption and / or why that portion of the first image 211 violates the condition 214. Similar to the previous example, the AI system 160 may then generate a second prompt 172 for the language model 170 that instructs the language model 170 to output the image editing prompt 215 based on the explanation 212.

[0091] The AI system 160 can then provide the image editing prompt 212 to the image editing model 204. The AI system 160 can also provide the first image 211 and / or a version of the first image 211 that includes a location indicator of a portion indicating a violation condition 214 of the image. The image editing model 204 can edit the first image 211 based on the image editing prompt 215 and output an edited version of the first image 211 as the second image 216.

[0092] In some implementations, the AI system 160 can evaluate the second image 216 using, for example, the image evaluation device 206 to ensure that the second image 216 meets the condition 214, as described above. If the second image 216 does not meet the condition 214, the AI system 160 can use similar techniques to generate another edited version of the first image 211. However, the AI system 160 can modify the prompt 172 to the language model 170 to ensure that the edited image meets the condition 214, or to increase the likelihood that the edited image meets the condition 214. In some examples, if the edited image does not meet the condition 214, the system can provide the AI system 160 with the original image caption, the original explanatory data, the image editing prompt, the edited image caption, and the explanation of the edited image, and the AI system 160 can modify the prompt 172 based on the provided information to increase the likelihood that the edited image will meet the condition 214.

[0093] If the second image 216 meets the condition, the AI system 160 can send the second image 216 to the client device 106. In some implementations, the digital component device 210 can use the second image 216 to generate a digital component and send the digital component to the client device 106. For example, the digital component device 210 can generate a digital component that depicts the second image 216 and includes a link to an electronic document and / or data / file, which enables the client device 106 to render the digital component. The AI system 160 can provide the digital component to the service device 110, and the service device 100 can distribute the digital component to the client device 106 in response to the component request 112, as described above.

[0094] Figure 3 is a flowchart of an example process 300 for generating personalized image advertisements. The operations of process 300 can be performed, for example, by Figure 1 the AI system 160 or another data processing device. The operations of process 300 can also be implemented as instructions stored on a computer-readable medium, which can be non-transitory. Executing the instructions by one or more data processing devices causes the one or more data processing devices to perform the operations of process 300.

[0095] The system receives data (302) indicating that a first image violates one or more conditions. For example, the system can evaluate the first image based on one or more conditions or receive data from another system indicating that the first image violates one or more conditions. In some implementations, the system provides the image and the conditions along with a request to a language model or a multimodal model to predict whether the image violates any of the conditions and, if so, output explanation data explaining why the image violates the conditions.

[0096] In response to receiving data indicating that the first image violates one or more conditions, the system generates an image editing prompt (304) for editing the first image. As described above, the system can use a chain of prompts to one or more machine learning models (e.g., a multimodal model and / or one or more language models) to generate the image editing prompt. The image editing prompt can instruct an image editing model to edit the first image such that the first image satisfies one or more conditions.

[0097] The system provides the image editing prompt and the first image as inputs to the image editing model (306). The image editing model can generate a second image by editing the first image based on the image editing prompt. The system receives the second image as the output of the image editing model.

[0098] The system provides the second image to one or more devices (310). For example, the system can provide the second image to the device that provided the first image to the system and / or to other devices, e.g., as an image digital component provided in response to a component request. For example, the system can be part of a service device 110 that distributes image digital components to client devices 106, as described above.

[0099] Figure 4 is a block diagram of an example computer system 400 that can be used to perform the operations described above. System 400 includes a processor 410, a memory 420, a storage device 430, and an input / output device 440. Each of the components 410, 420, 430, and 440 can be interconnected, for example, using a system bus 450. The processor 410 is capable of processing instructions for execution within the system 400. In one implementation, the processor 410 is a single-threaded processor. In another implementation, the processor 410 is a multi-threaded processor. The processor 410 is capable of processing instructions stored in the memory 420 or on the storage device 430.

[0100] The memory 420 stores information within the system 400. In one implementation, the memory 420 is a computer-readable medium. In one implementation, the memory 420 is a volatile memory unit. In another implementation, the memory 420 is a non-volatile memory unit.

[0101] The storage device 430 can provide large-capacity storage for the system 400. In one implementation, the storage device 430 is a computer-readable medium. In various different implementations, the storage device 430 can include, for example, a hard disk device, an optical disk device, a storage device shared by multiple computing devices via a network (e.g., a cloud storage device), or some other large-capacity storage device.

[0102] The input / output device 440 provides input / output operations for the system 400. In one implementation, the input / output device 440 can include one or more of a network interface device (e.g., an Ethernet card), a serial communication device (e.g., and an RS-232 port), and / or a wireless interface device (e.g., and an 802.11 card). In another implementation, the input / output device can include a drive device configured to receive input data and send output data to other devices (e.g., a keyboard, a printer, a display, and other peripheral devices 460). However, other implementations can also be used, such as mobile computing devices, mobile communication devices, set-top box television client devices, etc.

[0103] Although example processing systems have been described in Figure 4 , the described subject matter and implementations of the functional operations can be implemented in other types of digital electronic circuitry or in computer software, firmware, or hardware (including the structures disclosed in this specification and their structural equivalents) or in a combination of one or more of them.

[0104] An electronic document (which will be abbreviated as a document for brevity) does not necessarily correspond to a file. A document can be stored in a part of a file that stores other documents, in a single file dedicated to the document in question, or in multiple coordinated files.

[0105] For cases where the systems discussed here collect and / or use personal information about users, users can be provided with the opportunity to enable / disable or control programs or features that can collect and / or use personal information (e.g., information about a user's social network, social actions or activities, a user's preferences, or a user's current location). Additionally, certain data can be disposed of in one or more ways before it is stored or used so that personally identifiable information associated with the user is removed. For example, a user's identity can be anonymized so that the user's personally identifiable information cannot be determined, or a user's geographical location can be generalized (such as to the city, zip code, or state level) if location information is obtained so that the user's specific location cannot be determined.

[0106] The subject matter and embodiments of the operations described in this specification can be implemented in digital electronic circuitry, or in computer software, firmware, or hardware (including the structures disclosed in this specification and structural equivalents thereof), or in combinations of one or more of them. Embodiments of the subject matter described in this specification can be implemented as one or more computer programs, i.e., one or more modules of computer program instructions, encoded on a computer storage medium for execution by, or to control the operation of, a data processing apparatus. Alternatively or additionally, the program instructions can be encoded on an artificially generated propagated signal (e.g., a machine-generated electrical, optical, or electromagnetic signal) that is generated to encode information for transmission to a suitable receiver apparatus for execution by the data processing apparatus. A computer storage medium can be a computer-readable storage device, a computer-readable storage substrate, a random or serial access memory array or device, or a combination of one or more of them, or be included in a computer-readable storage device, a computer-readable storage substrate, a random or serial access memory array or device, or a combination of one or more of them. Moreover, although a computer storage medium is not a propagated signal, a computer storage medium can be a source or destination of computer program instructions encoded in an artificially generated propagated signal. A computer storage medium can also be one or more separate physical components or media (e.g., multiple CDs, disks, or other storage devices), or be included in one or more separate physical components or media.

[0107] The operations described in this specification can be implemented as operations performed by a data processing apparatus on data stored on one or more computer-readable storage devices or received from other sources.

[0108] The term “data processing apparatus” encompasses all kinds of devices, apparatus, and machines for processing data, including, by way of example, programmable processors, computers, system-on-a-chip, or multiple of the foregoing or combinations of the foregoing. The apparatus can include special purpose logic circuitry, e.g., an FPGA (field programmable gate array) or an ASIC (application specific integrated circuit). In addition to hardware, the apparatus can also include code that creates an execution environment for the computer programs in question, e.g., code that constitutes processor firmware, a protocol stack, a database management system, an operating system, a cross-platform runtime environment, a virtual machine, or a combination of one or more of them. The apparatus and the execution environment can implement various different computing model infrastructures, such as web services, distributed computing, and grid computing infrastructures.

[0109] This document relates to service devices. As used herein, a service device is one or more data processing devices that perform operations to facilitate the distribution of content over a network. The service device is depicted as a single box in a block diagram. However, while a service device can be a single apparatus or a single set of apparatuses, the present disclosure contemplates that a service device can also be a group of apparatuses or even multiple different systems that communicate to provide various content to client devices. For example, a service device can encompass one or more of a search system, a video streaming service, an audio streaming service, an email service, a navigation service, an advertising service, a gaming service, or any other service.

[0110] A computer program (also referred to as a program, software, software application, script, or code) can be written in any form of programming language (including compiled or interpreted languages, declarative or procedural languages), and it can be deployed in any form, including as a stand-alone program or as a module, component, subroutine, object, or other unit suitable for use in a computing environment. A computer program can, but need not, correspond to a file in a file system. The program can be stored in a part of a file that holds other programs or data (such as one or more scripts stored in a markup language document), in a single file dedicated to the program in question, or in multiple coordinated files (such as files that store one or more modules, subroutines, or portions of code). A computer program can be deployed to execute on one computer or on multiple computers distributed at one site or across multiple sites and interconnected by a communication network.

[0111] The processes and logical flows described in this specification can be performed by one or more programmable processors that execute one or more computer programs to perform actions by operating on input data and generating output. The processes and logical flows can also be performed by, and the apparatus can also be implemented as, special purpose logic circuitry, such as an FPGA (Field Programmable Gate Array) or an ASIC (Application Specific Integrated Circuit).

[0112] For example, processors suitable for the execution of a computer program include both general and special purpose microprocessors, as well as any one or more processors of any kind of digital computer. In general, a processor will receive instructions and data from a read only memory or a random access memory or both. The basic elements of a computer are a processor for performing actions in accordance with instructions and one or more memory devices for storing the instructions and data. In general, a computer will also include one or more mass storage devices for storing data, such as magnetic disks, magneto-optical disks, or optical disks, or operatively coupled to receive data from or transfer data to or both from the one or more mass storage devices. However, a computer need not have such devices. In addition, a computer may be embedded in another device (such as a mobile phone, a personal digital assistant (PDA), a mobile audio or video player, a game console, a global positioning system (GPS) receiver, or a portable storage device (such as a universal serial bus (USB) flash drive), to name just a few). Devices suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and memory devices, by way of example, including: semiconductor memory devices such as EPROM, EEPROM, and flash memory devices; magnetic disks such as internal hard disks or removable disks; magneto-optical disks; and CD-ROM and DVD-ROM disks. The processor and the memory may be supplemented by, or incorporated in, special purpose logic circuitry.

[0113] To provide for interaction with a user, embodiments of the subject matter described in this specification may be implemented on a computer having a display device (such as a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user and a keyboard and a pointing device (such as a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices may also be used to provide for interaction with the user; for example, feedback provided to the user may be any form of sensory feedback, such as visual feedback, auditory feedback, or tactile feedback; and input received from the user may be in any form, including acoustic, speech, or tactile input. Additionally, a computer may interact with the user by sending documents to and receiving documents from a device used by the user; for example, by sending a web page to a web browser on a client device of the user in response to a request received from the web browser.

[0114] Embodiments of the subject matter described in this specification can be implemented in a computing system that includes a backend component (e.g., as a data server), or that includes a middleware component (e.g., an application server), or that includes a frontend component (e.g., a client computer having a graphical user interface or a web browser through which a user can interact with an implementation of the subject matter described in this specification), or any combination of one or more such backend, middleware, or frontend components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include local area networks (“LANs”) and wide area networks (“WANs”), the Internet (e.g., the Internet), and peer-to-peer networks (e.g., ad hoc peer-to-peer networks).

[0115] The computing system can include clients and servers. The clients and servers are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other. In some embodiments, the server sends data (e.g., an HTML page) to the client device (e.g., for purposes of displaying data to and receiving user input from a user interacting with the client device). Data generated at the client device (e.g., the result of a user interaction) can be received at the server from the client device.

[0116] Although this specification contains many specific implementation details, these should not be construed as limitations on the scope of any invention or of what may be claimed, but rather as descriptions of features specific to particular embodiments of particular inventions. Certain features that are described in this specification in the context of separate embodiments can also be implemented in combination in a single embodiment. Conversely, various features that are described in the context of a single embodiment can also be implemented separately or in any suitable sub-combination in multiple embodiments. Moreover, although features may be described above as acting in certain combinations and even initially claimed as such, in some cases, one or more features from a claimed combination can be deleted from the combination, and the claimed combination may relate to a sub-combination or a variation of a sub-combination.

[0117] Similarly, although operations are depicted in the drawings in a particular order, this should not be understood to require that such operations be performed in the particular order shown or in sequential order, or that all illustrated operations be performed, to achieve desirable results. In certain circumstances, multitasking and parallel processing may be advantageous. Moreover, the separation of various system components described above in the embodiments should not be understood to require such separation in all embodiments, and it should be understood that the described program components and systems can generally be integrated together in a single software product or packaged into multiple software products.

[0118] Accordingly, specific embodiments of the subject matter have been described. Other embodiments are within the scope of the appended claims. In some cases, the acts recited in the claims can be performed in a different order and still achieve the desired result. Additionally, the processes depicted in the figures are not necessarily required to be in the particular order or sequential order shown to achieve the desired result. In certain implementations, multitasking and parallel processing may be advantageous.

Claims

1. A method, comprising: Receiving data indicating that a first image violates one or more conditions; In response to receiving the data indicating that the first image violates the one or more conditions, generating an image editing prompt for instructing an image editing model to edit the first image to meet the one or more conditions, the generating comprising: Generating explanation data indicating the position of the content in the image that violates the one or more conditions; and Using a language model to generate the image editing prompt based on the explanation data and the one or more conditions; Providing the image editing prompt and the first image as inputs to an image editing model; Receiving a second image as an output of the image editing model; and Providing the second image to one or more devices.

2. The method according to claim 1, wherein generating the explanation data comprises providing the image and the one or more conditions to a multimodal model that is trained to identify the positions in the image that violate the input conditions.

3. The method according to claim 1, wherein generating the explanation data comprises providing the image and the one or more conditions to a multimodal model that is trained to predict whether the image violates the input conditions and output data indicating the positions of the content in the image that is likely to violate at least one of the input conditions.

4. The method according to claim 1 or 2, wherein generating the explanation data comprises: Providing the first image to a first machine learning model that is trained to generate an image caption for the image; Receiving the image caption for the image from the first machine learning model; Providing the image caption and the one or more conditions to a second machine learning model that is trained to output explanation data for the image based on the input image caption and the input conditions; And Receiving the explanation data for the first image from the second machine learning model.

5. The method according to any one of the preceding claims, wherein the explanation data comprises a position indicator, and wherein the position indicator indicates the position of the content in the first image that is determined to violate at least one of the one or more conditions.

6. The method according to claim 5, wherein the position indicator comprises a bounding box drawn around the content in the first image that is determined to violate at least one of the one or more conditions.

7. The method according to claim 5, wherein the position indicator comprises coordinates defining a bounding box around the content in the first image that is determined to violate at least one of the one or more conditions.

8. The method according to any one of claims 5 to 7, wherein providing the image editing prompt and the first image as inputs to an image editing model comprises providing the position indicator to the image editing model.

9. The method according to any one of the preceding claims, wherein the explanation data comprises an explanation indicating why the first image violates the one or more conditions.

10. The method according to any one of the preceding claims, wherein the image editing hint comprises at least a part of the interpretation data.

11. The method according to any one of the preceding claims, wherein generating the image editing hint using the language model based on the interpretation data and the one or more conditions comprises using the interpretation data to generate a hint for the language model and providing the hint to the language model.

12. The method according to claim 11, wherein the hint comprises an instruction to instruct the language model to generate the image editing hint based on the interpretation data and data defining each condition violated by the first image, and wherein the interpretation data comprises the name of each condition violated by the first image.

13. The method according to claim 11 or 12, wherein generating the image editing hint comprises: obtaining a hint template adapted to the image editing model; and populating the hint with at least a part of the interpretation data, the at least a part of the interpretation data comprising the name of the condition determined to be violated by the first image, wherein the image editing hint output by the language model is adapted to the image editing model.

14. A system, comprising: one or more processors; and one or more storage devices storing instructions that, when executed by the one or more processors, cause the one or more processors to perform operations, the operations including: receiving data indicating that a first image violates one or more conditions; in response to receiving the data indicating that the first image violates the one or more conditions, generating an image editing hint instructing an image editing model to edit the first image to meet the one or more conditions, the generating including: generating interpretation data indicating the location of the content in the image that violates the one or more conditions; and using a language model to generate the image editing hint based on the interpretation data and the one or more conditions; providing the image editing hint and the first image as inputs to an image editing model; receiving a second image as an output of the image editing model; and providing the second image to one or more devices.

15. The system according to claim 14, wherein generating the interpretation data comprises providing the image and the one or more conditions to a multimodal model trained to identify the locations in an image that violate input conditions.

16. The system according to claim 14, wherein generating the interpretation data comprises providing the image and the one or more conditions to a multimodal model trained to predict whether an image violates input conditions and output data indicating the locations of the content in the image that is likely to violate at least one of the input conditions.

17. A computer-readable storage medium carrying instructions that, when executed by one or more processors, cause the one or more processors to perform operations, the operations including: receiving data indicating that a first image violates one or more conditions; In response to receiving the data indicating that the first image violates the one or more conditions, generating an image editing prompt for instructing an image editing model to edit the first image to meet the one or more conditions, the generating including: generating explanation data indicating the locations of the content in the image that violates the one or more conditions; and using a language model to generate the image editing prompt based on the explanation data and the one or more conditions; providing the image editing prompt and the first image as inputs to the image editing model; receiving a second image as an output of the image editing model; and providing the second image to one or more devices.

18. The computer-readable storage medium according to claim 17, wherein generating the explanation data includes providing the image and the one or more conditions to a multimodal model that is trained to identify the locations in the image that violate the input conditions.

19. A computer program product comprising instructions that, when executed by one or more processors, cause the one or more processors to perform operations, the operations including: receiving data indicating that a first image violates one or more conditions; In response to receiving the data indicating that the first image violates the one or more conditions, generating an image editing prompt for instructing an image editing model to edit the first image to meet the one or more conditions, the generating including: generating explanation data indicating the locations of the content in the image that violates the one or more conditions; and using a language model to generate the image editing prompt based on the explanation data and the one or more conditions; providing the image editing prompt and the first image as inputs to the image editing model; receiving a second image as an output of the image editing model; and providing the second image to one or more devices.

20. The computer program product according to claim 19, wherein generating the explanation data includes providing the image and the one or more conditions to a multimodal model that is trained to identify the locations in the image that violate the input conditions.

Citation Information

Cited By

  • Image editing system, method and equipment based on artificial intelligence and medium

    CN122134856A