Generating content items using generative neural networks and entity reward models

The content item generation system uses a generative neural network and entity reward models to efficiently produce customized content, addressing resource inefficiencies and personalization challenges in conventional systems, ensuring stylistic coherence and engagement.

WO2026054772A1PCT designated stage Publication Date: 2026-03-12GOOGLE LLC
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-09-06
Publication Date
2026-03-12

AI Technical Summary

Technical Problem

Conventional content generation systems require significant computational and memory resources to produce customized content items, often generating multiple generic items before finding one that meets an entity's needs or requiring post-generation refinements, and fail to personalize content effectively.

Method used

A content item generation system using a generative neural network and multiple entity reward models to directly produce customized content tailored to specific entities, reducing the need for multiple generic item generation and post-processing.

Benefits of technology

The system efficiently generates personalized content items, saving computational resources and ensuring stylistic coherence and engagement by aligning with entity-specific attributes, thus providing a more engaging user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2024045567_12032026_PF_FP_ABST
    Figure US2024045567_12032026_PF_FP_ABST
Patent Text Reader

Abstract

Methods, systems, and apparatus, including computer programs encoded on a computer storage medium, for generating content items using a generative neural network and a plurality of entity reward model. One of the methods include: receiving a request to generate a customized content item that is customized for a target entity from a plurality of entities; identifying a target entity reward model that corresponds to the target entity; and generating the customized content item, comprising using the generative neural network to perform a data generation process that is guided by the target entity reward model.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Attorney Docket No.: 56113-0771WO1

[0002] GENERATING CONTENT ITEMS USING GENERATIVE NEURAL NETWORKS AND ENTITY REWARD MODELS

[0003] BACKGROUND

[0004]

[0001] This specification relates to generating content items using neural networks. For example, the content items can include text data, image data, video data, audio data, or the like.

[0005]

[0002] Neural networks are machine learning models that employ one or more layers of nonlinear units to predict an output for a received input. Some neural networks include one or more hidden layers in addition to an output layer. The output of each hidden layer is used as input to another layer in the network, e.g., the next hidden layer or the output layer. Each layer of the netw ork generates an output from a received input in accordance with current values of a respective set of weights.

[0006] SUMMARY

[0007]

[0003] This specification describes a content item generation system implemented as computer programs on one or more computers in one or more locations that generates customized content items using a generative neural network and multiple entity reward models.

[0008]

[0004] Particular embodiments of the subject matter described in this specification can be implemented so as to realize one or more of the following advantages.

[0009]

[0005] The content item generation system described in this specification can generate customized content items for many different entities by using a generative neural network and multiple entity reward models that correspond respectively to the different entities. On the other hand, conventional systems may do this by either generating multiple different candidate content items for selection by an entity, or allowing the entity7to specify7postgeneration refinements to an initially generated generic content item.

[0010]

[0006] The capability of the described system to generate customized content items saves computational and memory resources that are otherwise required by those conventional systems, e.g., the execution of multiple runs of the generative neural netw ork to generate multiple content items that are generic, i.e., not customized, until a content item that satisfies the needs of an entity7is generated; or the further processing, e.g.. modification or adjustment, of such generic content items to obtain the customized content items. Attorney Docket No.: 56113-0771WO1

[0011]

[0007] The described system and techniques for generating customized content items are useful in many content item generation applications, e.g., conditional text, image, video, or audio generation tasks. As one example, the described system and techniques are useful in document collaboration or co-authoring applications - for example by generating a portion of a document that corresponds to one of the collaborators of the document while accommodating for the varying language styles of other collaborators which provide other portions of the same document, thereby generating a collaborative document that has stylistically coherent (and rather than disjoint) content.

[0012]

[0008] As another example, the described system and techniques for generating customized content items are useful in image or video generation applications - for example by generating different images or videos that each consistently show a particular subject instance of interest (e.g., a particular person, a particular animal, a particular car, a particular boat, etc.) and rather than various subject instances (e.g., various people, various animals, various cars, various boats, etc.), thereby showing visual content that is personalized for a recipient, providing a more engaging user experience.

[0013]

[0009] As another example, the described system and techniques for generating customized content items are useful in public document creation applications - for example by generating public documents that have content tailored to target readers of the public documents. For example, when generating blog posts for a social media app the target readers of which are sports enthusiasts, the described system and techniques can generate blog posts that include content tailored to sports enthusiasts, e.g., information about sports events, quotes from famous sports players, etc.

[0014]

[0010] The details of one or more embodiments of the subject matter of this specification are set forth in the accompanying drawings and the description below. Other features, aspects, and advantages of the subject matter will become apparent from the description, the drawings, and the claims.

[0015] BRIEF DESCRIPTION OF THE DRAWINGS

[0016]

[0011] FIG. 1 is a diagram of an example content item generation system.

[0017]

[0012] FIG. 2 is a flow diagram of an example process for generating a customized content item.

[0018]

[0013] FIG. 3 is a flow diagram of an example process for performing a data generation process. Attorney Docket No.: 56113-0771WO1

[0019]

[0014] FIG. 4 is a flow diagram of another example process for performing a data generation process.

[0020]

[0015] FIG. 5 is a flow diagram of another example process for performing a data generation process.

[0021]

[0016] Like reference numbers and designations in the various draw ings indicate like elements.

[0022] DETAILED DESCRIPTION

[0023]

[0017] FIG. 1 is a diagram of an example content item generation system 100. The content item generation system 100 is an example of a system implemented as computer programs on one or more computers in one or more locations, in which the systems, components, and techniques described below can be implemented.

[0024]

[0018] The content item generation system 100 is a system that generates customized content items 152 in response to received requests 102. The content item generation system 100 can generate any kind of customized content items 152, e.g., textual content items, image content items, video content items, audio content items, and so on.

[0025]

[0019] A content item is referred to as a “customized” content item that is customized for an entity when the content item includes attributes that are specific to the entity. An entity may be an individual, e.g., a person. Alternatively, an entity may be an organization, e.g., a named group of people, a web publisher, a website, and so on.

[0026]

[0020] In some cases, the content item generation system 100 can be a text generation system that generates text sequences, i.e., each customized content item 152 generated by the system is an output sequence of text that includes a sequence of text tokens from a vocabulary of text tokens that includes, e.g., one or more of characters, sub-words, words, punctuation marks, numbers, or other symbols that appear in a natural language or a computer language. For example, the system can generate text sequences in response to a request 102 submitted by a user of the system and provide the text sequences for presentation to the user which submitted the request 102.

[0027]

[0021] In some of these cases, the content item generation system 100 can receive a conditioning input as part of, or associated with, the request 102 and generate a customized content item 152 that is a response to the conditioning input.

[0028]

[0022] For example, the conditioning input can be an input sequence of text and the output sequence is another sequence of text, e.g.. a translation of the input sequence of text, a completion of the input sequence of text, a paraphrase of the input sequence of text, a Attorney Docket No.: 56113-0771WO1 response to a question posed in the input sequence, or a sequence of text that is about a topic specified by the input sequence of text. As another example, the conditioning input can be an input other than text, e.g., an image, and the output sequence can be text that describes the input.

[0029]

[0023] In these cases, an entity can have preferred attributes that should be included in the output sequences of text. Thus, in order for an output sequence of text to be customized for the entity, the output sequence of text should include entity -specific attributes, e.g.. it should adhere to the entity’s language style.

[0030]

[0024] For example, if the input sequence of text is a sequence of text in one language, and the output sequence of text is a piece of text in another language that is a predicted proper translation of the input sequence of text into the other language, the output sequence of text that is customized for an entity’ should align with the specific manner in which the entity communicates in the language of the input sequence of text.

[0031]

[0025] As another example, if the conditioning input includes an image, and the output sequence of text is a piece of text in another language that is a caption of the image or an answer to a question posed about the image, the output sequence of text that is customized for an entity should have entity-specific language style preferences or characteristics for captioning or the format of the answer.

[0032]

[0026] As a particular example, the content item generation system 100 can be part of a dialog system and the conditioning input can include audio or text from the most recent conversational turn submitted by a user of the dialog system during the dialog while the output sequence of text is the next turn in the conversation, e.g., either text or audio that is a response to the most recent conversational turn. Optionally, the conditioning input can also include one or more historical conversational turns that occurred earlier in the conversation. In this example, the output sequence of text that is customized for an entity (e.g., the user of the dialog system) should align with a specific manner in which the entity communicates in a more casual setting, e.g., includes specific slang, specific idioms, specific abbreviations, specific punctuations, etc.

[0033]

[0027] As another particular example, the content item generation system 100 can be part of a computer code generation system and the conditioning input can be a text description of a desired piece of code or a snippet of computer code in a programming language and the output sequence of text can be computer code, e.g., a snippet of code that is described by the conditioning input or a snippet of code that follows the conditioning input in a computer program. In this example, the output sequence of text that is customized for an entity should Attorney Docket No.: 56113-0771WO1 align with a specific manner in which the entity writes code, e.g., includes specific indentation structure, specific variable names, etc.

[0034]

[0028] In some cases, the content item generation system 100 can be an image or video generation system that generates images or videos that each have multiple frames (where each frame is an image) by generating images, e.g., either as sequences of pixels or through an iterative denoising process. For example, the content item generation system 100 can generate an image or a video conditioned on a conditioning input that includes a text description of the content of the image or the video.

[0035]

[0029] In these cases, an entity can have preferred attributes that should be included in the images generated by the system. Thus, in order for an image to be customized for the entity, the image should include entity-specific attributes, e.g., it should adhere to the entity's visual style that is represented by one or more entity-specific visual attributes.

[0036]

[0030] For example, an image is customized for an entity when it includes a specific color, a specific texture, a specific brightness, etc., that is specific to the entity. As another example, an image is customized for an entity when it includes a depiction of a specific object, a depiction of a specific watermark, a depiction of a specific logo, a depiction of a specific symbol, etc., that is specific to the entity7.

[0037]

[0031] In some cases, the content item generation system 100 can be an audio generation system that generates audio signals, e.g., each customized content item 152 is an output audio example that includes a sample of an audio wave at each of a sequence of output time steps that span a specified time window. For example, the output time steps can be arranged at regular intervals within the specified time window. The audio sample at a given output time step can be an amplitude value of the audio wave or an amplitude value that has been compressed, companded, or both. For example, the audio sample can be a raw amplitude value or a mu-law companded representation of the amplitude value.

[0038]

[0032] In these cases, an entity can have preferred attributes that should be included in the audio signals generated by the system. Thus, in order for an audio signal to be customized for the entity, the audio signal should include entity-specific attributes, e.g., it should adhere to the entity’s audio style that is represented by one or more entity-specific audio attributes.

[0039]

[0033] For example, an output audio example is customized for an entity when it has a specific acoustic property, e.g., a specific speaker identity, a specific recording condition (such as a specific level of reverberation, distortion, or background noise, etc.) that is specific to the entity. As another example, an output audio example is customized for an entity when it has a specific semantic property, e.g., a specific linguistic content (when the output audio Attorney Docket No.: 56113-0771WO1 example represents speech), or a specific melody or a specific rhythm (when the output audio example represents music) that is specific to the entity.

[0040]

[0034] In particular, the content item generation system 100 receives a request 102 for a customized content item 152 and, in response, generates the customized content item 152 using a generative neural network 110.

[0041]

[0035] The generative neural network 110 can be any appropriate generative neural network that has a set of generative neural network parameters and that can be used to generate a content item that includes data in a single modality or multiple modalities by performing a data generation process in accordance with the set of generative neural network parameters.

[0036] In some implementations, the generative neural network 110 can be a language model neural network that executes an auto-regressive token generation process to auto-regressively generate a customized content item 152, e.g., a sequence of text tokens, a sequence of pixel tokens, a sequence of audio tokens, a sequence of multi-modal tokens, e.g., text and pixel tokens, or the like, across multiple time steps, for example by generating one token at each time step conditioned on any tokens that have already been generated in previous time steps.

[0042]

[0037] The language model neural network can have any of a variety of Transformer-based neural network architectures, e.g., encoder-only Transformer architectures, encoder-decoder Transformer architectures, decoder-only Transformer architectures, other attention-based architectures, and so on.

[0043]

[0038] Examples of language model neural networks include those described in Cohn Raffel, et al. Exploring the limits of transfer learning with a unified text-to-text transformer. arXiv preprint arXiv: 1910.10683, 2019; Daniel Adiwardana, et al. Towards a human-like opendomain chatbot. CoRR, abs / 2001.09977, 2020; Tom B Brown, et al. Language models are few-shot learners. arXiv preprint arXiv:2005. 14165, 2020; Aakanksha Chowdhery, et al. PaLM: Scaling Language Modeling with Pathways, arXiv preprint arXiv:2204.02311; Rohan Anil, et al. Palm 2 technical report. arXiv preprint arXiv:2305.10403, 2023; Borsos, Zalan, et al. AudioIm: a language modeling approach to audio generation. IEEE / ACM Transactions on Audio, Speech, and Language Processing (2023); and Agostinelli, Andrea, et al. Musiclm: Generating music from text." arXiv preprint arXiv:2301.11325 (2023).

[0044]

[0039] In some implementations, the generative neural network 110 can be a diffusion model neural network that executes a reverse diffusion process to iteratively generate a customized content item 152, e.g., an image, a video, or an audio, across multiple reverse diffusion steps starting from random noise. Attorney Docket No.: 56113-0771WO1

[0045]

[0040] For example, the diffusion model neural network can generate an image by performing a reverse diffusion process to generate a diffusion output that includes or otherwise specifies a plurality of color values for pixels in the image arranged according to a specified order.

[0046]

[0041] As another example, the diffusion model neural network can generate an image by performing a reverse diffusion process to generate a diffusion output that includes or otherwise specifies a plurality of tokens that represent image patch embeddings of the image which can then be processed by a decoder neural network to generate the image.

[0047]

[0042] Examples of diffusion model neural networks include those described in Chitwan Saharia, et al. Photorealistic text-to-image diffusion models with deep language understanding. Advances in Neural Information Processing Systems, 35:36479-36494, 2022; Adily a Ramesh, et al. Hierarchical text-conditional image generation with clip latents. arXiv preprint arXiv: 2204.06125; and Robin Rombach, et al. High-resolution image synthesis with latent diffusion model, Proceedings of the IEEE / CVF conference on computer vision and pattern recognition. 2022.

[0048]

[0043] As another example, the diffusion model neural network can generate an image or a video that has multiple frames (where each frame is an image) by iteratively predicting masked tokens over a decoding process in a discrete token space, e.g., as described in Huiwen Chang, et al. Muse: Text-to-image generation via masked generative transformers. arXiv preprint arXiv:2301.00704, 2023. and Huiwen Chang, et al. Maskgit: Masked generative image transformer. arXiv preprint arXiv:2202.04200, 2022.

[0049]

[0044] The content item generation system 100 includes or accesses a plurality' of entity reward models 120A-N. The plurality of entity' reward models 120A-N correspond (or map) respectively to a plurality of entities. For example, in FIG. 1, a first entity reward model 120A corresponds to a first entity, a second entity reward model 120B corresponds to a second entity7, a third entity reward model 120C corresponds to a third entity, and so on.

[0050]

[0045] The plurality of entity7reward models 120A-N can be used together with the generative neural network 110 to generate the customized content items 152. As will be described in more detail below with reference to FIGS. 2-5. upon receiving a request 102 for a customized content item 152 submitted by7a user of the system, the content item generation system 100 can first identify, from among the plurality7of entity7reward models 120A-N, a target entity7reward model based on the request 102, and then use the generative neural network 110 to perform a data generation process that is guided by the target entity reward model to generate the customized content item 152 in response to the request 102. Attorney Docket No.: 56113-0771WO1

[0051]

[0046] The generative neural network 110 can have been trained, e g., based on optimizing a next token prediction objective (when it is configured as a language model neural network) or a score matching objective (when it is configured as a diffusion model neural network), independently of the plurality of entity reward models 120A-N. That is, an entity reward model generally was not used to guide the data generation process during training of the generative neural network 110. The content item generation system 100 can thus utilize the plurality of entity reward models 120A-N to enhance an “off-the-shelf generative neural network 110 to generate the customized content items 152.

[0052]

[0047] In some implementations, the content item generation system 100 locally stores the plurality of entity reward models 120A-N.

[0053]

[0048] In some other implementations, the content item generation system 100 does not locally store any of the plurality of entity reward models 120A-N. For example, each entity can store a corresponding entity reward model in a remote system, e.g., that runs on a client device, and the content item generation system 100 can access the corresponding entity reward model of the entity through an application programming interface (API) or another data interface made available by the remote system.

[0054]

[0049] For example, the API can be a secure API that provides a secure communication channel between the remote system and the content item generation system 100. As another example, the API can be an access protected API, such that the entity reward model, which is implemented in the remote system, is permission protected - for example, the entity reward model may be accessible to the content item generation system 100 only when an access request is granted by the remote system. In this example, the content item generation system 100 can generate content items that are customized to a given entity while ensuring the protection of the privacy and security of the data of the given entity.

[0055]

[0050] In some other implementations, the content item generation system 100 locally stores some of the plurality of entity rew ard models 120A-N that correspond to a subset of plurality of entities, while others of the plurality of entity rew ard models 120A-N that correspond to another subset of plurality of entities are stored at remote systems.

[0056]

[0051] Each entity rew ard model has a respective set of reward model parameters and is configured to process, in accordance with the respective set of reward model parameters, a rew ard model input that includes a final content item or an intermediate content item to compute a reward score indicating how likely the final content item or the intermediate content item is customized to an entity corresponding to the entity reward model. Optionally, Attorney Docket No.: 56113-0771WO1 when the content item generation system 100 also receives a conditioning input, the reward model input can also include the conditioning input.

[0057]

[0052] The final content item includes data generated as a result of a data generation process executed by the generative neural network 110.

[0058]

[0053] For example, the final content item can be the customized content item 152 that is generated as a result of an auto-regressive token generation process executed by a language model neural network, i.e.. after the last time step in the multiple time steps in the autoregressive token generation process.

[0059]

[0054] As another example, the final content item can be the customized content item 152 that is generated as a result of a reverse diffusion process executed by a diffusion model neural network, i.e., after the last reverse diffusion step in the multiple reverse diffusion steps in the reverse diffusion process.

[0060]

[0055] An intermediate content item includes data generated at any intermediate time point during (but before the end of) a data generation process executed by the generative neural network 110.

[0061]

[0056] For example, the intermediate content item can be an incomplete sequence of tokens included in the customized content item 152 that is generated as of an intermediate time step in the multiple time steps in the auto-regressive token generation process that is executed by the language model neural network. The sequence of tokens included in the intermediate content item is called “incomplete7’ because the language model neural network will continue to add additional tokens to the sequence of tokens in subsequent time steps in the autoregressive token generation process in order to provide the customized content item 152.

[0062]

[0057] As another example, the intermediate content item can be an intermediate representation of the customized content item 152 that is generated as of an intermediate reverse diffusion step in the multiple reverse diffusion steps in the reverse diffusion process that is executed by the diffusion model neural network. The intermediate representation of the customized content item 152 may be referred to as a noisy version of the customized content item 152 because the intermediate representation includes noise (relative to the customized content item 152) which will need to be removed in subsequent reverse diffusion steps in the reverse diffusion process in order to provide the customized content item 152.

[0063]

[0058] Each entity' reward model can have any of a variety' of architectures. For example, each entity reward model can be one of: a neural network, a decision tree model, a random forest model, a gradient boosting model, a linear regression model, a logistical regression model, a support vector machine (SVM) model, and so on. As another example, each of Attorney Docket No.: 56113-0771WO1 some or all of the plurality of entity reward models 120A-N can be initiated from the generative neural network 110, and then fine-tuned (e.g., using Low Rank Adaptation (LoRA) fine-tuning) or otherwise adapted to a corresponding entity by learning a set of entity-specific parameter values.

[0064]

[0059] In some implementations, the plurality of entity reward models 120A-N have the same architecture (but different values for the respective sets of reward model parameters), whereas, in other implementations, the plurality of entity reward models 120A-N have different architectures. For example, the first entity reward model 120 A can be a convolutional neural network while the second entity reward model 120B can be a feedforward neural network.

[0065]

[0060] As a particular example, similar to the generative neural network 110, an entity reward model can be implemented as a language model neural network configured to process an input that includes a final content item or an intermediate content item, and generate an output indicating the likelihood, as determined by the language model neural network, that the final content item or the intermediate content item is customized to an entity corresponding to the entity reward model.

[0066]

[0061] For example, the output of the entity reward model can be a value between zero and one, with zero indicating that the final content item or the intermediate content item is unlikely to be customized to the entity, and one indicating that the final content item or the intermediate content item being is likely to be customized to the entity.

[0067]

[0062] As another example, the output of the entity reward model can include tokens representing one of a predetermined list of classifications, e.g., “customized” or “generic”; or “customized” or “generic” or “hybrid”. The entity reward model can then output a predetermined score assigned to each classification generated by the language model neural network as the reward score.

[0068]

[0063] In some implementations, an entity reward model can be trained on a training dataset that includes a plurality of training content items. Each training content item is associated with a ground truth reward score indicating a likelihood that the training content item is customized to an entity.

[0069]

[0064] For example, the training dataset can include a customized corpus of content items that are customized for an entity7corresponding to the entity rew ard model, and the ground truth reward scores associated with the content items included in the customized corpus can indicate that the training content items are customized for an entity, e.g., can be a value of one. Attorney Docket No.: 56113-0771WO1

[0070]

[0065] A customized corpus of content items may include documents, presentations, news items, articles, blog posts, books, book reviews, magazines, magazine articles, text messages, e-mail messages, social media content, images, photos, audio recordings, video recordings, or any other type of information authored by either the entity or another preferred content provider identified by the entity.

[0071]

[0066] Optionally, the training dataset can include another corpus of content items that are not customized for the entity, e.g.. that are instead customized for other entities, and the ground truth reward scores associated with the content items included in the other, noncustomized corpus can indicate that the training content items are not customized for an entity , e.g., can be a value of zero.

[0072]

[0067] In these implementations, the entity reward model can be trained based on optimizing a classification loss function that measures, for each training content item, a difference between the ground truth reward score and a predicted reward score generated by the entity reward model based on processing the training content item.

[0073]

[0068] In effect, the entity reward model is trained to generate predicted reward scores for each training content item that match the corresponding ground truth reward scores associated with the content item - for example predicted reward scores that are closer to one for content items included in the customized corpus and predicted reward scores that are closer to zero for content items included in the non-customized corpus.

[0074]

[0069] FIG. 2 is a flow diagram of an example process 200 for generating a customized content item. For convenience, the process 200 will be described as being performed by a system of one or more computers located in one or more locations. For example, a content item generation system, e.g.. the content item generation system 100 depicted in FIG. 1, appropriately programmed in accordance with this specification, can perform the process 200.

[0075]

[0070] The system receives a request to generate a customized content item that is customized for a target entity' by using a generative neural network (step 202). The system can use the generative neural network to generate any kind of customized output data items, e.g., textual data items, image data items, video content items, audio data items, and so on.

[0071] The system identifies, from among a plurality of entity reward models that correspond respectively to a plurality' of entities, a target entity' reward model that corresponds to the target entity' (step 204). Specifically, the system identifies one of the plurality of entities as the target entity, and then uses the entity reward model that corresponds to the identified target entity as the target entity reward model. Attorney Docket No.: 56113-0771WO1

[0076]

[0072] The target entity can be identified in any of a variety of ways based on the request. In some cases, the system can receive a conditioning input as part of, or associated with, the request to generate the customized content item from a user. Such a conditioning input provides context for the customized content item. The system can then use the conditioning input to determine which one of the plurality' of entities should be the target entity7.

[0077]

[0073] For example, when the conditioning input includes data that identifies a recipient of the customized content item, e.g., the conditioning input is “Write a blog post for tennis enthusiasts” or “Make an image for a birthday card for Alice,” the system can use the recipient identified in the conditioning input to determine the target entity, e.g., a target entity7of a named group “tennis enthusiasts” or a target entity of an individual “Alice.”

[0074] As another example, the system can process the conditioning input using an entity classification model to generate an entity classification output that defines the target entity. For example, the entity' classification model can be one of: a neural network, a decision tree model, a random forest model, a gradient boosting model, a linear regression model, a logistical regression model, a support vector machine (SVM) model, and so on, and the entity classification output includes a respective score for each of the plurality of entities. The system can then determine the target entity based on the respective scores for the plurality of entities.

[0078]

[0075] In some cases, the system can obtain metadata associated with a user session between the user and the system. Such metadata includes information that characterizes the user, e.g.. information about the user’s identity, the user’s social network, social actions, or activities, profession, the user’s preferences, or the user’s current geographic location. The system can then use the metadata to determine which one of the plurality of entities should be the target entity. For example, the system can determine, as the target entity, one of the plurality of entities that is within a proximity of the user’s current geographic location, or that is a contact in the user’s social network.

[0079]

[0076] In these cases, a user may be provided with controls (e.g., user interface elements with which a user can interact) allowing the user to make an election as to both if and when systems, programs, or features described herein may enable collection of user information (e.g., information about a user’s social network, social actions, or activities, profession, a user’s preferences, or a user's current location), and if the user is sent content or communications from a server. In addition, certain data may be treated in one or more ways before it is stored or used, so that personally identifiable information is removed. For example, a user’s identity may be treated so that no personally identifiable information can be Attorney Docket No.: 56113-0771WO1 determined for the user, or a user’s geographic location may be generalized where location information is obtained (such as to a city; postal code, or state level), so that a particular location of a user cannot be determined. Thus, the user may have control over what information is collected about the user, how that information is used, and what information is provided to the user.

[0080]

[0077] The system generates the customized content item based on using the generative neural network to perform a data generation process that is guided by the target entity reward model (step 206). That is, the system uses both the generative neural network and the target entity7reward model to generate the customized content item in response to the request.

[0081]

[0078] Depending on the configuration of the generative neural network, there are many ways in which a target entity reward model-guided data generation process can be performed. A few examples of the target entity reward model-guided data generation process will be described below in FIGS 3-5.

[0082]

[0079] FIG. 3 is a flow diagram of an example process 300 for performing a data generation process. For convenience, the process 300 will be described as being performed by a system of one or more computers located in one or more locations. For example, a content item generation system, e.g., the content item generation system 100 depicted in FIG. 1, appropriately programmed in accordance with this specification, can perform the process 300.

[0083]

[0080] The system generates, using the generative neural network, a plurality7of candidate content items (step 302). That is, the system uses the same generative neural network to generate multiple different candidate content items in response to the same request. The generative neural network can be any appropriate generative neural network, e.g., a language model neural network or a diffusion model neural network. The multiple different candidate content items can be generated either in parallel or sequentially.

[0084]

[0081] For example, when configured as a language model neural network that operates auto- regressively, the system can use the same language model neural network to generate multiple different candidate content items in response to the same request, e.g.. by using beam search decoding from score distributions generated by the language model neural network, using a Sample-and-Rank decoding strategy, or using another decoding strategy that leverages the auto-regressive nature of the language model neural network.

[0085]

[0082] Once the plurality7of candidate content items are generated, for each of the plurality7of candidate content items, the system processes a reward model input that includes the candidate content item using the target entity reward model to generate a respective reward Attorney Docket No.: 56113-0771WO1 score (step 304). Optionally, the reward model input can also include the conditioning input. For each candidate content item, the respective reward score indicates a likelihood, determined by the target entity reward model, that the candidate content item is customized to the target entity corresponding to the target entity reward model.

[0086]

[0083] The system selects, as the customized content item, a candidate content item from the plurality of candidate content items in accordance with the respective reward scores (step 306). For example, the system can then select, as the customized content item to be provided in response to the request, the candidate customized content item with the highest respective reward score. In effect, the system selects the candidate content item that is most likely to be customized for the target entity.

[0087]

[0084] As another example, the system can maintain a threshold value that can be compared against the respective reward scores. In this example, the system can refrain from selecting a candidate customized content item if the respective reward score does not satisfy the threshold value (even if the candidate customized content item would otherwise have the highest respective reward score), and instead re-run the generative neural network to generate additional candidate customized content items.

[0088]

[0085] FIG. 4 is a flow diagram of another example process 400 for performing a data generation process. For convenience, the process 400 will be described as being performed by a system of one or more computers located in one or more locations. For example, a content item generation system, e.g., the content item generation system 100 depicted in FIG. 1, appropriately programmed in accordance with this specification, can perform the process 400.

[0089]

[0086] In the example of FIG. 4, the generative neural network can be a language model neural network that executes an auto-regressive token generation process to auto-regressively generate a customized content item, e.g., a sequence of text tokens, a sequence of pixel tokens, a sequence of audio tokens, a sequence of multi-modal tokens, e.g., text and pixel tokens, or the like, across multiple time steps, for example by generating one token at each time step conditioned on any tokens that have already been generated in previous time steps.

[0087] The system generates, using the language model neural network, a score distribution for a particular position in the customized content item that corresponds to a particular time step in the multiple time steps (step 402). To generate the score distribution for the particular position, the language model neural network processes at least the most recently selected token, i.e., the token that is selected for the immediately preceding position in the customized content item that precedes the particular position of the particular token. Attorney Docket No.: 56113-0771WO1

[0090]

[0088] The score distribution assigns a respective score to each token in a vocabulary' of tokens. For example, the score distribution can be a probability distribution, and the respective score assigned to each token can be a respective probability score.

[0091]

[0089] The vocabulary' of tokens can include any of a variety7of tokens that represent text symbols or other symbols. For example, the vocabulary' of tokens can include one or more of: characters, sub-words, words, punctuation marks, numbers, or other symbols that appear in a corpus of natural language text and / or computer code.

[0092]

[0090] Additionally or alternatively, the vocabulary' of tokens can include tokens that can represent data other than text. For example, the vocabulary' of tokens can include image tokens that represent a discrete set of image patch embeddings of an image that can be generated by an image encoder neural network based on processing the image patches of the image. As another example, the vocabulary of tokens can include audio tokens that represent code vectors in a codebook of a quantizer, e.g., a residual vector quantizer.

[0093]

[0091] The system selects at least a subset of tokens from the vocabulary of tokens in accordance with the score distribution (step 404). For example, the system can select the top- k highest-scoring tokens or can sample, e.g., using nucleus sampling or another sampling technique, one or more tokens based on the score distribution. As another example, the system can select all of the tokens from the vocabulary' for inclusion in the subset. As yet another example, the system can select tokens having scores that are above a predetermined threshold.

[0094]

[0092] For each of the tokens in the subset, the system processes a reward model input that includes at least the token using the target entity' reward model to generate a respective reward score (step 406). Optionally, the reward model input can also include any tokens that have already been generated in previous time steps that precede the particular time step. Further optionally, the reward model input can also include the conditioning input.

[0095]

[0093] Thus, the reward model input can include an intermediate content item, e.g., an incomplete sequence of tokens, that is generated as of the particular time step in the multiple time steps in the auto-regressive token generation process that is executed by the language model neural network.

[0096]

[0094] For each token, the respective rew ard score indicates a likelihood, determined by7the target entity7rew ard model, that an intermediate content item that includes the token, and, optionally, any previously generated tokens is customized to the target entity corresponding to the target entity reward model. Attorney Docket No.: 56113-0771WO1

[0097]

[0095] The system selects, as a particular token to occupy the particular position in the customized content item, a token from the subset of tokens in accordance with the respective reward scores (step 408). In some implementations, the system selects, as the particular token to occupy the particular position, the token that yields the highest reward score among all tokens in the subset.

[0098]

[0096] In some implementations, for each of the tokens in the subset, the system computes a respective weighted score for the token based on (i) the respective score that is included in the score distribution generated by using the language model neural network and that is assigned to token and (ii) the respective reward score generated by using the target entity reward model. For example, for each token, the respective weighted score can be computed by multiplying the respective score included in the score distribution with the respective reward score.

[0099]

[0097] In these implementations, the system can then select, as the particular token to occupy the particular position, the token with the highest respective weighted score. In effect, the system selects the token that is most likely to be customized for the target entity to add to the incomplete sequence of tokens that has been generated so far.

[0100]

[0098] By repeatedly performing the process 400 for some or all of the multiple time steps in the auto-regressive token generation process to generate the respective tokens to occupy the multiple positions, the system can generate the customized content item which includes respective tokens to be provided in response to the request. For example, the process 400 can be performed at predetermined (e.g., fixed) intervals, during the multiple time steps, e.g., performed every 5 time steps, 10 time steps, 20 time steps, and so on.

[0101]

[0099] FIG. 5 is a flow diagram of another example process 500 for performing a generation process. For convenience, the process 500 will be described as being performed by a system of one or more computers located in one or more locations. For example, a content item generation system, e.g., the content item generation system 100 depicted in FIG. 1, appropriately programmed in accordance with this specification, can perform the process 500.

[0102]

[0100] In the example of FIG. 5, the generative neural network can be a diffusion model neural network that executes a reverse diffusion process to iteratively generate a customized content item, e.g., an image, a video, or an audio, across multiple reverse diffusion steps starting from random noise.

[0103]

[0101] The system obtains an intermediate representation of the customized content item at a particular reverse diffusion step in the multiple reverse diffusion steps (step 502). If the Attorney Docket No.: 56113-0771WO1 particular reverse diffusion step is the first reverse diffusion step in the reverse diffusion process, the intermediate representation is an initial (e.g., a randomly initialized) intermediate representation . For any subsequent reverse diffusion step, the intermediate representation is the updated intermediate representation that has been generated in the immediately preceding reverse diffusion step.

[0104]

[0102] The system processes a diffusion model input that includes at least the intennediate representation using the diffusion model neural network to generate a diffusion model output (step 504). Optionally, the diffusion model input can also include the conditioning input. Further optionally, the diffusion model input can also include a timestep index that corresponds to the particular reverse diffusion step in the multiple reverse diffusion steps. In some implementations, the diffusion model output can be an estimate of the noise that needs to be, e.g., added to the customized content item being generated by the system, to arrive at the intermediate representation of the customized content item.

[0105]

[0103] The system processes at least the intermediate representation using the target entityreward model to generate a reward score (step 506). In some implementations, the target entity reward model can be configured as a classifier model, and the reward score can be in the form of a classification output that indicates whether the intermediate representation is customized to the target entity or not customized to the target entity.

[0106]

[0104] The system updates, based on the diffusion model output and the reward score, the intermediate representation to generate an updated intermediate representation of the customized content item at the particular reverse diffusion step (step 508). In some implementations, the updated intermediate representation can be generated by the system by computing, with respect to the intermediate representation, gradients of the reward score; and modifying the intermediate representation using the computed gradients of the reward score. Then, the system can determine the updated intermediate representation based on the modified intermediate representation.

[0107]

[0105] For example, the updated intermediate representation can be determined by: '( / x + ,$£ Vr(log p^(? / |. ), £) where represents the target entity- reward model (a classifier model), peand Jeare parameters included in the diffusion model output that defines the mean and diagonal covariance matrix, respectively, of a noise distribution, e.g., a diagonal Gaussian distribution, represents the gradients of the reward score (a classification output), xt Attorney Docket No.: 56113-0771WO1 is the intermediate representation, xt-1is the updated intermediate representation, and s is the gradient scale.

[0108]

[0106] When the particular reverse diffusion step is not the last reverse diffusion step in the reverse diffusion process, the system can perform a next iteration of the process 500, where the updated intermediate representation will be used as the intermediate representation in the next iteration of the process 500.

[0109]

[0107] Alternatively, when the particular reverse diffusion step is the last reverse diffusion step, the system can generate the customized content item based on the updated intermediate representation. For example, the system can process the updated intermediate representation using a decoder neural network to generate a decoder output and use the decoder output as the customized content item. As another example, the updated intermediate representation at the last reverse diffusion step corresponds to the customized content item.

[0110]

[0108] This specification uses the term “configured” in connection with systems and computer program components. For a system of one or more computers to be configured to perform particular operations or actions means that the system has installed on it software, firmware, hardware, or a combination of them that in operation cause the system to perform the operations or actions. For one or more computer programs to be configured to perform particular operations or actions means that the one or more programs include instructions that, when executed by data processing apparatus, cause the apparatus to perform the operations or actions.

[0111]

[0109] Embodiments of the subject matter and the functional operations described in this specification can be implemented in digital electronic circuitry, in tangibly-embodied computer software or firmware, in computer hardware, including the structures disclosed in this specification and their structural equivalents, or in combinations of one or more of them. Embodiments of the subject matter described in this specification can be implemented as one or more computer programs, i.e., one or more modules of computer program instructions encoded on a tangible non transitory storage medium for execution by, or to control the operation of, data processing apparatus. The computer storage medium can be a machine- readable storage device, a machine-readable storage substrate, a random or serial access memory device, or a combination of one or more of them. Alternatively or in addition, the program instructions can be encoded on an artificially generated propagated signal, e.g., a machine-generated electrical, optical, or electromagnetic signal, that is generated to encode information for transmission to suitable receiver apparatus for execution by a data processing apparatus. Attorney Docket No.: 56113-0771WO1

[0112]

[0110] The term “data processing apparatus” refers to data processing hardware and encompasses all kinds of apparatus, devices, and machines for processing data, including by way of example a programmable processor, a computer, or multiple processors or computers. The apparatus can also be, or further include, special purpose logic circuitry, e.g., an FPGA (field programmable gate array) or an ASIC (application specific integrated circuit). The apparatus can optionally include, in addition to hardware, code that creates an execution environment for computer programs, e.g.. code that constitutes processor firmware, a protocol stack, a database management system, an operating system, or a combination of one or more of them.

[0113]

[0111] A computer program, which may also be referred to or described as a program, software, a software application, an app. a module, a software module, a script, or code, can be written in any form of programming language, including compiled or interpreted languages, or declarative or procedural languages; and it can be deployed in any form, including as a stand alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment. A program may, but need not, correspond to a file in a file system. A program can be stored in a portion of a file that holds other programs or data, e.g., one or more scripts stored in a markup language document, in a single file dedicated to the program in question, or in multiple coordinated files, e.g., files that store one or more modules, sub programs, or portions of code. A computer program can be deployed to be executed on one computer or on multiple computers that are located at one site or distributed across multiple sites and interconnected by a data communication network.

[0114]

[0112] In this specification, the term “database” is used broadly to refer to any collection of data: the data does not need to be structured in any particular way, or structured at all, and it can be stored on storage devices in one or more locations. Thus, for example, the index database can include multiple collections of data, each of which may be organized and accessed differently.

[0115]

[0113] Similarly, in this specification the term “engine” is used broadly to refer to a software-based system, subsystem, or process that is programmed to perform one or more specific functions. Generally, an engine will be implemented as one or more software modules or components, installed on one or more computers in one or more locations. In some cases, one or more computers will be dedicated to a particular engine; in other cases, multiple engines can be installed and running on the same computer or computers.

[0116]

[0114] The processes and logic flows described in this specification can be performed by one or more programmable computers executing one or more computer programs to perform Attorney Docket No.: 56113-0771WO1 functions by operating on input data and generating output. The processes and logic flows can also be performed by special purpose logic circuitry, e.g.. an FPGA or an ASIC, or by a combination of special purpose logic circuitry and one or more programmed computers.

[0117]

[0115] Computers suitable for the execution of a computer program can be based on general or special purpose microprocessors or both, or any other kind of central processing unit. Generally, a central processing unit will receive instructions and data from a read only memory or a random access memory or both. The essential elements of a computer are a central processing unit for performing or executing instructions and one or more memory devices for storing instructions and data. The central processing unit and the memory' can be supplemented by, or incorporated in, special purpose logic circuitry. Generally, a computer will also include, or be operatively coupled to receive data from or transfer data to. or both, one or more mass storage devices for storing data, e.g., magnetic, magneto optical disks, or optical disks. However, a computer need not have such devices. Moreover, a computer can be embedded in another device, e.g.. a mobile telephone, a personal digital assistant (PDA), a mobile audio or video player, a game console, a Global Positioning System (GPS) receiver, or a portable storage device, e.g., a universal serial bus (USB) flash drive, to name just a few.

[0118]

[0116] Computer readable media suitable for storing computer program instructions and data include all forms of non volatile memory', media and memory devices, including by way of example semiconductor memory devices, e.g., EPROM, EEPROM, and flash memory devices; magnetic disks, e.g., internal hard disks or removable disks; magneto optical disks; and CD ROM and DVD-ROM disks.

[0119]

[0117] To provide for interaction with a user, embodiments of the subject matter described in this specification can be implemented on a computer having a display device, e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor, for displaying information to the user and a keyboard and a pointing device, e.g., a mouse or a trackball, by which the user can provide input to the computer. Other kinds of devices can be used to provide for interaction with a user as well; for example, feedback provided to the user can be any form of sensory feedback, e.g., visual feedback, auditory feedback, or tactile feedback; and input from the user can be received in any form, including acoustic, speech, or tactile input. In addition, a computer can interact with a user by sending documents to and receiving documents from a device that is used by the user; for example, by sending web pages to a web browser on a user’s device in response to requests received from the web browser. Also, a computer can interact with a user by sending text messages or other forms of message to a personal device, Attorney Docket No.: 56113-0771WO1 e.g., a smartphone that is running a messaging application, and receiving responsive messages from the user in return.

[0120]

[0118] Data processing apparatus for implementing machine learning models can also include, for example, special-purpose hardware accelerator units for processing common and compute-intensive parts of machine learning training or production, i.e., inference, workloads.

[0121]

[0119] Machine learning models can be implemented and deployed using a machine learning framework, e.g., a TensorFlow framework or a Jax framework.

[0122]

[0120] Embodiments of the subject matter described in this specification can be implemented in a computing system that includes a back end component, e.g., as a data server, or that includes a middleware component, e.g., an application server, or that includes a front end component, e.g., a client computer having a graphical user interface, a web browser, or an app through which a user can interact with an implementation of the subject matter described in this specification, or any combination of one or more such back end, middleware, or front end components. The components of the system can be interconnected by any form or medium of digital data communication, e.g., a communication network. Examples of communication networks include a local area network (LAN) and a wide area network (WAN), e.g., the Internet.

[0123]

[0121] The computing system can include clients and servers. A client and server are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other. In some embodiments, a server transmits data, e.g., an HTML page, to a user device, e.g.. for purposes of displaying data to and receiving user input from a user interacting with the device, which acts as a client. Data generated at the user device, e.g., a result of the user interaction, can be received at the server from the device.

[0124]

[0122] While this specification contains many specific implementation details, these should not be construed as limitations on the scope of any invention or on the scope of what may be claimed, but rather as descriptions of features that may be specific to particular embodiments of particular inventions. Certain features that are described in this specification in the context of separate embodiments can also be implemented in combination in a single embodiment. Conversely, various features that are described in the context of a single embodiment can also be implemented in multiple embodiments separately or in any suitable subcombination. Moreover, although features may be described above as acting in certain combinations and Attorney Docket No.: 56113-0771WO1 even initially be claimed as such, one or more features from a claimed combination can in some cases be excised from the combination, and the claimed combination may be directed to a subcombination or variation of a subcombination.

[0125]

[0123] Similarly, while operations are depicted in the drawings and recited in the claims in a particular order, this should not be understood as requiring that such operations be performed in the particular order shown or in sequential order, or that all illustrated operations be performed, to achieve desirable results. In certain circumstances, multitasking and parallel processing may be advantageous. Moreover, the separation of various system modules and components in the embodiments described above should not be understood as requiring such separation in all embodiments, and it should be understood that the described program components and systems can generally be integrated together in a single software product or packaged into multiple software products.

[0126]

[0124] Particular embodiments of the subject matter have been described. Other embodiments are within the scope of the following claims. For example, the actions recited in the claims can be performed in a different order and still achieve desirable results. As one example, the processes depicted in the accompanying figures do not necessarily require the particular order shown, or sequential order, to achieve desirable results. In some cases, multitasking and parallel processing may be advantageous.

[0127]

[0125] What is claimed is:

Claims

Attorney Docket No.: 56113-0771WO1CLAIMS1. A computed-implemented method comprising: receiving a request to generate, using a generative neural network, a customized content item that is customized for a target entity from a plurality of entities, wherein the generative neural network has a set of generative neural network parameters and is configured to generate a final content item in accordance with the set of generative neural network parameters; identifying a target entity reward model that corresponds to the target entity, wherein the target entity reward model corresponding to the target entity has a respective set of reward model parameters and is configured to process the final content item in accordance with the respective set of reward model parameters to generate a reward score that indicates a likelihood that the final content item is customized to the target entity; and generating the customized content item, comprising using the generative neural network to perform a data generation process that is guided by the target entity reward model.

2. The method of claim 1, wherein receiving the request to generate the customized content item comprises receiving a conditioning input as part of or in association with the request, and wherein identifying the target entity reward model comprises identifying the target entity reward model based on the conditioning input.

3. The method of claim 2, wherein identify ing the target entity reward model based on the conditioning input comprises processing the conditioning input using an entity classification model to generate an entity classification output that defines the target entity.

4. The method of claim 2, wherein receiving the request to generate the customized content item comprises receiving data identifying the target entity, and wherein identifying the target entity reward model comprises identifying the target entity reward model in accordance wi th the data.

5. The method of any one of claims 1-4, wherein the target entity reward model and the generative neural network are hosted on a same server.

6. The method of any one of claims 1-4, wherein the target entity reward models and the generative neural network are hosted on different servers.

7. The method of any one of claims 1-4, wherein the target entity reward model is hosted on a client device and the generative neural network is hosted on a server.Attorney Docket No.: 56113-0771WO18. The method of any one of claims 6-7, wherein performing the generation process that is guided by the target entity reward model comprises using the target entity reward model through an access protected application programming interface (API).

9. The method of any one of claims 1-8, wherein generating the customized content item based on using the generative neural network comprises: generating, using the generative neural network, a plurality of candidate content items; for each of the plurality of candidate content items: processing the candidate content item using the target entity reward model to generate a respective reward score; and selecting, as the customized content item, a candidate content item from the plurality of candidate content items in accordance with the respective reward scores.

10. The method of any one of claims 1-9, wherein the target entity reward model is trained on a reward model training dataset that comprises a plurality of training content items and, for each of the plurality of training content items, a ground truth reward score.

11. A computed-implemented method comprising: receiving a request to generate, using a generative neural network, a customized content item that is customized for a target entity from a plurality of entities, wherein the generative neural network has a set of generative neural network parameters and is configured to generate an intermediate content item in accordance with the set of generative neural network parameters; identifying a target entity reward model that corresponds to the target entity, wherein the target entity reward model corresponding to the target entity- has a respective set of reward model parameters and is configured to process the intermediate content item in accordance with the respective set of reward model parameters to generate a reward score that indicates a likelihood that the intermediate content item is customized to the target entity; and generating the customized content item, comprising using the generative neural network to perform a data generation process that is guided by the target entity reward model.

12. The method of claim 11, wherein the generative neural network is a language model neural network, wherein the generation process is an auto-regressive token generation process, and wherein generating the customized content item based on using the generative neural network comprises, at a particular position in the customized content item: generating, using the language model neural network, a score distribution that assignsAttorney Docket No.: 56113-0771WO1 a respective score to each token in a vocabulary of tokens; selecting a subset of tokens from the vocabulary of tokens in accordance with the score distribution; for each of the tokens in the subset: processing at least the token using the target entity reward model to generate a respective reward score; and selecting, as a particular token to occupy the particular position, a token from the subset of tokens in accordance with the respective reward scores.

13. The method of any one of claims 11-12, wherein the generative neural network is a diffusion model neural network, wherein the generation process is a reverse diffusion process, and wherein generating the customized content item based on using the generative neural network comprises, at a particular reverse diffusion step: obtaining an intermediate representation of the customized content item at the particular reverse diffusion step; processing at least the intermediate representation using the diffusion model neural network to generate a diffusion model output; processing at least the intermediate representation using the target entity reward model to generate a reward score; and updating, based on the diffusion model output and the reward score, the intermediate representation to generate an updated intermediate representation of the customized content item at the particular reverse diffusion step.

14. The method of any one of claims 11-13, wherein the target entity reward model is trained on a reward model training dataset that comprises a plurality of training content items and, for each of the plurality of training content items, a ground truth reward score.

15. The method of any one of claims 11-14, wherein receiving the request to generate the customized content item comprises receiving a conditioning input as part of or in association with the request, and wherein identifying the target entity reward model comprises identifying the target entity reward model based on the conditioning input.

16. The method of claim 15, w herein identifying the target entity rew ard model based on the conditioning input comprises processing the conditioning input using an entity classification model to generate an entity classification output that defines the target entity.

17. The method of claim 15, wherein receiving the request to generate the customized content item comprises receiving data identifying the target entity, and wherein identifyingAttorney Docket No.: 56113-0771WO1 the target entity reward model comprises identifying the target entity reward model in accordance with the data.

18. The method of any one of claims 11-17, wherein the target entity reward model and the generative neural network are hosted on a same server.

19. The method of any one of claims 11-17, wherein the target entity reward models and the generative neural network are hosted on different servers.

20. The method of any one of claims 11-17, wherein the target entity reward model is hosted on a client device and the generative neural network is hosted on a server.

21. The method of any one of claims 19-20, wherein performing the generation process that is guided by the target entity reward model comprises using the target entity reward model through an access protected application programming interface (API).

22. A system comprising one or more computers and one or more storage devices storing instructions that, when executed by the one or more computers, cause the one or more computers to perform the respective operations of any one of the methods of any of the preceding claims 1-21.

23. One or more computer storage media storing instructions that, when executed by one or more computers, cause the one or more computers to perform the respective operations of any one of the methods of any of the preceding claims 1-21.

Citation Information

Patent Citations

  • System and method for a personalized search and discovery engine

    US20210174164A1

  • Deep generation of user-customized items

    US20210192594A1