Systems and methods for multi-reward reinforcement learning framework for text-to-image generation
Patent Information
- Application Number
- EP2024820871
- Authority / Receiving Office
- EP · EP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-01-10
- Filing Date
- 2024-11-13
- Publication Date
- 2025-08-27
AI Technical Summary
Existing image generation models struggle to generate images that satisfy performance metrics due to insufficient contextual information in textual prompts, leading to a loss of focus on the original prompt and degradation of image quality.
A multi-reward reinforcement learning framework is employed to jointly train a prompt expansion model and an image generation model, optimizing both models simultaneously to enhance the contextual relevance and quality of generated images by incorporating multiple reward criteria.
The framework improves image generation by maintaining focus on the original textual prompt while enriching it with additional context, resulting in higher-quality images that meet multiple performance metrics such as visual clarity, resolution, and aesthetic appeal.
Smart Images

Figure US2024055734_17072025_PF_FP_ABST
Abstract
Description
SYSTEMS AND METHODS FOR MULTI-REWARD REINFORCEMENT LEARNING FRAMEWORK FOR TEXT-TO-IMAGE GENERATIONCROSS-REFERENCE TO RELATED APPLICATIONS
[0001] The present application claims priority to US Provisional Application No.: 63 / 619,632, entitled “MULTI-REWARD REINFORCEMENT LEARNING FRAMEWORK FOR TEXT-TO-IMAGE GENERATION,” filed on January 10, 2024, the disclosure of which is incorporated by reference herein in its entirety.FIELD
[0002] The present disclosure relates generally to machine-learning processes and machine-learned devices and systems. More particularly, the present disclosure relates to a multi-reward reinforcement learning framework for text-to-image generation.BACKGROUND
[0003] A computer can receive input(s). The computer can execute instructions to process the input(s) to generate output(s) using a parameterized model. The computer can obtain feedback on its performance in generating the outputs with the model. The computer can generate feedback by evaluating its performance. The computer can receive feedback from an external source. The computer can update parameters of the model based on the feedback to improve its performance. In this manner, the computer can iteratively “learn” to generate the desired outputs. The resulting model is often referred to as a machine-learned model.BRIEF DESCRIPTION
[0004] Aspects and advantages of the invention in accordance with the present disclosure will be set forth in part in the following description, or can be obvious from the description, or can be learned through the practice of the technology.
[0005] In accordance with one embodiment, a method for a multi-reward reinforcement learning framework for text-to-image generation is provided. The method includes training, in tandem, a prompt expansion model (prompt expansion model) and an image generation model (image generation models) using a multi-reward reinforcement learning model by: processing, by the prompt expansion model, a training query and training context data to generate an expanded training query; generating, by the image generation model, a training set of image data based on the expanded training query; generating a set of reward scores for each image datum within the training set of image data using a set of reward models, wherein generating the set of reward scores comprises generating at least one reward score for each reward criterion within a plurality of reward criteria; and adjusting weights and biases associated with the plurality of reward criteria based on the set of reward scores using the multi-reward reinforcement learning model.
[0006] In accordance with another embodiment, a system for multi-reward reinforcement learning framework for text-to-image generation is provided. The system includes one or more processors; and one or more transitory or non-transitory computer-readable media storing instructions that are executable to cause the one or more processors to perform operations, the operations comprising: obtaining, by the one or more processors, input data comprising a user query and query context data; training, in tandem, a prompt expansion model (prompt expansion model) and an image generation model using a multi-reward reinforcement learning model, wherein training the prompt expansion model and the image generation model comprises: processing, by the prompt expansion model, a training query and training context data to generate an expanded training query; generating, using the image generation model, a training set of image data based on the expanded training query; generating a set of reward scores for each image datum within the training set of image data using a set of reward models, wherein generating the set of reward scores comprises generating at least one reward score for each reward criterion within a plurality of reward criteria; selecting a subset of image data from the training set of image data as a function of the set of reward scores using a non-dominated sorting algorithm; and minimizing weights and biases associated with each reward criterion that is not represented within the subset of image data using the multi-reward reinforcement learning model; and generating, using the trained prompt expansion model and the trained image generation model, image data based on the user query and the query context data.
[0007] In accordance with one embodiment, a method for a multi-reward reinforcement learning framework for text-to-image generation is provided. The method includes training, in tandem, a prompt expansion model (prompt expansion model) and an image generation model using a multi-reward reinforcement learning model by: processing, by the prompt expansion model, a training query and training context data to generate an expanded training query; generating, by the image generation model, a training set of image data based on the expanded training query; and training, in tandem, the prompt expansion model and the image generation model using the multi-reward reinforcement learning model based on the training set of image data.
[0008] These and other features, aspects and advantages of the present invention will become better understood with reference to the following description and appended claims. The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments of the technology and, together with the description, serve to explain the principles of the technology.BRIEF DESCRIPTION OF THE DRAWINGS
[0009] A full and enabling disclosure of the present invention, including the best mode of making and using the present systems and methods, directed to one of ordinary skill in the art, is set forth in the specification, which makes reference to the appended figures, in which:
[0010] Figures 1A-B depict a diagram of an exemplary framework for multireward reinforcement learning for text-to-image generative models according to exemplary embodiments of the present disclosure;
[0011] Figure 2 depicts a block diagram of an exemplary system for generating image data using an image generation model according to exemplary embodiments of the present disclosure;
[0012] Figure 3 depicts a flow diagram of an exemplary embodiment of a method for multi-reward reinforcement learning for text-to-image generative models according to exemplary embodiments of the present disclosure;
[0013] Figure 4 is a flow chart diagram illustrating an example method for training a machine-learned model according to example implementations of aspects of the present disclosure;
[0014] Figure 5 is a block diagram of an example processing flow for using machine-learned model(s) to process input(s) to generate output(s) according to example implementations of aspects of the present disclosure;
[0015] Figure 6 is a block diagram of an example sequence processing model according to example implementations of aspects of the present disclosure;
[0016] Figure 7 is a block diagram of an example technique for populating an example input sequence for processing by a sequence processing model according to example implementations of aspects of the present disclosure;
[0017] Figure 8 is a block diagram of an example model development platform according to example implementations of aspects of the present disclosure;
[0018] Figure 9 is a block diagram of an example training workflow for training a machine-learned model according to example implementations of aspects of the present disclosure;
[0019] Figure 10 is a block diagram of an inference system for operating one or more machine-learned model(s) to perform inference according to example implementations of aspects of the present disclosure;
[0020] Figure 11 is a block diagram of an example networked computing system according to example implementations of aspects of the present disclosure;
[0021] Figure 12 is a block diagram of an example computing device according to example implementations of aspects of the present disclosure; and
[0022] Figure 13 is a block diagram of an example computing device according to example implementations of aspects of the present disclosure.DETAILED DESCRIPTION
[0023] Generally, the present disclosure is directed to a multi-reward reinforcement learning framework for image generation models. More particularly, the present disclosure provides for a means for training an image generation models and prompt expansion model prompt expansion model in tandem using a multireward reinforcement learning model. Training the image generation models andprompt expansion model in tandem can provide for improved fine tuning of both the image generation models and the prompt expansion model.
[0024] The present disclosure provides for the use of a batch-wise Pareto optimal selection. The system can identify optimal trade-offs between a number of rewards during the training phase. The system can utilize a joint optimization approach for the image generation models and the prompt expansion model to facilitate the generation of expanded text prompts. The joint optimization approach can provide for improved image generation while re-focusing the image generation models on the original textual prompt. This can be performed by assigning and optimizing rewards to both the original textual prompt and the expanded text prompts. This allows the image generation models to generate images that utilize the expanded prompt while keeping the original text prompt within the context window for image generation.
[0025] The present disclosure provides for a number of technical benefits to technical problems. By way of example, existing image generation models have encountered several challenges in generating images that satisfy performance metrics. This is true, at least partially, because the provided textual prompts usually lack sufficient context for the image generation models to generate images that satisfy the performance metrics such as image quality, resolution, clarity, or other performance metrics. To counteract this issue, prompt expansion models have been employed to enrich the textual prompts with additional details for use by the image generation models for image generation. Existing methods using prompt expansion model provide for a number of technical challenges. Existing systems simply provide an expanded prompt to the image generation models, however, as discussed, providing solely an expanded prompt can lead to a loss of focus on the original prompt by the image generation models. This results in the generation of images that fail to satisfy performance metrics.
[0026] The present disclosure can prevent over-optimization and degradation of metrics resulting from training with simple aggregation of multiple rewards by jointly optimizing multiple rewards in using the multi-reward reinforcement learning model. By adjusting the image generation models and prompt expansion model at the same time, the system can provide for improved tuning of both the prompt expansion model as well as the image generation models. This can provide for an improved imagegeneration model that outperforms baseline text-to-image methods across a number of quality criteria metrics. The quality criteria metrics can include, for example, visual clarity, contrast, pixelation, resolution, aesthetics, human preferences, image sentiment, or text-image alignment.
[0027] The improvements associated with the systems and methods discussed herein can be further understood with reference to the figures. Reference now is made to the figures, which provide example arrangements of computing systems, model structures, and data flows for illustration purposes only.
[0028] Referring now to the drawings, Figure 1 illustrates an exemplary block diagram of a system for multi-reward reinforcement learning for image generation models. System 100 can include a user query 102a-b, query context data 104a-b, prompt expansion model (prompt expansion model) 106, expanded user query 108a-b, image generation model (image generation models) 110, training set of image data 112, multi-reward reinforcement learning model 114, plurality of reward criteria 116a-d, weights and biases 118, set of reward scores 120a-d, set of reward models 122a-d, image data 124, and the like.
[0029] The operations include training, in tandem, a prompt expansion model 106 (prompt expansion model) and an image generation model 110 using a multi -reward reinforcement learning model 114. Training the prompt expansion model 106 and image generation model 110 in tandem means that both models are optimized or finetuned simultaneously. Training the models in tandem helps both models to work in harmony without over-optimization and degradation of metrics. This is done to ensure the generated text of the prompt expansion model 106 is directly in line with how the image generation model 110 will interpret the text. By training the prompt expansion model 106 and the image generation model 110 in tandem, both models receive feedback from the multi-reward reinforcement learning model 114 regarding their actions within the environment.
[0030] Jointly training the prompt expansion model (prompt expansion model) 106 and the image generation model (image generation models) 110 can include iteratively updating parameters such as weights, biases, and coefficients based on feedback from the plurality of reward criteria 116a-d. Both the prompt expansionmodel 106 and the image generation model 110 can be optimized simultaneously to ensure they work cohesively.
[0031] In an embodiment, updating the parameters can include employing techniques like gradient descent or least-squares processes, allowing for the adjustment of weights and biases based on the feedback received (i.e., set of reward scores 120a-d). The updates to the parameters can be performed iteratively to ensure that both agents' (i.e., image generation model 110 and the prompt expansion model 106) parameters are refined in response to the rewards and errors encountered during training.
[0032] Joint training of the prompt expansion model 106 and image generation model 110 can be iteratively repeated until either the available training data is exhausted, or a convergence test is achieved. As used in the current disclosure, a convergence test can be used to evaluate whether the models have reached an acceptable level of accuracy. This can be achieved by evaluating successive error values. If these differences in the error values fall below a defined threshold, it indicates that the models have stabilized. Alternatively, the error values can be compared to a predetermined threshold to determine if further training is warranted.
[0033] Joint training of the prompt expansion model 106 and image generation model 110 includes obtaining input data comprising a user query 102a-b. As used in the current disclosure, a user query 102a refers to a request for information related to the generation of an image. The user query 102a can include a description of the desired image data. The user query 102a can include information related to subject matter, aesthetics, setting, background, lighting, tone, emotions, art style, and the like. The user query 102a can be received from a user through various channels such as online forms, customer service emails, or live chat systems. Users can submit their queries by entering text or selecting options that describe their request or issue. The descriptions within the user query 102a can range from generalized inquiries to highly specific requests. For example, a generalized user query can simply state “a car” or “landscape,” while a more detailed user query 102a can contain a description such as “a pink sports car parked in front of a beautiful botanical garden of pink flowers.”
[0034] In an embodiment, a user query 102a can include a training query 102b. As used in the current disclosure, the training query 102b is an exemplary user querydesigned to provide input data for training a model. The training query 102b can be used as training data to help train or fine-tune the prompt expansion model 106, image generation model 110, or any other model discussed herein in. The training query 102b can be used to help the agents understand the relationship between textual input and visual output or a textual input and an expanded textual input. In some embodiments, the training query 102b can be the same or substantially similar to the user query 102a. This can mean that the training query 102b includes exemplary descriptions of a desired image.
[0035] System 100 can receive query context data 104a associated with the user query 102a-b. Query context data 104a refers to the contextual information associated with the user or the user query 102a-b. Query context data 104a provides additional insight into the circumstances surrounding the user query 102a-b. Query context data 104a can be used to interpret user queries 102a-b by considering the broader context in which they arise. Query context data 104a can include a range of information derived from the user’s past interactions with the system 100.
[0036] Query context data 104a can include information about the user’s previous interactions with the system 100. This can include previous queries submitted by the user, links that have been clicked, pages that have been visited, user feedback related to previous images, previous user queries 102a-d, and other relevant actions or behaviors exhibited within system 100. For example, suppose a user has submitted an inquiry about an “animated raccoon outfitted in military gear.” Query context data 104a can contain details about the user’s historical feedback about previously generated images. By analyzing these previous interactions, the system can tailor the generated image data to more precisely reflect the user’s current context.
[0037] In an embodiment, query context data 104a includes training context data 104b. As used in the current disclosure, the training context data 104b is exemplary contextual data that is used for training a model. The model can include but is not limited to the prompt expansion model 106 and image generation model 110. Training context data 104b is designed to provide the models with exemplary contextual insights associated with user interactions. The training context data 104b can include historical versions of query context data 104a associated with user interactions. Training context data 104b can include pairs of submitted queries and image data,feedback on generated images, and patterns in user behavior. Training context data 104b can be generated and optimized for use as training data.
[0038] With continued reference to Figures 1A-B, system 100 includes prompt expansion model (prompt expansion model) 106. As used in the current disclosure, the prompt expansion model 106 is a model that is configured to generate an expanded user query 108a-b. prompt expansion model 106 can be consistent with the machine-learning models described herein below in Figures 4-13. The prompt expansion model 106 operates by predicting and generating expanded versions of the user query 102a-b based on the query context data 104a-b. The prompt expansion model 106 evaluates the query context data 104a-b to predict and generate expanded versions of the user query that align more closely with the user’s contextual data. For instance, if the query context data 104a-b indicates that the user previously expressed a preference for vibrant colors in their queries, the prompt expansion model 106 can expand the user query 102a-b to include descriptors that reflect this preference.
[0039] In an embodiment, the prompt expansion model 106 can include a neural network architecture. The prompt expansion model 106 can include multiple layers of interconnected nodes, or neurons, which are configured to process data in a hierarchical manner. Each layer of the neural network can be responsible for different aspects of the input, enabling the prompt expansion model 106 to learn complex patterns and relationships within the data. The user query 102a-b can be processed using these layers where the neural network analyzes the text and identifies key components that can be expanded upon.
[0040] The nodes in the prompt expansion model 106 can be organized in a structured network, such as a convolutional neural network, which includes an input layer of nodes, one or more intermediate layers, and an output layer of nodes. During the training of the prompt expansion model 106, connections between these nodes can be established by applying elements from the training dataset to the nodes. This can include using the multi-reward reinforcement learning model 114 to adjust the connections and weights between nodes in adjacent layers based on one or more reward criteria 116a-d. The adjustments between the connections and weights between nodes can be done with the goal of optimizing the prompt expansion model 106 to produce the desired outputs.
[0041] With continued reference to Figures 1A-B, jointly training the prompt expansion model 106 and image generation model 110 includes processing the training query 102b and training context data 104b to generate an expanded training query 108b. The expanded user query 108a-b is a user query 102a-b that has been augmented to include more details and contextual information. This expanded user query 108a-b can be used as an input for the image generation model 110. By incorporating additional descriptive elements, the expanded user query 108a-b captures the user's intent more effectively. This can include adding adjectives to specify colors, incorporating details about the setting, or inferring emotional context based on the user’s history.
[0042] The prompt expansion model 106 can be configured to expand the text of the user query 102a-b by providing a more detailed and contextually relevant description of the subject matter of the query. The prompt expansion model 106 is configured to elaborate on the user query 102a-b by incorporating additional descriptive elements, such as colors, background details, emotional context, or details related to the subject matter of the user query 102a-b. Additionally, the prompt expansion model 106 can augment the user query 102a-b by incorporating the plurality of reward criteria 116a-d into the expanded user query 108a-b. For instance, if the original user query 102a-b emphasizes a specific mood, the prompt expansion model 106 can augment it to include descriptive adjectives and contextual elements that align with the image sentiment reward criteria 116b.
[0043] The expanded user query 108a-b can elaborate on the user query 102a-b by adding additional details. In a non-limiting example, if a user query 102a-b states "provide an image of a fast car," the expanded user query 108a-b can include an augmented query that specifies "a sleek, lime green sports car parked near a beach during sunset." This augmented query includes additional details, such as color, style, setting, and emotional context.
[0044] The augmented text of the expanded user query 108a-b can be generated based on an analysis of the user query 102a-b and query context data 104a-b. The prompt expansion model 106 can evaluate query context data 104a-b such as insights from previous interactions, user preferences, and other contextual factors to determine which details should be added to the query. This means that expanded user query108a-b can vary significantly based on the user's history and interests. For example, a user who frequently requests nature scenes can receive an expanded user query 108a- b that emphasizes natural elements, such as "a tranquil forest with a flowing river" instead of a more generic description.
[0045] In an embodiment, the expanded user query 108a can include an expanded training query 108. The expanded training query 108b is an output of the prompt expansion model 106 during the training phase. The expanded training query 108b is configured to serve as a reference point for training both the prompt expansion model 106 and the image generation model 110. The expanded training query 108b is generated from the training query 102b and the training context data 104b during the joint training process of the prompt expansion model 106 and the image generation model 110. The expanded training query 108b can serve multiple purposes in the training process. The expanded training query 108b can provide a clear target for the image generation model 110 during its learning phase, helping the model understand the kinds of visual elements to include in its outputs. Additionally, the expanded training query 108b can reinforce the importance of context and detail in query expansion, allowing the model to learn from training examples as it adjusts its parameters. The expanded training query 108b can allow for the evaluation of both the prompt expansion model 106 and the image generation model 110 through error analysis.
[0046] With continued reference to Figures 1 A-B, system 100 includes the image generation model 110. As used in the current disclosure, the image generation model is a model that is designed to generate images based on textual inputs and contextual data. The image generation model 110 is configured to interpret textual data to generate image data. The generated image data can include but is not limited to the training set of image data 112, the subset of image data 202, the Pareto-optimal set 204, the image data 124, and any other image data discussed herein. In some embodiments, the image generation model 110 can include a large language model that is tailored for the image generation. The LLM can use several natural language processing techniques to convert textual prompts into image data. The LLM can be used to generate image data that reflects the textual input. The image generation model can include any of the machine learning models, natural language processingmodels, image processing models, large language models, and the like that are discussed herein below in Figures 4-13.
[0047] With continued reference to Figures 1A-B, jointly training the prompt expansion model 106 and image generation model 110 includes generating a training set of image data 112 based on the expanded training query 108b. As used in the current disclosure, the training set of image data 112 is a collection of images that are used to jointly train or optimize the prompt expansion model 106 and image generation model 110. The training set of image data 112 includes a plurality of images generated by the image generation model 110 based on the expanded training query 108b. This can include generating multiple iterations of image data using a single expanded training query 108b. Alternatively or additionally, image generation model 110 can generate multiple iterations of image data using multiple expanded training queries 108b.
[0048] Generating the training sets of image data 112 can be achieved by employing techniques such as stochastic sampling or controlled randomness. Stochastic sampling introduces a level of variance within the image generation process. This variance can come in the form of adjusting parameters, such as color saturation, lighting, perspective, composition, and the like. This variance can help the image generation model 110 to produce different iterations of image data based on the same or similar expanded training query 108b.
[0049] In an embodiment, each image datum within the training set of image data 112, can be tailored to emphasize one or more reward criteria 116a-d. Each image datum produced by the image generation model 110 can be conceptualized as the result of a decision-making process that weighs these multiple reward criteria 116a-d. For example, when the expanded training query 108b includes details about an image of an “urban landscape,” the image generation model 110 can prioritize the reward criterion 116c associated with aesthetic quality in one iteration, leading to vibrant colors and dramatic lighting. In another iteration, the image generation model 110 can be configured to focus on the text-to-image alignment criterion 116a. By optimizing parameters associated with the reward criteria 116a-d, the image generation model 110 can create images that emphasize or optimize different reward criteria 116a-d.
[0050] Optimization of the parameters associated with the reward criteria 116a-d can enable the image generation model 110 to optimize the image generation process for specific outcomes. The ability to emphasize one or more reward criteria 116a-d can allow the image generation model 110 to produce images that create a balanced representation that fulfills multiple reward criteria 116a-d simultaneously. This can include identifying optimal trade-offs between conflicting reward criteria 116a-d. Optimizing the image generation model 110 for one reward criterion can negatively impact a second reward criterion. In such cases, optimization of the image generation model 110 can include employing strategies to find the most effective balance. By employing techniques such as multi -objective optimization, the image generation model 110 can iteratively adjust the parameters associated with each reward criterion 116a-d. The multi-objective optimization techniques are discussed in greater detail below.
[0051] With continued reference to Figures 1A-B, the operations can further include evaluating the plurality of image data according to the plurality of reward criteria 116a-d. As used in the current disclosure, the plurality of reward criteria 116a- d are specific metrics or objectives utilized in reinforcement learning to evaluate the performance of an agent, such as the image generation model 110 or the prompt expansion model 106. These reward criteria 116a-d can be used to guide the training process by providing feedback on the quality of generated outputs. The reward criteria 116a-d can be used to provide measurable objectives that guide the behavior of agents. This feedback is used to help the image generation model 110 or the prompt expansion model 106 learn which actions within the environment to prioritize. The plurality of reward criteria 116a-d can be used to help assess the quality and effectiveness of the generated outputs (i.e., image data and expanded user prompt). This assessment is subsequently used to steer the behaviors of the agents toward desired outcomes.
[0052] Each criterion of the plurality of reward criteria 116a-d can serve as a distinct lens through which the generated images or expanded prompts can be evaluated. The agents receive rewards or penalties based on their actions. These rewards or penalties are used to influence the agents regarding the success or failure of those actions. For the image generation model 110 or the prompt expansion model106 defining clear reward criteria 116a-d enables the models to quantify a “good” image or expanded prompt versus a “less effective” one. This evaluation process can help the models identify which features to prioritize or modify in future iterations.
[0053] Exemplary embodiments of the plurality of reward criteria 116a-d can include text-image alignment reward criteria 116a. As used in the current disclosure, the text-image alignment reward criteria 116a refers to a set of evaluation metrics used to assess how well a generated image corresponds to a given textual description or prompt. Text-image alignment reward criteria 116a can be used to evaluate the relationship between generated images and their associated textual prompts. These textual prompts can include the user query 102a-b or the expanded user query 108a-b. Text-image alignment reward criteria 116a can include considerations of the overall context and tone of the text. The text-image alignment reward criteria 116a can evaluate the textual input and the generated image within the context of the query context data 104a-b. For example, if the textual prompt describes an "animated frog dancing on a lily pad," the image should include identifiable features such as the frog, the dancing posture, and the lily pad. Text-image alignment reward criteria 116a can be used to ensure that the textually described elements and actions are visually present.
[0054] In an embodiment, the text-image alignment reward criteria 116a can include a first text-image alignment reward criteria 116a and a second text-image alignment reward criteria 116a. The first text-image alignment reward criteria 116a can be associated with the relationship between the generated image and the original user query 102a-b. In contrast, a second text-image alignment reward criteria 116a can be associated with the expanded user query 108a-b. The second text-image alignment reward criteria 116a can be used to assess how effectively the generated image aligns with the broader context of the expanded user query 108a-b. The first and second text-image alignment reward criteria 116a can be used to ensure that the generated image captures the key elements and themes explicitly mentioned in the written text of the user query 102a-b and the expanded user query 108a-b.
[0055] Exemplary embodiments of the plurality of reward criteria 116a-d can include one or more image sentiment reward criteria 116b. As used in the current disclosure, the image sentiment reward criteria 116b are used to evaluate theemotional tone and mood conveyed by generated images in relation to their corresponding textual descriptions or prompts. The image sentiment reward criteria 116b can be used to ensure that the generated image data reflects the emotional context that was described in the textual prompts. In some cases, image sentiment reward criteria 116b can be an assessment of specific visual cues that contribute to the overall sentiment of an image. These cues can include factors such as color palettes, facial expressions, and composition. For instance, vibrant colors and cheerful facial expressions are generally considered to be associated with feelings of joy, while muted tones and solemn postures can suggest sadness or contemplation.
[0056] Exemplary embodiments of the plurality of reward criteria 116a-d can include one or more aesthetic reward criteria 116c. As used in the current disclosure, the aesthetic reward criteria 116c are reward criteria that are used to assess the visual quality and artistic appeal of generated images. The aesthetic reward criteria 116c can focus on various elements that contribute to the overall appearance and attractiveness of an image. This can include considerations such as composition, color harmony, lighting, detail and texture, art style, and the like. The aesthetic reward criteria 116c can include an evaluation of how well the elements within the image are arranged and balanced or the use of color within the image.
[0057] In an embodiment, the aesthetic reward criteria 116c can be derived from historical ratings of the aesthetic qualities of real images. System 100 can generate the aesthetic reward criteria 116c from a number of subjective assessments of the aesthetic appeal of image data. These assessments can be analyzed to identify the patterns in how aesthetic qualities are perceived. These annotated datasets can be used to train a machine-learning model such as the third reward model 122c or any other model disclosed herein. By applying techniques such as supervised learning, the machine-learning model can develop an understanding of factors such as composition, color harmony, and texture that resonate with viewers. This trained model can then be utilized to generate the aesthetic reward criteria 116c.
[0058] Exemplary embodiments of the plurality of reward criteria 116a-d can include one or more human preference reward criteria 116d. Human preference reward criteria 116d can be used to evaluate generated images based on a dataset composed of human feedback. This can be done to ensure that the output alignsclosely with what individuals find appealing or desirable. Generating the human preference reward criteria 116d can involve the collection of user feedback through surveys, ratings, or other interactive methods. This feedback can then be aggregated to create a dataset that reflects the diverse preferences of the audience. This dataset will include a plurality of text-image pairs that reflect the preferences of the users. Once this annotated dataset is established, machine learning models, such as the fourth reward model 122d, can be applied to analyze the data and extract meaningful insights about user preferences. The fourth reward model 122d can be used to identify and score key features (i.e., color schemes, compositions, and subject matter) based on the annotated dataset. In some embodiments, this annotated dataset can be used as training data by the fourth reward model 122d.
[0059] With continued reference to Figures 1A-B, the operations can further include generating a set of reward scores 120a-d for each image datum within the training set of image data 112 using a set of reward models 122a-d. As used in the current disclosure, a reward score is a quantitative measure that reflects the effectiveness of an action or decision made by an agent in achieving specific objectives, such as the plurality of reward criteria 116a-d. The agents are configured to optimize their behavior by receiving feedback based on the plurality of reward criteria 116a-d. This feedback is quantified by the set of reward scores 120a-d. Each reward score of the set of reward scores 120a-d represents a quantification of a distinct aspect of the behavior of the agents within the environment. For example, a reward score can reflect how well a generated image aligns with a textual prompt.
[0060] The set of reward scores 120a-d can facilitate the refinement of the decision-making processes of the agents within the environment. By receiving quantitative feedback on multiple reward criteria 116a-d, the agents can learn to identify patterns and make informed choices that enhance their overall effectiveness. This encourages the agents to optimize their strategies based on their impact on the set of reward scores 120a-d. The refinement encourages the agents to capitalize on successful strategies that yield higher scores.
[0061] In an embodiment, generating the set of reward scores 120a-d includes producing at least one distinct reward score 120 for each reward criterion within the plurality of reward criteria 116a-d. The first reward score 120a is associated with thetext-image alignment reward criteria 116a. The first reward score 120a is a quantification of how effectively the generated image corresponds to the textual prompt (i.e., user query 102a-b or the expanded user query 108a-b). The second reward score 120b corresponds to the image sentiment reward criteria 116b. The second reward score 120b represents a quantification of the alignment of the emotional tone conveyed by the generated images with the textual input. The third reward score 120c is associated with the aesthetic reward criteria 116c. The third reward score 120c is a quantification of the visual quality and artistic appeal of the generated images. The fourth reward score 120d is associated with the human preference reward criteria 116d. The fourth reward score 120d is a quantification of how well the generated images align with user preferences and tastes.
[0062] Each reward score within the set of reward scores 120a-d can be generated using a reward model of the set of reward models 122a-d. Each reward model of the set of reward models 122a-d can be tailored to a specific reward criterion 116a-d within the plurality of reward criteria 116a-d. This can enable the model to analyze and score the images effectively. Each reward model within the set of reward models 122a-d can be the same or substantially similar to the machine learning models, natural language processing models, image processing models, large language models, and the like that are discussed herein below in Figures 4-13. For instance, the first reward model 122a is associated with the text-image alignment reward criteria 116a. Thus, the first reward model 122a can employ one or more natural language processing techniques and image processing techniques to evaluate how well a generated image corresponds to its accompanying textual prompt. The first reward model 122a can analyze semantic relationships and visual features to produce the first reward score 120a.
[0063] Similarly, a second reward model 122b can be dedicated to the image sentiment reward criteria 116b. The second reward model 122b can employ machine learning algorithms trained on sentiment analysis to gauge the emotional tone of the images. The second reward model 122b can evaluate features of the image such as colors, expressions, and overall composition to quantify how well the images align with the desired sentiment of the textual prompt. The third reward model 122c can be associated with the aesthetic reward criteria 116c. The third reward model 122c canuse an image processing model to evaluate the artistic quality of the image data. The fourth reward model 122d is associated with the human preference reward criteria 116d. The fourth reward model 122d can be trained on a dataset composed of user feedback and preference data to inform its scoring.
[0064] With continued reference to Figures 1A-B, the operations can further include adjusting weights and biases 118 associated with the plurality of reward criteria 116a-d based on the set of reward scores 122a-d using the multi-reward reinforcement learning model 114. As used in the current disclosure, the multi-reward reinforcement learning model 114 is a model that is designed to optimize the performance of agents in generating outputs based on multiple reward criteria. The multi-reward reinforcement learning model 114 is configured to adjust the biases and weights 118 associated with the plurality of the plurality of reward criteria 116a-d based on the set of reward scores 122a-d. The multi-reward reinforcement learning model 114 is configured to account for the performance of multiple agents across multiple reward criteria 116a-d. The multi-reward reinforcement learning model 114 can be consistent with the machine-learning models described herein below in Figures 4-13. In an embodiment, the multi-reward reinforcement learning model 114 can be configured to identify optimal trade-offs between the plurality of reward criteria 116a- d. Based on these optimal trade-offs, the multi-reward reinforcement learning model 114 can optimize the biases and weights 118 associated with the plurality of reward criteria 116a-d.
[0065] The multi-reward reinforcement learning model 114 can be used to finetune the biases and weights 118 that determine the relative importance of each reward criterion. Weights can be used the degree of influence that each reward criterion has on the final decision-making process. Bias is a term associated with the adjustments that can shift the output in a desired direction. The adjustment of these weights and biases 118 can be informed by the set of reward scores 120a-d generated from the reward models 122a-d. By analyzing the performance of the agents based on these scores, the multi-reward reinforcement learning model 114 can identify which reward criteria 116a-d are contributing positively to achieving the desired outcomes and which can require recalibration. For example, if the text-image alignment reward score 120a consistently yields high scores while the aesthetic reward score 120c fallsshort, the multi-reward reinforcement learning model 114 can maximize the weight assigned to the text-image alignment reward criterion 116a while minimizing the weight for the aesthetic reward criteria 116c. Adjusting weights and biases 118 can include an iterative learning process where the multi-reward reinforcement learning model 114 evaluates the effects of these adjustments on subsequent sets of reward scores 120a-d. By employing techniques such as gradient descent or other optimization algorithms, the multi-reward reinforcement learning model 114 can systematically refine its parameters to enhance performance.
[0066] With continued reference to FIG IB, once training of the prompt expansion model 106 and the image generation model 110 has been completed, system 100 can be used to generate image data 124 based on the expanded user query 108a. Generating the image data 124 can include converting the user query 108a into the expanded user query 108a using the trained prompt expansion model 106, as described herein above. The expanded user query 108a can be used as an input into the trained image generation model 110. The image generation model 110 is then configured to generate image data 124 which is in line with the reward criteria 116a- d. This can be done using any of the processes for image generation discussed herein.
[0067] In some embodiments, the prompt expansion model 106 and the image generation model 110 can be frozen once the training process is complete. During the training phase, both of the agents can continually adjust their behavior based on the feedback received from the set of reward scores 122a-d. Once the agents have sufficiently learned the optimal strategies, a convergence test can be employed to evaluate whether the agents' performance has stabilized and reached an acceptable level of consistency in their outputs. Once the agents pass this convergence test, indicating that their learning has plateaued and further adjustments are unlikely to yield significant improvements, they can be frozen. Freezing the agents can include locking their weights and biases in place, effectively halting any further updates to their parameters.
[0068] Referring now to FIG 2, an exemplary depiction of a block diagram of an exemplary system for generating image data using an image generation model according to exemplary embodiments of the present disclosure. Figure 2 includes asubset of image data 202, Pareto-optimal set 204, non-dominated sorting algorithm 206, policy gradient update 208, and the like.
[0069] In an embodiment, training the prompt expansion model 106 and the image generation model 110 includes selecting a subset of image data 202 from the training set of image data 112 as a function of the set of reward scores 120a-d using a non-dominated sorting algorithm 206. As used in the current disclosure, the subset of image data 202 is a collection of images selected from the broader training set of image data 112. The subset of image data 202 can be chosen based on the performance of the image generation model 110 as quantified by the set of reward scores 120a-d. The subset of image data 202 can be characterized by its representation of the optimal trade-offs among the various reward scores 120a-d.
[0070] In an embodiment, the subset of image data 202 can be represented as a pareto-optimal set 204. As used in the current disclosure, a Pareto-optimal set 204 refers to a collection of solutions in a multi-objective optimization problem where no individual solution can be improved in one criterion without degrading performance in another. The Pareto-optimal set 204 can represent a subset of image data 202 that displays the optimal trade-offs among various reward scores 120a-d. The image data within the Pareto-optimal set 204 can be considered optimal because it reflects a balanced performance across two or more reward criteria 116a-d. In a non-limiting example, assume a first image excels in a first reward criterion but is weaker in a second reward criterion. A second image displays the opposite strengths, both the first image and the second image can be selected to be a part of the Pareto-optimal set 204.
[0071] Optimal trade-offs among the various reward scores 120a-d can refer to identifying the best possible outcomes across multiple competing reward criteria 116a-d. Each reward score with the set of reward scores 120a-d can represent a quantification of the performance of the prompt expansion model 106 and the image generation model 110 while generating image data. One of the key challenges in identifying and producing the optimal trade-offs is managing the competing reward scores 120. Optimal tradeoffs between multiple competing reward criteria 116a-d can be identified when the improvement of one reward score often leads to a decline in another reward score.
[0072] The non-dominated sorting algorithm 206 can be used to identify the optimal trade-offs between the image data based on their respective reward scores 120a-d. The non-dominated sorting algorithm 206 can be used to analyze how changes in one reward score impact the remainder of the reward scores. As used in the current disclosure, the non-dominated sorting algorithm 206 is a technique used in multi-objective optimization to categorize and rank solutions based on their performance across several reward criteria. The non-dominated sorting algorithm 206 is configured to identify one or more non-dominated solutions within the training set of image data 112. An image can be deemed non-dominated if there is no other image that performs better across two or more reward criteria. In a non-limiting example, if Image A excels in a first reward criterion while Image B is superior in a second reward criterion, neither image dominates the other, as they both have strengths in different areas.
[0073] The non-dominated sorting algorithm 206 can be configured to evaluate the relationship between the images within the training set of image data 112. This can include classifying the images based on their overall performance as reflected by the set of reward scores 120a-d. Exemplary categories can include but are not limited to non-dominated categories, text-image alignment categories, image sentiment categories, aesthetic categories, human preference categories, and the like. The nondominated category can be a category that is used to represent the highest-performing solutions across a defined group of reward criteria 116a-d. Subsequent categories can be reserved for images that perform well across one or more reward criteria 116a-d.
[0074] Once the images have been ranked or classified using the non-dominated sorting algorithm 206, the non-dominated sorting algorithm 206 can select the subset of image data 202 based on these identified trade-offs. The subset of image data 202 can be identified from the non-dominated category. This category can represent the most optimal tradeoffs between the image data.
[0075] With continued reference to FIG 2, training the prompt expansion model 106 and the image generation model 110 can include adjusting the weights and biases 118 associated with the plurality of reward criteria 116a-d based on the subset of image data 202 using the multi-reward reinforcement learning model 114. The multireward reinforcement learning model 114 can be configured to identify specificweights and biases 118 that need to be adjusted to fine-tune the prompt expansion model 106 and the image generation model 110 based on the subset of image data 202 or the Pareto-optimal set 204. The weights and biases 118 can be fine-tuned based on the identified optimal trade-offs between various reward scores. The multi-reward reinforcement learning model 114 can be configured to examine the characteristics of these images within the subset of image data 202 or the Pareto-optimal set 204. Based on this examination the multi -reward reinforcement learning model 114 can determine which features contribute to the most favorable results. In a non-limiting example, if the images that are represented within the Pareto-optimal set 204 demonstrate strong alignment with prompts and effective emotional expression, the multi-reward reinforcement learning model 114 can adjust the weights associated with these dimensions. This can be done with the goal of improving the generated image data across later iterations.
[0076] With continued reference to FIG 2, training the prompt expansion model 106 and the image generation model 110 can include determining a policy gradient update 208 as a function of the subset of image data 202. As used in the current disclosure, the policy gradient update 208 is used to optimize the policies of the agents. The policy gradient update 208 can be used to help the agents optimize their interaction within the environment to maximize / minimize the set of reward scores 120a-d. The policy gradient update 208 can be created based on a series of evaluations of the performance of the model against the subset of image data 202 or the Pareto-optimal set 204. The subset of image data 202 can be used as a point of reference when making a determination of how well the agents' actions align with desired outcomes. By analyzing the reward scores 116a-d associated with the subset of image data 202 the multi-reward reinforcement learning model 114 can make a determination of the gradients that indicate how changes to the policy would impact future performance. In an embodiment, the multi-reward reinforcement learning model 114 can evaluate the reward scores 116a-d for actions taken based on the subset of image data 202. The gradients can be generated based on each image datum’s contribution to the set of reward scores 116a-d.
[0077] Once the gradients are computed, the policy gradient update 208 is applied to adjust the policy parameters in the direction that maximizes / minimizes theidentified reward criteria 116a-d. These adjustments can be made to encourage the agents to explore and select actions that lead to high-reward outcomes. Additionally, by optimizing the biases and weights associated with the reward criteria, the agents can learn to make the optimal trade-offs between competing reward criteria. This process can be performed iteratively, meaning that the policy gradient can be updated until the agents can successfully pass a convergence test.
[0078] In an embodiment, determining the policy gradient update 208 can involve minimizing the reward scores 116a-d associated with each reward criterion that is not adequately represented within the subset of image data 202 or the Pareto-optimal set 204. By identifying these underrepresented reward scores, the multi-reward reinforcement learning model 114 can develop a targeted policy focused on improving the performance across the unrepresented reward criteria. To achieve this minimization, the multi-reward reinforcement learning model 114 can employ a gradient descent approach, incorporating the gradients of the underrepresented reward scores 120a-d into the overall policy gradient update. By actively minimizing these scores, the agents learn to address performance gaps.
[0079] In an additional embodiment, determining the policy gradient update 208 can also include maximizing the reward scores 120a-d associated with each reward criterion 116a-d that is present within the subset of image data 202. This can be done to encourage the agents to capitalize on their strengths by reinforcing the positive behaviors that lead to favorable reward scores 120a-d. To achieve this maximization, the multi-reward reinforcement learning model 114 can identify the reward scores 120a-d that correspond to the reward criteria reflected in the subset of image data 202. By analyzing the successful image outputs, the multi-reward reinforcement learning model 114 can ascertain which features and actions yield higher rewards. This understanding allows the multi-reward reinforcement learning model 114 to adjust its policy parameters in a way that favors actions leading to similar successful outcomes in future iterations. In some cases, the process of maximizing these reward scores 120a-d can include applying techniques such as stochastic gradient ascent. By calculating the gradients of the represented reward scores 120a-d and integrating these into the policy gradient update, the multi-reward reinforcement learning model 114can make informed adjustments that amplify the probability of producing desirable outputs.
[0080] Referring now to Figure 3 depicts a flow diagram of a method 300 to perform multi-reward reinforcement learning for text-to-image generative models according to exemplary embodiments of the present disclosure. The method 300 can be performed by processing logic that can include hardware (e.g., processing device, circuitry, dedicated logic, programmable logic, microcode, hardware of a device, integrated circuit, etc.), software (e.g., instructions run or executed on a processing device), or a combination thereof. In some embodiments, method is performed by a server computing system (e.g., server computing system 60) or client computing system (e.g., client computing device 50). Although shown in a particular sequence or order, unless otherwise specified, the order of the processes can be modified. Thus, the illustrated embodiments should be understood only as examples, and the illustrated processes can be performed in a different order, and some processes can be performed in parallel. Additionally, one or more processors can be omitted in various embodiments. Thus, not all processes are required in every embodiment. Other process flows are possible.
[0081] At operation 302, processing logic can train, in tandem, a prompt expansion model (prompt expansion model) and an image generation model (image generation models) using a multi-reward reinforcement learning model. Joint training of the prompt expansion model and the image generation models includes iteratively updating the parameters of the models based on feedback from the plurality of reward criteria. This feedback is used to help the image generation models and the prompt expansion model learn which actions within the environment to prioritize based on the quality and effectiveness of the generated outputs (i.e., image data and expanded user prompt).
[0082] At operation 304, training the prompt expansion model and the image generation model includes processing, by the prompt expansion model, a training query and training context data to generate an expanded training query. The prompt expansion model is configured to augment the training query with additional contextual and descriptive elements to generate the expanded training query. Both the training query and the training context data are exemplary representations of the userquery and the query context data that are used as training data for the prompt expansion model. Similar to the training query and training context data, the expanded training query is an exemplary representation of the expanded user query that can be used to train the image generation models.
[0083] At operation 306, training the prompt expansion model and the image generation model includes generating, by the image generation model, a training set of image data based on the expanded training query. The image generation models is configured to generate the training set of image data from the textual input of the expanded training query. The training set of image data is a collection of generated images that are exemplary representations of potential outputs of the image generation models. The training set of image data may represent optimal trade-offs between two or more reward criteria.
[0084] At operation 308, training the prompt expansion model and the image generation model includes generating a set of reward scores for each image datum within the training set of image data using a set of reward models, wherein generating the set of reward scores includes generating at least one reward score for each reward criterion within a plurality of reward criteria. In an embodiment, the plurality of reward criteria includes one or more text-image alignment reward criteria, one or more image sentiment reward criteria, one or more aesthetic reward criteria, or one or more human preference reward criteria. In some cases, the one or more text-image alignment reward criteria can include: a first text-image alignment reward criteria associated with the training query; and a second text-image alignment reward criteria associated with the expanded training query.
[0085] The quality of the expanded training query and the training set of image data may be represented by the set of reward scores. The set of reward scores are quantifications of how well the expanded training query and the training set of image data align with the set of reward criteria. These reward scores may be used as feedback to both the prompt expansion model and the image generation models, respectively.
[0086] At operation 310, training the prompt expansion model and the image generation model includes adjusting weights and biases associated with the plurality of reward criteria based on the set of reward scores using the multi-rewardreinforcement learning model. Both the prompt expansion model and the image generation models can be fine-tuned based on the optimal trade-offs between various reward scores that are present within the training set of image data. The multi-reward reinforcement learning model is used to examine the characteristics of the training set of image data to determine which features or actions contribute to the most favorable results. Based on these optimal trade-offs, the multi-reward reinforcement learning model may simultaneously adjust the weights and biases associated with the reward criteria for both the prompt expansion model and the image generation models.
[0087] In some implementations, processing logic can obtain input data including a user query and query context data. In some implementations, processing logic can generate, using the trained prompt expansion model and the trained image generation model, image data based on the user query and the query context data. For instance, processing logic can generate the image data by processing, by the trained prompt expansion model, the user query and the query context data to generate an expanded user query. For instance, processing logic can generate the image data by generating, using the trained image generation model, the image data based on the expanded user query. In some instances, the expanded user query can be generated by incorporating the plurality of reward criteria into the expanded user query.
[0088] In some implementations, processing logic can train the prompt expansion model and the image generation model includes selecting a subset of image data from the training set of image data as a function of the set of reward scores using a nondominated sorting algorithm. In some implementations, processing logic can adjust the weights and biases associated with the plurality of reward criteria based on the subset of image data using the multi-reward reinforcement learning model. In some implementations, the subset of image data can include a Pareto-optimal set associated with the set of reward scores.
[0089] In some embodiments, training the prompt expansion model and the image generation model can include determining, by the computing system, a policy gradient update as a function of the subset of image data. Training the prompt expansion model and image generation model can include adjusting the weights and biases associated with the plurality of reward criteria using the multi-reward reinforcement learning model as a function of the policy gradient update. In somecases, determining the policy gradient update can include minimizing the reward scores associated with each reward criterion that is not represented within the subset of image data. Additionally, determining the policy gradient update can include maximizing the reward scores associated with each reward criterion that is represented within the subset of image data. Furthermore, the policy gradient update can include maximizing one or more text-image alignment reward criteria.
[0090] Figure 4 depicts a flowchart of a method 400 for training one or more machine-learned models according to aspects of the present disclosure. For instance, an example machine-learned model can include the prompt expansion model 106, image generation model 110, multi-reward reinforcement learning model 114, Set of Reward Models 122a-d, along with any other model or algorithm mentioned here.
[0091] The method 400 can be performed by processing logic that can include hardware (e.g., processing device, circuitry, dedicated logic, programmable logic, microcode, hardware of a device, integrated circuit, etc.), software (e.g., instructions run or executed on a processing device), or a combination thereof. In some embodiments, method is performed by a server computing system (e.g., server computing system 60) or client computing system (e.g., client computing device 50). Although shown in a particular sequence or order, unless otherwise specified, the order of the processes can be modified. Thus, the illustrated embodiments should be understood only as examples, and the illustrated processes can be performed in a different order, and some processes can be performed in parallel. Additionally, one or more processors can be omitted in various embodiments. Thus, not all processes are required in every embodiment. Other process flows are possible.
[0092] At operation 402, processing logic can obtain a training instance. A set of training data can include a plurality of training instances divided between multiple datasets (e.g., a training dataset, a validation dataset, or testing dataset). A training instance can be labeled or unlabeled. Although referred to in example method 400 as a “training” instance, it is to be understood that runtime inferences can form training instances when a model is trained using an evaluation of the model’s performance on that runtime instance (e.g., online training / learning). Example data types for the training instance and various tasks associated therewith are described throughout the present disclosure.
[0093] At operation 404, processing logic can process, using one or more machine-learned models, the training instance to generate an output. The output can be directly obtained from the one or more machine-learned models or can be a downstream result of a chain of processing operations that includes an output of the one or more machine-learned models.
[0094] At operation 406, processing logic can receive an evaluation signal associated with the output. The evaluation signal can be obtained using a loss function. Various determinations of loss can be used, such as mean squared error, likelihood loss, cross entropy loss, hinge loss, contrastive loss, or various other loss functions. The evaluation signal can be computed using known ground-truth labels (e.g., supervised learning), predicted or estimated labels (e.g., semi- or self-supervised learning), or without labels (e.g., unsupervised learning). The evaluation signal can be a reward (e.g., for reinforcement learning). The reward can be computed using a machine-learned reward model configured to generate rewards based on output(s) received. The reward can be computed using feedback data describing human feedback on the output(s).
[0095] At operation 408, processing logic can update the machine-learned model using the evaluation signal. For example, values for parameters of the machine- learned model(s) can be learned, in some embodiments, using various training or learning techniques, such as, for example, backwards propagation. For example, the evaluation signal can be backpropagated from the output (or another source of the evaluation signal) through the machine-learned model(s) to update one or more parameters of the model(s) (e.g., based on a gradient of the evaluation signal with respect to the parameter value(s)). For example, system(s) containing one or more machine-learned models can be trained in an end-to-end manner. Gradient descent techniques can be used to iteratively update the parameters over a number of training iterations. In some implementations, performing backwards propagation of errors can include performing truncated backpropagation through time. Example method 400 can include implementing a number of generalization techniques (e.g., weight decays, dropouts, etc.) to improve the generalization capability of the models being trained.
[0096] In some implementations, example method 400 can be implemented for training a machine-learned model from an initialized state to a fully trained state (e.g.,when the model exhibits a desired performance profile, such as based on accuracy, precision, recall, etc.).
[0097] In some implementations, example method 400 can be implemented for particular stages of a training procedure. For instance, in some implementations, example method 400 can be implemented for pre-training a machine-learned model. Pre-training can include, for instance, large-scale training over potentially noisy data to achieve a broad base of performance levels across a variety of tasks / data types.
[0098] In some implementations, example method 400 can be implemented for fine-tuning a machine-learned model. Fine-tuning can include, for instance, smaller- scale training on higher-quality (e.g., labeled, curated, etc.) data. Fine-tuning can affect all or a portion of the parameters of a machine-learned model. For example, various portions of the machine-learned model can be “frozen” for certain training stages. For example, parameters associated with an embedding space can be “frozen” during fine-tuning (e.g., to retain information learned from a broader domain(s) than present in the fine-tuning dataset(s)). In some implementations, example method 400 uses adapter modules. Adapters can be small trainable layers that are inserted between pre-existing layers of a pre-trained model. During the fine-tuning process, the original parameters of the pre-trained model are typically frozen, and only the parameters of the adapters are updated.
[0099] In some implementations, example method 400 can be implemented to execute parameter-efficient fine-tuning methods, such as Layerwise Optimization of Residuals (LoRA). LoRA can refine pre-trained models with minimal adjustments to the original parameters. This can be achieved by introducing trainable low-rank matrices that modify the behavior of the pre-trained weights without directly altering them. In some implementations, during fine-tuning, only these auxiliary matrices are updated, which significantly reduces the number of parameters that are trained.
[0100] An example fine-tuning approach includes reinforcement learning. Reinforcement learning can be based on user feedback on model performance during use.
[0101] Figure 5 is a block diagram of an example processing flow for using machine-learned model(s) 1 to process input(s) 2 to generate output(s) 3.
[0102] Machine-learned model(s) 1 can be or include one or multiple machine- learned models or model components. Example machine-learned models can include neural networks (e.g., deep neural networks). Example machine-learned models can include non-linear models or linear models. Example machine-learned models can use other architectures in lieu of or in addition to neural networks. Example machine- learned models can include decision tree-based models, support vector machines, hidden Markov models, Bayesian networks, linear regression models, k-means clustering models, etc.
[0103] Machine-learned model(s) 1 can be or include, or otherwise be representative of any one or more of the machine-learned models described above with respect to the preceding figures. For example, machine-learned model(s) 1 can be or include, or otherwise be representative of any one or more of prompt expansion model 106, image generation model 110, multi-reward reinforcement learning model 114, Set of Reward Models 122a-d, along with any other model or algorithm mentioned here. Although various features, variations, and implementations described below are described with respect to machine-learned model(s) 1, it is to be understood that such features, variations, and implementations are to be understood as described with respect to each of prompt expansion model 106, image generation model 110, multi-reward reinforcement learning model 114, Set of Reward Models 122a-d, along with any other model, algorithm, or machine-learned component described herein.
[0104] Example neural networks can include feed-forward neural networks, recurrent neural networks (RNNs), including long short-term memory (LSTM) based recurrent neural networks, convolutional neural networks (CNNs), diffusion models, generative-adversarial networks, or other forms of neural networks. Example neural networks can be deep neural networks. Some example machine-learned models can leverage an attention mechanism such as self-attention. For example, some example machine-learned models can include multi-headed self-attention models.
[0105] Machine-learned model(s) 1 can include a single or multiple instances of the same model configured to operate on data from input(s) 2. Machine-learned model(s) 1 can include multiple different models or multiple different model portions configured to operate on data from input(s) 2.
[0106] Machine-learned model(s) 1 can include an ensemble of different models that can cooperatively interact to process data from input(s) 2. For example, a model ensemble can include multiple models that have different attributes (e.g., different architectures, trained with different recipes, etc.). The ensemble can output an overall output based on the individual outputs of the constituent models. In this manner, for instance, the diverse constituent models can work together to provide system-level robustness by effectively aggregating over individual strengths and weaknesses of any given model. The respective individual outputs can be combined in a weighted combination, using a voting or routing mechanism, or a learned output layer (e.g., one or more feedforward or fully connected layers).
[0107] Machine-learned model(s) 1 can employ a mixture-of-experts structure. See, e.g., Zhou et al., Mixture-of-Experts with Expert Choice Routing, ARXIV:2202.09368V2 (Oct. 14, 2022). For example, different portions of a model can learn (explicitly or implicitly) different expertise areas, with pathways through the model being selected by a learned routing mechanism that engages the appropriate expert for a given input (e.g., a given portion of an input, such as on a per-token basis). For example, a feedforward network can be sparsely activated for a given portion of an input based on an output of a routing mechanism that processes the portion of the input. In this manner, for instance, the group of activated weights can form an “expert” that is selected by the router. On each forward pass, only a subset of the total model weights can be engaged, thereby decreasing the quantity of operations performed for processing a given input compared to a densely activated model. In this manner, for instance, the expressive and interpretive power of a high-parameter-count model can be achieved with more compute-efficient forward passes.
[0108] Input(s) 2 can generally include or otherwise represent various types of data. Input(s) 2 can include one type or many different types of data. Output(s) 3 can be data of the same type(s) or of different types of data as compared to input(s) 2. Output(s) 3 can include one type or many different types of data.
[0109] Example data types for input(s) 2 or output(s) 3 include natural language text data, software code data (e.g., source code, object code, machine code, or any other form of computer-readable instructions or programming languages), machine code data (e.g., binary code, assembly code, or other forms of machine-readableinstructions that can be executed directly by a computer's central processing unit), assembly code data (e.g., low-level programming languages that use symbolic representations of machine code instructions to program a processing unit), genetic data or other chemical or biochemical data, image data, audio data, audiovisual data, haptic data, biometric data, medical data, financial data, statistical data, geographical data, astronomical data, historical data, sensor data generally (e.g., digital or analog values, such as voltage or other absolute or relative level measurement values from a real or artificial input, such as from an audio sensor, light sensor, displacement sensor, etc.), and the like. Data can be raw or processed and can be in any format or schema.
[0110] In multimodal inputs 2 or outputs 3, example combinations of data types include image data and audio data, image data and natural language data, natural language data and software code data, image data and biometric data, sensor data and medical data, etc. It is to be understood that any combination of data types in an input 2 or an output 3 can be present.
[0111] An example input 2 can include one or multiple data types, such as the example data types noted above. An example output 3 can include one or multiple data types, such as the example data types noted above. The data type(s) of input 2 can be the same as or different from the data type(s) of output 3. It is to be understood that the example data types noted above are provided for illustrative purposes only. Data types contemplated within the scope of the present disclosure are not limited to those examples noted above.
[0112] Figure 6 is a block diagram of an example implementation of an example machine-learned model configured to process sequences of information. For instance, an example implementation of machine-learned model(s) 1 can include machine- learned sequence processing model(s) 4. An example system can pass input(s) 2 to sequence processing model(s) 4. Sequence processing model(s) 4 can include one or more machine-learned components. Sequence processing model(s) 4 can process the data from input(s) 2 to obtain an input sequence 5. Input sequence 5 can include one or more input elements 5-1, 5-2, . . . , 5-A7, etc. obtained from input(s) 2. Sequence processing model 4 can process input sequence 5 using prediction layer(s) 6 to generate an output sequence 7. Output sequence 7 can include one or more outputelements 7-1, 7-2, . . . , 7-7V, etc. generated based on input sequence 5. The system can generate output(s) 3 based on output sequence 7.
[0113] Sequence processing model(s) 4 can include one or multiple machine- learned model components configured to ingest, generate, or otherwise reason over sequences of information. For example, some example sequence processing models in the text domain are referred to as “Large Language Models,” or LLMs. See, e.g., PaLM 2 Technical Report, GOOGLE, https: / / ai.google / static / documents / palm2techreport.pdf (n.d.). Other example sequence processing models can operate in other domains, such as image domains, see, e.g., Dosovitskiy et al., An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale, ARXIV:2010.11929v2 (Jun. 3, 2021), audio domains, see, e.g., Agostinelli et al., MusicLM: Generating Music From Text, ARXIV:2301.11325vl (Jan. 26, 2023), biochemical domains, see, e.g., Jumper et al., Highly accurate protein structure prediction with AlphaFold, 596 Nature 583 (Aug. 26, 2021), by way of example. Sequence processing model(s) 4 can process one or multiple types of data simultaneously. Sequence processing model(s) 4 can include relatively large models (e.g., more parameters, computationally expensive, etc.), relatively small models (e.g., fewer parameters, computationally lightweight, etc.), or both.
[0114] In general, sequence processing model(s) 4 can obtain input sequence 5 using data from input(s) 2. For instance, input sequence 5 can include a representation of data from input(s) 2 in a format understood by sequence processing model(s) 4. One or more machine-learned components of sequence processing model(s) 4 can ingest the data from input(s) 2, parse the data into pieces compatible with the processing architectures of sequence processing model(s) 4 (e.g., via “tokenization”), and project the pieces into an input space associated with prediction layer(s) 6 (e.g., via “embedding”).
[0115] Sequence processing model(s) 4 can ingest the data from input(s) 2 and parse the data into a sequence of elements to obtain input sequence 5. For example, a portion of input data from input(s) 2 can be broken down into pieces that collectively represent the content of the portion of the input data. The pieces can provide the elements of the sequence.
[0116] Elements 5-1, 5-2, . . . , 5-M can represent, in some cases, building blocks for capturing or expressing meaningful information in a particular data domain. For instance, the elements can describe “atomic units” across one or more domains. For example, for textual input source(s), the elements can correspond to groups of one or more words or sub-word components, such as sets of one or more characters.
[0117] For example, elements 5-1, 5-2, . . . , 5-M can represent tokens obtained using a tokenizer. For instance, a tokenizer can process a given portion of an input source and output a series of tokens (e.g., corresponding to input elements 5-1, 5-2, . . . , 5-M) that represent the portion of the input source. Various approaches to tokenization can be used. For instance, textual input source(s) can be tokenized using a byte-pair encoding (BPE) technique. See, e.g., Kudo et al., SentencePiece: A simple and language independent subword tokenizer and detokenizer for Neural Text Processing, PROCEEDINGS OF THE 2018 CONFERENCE ON EMPIRICAL METHODS IN NATURAL LANGUAGE PROCESSING (System Demonstrations), pages 66-71 (October 31-November 4, 2018), https: / / aclanthology.org / D18-2012.pdf. Image-based input source(s) can be tokenized by extracting and serializing patches from an image.
[0118] In general, arbitrary data types can be serialized and processed into input sequence 5. It is to be understood that element(s) 5-1, 5-2, . . . , 5-M depicted in Figure 6 can be the tokens or can be the embedded representations thereof.
[0119] Prediction layer(s) 6 can predict one or more output elements 7-1, 7-2, . . . , 7-V based on the input elements. Prediction layer(s) 6 can include one or more machine-learned model architectures, such as one or more layers of learned parameters that manipulate and transform the input(s) to extract higher-order meaning from, and relationships between, input element(s) 5-1, 5-2, . . . , 5-M. In this manner, for instance, example prediction layer(s) 6 can predict new output element(s) in view of the context provided by input sequence 5.
[0120] Prediction layer(s) 6 can evaluate associations between portions of input sequence 5 and a particular output element. These associations can inform a prediction of the likelihood that a particular output follows the input context. For example, consider the textual snippet, “The carpenter’s toolbox was small and heavy. It was full of ” Example prediction layer(s) 6 can identify that “It” refers back to “toolbox” by determining a relationship between the respective embeddings. Exampleprediction layer(s) 6 can also link “It” to the attributes of the toolbox, such as “small” and “heavy.” Based on these associations, prediction layer(s) 6 can, for instance, assign a higher probability to the word “nails” than to the word “sawdust.”
[0121] A transformer is an example architecture that can be used in prediction layer(s) 4. See, e.g., Vaswani et al., Atention Is All You Need, ARXIV: 1706.03762v7 (Aug. 2, 2023). A transformer is an example of a machine-learned model architecture that uses an attention mechanism to compute associations between items within a context window. The context window can include a sequence that contains input sequence 5 and potentially one or more output element(s) 7-1, 7-2, . . . , 1-N. A transformer block can include one or more attention layer(s) and one or more postattention layer(s) (e.g., feedforward layer(s), such as a multi-layer perceptron).
[0122] Prediction layer(s) 6 can include other machine-learned model architectures in addition to or in lieu of transformer-based architectures. For example, recurrent neural networks (RNNs) and long short-term memory (LSTM) models can also be used, as well as convolutional neural networks (CNNs). In general, prediction layer(s) 6 can leverage various kinds of artificial neural networks that can understand or generate sequences of information.
[0123] Output sequence 7 can include or otherwise represent the same or different data types as input sequence 5. For instance, input sequence 5 can represent textual data, and output sequence 7 can represent textual data. Input sequence 5 can represent image, audio, or audiovisual data, and output sequence 7 can represent textual data (e.g., describing the image, audio, or audiovisual data). It is to be understood that prediction layer(s) 6, and any other interstitial model components of sequence processing model(s) 4, can be configured to receive a variety of data types in input sequence(s) 5 and output a variety of data types in output sequence(s) 7.
[0124] Output sequence 7 can have various relationships to input sequence 5. Output sequence 7 can be a continuation of input sequence 5. Output sequence 7 can be complementary to input sequence 5. Output sequence 7 can translate, transform, augment, or otherwise modify input sequence 5. Output sequence 7 can answer, evaluate, confirm, or otherwise respond to input sequence 5. Output sequence 7 can implement (or describe instructions for implementing) an instruction provided via input sequence 5.
[0125] Output sequence 7 can be generated autoregressively. For instance, for some applications, an output of one or more prediction layer(s) 6 can be passed through one or more output layers (e.g., softmax layer) to obtain a probability distribution over an output vocabulary (e.g., a textual or symbolic vocabulary) conditioned on a set of input elements in a context window. In this manner, for instance, output sequence 7 can be autoregressively generated by sampling a likely next output element, adding that element to the context window, and re-generating the probability distribution based on the updated context window, and sampling a likely next output element, and so forth.
[0126] Output sequence 7 can also be generated non-autoregressively. For instance, multiple output elements of output sequence 7 can be predicted together without explicit sequential conditioning on each other. See, e.g., Saharia et al., Non- Autoregressive Machine Translation with Latent Alignments, ARXIV:2004.07437V3 (Nov. 16, 2020).
[0127] Output sequence 7 can include one or multiple portions or elements. In an example content generation configuration, output sequence 7 can include multiple elements corresponding to multiple portions of a generated output sequence (e.g., a textual sentence, values of a discretized waveform, computer code, etc.). In an example classification configuration, output sequence 7 can include a single element associated with a classification output. For instance, an output “vocabulary” can include a set of classes into which an input sequence is to be classified. For instance, a vision transformer block can pass latent state information to a multilayer perceptron that outputs a likely class value associated with an input image.
[0128] Figure 7 is a block diagram of an example technique for populating an example input sequence 8. Input sequence 8 can include various functional elements that form part of the model infrastructure, such as an element 8-0 obtained from a task indicator 9 that signals to any model(s) that process input sequence 8 that a particular task is being performed (e.g., to help adapt a performance of the model(s) to that particular task). Input sequence 8 can include various data elements from different data modalities. For instance, an input modality 10-1 can include one modality of data. A data-to- sequence model 11-1 can process data from input modality 10-1 to project the data into a format compatible with input sequence 8 (e.g., one or morevectors dimensioned according to the dimensions of input sequence 8) to obtain elements 8-1, 8-2, 8-3. Another input modality 10-2 can include a different modality of data. A data-to-sequence model 11-2 can project data from input modality 10-2 into a format compatible with input sequence 8 to obtain elements 8-4, 8-5, 8-6. Another input modality 10-3 can include yet another different modality of data. A data-to- sequence model 11-3 can project data from input modality 10-3 into a format compatible with input sequence 8 to obtain elements 8-7, 8-8, 8-9.
[0129] Input sequence 8 can be the same as or different from input sequence 5. Input sequence 8 can be a multimodal input sequence that contains elements that represent data from different modalities using a common dimensional representation. For instance, an embedding space can have P dimensions. Input sequence 8 can be configured to contain a plurality of elements that have P dimensions. In this manner, for instance, example implementations can facilitate information extraction and reasoning across diverse data modalities by projecting data into elements in the same embedding space for comparison, combination, or other computations therebetween.
[0130] For example, elements 8-0, . . . , 8-9 can indicate particular locations within a multidimensional embedding space. Some elements can map to a set of discrete locations in the embedding space. For instance, elements that correspond to discrete members of a predetermined vocabulary of tokens can map to discrete locations in the embedding space that are associated with those tokens. Other elements can be continuously distributed across the embedding space. For instance, some data types can be broken down into continuously defined portions (e.g., image patches) that can be described using continuously distributed locations within the embedding space.
[0131] In some implementations, the expressive power of the embedding space cannot be limited to meanings associated with any particular set of tokens or other building blocks. For example, a continuous embedding space can encode a spectrum of high-order information. An individual piece of information (e.g., a token) can map to a particular point in that space: for instance, a token for the word “dog” can be projected to an embedded value that points to a particular location in the embedding space associated with canine-related information. Similarly, an image patch of an image of a dog on grass can also be projected into the embedding space. In someimplementations, the projection of the image of the dog can be similar to the projection of the word “dog” while also having similarity to a projection of the word “grass,” while potentially being different from both. In some implementations, the projection of the image patch cannot exactly align with any single projection of a single word. In some implementations, the projection of the image patch can align with a combination of the projections of the words “dog” and “grass.” In this manner, for instance, a high-order embedding space can encode information that can be independent of data modalities in which the information is expressed.
[0132] Task indicator 9 can include a model or model component configured to identify a task being performed and inject, into input sequence 8, an input value represented by element 8-0 that signals which task is being performed. For instance, the input value can be provided as a data type associated with an input modality and projected along with that input modality (e.g., the input value can be a textual task label that is embedded along with other textual data in the input; the input value can be a pixel-based representation of a task that is embedded along with other image data in the input; etc.). The input value can be provided as a data type that differs from or is at least independent from other input(s). For instance, the input value represented by element 8-0 can be learned within a continuous embedding space.
[0133] Input modalities 10-1, 10-2, and 10-3 can be associated with various different data types (e.g., as described above with respect to input(s) 2 and output(s) 3).
[0134] Data-to-sequence models 11-1, 11-2, and 11-3 can be the same or different from each other. Data-to-sequence models 11-1, 11-2, and 11-3 can be adapted to each respective input modality 10-1, 10-2, and 10-3. For example, a textual data-to- sequence model can subdivide a portion of input text and project the subdivisions into element(s) in input sequence 8 (e.g., elements 8-1, 8-2, 8-3, etc.). An image data-to- sequence model can subdivide an input image and project the subdivisions into element(s) in input sequence 8 (e.g., elements 8-4, 8-5, 8-6, etc.). An arbitrary datatype data-to-sequence model can subdivide an input of that arbitrary datatype and project the subdivisions into element(s) in input sequence 8 (e.g., elements 8-7, 8-8, 8- 9, etc.).
[0135] Data-to-sequence models 11-1, 11-2, and 11-3 can form part of machine- learned sequence processing model(s) 4. Data-to-sequence models 11-1, 11-2, and 11- 3 can be jointly trained with or trained independently from machine-learned sequence processing model(s) 4. Data-to-sequence models 11-1, 11-2, and 11-3 can be trained end-to-end with machine-learned sequence processing model(s) 4.
[0136] Figure 8 is a block diagram of an example model development platform 12 that can facilitate creation, adaptation, and refinement of example machine-learned models (e.g., machine-learned model(s) 1, sequence processing model(s) 4, etc.). Model development platform 12 can provide a number of different toolkits that developer systems can employ in the development of new or adapted machine-learned models.
[0137] Model development platform 12 can provide one or more model libraries13 containing building blocks for new models. Model libraries 13 can include one or more pre-trained foundational models 13-1, which can provide a backbone of processing power across various tasks. Model libraries 13 can include one or more pre-trained expert models 13-2, which can be focused on performance in particular domains of expertise. Model libraries 13 can include various model primitives 13-3, which can provide low-level architectures or components (optionally pre-trained), which can be assembled in various arrangements as desired. Model primitives 13-3 can include a library of pre-trained adapters or LoRA modules that can adapt a baseline foundational model to align its outputs with a desired performance profile, augment model capabilities (e.g., to adapt to a different input modality, etc.), and the like.
[0138] Model development platform 12 can receive selections of various model components 14. Model development platform 12 can pass selected model components14 to a workbench 15 that combines selected model components 14 into a development model 16.
[0139] Workbench 15 can facilitate further refinement and adaptation of development model 16 by leveraging a number of different toolkits integrated with model development platform 12. For example, workbench 15 can facilitate alignment of the development model 16 with a desired performance profile on various tasks using a model alignment toolkit 17.
[0140] Model alignment toolkit 17 can provide a number of tools for causing development model 16 to generate outputs aligned with desired behavioral characteristics. Alignment can include increasing an accuracy, precision, recall, etc. of model outputs. Alignment can include enforcing output styles, schema, or other preferential characteristics of model outputs. Alignment can be general or domainspecific. For instance, a pre-trained foundational model 13-1 can begin with an initial level of performance across multiple domains. Alignment of the pre-trained foundational model 13-1 can include improving a performance in a particular domain of information or tasks (e.g., even at the expense of performance in another domain of information or tasks).
[0141] Model alignment toolkit 17 can integrate one or more dataset(s) 17-1 for aligning development model 16. Curated dataset(s) 17-1 can include labeled or unlabeled training data. Dataset(s) 17-1 can be obtained from public domain datasets. Dataset(s) 17-1 can be obtained from private datasets associated with one or more developer system(s) for the alignment of bespoke machine-learned model(s) customized for private use-cases.
[0142] Pre-training pipelines 17-2 can include a machine-learned model training workflow configured to update development model 16 over large-scale, potentially noisy datasets. For example, pre-training can leverage unsupervised learning techniques (e.g., de-noising, etc.) to process large numbers of training instances to update model parameters from an initialized state and achieve a desired baseline performance. Pre-training pipelines 17-2 can leverage unlabeled datasets in dataset(s) 17-1 to perform pre-training. Workbench 15 can implement a pre-training pipeline 17- 2 to pre-train development model 16.
[0143] Fine-tuning pipelines 17-3 can include a machine-learned model training workflow configured to refine the model parameters of development model 16 with higher-quality data. Fine-tuning pipelines 17-3 can update development model 16 by conducting supervised training with labeled dataset(s) in dataset(s) 17-1. Fine-tuning pipelines 17-3 can update development model 16 by conducting reinforcement learning using reward signals from user feedback signals. Workbench 15 can implement a fine-tuning pipeline 17-3 to fine-tune development model 16.
[0144] Prompt libraries 17-4 can include sets of inputs configured to induce behavior aligned with desired performance criteria. Prompt libraries 17-4 can include few-shot prompts (e.g., inputs providing examples of desired model outputs for prepending to a desired runtime query), chain-of-thought prompts (e.g., inputs providing step-by-step reasoning within the exemplars to facilitate thorough reasoning by the model), and the like.
[0145] Example prompts can be retrieved from an available repository of prompt libraries 17-4. Example prompts can be contributed by one or more developer systems using workbench 15.
[0146] In some implementations, pre-trained or fine-tuned models can achieve satisfactory performance without exemplars in the inputs. For instance, zero-shot prompts can include inputs that lack exemplars. Zero-shot prompts can be within a domain within a training dataset or outside of the training domain(s).
[0147] Prompt libraries 17-4 can include one or more prompt engineering tools. Prompt engineering tools can provide workflows for retrieving or learning optimized prompt values. Prompt engineering tools can facilitate directly learning prompt values (e.g., input element values) based on one or more training iterations. Workbench 15 can implement prompt engineering tools in development model 16.
[0148] Prompt libraries 17-4 can include pipelines for prompt generation. For example, inputs can be generated using development model 16 itself or other machine-learned models. In this manner, for instance, a first model can process information about a task and output an input for a second model to process in order to perform a step of the task. The second model can be the same as or different from the first model. Workbench 15 can implement prompt generation pipelines in development model 16.
[0149] Prompt libraries 17-4 can include pipelines for context injection. For instance, a performance of development model 16 on a particular task can improve if provided with additional context for performing the task. Prompt libraries 17-4 can include software components configured to identify desired context, retrieve the context from an external source (e.g., a database, a sensor, etc.), and add the context to the input prompt. Workbench 15 can implement context injection pipelines in development model 16.
[0150] Although various training examples described herein with respect to model development platform 12 refer to “pre-training” and “fine-tuning,” it is to be understood that model alignment toolkit 17 can generally support a wide variety of training techniques adapted for training a wide variety of machine-learned models. Example training techniques can correspond to the example training method 400 described above.
[0151] Model development platform 12 can include a model plugin toolkit 18. Model plugin toolkit 18 can include a variety of tools configured for augmenting the functionality of a machine-learned model by integrating the machine-learned model with other systems, devices, and software components. For instance, a machine- learned model can use tools to increase performance quality where appropriate. For instance, deterministic tasks can be offloaded to dedicated tools in lieu of probabilistically performing the task with an increased risk of error. For instance, instead of autoregressively predicting the solution to a system of equations, a machine-learned model can recognize a tool to call for obtaining the solution and pass the system of equations to the appropriate tool. The tool can be a traditional system of equations solver that can operate deterministically to resolve the system of equations. The output of the tool can be returned in response to the original query. In this manner, tool use can allow some example models to focus on the strengths of machine-learned models — e.g., understanding an intent in an unstructured request for a task — while augmenting the performance of the model by offloading certain tasks to a more focused tool for rote application of deterministic algorithms to a well-defined problem.
[0152] Model plugin toolkit 18 can include validation tools 18-1. Validation tools 18-1 can include tools that can parse and confirm output(s) of a machine-learned model. Validation tools 18-1 can include engineered heuristics that establish certain thresholds applied to model outputs. For example, validation tools 18-1 can ground the outputs of machine-learned models to structured data sources (e.g., to mitigate “hallucinations”).
[0153] Model plugin toolkit 18 can include tooling packages 18-2 for implementing one or more tools that can include scripts or other executable code that can be executed alongside development model 16. Tooling packages 18-2 can includeone or more inputs configured to cause machine-learned model(s) to implement the tools (e.g., few-shot prompts that induce a model to output tool calls in the proper syntax, etc.). Tooling packages 18-2 can include, for instance, fine-tuning training data for training a model to use a tool.
[0154] Model plugin toolkit 18 can include interfaces for calling external application programming interfaces (APIs) 18-3. For instance, in addition to or in lieu of implementing tool calls or tool code directly with development model 16, development model 16 can be aligned to output instructions that initiate API calls to send or obtain data via external systems.
[0155] Model plugin toolkit 18 can integrate with prompt libraries 17-4 to build a catalog of available tools for use with development model 16. For instance, a model can receive, in an input, a catalog of available tools, and the model can generate an output that selects a tool from the available tools and initiates a tool call for using the tool.
[0156] Model development platform 12 can include a computational optimization toolkit 19 for optimizing a computational performance of development model 16. For instance, tools for model compression 19-1 can allow development model 16 to be reduced in size while maintaining a desired level of performance. For instance, model compression 19-1 can include quantization workflows, weight pruning and sparsification techniques, etc. Tools for hardware acceleration 19-2 can facilitate the configuration of the model storage and execution formats to operate optimally on different hardware resources. For instance, hardware acceleration 19-2 can include tools for optimally sharding models for distributed processing over multiple processing units for increased bandwidth, lower unified memory requirements, etc. Tools for distillation 19-3 can provide for the training of lighter- weight models based on the knowledge encoded in development model 16. For instance, development model 16 can be a highly performant, large machine-learned model optimized using model development platform 12. To obtain a lightweight model for running in resource-constrained environments, a smaller model can be a “student model” that learns to imitate development model 16 as a “teacher model.” In this manner, for instance, the investment in learning the parameters and configurations of developmentmodel 16 can be efficiently transferred to a smaller model for more efficient inference.
[0157] Workbench 15 can implement one, multiple, or none of the toolkits implemented in model development platform 12. Workbench 15 can output an output model 20 based on development model 16. Output model 20 can be a deployment version of development model 16. Output model 20 can be a development or training checkpoint of development model 16. Output model 20 can be a distilled, compressed, or otherwise optimized version of development model 16.
[0158] Figure 9 is a block diagram of an example training flow for training a machine-learned development model 16. One or more portion(s) of the example training flow can be implemented by a computing system that includes one or more computing devices such as, for example, computing systems described with reference to the other figures. Each respective portion of the example training flow can be performed by any (or any combination) of one or more computing devices. Moreover, one or more portion(s) of the example training flow can be implemented on the hardware components of the device(s) described herein, for example, to train one or more systems or models. Figure 9 depicts elements performed in a particular order for purposes of illustration and discussion. Those of ordinary skill in the art, using the disclosures provided herein, will understand that the elements of any of the methods discussed herein can be adapted, rearranged, expanded, omitted, combined, or modified in various ways without deviating from the scope of the present disclosure. Figure 9 is described with reference to elements / terms described with respect to other systems and figures for exemplary illustrated purposes and is not meant to be limiting. One or more portions of the example training flow can be performed additionally, or alternatively, by other systems.
[0159] Initially, development model 16 can persist in an initial state as an initialized model 21. Development model 16 can be initialized with weight values. Initial weight values can be random or based on an initialization schema. Initial weight values can be based on prior pre-training for the same or for a different model.
[0160] Initialized model 21 can undergo pre-training in a pre-training stage 22. Pre-training stage 22 can be implemented using one or more pre-training pipelines 17- 2 over data from dataset(s) 17-1. Pre-training can be omitted, for example, ifinitialized model 21 is already pre-trained (e.g., development model 16 contains, is, or is based on a pre-trained foundational model or an expert model).
[0161] Pre-trained model 23 can then be a new version of development model 16, which can persist as development model 16 or as a new development model. Pretrained model 23 can be the initial state if development model 16 was already pretrained. Pre-trained model 23 can undergo fine-tuning in a fine-tuning stage 24. Fine- tuning stage 24 can be implemented using one or more fine-tuning pipelines 17-3 over data from dataset(s) 17-1. Fine-tuning can be omitted, for example, if a pre-trained model has satisfactory performance, if the model was already fine-tuned, or if other tuning approaches are preferred.
[0162] Fine-tuned model 29 can then be a new version of development model 16, which can persist as development model 16 or as a new development model. Finetuned model 29 can be the initial state if development model 16 was already finetuned. Fine-tuned model 29 can undergo refinement with user feedback 26. For instance, refinement with user feedback 26 can include reinforcement learning, optionally based on human feedback from human users of fine-tuned model 25. As reinforcement learning can be a form of fine-tuning, it is to be understood that finetuning stage 24 can subsume the stage for refining with user feedback 26. Refinement with user feedback 26 can produce a refined model 27. Refined model 27 can be output to downstream system(s) 28 for deployment or further development.
[0163] In some implementations, computational optimization operations can be applied before, during, or after each stage. For instance, initialized model 21 can undergo computational optimization 29-1 (e.g., using computational optimization toolkit 19) before pre-training stage 22. Pre-trained model 23 can undergo computational optimization 29-2 (e.g., using computational optimization toolkit 19) before fine-tuning stage 24. Fine-tuned model 25 can undergo computational optimization 29-3 (e.g., using computational optimization toolkit 19) before refinement with user feedback 26. Refined model 27 can undergo computational optimization 29-4 (e.g., using computational optimization toolkit 19) before output to downstream system(s) 28. Computational optimization(s) 29-1, . . . , 29-4 can all be the same, all be different, or include at least some different optimization techniques.
[0164] Figure 10 is a block diagram of an inference system for operating one or more machine-learned model(s) 1 to perform inference (e.g., for training, for deployment, etc.). A model host 31 can receive machine-learned model(s) 1. Model host 31 can host one or more model instance(s) 31-1, which can be one or multiple instances of one or multiple models. Model host 31 can host model instance(s) 31-1 using available computer resources 31-2 associated with model host 31.
[0165] Model host 31 can perform inference on behalf of one or more client(s) 32. Client(s) 32 can transmit an input request 33 to model host 31. Using input request 33, model host 31 can obtain input(s) 2 for input to machine-learned model(s) 1. Machine-learned model(s) 1 can process input(s) 2 to generate output(s) 3. Using output(s) 3, model host 31 can return an output payload 34 for responding to input request 33 from client(s) 32. Output payload 34 can include or be based on output(s) 3.
[0166] Model host 31 can leverage various other resources and tools to augment the inference task. For instance, model host 31 can communicate with tool interfaces 35 to facilitate tool use by model instance(s) 31-1. Tool interfaces 35 can include local or remote APIs. Tool interfaces 35 can include integrated scripts or other software functionality. Model host 31 can engage online learning interface(s) 36 to facilitate ongoing improvements to machine-learned model(s) 1. For instance, online learning interface(s) 36 can be used within reinforcement learning loops to retrieve user feedback on inferences served by model host 31. Model host 31 can access runtime data source(s) 37 for augmenting input(s) 2 with additional contextual information. For instance, runtime data source(s) 37 can include a knowledge graph 37-1 that facilitates structured information retrieval for information associated with input request(s) 33 (e.g., a search engine service). Runtime data source(s) 37 can include public or private, external, or local database(s) 37-2 that can store information associated with input request(s) 33 for augmenting input(s) 2. Runtime data source(s) 37 can include account data 37-3 which can be retrieved in association with a user account corresponding to a client 32 for customizing the behavior of model host 31 accordingly.
[0167] Model host 31 can be implemented by one or multiple computing devices or systems. Client(s) 2 can be implemented by one or multiple computing devices or systems, which can include computing devices or systems shared with model host 31.
[0168] For example, model host 31 can operate on a server system that provides a machine-learning service to client device(s) that operate client(s) 32 (e.g., over a local or wide-area network). Client device(s) can be end-user devices used by individuals. Client device(s) can be server systems that operate client(s) 32 to provide various functionality as a service to downstream end-user devices.
[0169] In some implementations, model host 31 can operate on a same device or system as client(s) 32. Model host 31 can be a machine-learning service that runs on- device to provide machine-learning functionality to one or multiple applications operating on a client device, which can include an application implementing client(s) 32. Model host 31 can be a part of a same application as client(s) 32. For instance, model host 31 can be a subroutine or method implemented by one part of an application, and client(s) 32 can be another subroutine or method that engages model host 31 to perform inference functions within the application. It is to be understood that model host 31 and client(s) 32 can have various different configurations.
[0170] Model instance(s) 31-1 can include one or more machine-learned models that are available for performing inference. Model instance(s) 31-1 can include weights or other model components that are stored in persistent storage, temporarily cached, or loaded into high-speed memory. Model instance(s) 31-1 can include multiple instance(s) of the same model (e.g., for parallel execution of more requests on the same model). Model instance(s) 31-1 can include instance(s) of different model(s). Model instance(s) 31-1 can include cached intermediate states of active or inactive model(s) used to accelerate inference of those models. For instance, an inference session with a particular model can generate significant amounts of computational results that can be re-used for future inference runs (e.g., using a KV cache for transformer-based models). These computational results can be saved in association with that inference session so that session can be executed more efficiently when resumed.
[0171] Compute resource(s) 31-2 can include one or more processors (central processing units, graphical processing units, tensor processing units, machine-learningaccelerators, etc.) connected to one or more memory devices. Compute resource(s) 31-2 can include a dynamic pool of available resources shared with other processes. Compute resource(s) 31-2 can include memory devices large enough to fit an entire model instance in a single memory instance. Compute resource(s) 31-2 can also shard model instance(s) across multiple memory devices (e.g., using data parallelization or tensor parallelization, etc.). This can be done to increase parallelization or to execute a large model using multiple memory devices which individually cannot be able to fit the entire model into memory.
[0172] Input request 33 can include data for input(s) 2. Model host 31 can process input request 33 to obtain input(s) 2. Input(s) 2 can be obtained directly from input request 33 or can be retrieved using input request 33. Input request 33 can be submitted to model host 31 via an API.
[0173] Model host 31 can perform inference over batches of input requests 33 in parallel. For instance, a model instance 31-1 can be configured with an input structure that has a batch dimension. Separate input(s) 2 can be distributed across the batch dimension (e.g., rows of an array). The separate input(s) 2 can include completely different contexts. The separate input(s) 2 can be multiple inference steps of the same task. The separate input(s) 2 can be staggered in an input structure, such that any given inference cycle can be operating on different portions of the respective input(s) 2. In this manner, for instance, model host 31 can perform inference on the batch in parallel, such that output(s) 3 can also contain the batch dimension and return the inference results for the batched input(s) 2 in parallel. In this manner, for instance, batches of input request(s) 33 can be processed in parallel for higher throughput of output payload(s) 34.
[0174] Output payload 34 can include or be based on output(s) 3 from machine- learned model(s) 1. Model host 31 can process output(s) 3 to obtain output payload 34. This can include chaining multiple rounds of inference (e.g., iteratively, recursively, across the same model(s) or different model(s)) to arrive at a final output for a task to be returned in output payload 34. Output payload 34 can be transmitted to client(s) 32 via an API.
[0175] Online learning interface(s) 36 can facilitate reinforcement learning of machine-learned model(s) 1. Online learning interface(s) 36 can facilitatereinforcement learning with human feedback (RLHF). Online learning interface(s) 36 can facilitate federated learning of machine-learned model(s) 1.
[0176] Model host 31 can access a library of pre-trained adapters or LoRA modules that can adapt a baseline model to align its outputs with a desired performance profile, augment model capabilities (e.g., to adapt to a different input modality, etc.), and the like. For instance, model host 31 can receive an input request to load a customized model, and model host 31 can retrieve one or more components to adapt a baseline model to the custom profile. Model host 31 can determine that a particular functionality is needed for a particular task (e.g., based on an output of a model that preprocesses an input) and retrieve a pre-trained component accordingly.
[0177] Model host 31 can execute machine-learned model(s) 1 to perform inference for various tasks using various types of data. For example, various different input(s) 2 and output(s) 3 can be used for various different tasks. In some implementations, input(s) 2 can be or otherwise represent image data. Machine- learned model(s) 1 can process the image data to generate an output. As an example, machine-learned model(s) 1 can process the image data to generate an image recognition output (e.g., a recognition of the image data, a latent embedding of the image data, an encoded representation of the image data, a hash of the image data, etc.). As another example, machine-learned model(s) 1 can process the image data to generate an image segmentation output. As another example, machine-learned model(s) 1 can process the image data to generate an image classification output. As another example, machine-learned model(s) 1 can process the image data to generate an image data modification output (e.g., an alteration of the image data, etc.). As another example, machine-learned model(s) 1 can process the image data to generate an encoded image data output (e.g., an encoded or compressed representation of the image data, etc.). As another example, machine-learned model(s) 1 can process the image data to generate an upscaled image data output. As another example, machine- learned model(s) 1 can process the image data to generate a prediction output.
[0178] In some implementations, the task is a computer vision task. In some cases, input(s) 2 includes pixel data for one or more images and the task is an image processing task. For example, the image processing task can be image classification, where the output is a set of scores, each score corresponding to a different object classand representing the likelihood that the one or more images depict an object belonging to the object class. The image processing task can be object detection, where the image processing output identifies one or more regions in the one or more images and, for each region, a likelihood that region depicts an object of interest. As another example, the image processing task can be image segmentation, where the image processing output defines, for each pixel in the one or more images, a respective likelihood for each category in a predetermined set of categories. For example, the set of categories can be foreground and background. As another example, the set of categories can be object classes. As another example, the image processing task can be depth estimation, where the image processing output defines, for each pixel in the one or more images, a respective depth value. As another example, the image processing task can be motion estimation, where the network input includes multiple images, and the image processing output defines, for each pixel of one of the input images, a motion of the scene depicted at the pixel between the images in the network input.
[0179] In some implementations, input(s) 2 can be or otherwise represent natural language data. Machine-learned model(s) 1 can process the natural language data to generate an output. As an example, machine-learned model(s) 1 can process the natural language data to generate a language encoding output. As another example, machine-learned model(s) 1 can process the natural language data to generate a latent text embedding output. As another example, machine-learned model(s) 1 can process the natural language data to generate a translation output. As another example, machine-learned model(s) 1 can process the natural language data to generate a classification output. As another example, machine-learned model(s) 1 can process the natural language data to generate a textual segmentation output. As another example, machine-learned model(s) 1 can process the natural language data to generate a semantic intent output. As another example, machine-learned model(s) 1 can process the natural language data to generate an upscaled text or natural language output (e.g., text or natural language data that is higher quality than the input text or natural language, etc.). As another example, machine-learned model(s) 1 can process the natural language data to generate a prediction output (e.g., one or more predicted next portions of natural language content).
[0180] In some implementations, input(s) 2 can be or otherwise represent speech data (e.g., data describing spoken natural language, such as audio data, textual data, etc.). Machine-learned model(s) 1 can process the speech data to generate an output. As an example, machine-learned model(s) 1 can process the speech data to generate a speech recognition output. As another example, machine-learned model(s) 1 can process the speech data to generate a speech translation output. As another example, machine-learned model(s) 1 can process the speech data to generate a latent embedding output. As another example, machine-learned model(s) 1 can process the speech data to generate an encoded speech output (e.g., an encoded or compressed representation of the speech data, etc.). As another example, machine-learned model(s) 1 can process the speech data to generate an upscaled speech output (e.g., speech data that is higher quality than the input speech data, etc.). As another example, machine-learned model(s) 1 can process the speech data to generate a textual representation output (e.g., a textual representation of the input speech data, etc.). As another example, machine-learned model(s) 1 can process the speech data to generate a prediction output.
[0181] In some implementations, input(s) 2 can be or otherwise represent latent encoding data (e.g., a latent space representation of an input, etc.). Machine-learned model(s) 1 can process the latent encoding data to generate an output. As an example, machine-learned model(s) 1 can process the latent encoding data to generate a recognition output. As another example, machine-learned model(s) 1 can process the latent encoding data to generate a reconstruction output. As another example, machine-learned model(s) 1 can process the latent encoding data to generate a search output. As another example, machine-learned model(s) 1 can process the latent encoding data to generate a reclustering output. As another example, machine-learned model(s) 1 can process the latent encoding data to generate a prediction output.
[0182] In some implementations, input(s) 2 can be or otherwise represent statistical data. Statistical data can be, represent, or otherwise include data computed or calculated from some other data source. Machine-learned model(s) 1 can process the statistical data to generate an output. As an example, machine-learned model(s) 1 can process the statistical data to generate a recognition output. As another example, machine-learned model(s) 1 can process the statistical data to generate a predictionoutput. As another example, machine-learned model(s) 1 can process the statistical data to generate a classification output. As another example, machine-learned model(s) 1 can process the statistical data to generate a segmentation output. As another example, machine-learned model(s) 1 can process the statistical data to generate a visualization output. As another example, machine-learned model(s) 1 can process the statistical data to generate a diagnostic output.
[0183] In some implementations, input(s) 2 can be or otherwise represent sensor data. Machine-learned model(s) 1 can process the sensor data to generate an output. As an example, machine-learned model(s) 1 can process the sensor data to generate a recognition output. As another example, machine-learned model(s) 1 can process the sensor data to generate a prediction output. As another example, machine-learned model(s) 1 can process the sensor data to generate a classification output. As another example, machine-learned model(s) 1 can process the sensor data to generate a segmentation output. As another example, machine-learned model(s) 1 can process the sensor data to generate a visualization output. As another example, machine- learned model(s) 1 can process the sensor data to generate a diagnostic output. As another example, machine-learned model(s) 1 can process the sensor data to generate a detection output.
[0184] In some implementations, machine-learned model(s) 1 can be configured to perform a task that includes encoding input data for reliable or efficient transmission or storage (or corresponding decoding). For example, the task can be an audio compression task. The input can include audio data and the output can comprise compressed audio data. In another example, the input includes visual data (e.g., one or more images or videos), the output comprises compressed visual data, and the task is a visual data compression task. In another example, the task can comprise generating an embedding for input data (e.g., input audio or visual data). In some cases, the input includes audio data representing a spoken utterance and the task is a speech recognition task. The output can comprise a text output which is mapped to the spoken utterance. In some cases, the task comprises encrypting or decrypting input data. In some cases, the task comprises a microprocessor performance task, such as branch prediction or memory address translation.
[0185] In some implementations, the task is a generative task, and machine- learned model(s) 1 can be configured to output content generated in view of input(s) 2. For instance, input(s) 2 can be or otherwise represent data of one or more modalities that encodes context for generating additional content.
[0186] In some implementations, the task can be a text completion task. Machine- learned model(s) 1 can be configured to process input(s) 2 that represent textual data and to generate output(s) 3 that represent additional textual data that completes a textual sequence that includes input(s) 2. For instance, machine-learned model(s) 1 can be configured to generate output(s) 3 to complete a sentence, paragraph, or portion of text that follows from a portion of text represented by input(s) 2.
[0187] In some implementations, the task can be an instruction following task. Machine-learned model(s) 1 can be configured to process input(s) 2 that represent instructions to perform a function and to generate output(s) 3 that advance a goal of satisfying the instruction function (e.g., at least a step of a multi-step procedure to perform the function). Output(s) 3 can represent data of the same or of a different modality as input(s) 2. For instance, input(s) 2 can represent textual data (e.g., natural language instructions for a task to be performed) and machine-learned model(s) 1 can process input(s) 2 to generate output(s) 3 that represent textual data responsive to the instructions (e.g., natural language responses, programming language responses, machine language responses, etc.). Input(s) 2 can represent image data (e.g., imagebased instructions for a task to be performed, optionally accompanied by textual instructions) and machine-learned model(s) 1 can process input(s) 2 to generate output(s) 3 that represent textual data responsive to the instructions (e.g., natural language responses, programming language responses, machine language responses, etc.). One or more output(s) 3 can be iteratively or recursively generated to sequentially process and accomplish steps toward accomplishing the requested functionality. For instance, an initial output can be executed by an external system or be processed by machine-learned model(s) 1 to complete an initial step of performing a function. Multiple steps can be performed, with a final output being obtained that is responsive to the initial instructions.
[0188] In some implementations, the task can be a question answering task. Machine-learned model(s) 1 can be configured to process input(s) 2 that represent aquestion to answer and to generate output(s) 3 that advance a goal of returning an answer to the question (e.g., at least a step of a multi-step procedure to perform the function). Output(s) 3 can represent data of the same or of a different modality as input(s) 2. For instance, input(s) 2 can represent textual data (e.g., natural language instructions for a task to be performed) and machine-learned model(s) 1 can process input(s) 2 to generate output(s) 3 that represent textual data responsive to the question (e.g., natural language responses, programming language responses, machine language responses, etc.). Input(s) 2 can represent image data (e.g., image-based instructions for a task to be performed, optionally accompanied by textual instructions) and machine-learned model(s) 1 can process input(s) 2 to generate output(s) 3 that represent textual data responsive to the question (e.g., natural language responses, programming language responses, machine language responses, etc.). One or more output(s) 3 can be iteratively or recursively generated to sequentially process and accomplish steps toward answering the question. For instance, an initial output can be executed by an external system or be processed by machine-learned model(s) 1 to complete an initial step of obtaining an answer to the question (e.g., querying a database, performing a computation, executing a script, etc.). Multiple steps can be performed, with a final output being obtained that is responsive to the question.
[0189] In some implementations, the task can be an image generation task. Machine-learned model(s) 1 can be configured to process input(s) 2 that represent context regarding a desired portion of image content. The context can include text data, image data, audio data, etc. Machine-learned model(s) 1 can be configured to generate output(s) 3 that represent image data that depicts imagery related to the context. For instance, machine-learned model(s) 1 can be configured to generate pixel data of an image. Values for channel(s) associated with the pixels in the pixel data can be selected based on the context (e.g., based on a probability determined based on the context).
[0190] In some implementations, the task can be an audio generation task. Machine-learned model(s) 1 can be configured to process input(s) 2 that represent context regarding a desired portion of audio content. The context can include text data, image data, audio data, etc. Machine-learned model(s) 1 can be configured togenerate output(s) 3 that represent audio data related to the context. For instance, machine-learned model(s) 1 can be configured to generate waveform data in the form of an image (e.g., a spectrogram). Values for channel(s) associated with pixels of the image can be selected based on the context. Machine-learned model(s) 1 can be configured to generate waveform data in the form of a sequence of discrete samples of a continuous waveform. Values of the sequence can be selected based on the context (e.g., based on a probability determined based on the context).
[0191] In some implementations, the task can be a data generation task. Machine- learned model(s) 1 can be configured to process input(s) 2 that represent context regarding a desired portion of data (e.g., data from various data domains, such as sensor data, image data, multimodal data, statistical data, etc.). The desired data can be, for instance, synthetic data for training other machine-learned models. The context can include arbitrary data type(s). Machine-learned model(s) 1 can be configured to generate output(s) 3 that represent data that aligns with the desired data. For instance, machine-learned model(s) 1 can be configured to generate data values for populating a dataset. Values for the data object(s) can be selected based on the context (e.g., based on a probability determined based on the context).
[0192] Figure 11 is a block diagram of an example networked computing system that can perform aspects of example implementations of the present disclosure. The system can include a number of computing devices and systems that are communicatively coupled over a network 49. An example computing device 50 is described to provide an example of a computing device that can perform any aspect of the present disclosure (e.g., implementing model host 31, client(s) 32, or both). An example server computing system 60 is described as an example of a server computing system that can perform any aspect of the present disclosure (e.g., implementing model host 31, client(s) 32, or both). Computing device 50 and server computing system(s) 60 can cooperatively interact (e.g., over network 49) to perform any aspect of the present disclosure (e.g., implementing model host 31, client(s) 32, or both). Model development platform system 70 is an example system that can host or serve model development platform(s) 12 for development of machine-learned models. Third-party system(s) 80 are example system(s) with which any of computing device 50, server computing system(s) 60, or model development platform system(s) 70 caninteract in the performance of various aspects of the present disclosure (e.g., engaging third-party tools, accessing third-party databases or other resources, etc.).
[0193] Network 49 can be any type of communications network, such as a local area network (e.g., intranet), wide area network (e.g., Internet), or some combination thereof and can include any number of wired or wireless links. In general, communication over network 49 can be carried via any type of wired or wireless connection, using a wide variety of communication protocols (e.g., TCP / IP, HTTP, SMTP, FTP), encodings or formats (e.g., HTML, XML), or protection schemes (e.g., VPN, secure HTTP, SSL). Network 49 can also be implemented via a system bus. For instance, one or more devices or systems of Figure 11 can be co-located with, contained by, or otherwise integrated into one or more other devices or systems.
[0194] Computing device 50 can be any type of computing device, such as, for example, a personal computing device (e.g., laptop or desktop), a mobile computing device (e.g., smartphone or tablet), a gaming console or controller, a wearable computing device, an embedded computing device, a server computing device, a virtual machine operating on a host device, or any other type of computing device. Computing device 50 can be a client computing device. Computing device 50 can be an end-user computing device. Computing device 50 can be a computing device of a service provided that provides a service to an end user (who can use another computing device to interact with computing device 50).
[0195] Computing device 50 can include one or more processors 51 and a memory 52. Processor(s) 51 can be any suitable processing device (e.g., a processor core, a microprocessor, an ASIC, an FPGA, a controller, a microcontroller, etc.) and can be one processor or a plurality of processors that are operatively connected. Memory 52 can include one or more non-transitory computer-readable storage media, such as HBM, RAM, ROM, EEPROM, EPROM, flash memory devices, magnetic disks, etc., and combinations thereof. Memory 52 can store data 53 and instructions 54 which can be executed by processor(s) 51 to cause computing device 50 to perform operations. The operations can implement any one or multiple features described herein. The operations can implement example methods and techniques described herein.
[0196] Computing device 50 can also include one or more input components that receive user input. For example, a user input component can be a touch-sensitive component (e.g., a touch-sensitive display screen or a touch pad) that is sensitive to the touch of a user input object (e.g., a finger or a stylus). The touch-sensitive component can serve to implement a virtual keyboard. Other example user input components include a microphone, camera, LIDAR, a physical keyboard or other buttons, or other means by which a user can provide user input.
[0197] Computing device 50 can store or include one or more machine-learned models 55. Machine-learned models 55 can include one or more machine-learned model(s) 1, such as a sequence processing model 4. Machine-learned models 55 can include one or multiple model instance(s) 31-1. Machine-learned model(s) 55 can be received from server computing system(s) 60, model development platform system 70, third party system(s) 80 (e.g., an application distribution platform), or developed locally on computing device 50. Machine-learned model(s) 55 can be loaded into memory 52 and used or otherwise implemented by processor(s) 51. Computing device 50 can implement multiple parallel instances of machine-learned model(s) 55.
[0198] Server computing system(s) 60 can include one or more processors 61 and a memory 62. Processor(s) 61 can be any suitable processing device (e.g., a processor core, a microprocessor, an ASIC, an FPGA, a controller, a microcontroller, etc.) and can be one processor or a plurality of processors that are operatively connected. Memory 62 can include one or more non-transitory computer-readable storage media, such as HBM, RAM, ROM, EEPROM, EPROM, flash memory devices, magnetic disks, etc., and combinations thereof. Memory 62 can store data 63 and instructions 64 which can be executed by processor(s) 61 to cause server computing system(s) 60 to perform operations. The operations can implement any one or multiple features described herein. The operations can implement example methods and techniques described herein.
[0199] In some implementations, server computing system 60 includes or is otherwise implemented by one or multiple server computing devices. In instances in which server computing system 60 includes multiple server computing devices, such server computing devices can operate according to sequential computing architectures, parallel computing architectures, or some combination thereof.
[0200] Server computing system 60 can store or otherwise include one or more machine-learned models 65. Machine-learned model(s) 65 can be the same as or different from machine-learned model(s) 55. Machine-learned models 65 can include one or more machine-learned model(s) 1, such as a sequence processing model 4. Machine-learned models 65 can include one or multiple model instance(s) 31-1. Machine-learned model(s) 65 can be received from computing device 50, model development platform system 70, third party system(s) 80, or developed locally on server computing system(s) 60. Machine-learned model(s) 65 can be loaded into memory 62 and used or otherwise implemented by processor(s) 61. Server computing system(s) 60 can implement multiple parallel instances of machine-learned model(s) 65.
[0201] In an example configuration, machine-learned models 65 can be included in or otherwise stored and implemented by server computing system 60 to establish a client-server relationship with computing device 50 for serving model inferences. For instance, server computing system(s) 60 can implement model host 31 on behalf of client(s) 32 on computing device 50. For instance, machine-learned models 65 can be implemented by server computing system 60 as a portion of a web service (e.g., remote machine-learned model hosting service, such as an online interface for performing machine-learned model operations over a network on server computing system(s) 60). For instance, server computing system(s) 60 can communicate with computing device 50 over a local intranet or internet connection. For instance, computing device 50 can be a workstation or endpoint in communication with server computing system(s) 60, with implementation of machine-learned models 65 being managed by server computing system(s) 60 to remotely perform inference (e.g., for runtime or training operations), with output(s) returned (e.g., cast, streamed, etc.) to computing device 50. Machine-learned models 65 can work cooperatively or interoperatively with machine-learned models 55 on computing device 50 to perform various tasks.
[0202] Model development platform system(s) 70 can include one or more processors 71 and a memory 72. Processor(s) 71 can be any suitable processing device (e.g., a processor core, a microprocessor, an ASIC, an FPGA, a controller, a microcontroller, etc.) and can be one processor or a plurality of processors that areoperatively connected. Memory 72 can include one or more non-transitory computer- readable storage media, such as HBM, RAM, ROM, EEPROM, EPROM, flash memory devices, magnetic disks, etc., and combinations thereof. Memory 72 can store data 73 and instructions 74 which can be executed by processor(s) 71 to cause model development platform system(s) 70 to perform operations. The operations can implement any one or multiple features described herein. The operations can implement example methods and techniques described herein. Example operations include the functionality described herein with respect to model development platform 12. This and other functionality can be implemented by developer tool(s) 75.
[0203] Third-party system(s) 80 can include one or more processors 81 and a memory 82. Processor(s) 81 can be any suitable processing device (e.g., a processor core, a microprocessor, an ASIC, an FPGA, a controller, a microcontroller, etc.) and can be one processor or a plurality of processors that are operatively connected. Memory 82 can include one or more non-transitory computer-readable storage media, such as HBM, RAM, ROM, EEPROM, EPROM, flash memory devices, magnetic disks, etc., and combinations thereof. Memory 82 can store data 83 and instructions 84 which can be executed by processor(s) 81 to cause third-party system(s) 80 to perform operations. The operations can implement any one or multiple features described herein. The operations can implement example methods and techniques described herein. Example operations include the functionality described herein with respect to tools and other external resources called when training or performing inference with machine-learned model(s) 1, 4, 16, 20, 55, 65, etc. (e.g., third-party resource(s) 85).
[0204] Figure 11 illustrates one example arrangement of computing systems that can be used to implement the present disclosure. Other computing system configurations can be used as well. For example, in some implementations, one or both of computing system 50 or server computing system(s) 60 can implement all or a portion of the operations of model development platform system 70. For example, computing system 50 or server computing system(s) 60 can implement developer tool(s) 75 (or extensions thereof) to develop, update / train, or refine machine-learned models 1, 4, 16, 20, 55, 65, etc. using one or more techniques described herein with respect to model alignment toolkit 17. In this manner, for instance, computing system50 or server computing system(s) 60 can develop, update / train, or refine machine- learned models based on local datasets (e.g., for model personalization / customization, as permitted by user data preference selections).
[0205] Figure 12 is a block diagram of an example computing device 98 that performs according to example embodiments of the present disclosure. Computing device 98 can be a user computing device or a server computing device (e.g., computing device 50, server computing system(s) 60, etc.). Computing device 98 can implement model host 31. For instance, computing device 98 can include a number of applications (e.g., applications 1 through N). Each application can contain its own machine learning library and machine-learned model(s). For example, each application can include a machine-learned model. Example applications include a text messaging application, an email application, a dictation application, a virtual keyboard application, a browser application, etc. As illustrated in Figure 12, each application can communicate with a number of other components of the computing device, such as, for example, one or more sensors, a context manager, a device state component, or additional components. In some implementations, each application can communicate with each device component using an API (e.g., a public API). In some implementations, the API used by each application is specific to that application.
[0206] Figure 13 is a block diagram of an example computing device 99 that performs according to example embodiments of the present disclosure. Computing device 99 can be the same as or different from computing device 98. Computing device 99 can be a user computing device or a server computing device (e.g., computing device 50, server computing system(s) 60, etc.). Computing device 98 can implement model host 31. For instance, computing device 99 can include a number of applications (e.g., applications 1 through N). Each application can be in communication with a central intelligence layer. Example applications include a text messaging application, an email application, a dictation application, a virtual keyboard application, a browser application, etc. In some implementations, each application can communicate with the central intelligence layer (and model(s) stored therein) using an API (e.g., a common API across all applications).
[0207] The central intelligence layer can include a number of machine-learned models. For example, as illustrated in Figure 13, a respective machine-learned modelcan be provided for each application and managed by the central intelligence layer. In other implementations, two or more applications can share a single machine-learned model. For example, in some implementations, the central intelligence layer can provide a single model for all of the applications. In some implementations, the central intelligence layer is included within or otherwise implemented by an operating system of computing device 99.
[0208] The central intelligence layer can communicate with a central device data layer. The central device data layer can be a centralized repository of data for computing device 99. As illustrated in Figure 13, the central device data layer can communicate with a number of other components of the computing device, such as, for example, one or more sensors, a context manager, a device state component, or additional components. In some implementations, the central device data layer can communicate with each device component using an API (e.g., a private API).
[0209] The technology discussed herein makes reference to servers, databases, software applications, and other computer-based systems, as well as actions taken and information sent to and from such systems. The inherent flexibility of computer-based systems allows for a great variety of possible configurations, combinations, and divisions of tasks and functionality between and among components. For instance, processes discussed herein can be implemented using a single device or component or multiple devices or components working in combination. Databases and applications can be implemented on a single system or distributed across multiple systems. Distributed components can operate sequentially or in parallel.
[0210] While the present subject matter has been described in detail with respect to various specific example embodiments thereof, each example is provided by way of explanation, not limitation of the disclosure. Those skilled in the art, upon attaining an understanding of the foregoing, can readily produce alterations to, variations of, and equivalents to such embodiments. Accordingly, the subject disclosure does not preclude inclusion of such modifications, variations or additions to the present subject matter as would be readily apparent to one of ordinary skill in the art. For instance, features illustrated or described as part of one embodiment can be used with another embodiment to yield a still further embodiment. Thus, it is intended that the present disclosure cover such alterations, variations, and equivalents.
[0211] Aspects of the disclosure have been described in terms of illustrative embodiments thereof. Any and all features in the following claims can be combined or rearranged in any way possible, including combinations of claims not explicitly enumerated in combination together, as the example claim dependencies listed herein should not be read as limiting the scope of possible combinations of features disclosed herein. Accordingly, the scope of the present disclosure is by way of example rather than by way of limitation, and the subject disclosure does not preclude inclusion of such modifications, variations or additions to the present subject matter as would be readily apparent to one of ordinary skill in the art. Moreover, terms are described herein using lists of example elements joined by conjunctions such as “and,” “or,” “but,” etc. It should be understood that such conjunctions are provided for explanatory purposes only. Clauses and other sequences of items joined by a particular conjunction such as “or,” for example, can refer to “and / or,” “at least one of,” “any combination of’ example elements listed therein, etc. Terms such as “based on” should be understood as “based at least in part on.”
[0212] The term “can” should be understood as referring to a possibility of a feature in various implementations and not as prescribing an ability that is necessarily present in every implementation. For example, the phrase “X can perform Y” should be understood as indicating that, in various implementations, X has the potential to be configured to perform Y, and not as indicating that in every instance X must always be able to perform Y. It should be understood that, in various implementations, X can be unable to perform Y and remain within the scope of the present disclosure.
[0213] The term “can” should be understood as referring to a possibility of a feature in various implementations and not as prescribing an ability that is necessarily present in every implementation. For example, the phrase “X can perform Y” should be understood as indicating that, in various implementations, X has the potential to be configured to perform Y, and not as indicating that in every instance X must always be able to perform Y. It should be understood that, in various implementations, X can be unable to perform Y and remain within the scope of the present disclosure.
[0214] This written description uses examples to disclose the invention, including the best mode, and also to enable any person skilled in the art to practice the invention, including making and using any devices or systems and performing anyincorporated methods. The patentable scope of the invention is defined by the claims, and can include other examples that occur to those skilled in the art. Such other examples are intended to be within the scope of the claims if they include structural elements that do not differ from the literal language of the claims, or if they include equivalent structural elements with insubstantial differences from the literal language of the claims.
Claims
WHAT IS CLAIMED IS:
1. A computer-implemented method, comprising: training, in tandem, a prompt expansion model and an image generation model using a multi-reward reinforcement learning model by: processing, by the prompt expansion model, a training query and training context data to generate an expanded training query; generating, by the image generation model, a training set of image data based on the expanded training query; generating a set of reward scores for each image datum within the training set of image data using a set of reward models, wherein generating the set of reward scores comprises generating at least one reward score for each reward criterion within a plurality of reward criteria; and adjusting weights and biases associated with the plurality of reward criteria based on the set of reward scores using the multi-reward reinforcement learning model.
2. The computer-implemented method of claim 1, comprising: obtaining, by a computing system, input data comprising a user query and query context data; generating, using the trained prompt expansion model and the trained image generation model, image data based on the user query and the query context data by: processing, by the trained prompt expansion model, the user query and the query context data to generate an expanded user query; and generating, using the trained image generation model, the image data based on the expanded user query.
3. The computer-implemented method of claim 2, wherein generating the expanded user query comprises incorporating the plurality of reward criteria into the expanded user query.
4. The computer-implemented method of any preceding claim, wherein the plurality of reward criteria comprises one or more text-image alignment reward criteria.
5. The computer-implemented method of one of claims 2 and 4, wherein the one or more text-image alignment reward criteria comprises: a first text-image alignment reward criteria associated with the training query; and a second text-image alignment reward criteria associated with the expanded training query.
6. The computer-implemented method of any preceding claim, wherein the plurality of reward criteria comprises one or more image sentiment reward criteria.
7. The computer-implemented method of any preceding claim, wherein the plurality of reward criteria comprises one or more aesthetic reward criteria.
8. The computer-implemented method of any preceding claim, wherein the plurality of reward criteria comprises one or more human preference reward criteria.
9. The computer-implemented method of any preceding claim, wherein training the prompt expansion model and the image generation model comprises: selecting a subset of image data from the training set of image data as a function of the set of reward scores using a non-dominated sorting algorithm; and adjusting the weights and biases associated with the plurality of reward criteria based on the subset of image data using the multi-reward reinforcement learning model.
10. The computer-implemented method of any preceding claim, wherein the subset of image data comprises a pareto-optimal set associated with the set of reward scores.
11. The computer-implemented method of one of claims 9 and 10, wherein training the prompt expansion model and the image generation model comprises: determining, by a computing system, a policy gradient update as a function of the subset of image data; and adjusting the weights and biases associated with the plurality of reward criteria using the multi-reward reinforcement learning model as a function of the policy gradient update.
12. The computer-implemented method of claim 11, wherein determining the policy gradient update further comprises minimizing the reward scores associated with each reward criterion that is not represented within the subset of image data.
13. The computer-implemented method of one of claims 11 or 12, wherein determining the policy gradient update further comprises maximizing the reward scores associated with each reward criterion that is represented within the subset of image data.
14. The computer-implemented method of any of claims 11 to 13, wherein determining the policy gradient update comprises maximizing one or more text-image alignment reward criteria.
15. A computing system, comprising: one or more processors; and one or more transitory or non-transitory computer-readable media storing instructions that are executable to cause the one or more processors to perform operations, the operations comprising: obtaining, by the one or more processors, input data comprising a user query and query context data;training, in tandem, a prompt expansion model and an image generation model using a multi-reward reinforcement learning model, wherein training the prompt expansion model and the image generation model comprises: processing, by the prompt expansion model, a training query and training context data to generate an expanded training query; generating, using the image generation model, a training set of image data based on the expanded training query; generating a set of reward scores for each image datum within the training set of image data using a set of reward models, wherein generating the set of reward scores comprises generating at least one reward score for each reward criterion within a plurality of reward criteria; selecting a subset of image data from the training set of image data as a function of the set of reward scores using a non-dominated sorting algorithm; and minimizing weights and biases associated with each reward criterion that is not represented within the subset of image data using the multi-reward reinforcement learning model; and generating, using the trained prompt expansion model and the trained image generation model, image data based on the user query and the query context data.
16. The computing system of claim 15, wherein generating the image data further comprises: processing, by the trained prompt expansion model, the user query and the query context data to generate an expanded user query; and generating, by the trained image generation model, the image data based on the expanded user query.
17. The computing system of claim 16, wherein generating the expanded user query comprises incorporating the plurality of reward criteria into the expanded user query.
18. The computing system of any of claims 15 to 17, wherein training the prompt expansion model and the image generation model further comprises: determining, by the computing system, a policy gradient update as a function of the subset of image data; and adjusting the weights and biases associated with the plurality of reward criteria using the multi-reward reinforcement learning model as a function of the policy gradient update.
19. The computing system of claim 18, wherein determining the policy gradient update further comprises maximizing the reward scores associated with each reward criterion that is represented within the subset of image data.
20. A computer-implemented method, comprising: training, in tandem, a prompt expansion model (prompt expansion model) and an image generation model using a multi-reward reinforcement learning model by: processing, by the prompt expansion model, a training query and training context data to generate an expanded training query; generating, by the image generation model, a training set of image data based on the expanded training query; and training, in tandem, the prompt expansion model and the image generation model using the multi-reward reinforcement learning model based on the training set of image data.