System and method of a multi-reward reinforcement learning framework for generating images from text.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2024-11-13
- Publication Date
- 2026-08-13
AI Technical Summary
【0008】 本発明のこれら及び他の特徴、態様及び利点は、以下の説明及び添付の特許請求の範囲を参照して、よりよく理解されるようになる。本明細書に組み込まれ、本明細書の部分を構成する添付図面は、技術の実施形態を説明し、説明と併せて、技術の原理を説明するのに役立つ。
Smart Images

Figure 0007904996000001 
Figure 0007904996000002 
Figure 0007904996000003
Abstract
Description
[Technical Field]
[0001] This disclosure generally relates to machine learning processing, trained devices, and systems. More specifically, this disclosure relates to a multi-reward reinforcement learning framework for generating images from text.
[0002] Cross-reference of related applications This application claims priority to U.S. Provisional Patent Application No. 63 / 619,632, titled "MULTI-REWARD REINFORCEMENT LEARNING FRAMEWORK FOR TEXT-TO-IMAGE GENERATION," filed on 10 January 2024, the disclosures of which are incorporated by reference in their entirety. [Background technology]
[0003] A computer can receive input(s). A computer can execute instructions to process the input(s) to produce output(s) using a parameterized model. A computer can obtain feedback on its performance when producing output using the model. A computer can generate feedback by evaluating its performance. A computer can receive feedback from external sources. A computer can improve the model's performance by updating its parameters based on the feedback. In this way, a computer can iteratively "learn" to produce the desired output. The resulting model is often called a machine learning model. [Overview of the Initiative] [Means for solving the problem]
[0004] Aspects and advantages of the present invention as disclosed herein are either partially shown in the following description, or are obvious from the description, or can be acquired through the practice of the art.
[0005] According to one embodiment, a method is provided for a multi-reward reinforcement learning framework for text-to-image generation. The method involves training a prompt augmentation model and an image generation model in parallel using a multi-reward reinforcement learning model, the method comprising: processing training queries and training context data by the prompt augmentation model to generate augmented training queries; generating a training set of image data based on the augmented training queries by the image generation model; generating a set of reward scores for each image data in the training set of image data using a set of reward models, wherein generating the set of reward scores includes generating at least one reward score for each reward criterion in a plurality of reward criteria; and training by adjusting the weights and biases associated with the plurality of reward criteria based on the set of reward scores using the multi-reward reinforcement learning model.
[0006] According to another embodiment, a system for a multi-reward reinforcement learning framework for text-to-image generation is provided. The system comprises one or more processors and one or more temporary or non-temporary computer-readable media for storing executable instructions for causing one or more processors to perform an operation, wherein the operation includes one or more processors acquiring input data including user queries and query context data, and training a prompt augmentation model and an image generation model in parallel using a multi-reward reinforcement learning model, wherein training the prompt augmentation model and the image generation model involves the prompt augmentation model processing the training queries and training context data to generate an augmented training query, and the image generation model using the augmented training query to generate a tray of image data. The method includes generating a training set, generating a set of reward scores for each image data in the training set of image data using a set of reward models, wherein generating the set of reward scores includes generating at least one reward score for each reward criterion in a plurality of reward criteria, selecting a subset of image data from the training set of image data as a function of the set of reward scores using a non-dominant sorting algorithm, minimizing the weights and biases associated with each reward criterion not represented in the subset of image data using a multi-reward reinforcement learning model, and generating image data based on user queries and query context data using a trained prompt extension model and a trained image generation model.
[0007] According to one embodiment, a method is provided for a multi-reward reinforcement learning framework for text-to-image generation. The method includes: training a prompt augmentation model and an image generation model in parallel using a multi-reward reinforcement learning model; processing training queries and training context data by the prompt augmentation model to generate augmented training queries; generating a training set of image data by the image generation model based on the augmented training queries; and training the prompt augmentation model and the image generation model in parallel using the multi-reward reinforcement learning model based on the training set of image data.
[0008] These and other features, aspects and advantages of the present invention will be better understood by referring to the following description and the appended claims. The appended drawings incorporated herein and constituting parts thereof illustrate embodiments of the art and, together with the description, help to illustrate the principles of the art.
[0009] As those skilled in the art will understand, the full and implementable disclosure of the present invention, namely the best embodiments including methods for creating and using the systems and methods of the present invention, is described in the specification with reference to the accompanying drawings. The contents of the drawings are as follows: [Brief explanation of the drawing]
[0010] [Figure 1A] The diagram shows an exemplary framework of multi-reward reinforcement learning for a model that generates images from text, according to an exemplary embodiment of the present disclosure. [Figure 1B] The diagram shows an exemplary framework of multi-reward reinforcement learning for a model that generates images from text, according to an exemplary embodiment of the present disclosure. [Figure 2] A block diagram of an exemplary system for generating image data using an image generation model, according to an exemplary embodiment of the present disclosure, is shown. [Figure 3]This figure shows a flowchart of an exemplary embodiment of a multi-reward reinforcement learning method for a model that generates images from text, according to an exemplary embodiment of the present disclosure. [Figure 4] This flowchart illustrates an exemplary method for training a machine learning model according to exemplary embodiments of the aspects of this disclosure. [Figure 5] This is a block diagram of an exemplary processing flow for processing inputs(multiple) and generating outputs(multiple) using machine learning models(multiple) according to exemplary embodiments of the aspects of this disclosure. [Figure 6] This is a block diagram of an exemplary sequence processing model according to an exemplary embodiment of the aspects of the present disclosure. [Figure 7] This is an exemplary block diagram of the technique for filling in an exemplary input sequence for processing by a sequence processing model according to an exemplary embodiment of an aspect of the present disclosure. [Figure 8] This is a block diagram of an exemplary model development platform according to an exemplary embodiment of the aspects of the present disclosure. [Figure 9] This is an exemplary block diagram for training a machine learning model according to an exemplary embodiment of the aspects of the present disclosure. [Figure 10] This is a block diagram of an inference system for operating one or more machine learning models to perform inference, according to an exemplary embodiment of an aspect of the present disclosure. [Figure 11] This is a block diagram of an exemplary network computing system according to an exemplary embodiment of the aspects of the present disclosure. [Figure 12] This is a block diagram of an exemplary computing device according to an exemplary embodiment of the aspects of the present disclosure. [Figure 13] Block diagram showing an exemplary computing device according to an exemplary embodiment of the aspects of this disclosure. [Modes for carrying out the invention]
[0011] In general, this disclosure covers a multi-reward reinforcement learning framework for image generation models. More specifically, this disclosure provides a means for training image generation models and prompt extension models in parallel using a multi-reward reinforcement learning model. Training image generation models and prompt extension models in parallel enables more advanced fine-tuning of both the image generation models and prompt extension models.
[0012] This disclosure provides the use of batch-based Pareto optimal selection. The system can identify the optimal trade-off between several rewards during the training phase. The system can utilize a co-optimization technique for an image generation model and a prompt extension model to facilitate the generation of augmented text prompts. The co-optimization technique enables the provision of more advanced image generation while readjusting the image generation model's focus to the original text prompt. This can be achieved by rewarding both the original and augmented text prompts and optimizing them. This allows the image generation model to generate images that utilize the augmented prompt while retaining the original text prompt in the context window for image generation.
[0013] The present disclosure provides several technical advantages. By way of example, existing image generation models have faced several challenges in generating images that meet performance metrics. This is at least partially because the provided text prompts typically lack sufficient context for the image generation model to generate an image that meets image quality, resolution, sharpness, or other performance metrics. To address this problem, prompt extension models have been used to add additional details to the text prompts used by an image generation model for image generation. Existing methods of using prompt extension models have several technical challenges. Existing systems simply provide the extended prompt to the image generation model, but as described above, providing only the extended prompt can cause the image generation model to lose focus on the original prompt. This results in the generation of images that do not meet the performance metrics.
[0014] The present disclosure can prevent over-optimization and degradation of metrics resulting from simply aggregating and training multiple rewards by jointly optimizing multiple rewards when using a multi-reward reinforcement learning model. By simultaneously tuning an image generation model and a prompt extension model, the system can provide a higher degree of tuning in both the prompt extension model and the image generation model. This can provide a more advanced image generation model that outperforms baseline text-to-image generation techniques across several quality criteria metrics. The quality criteria metrics can include, for example, visual clarity, contrast, pixelation, resolution, aesthetic elements, human preferences, image sentiment, or alignment between text and image.
[0015] The improvements associated with the systems and methods discussed herein can be further understood with reference to the figures. Referring now to the drawings, the figures provide an exemplary arrangement of computing systems, model structures, and data flows for illustrative purposes only.
[0016] Referring to the diagram, Figure 1 shows an exemplary block diagram of a system for multi-reward reinforcement learning of an image generation model. System 100 includes user queries 102a-102b, query context data 104a-104b, prompt extension model 106, extended user queries 108a-108b, image generation model (multiple image generation models) 110, training set of image data 112, multi-reward reinforcement learning model 114, multiple reward criteria 116a-116d, weights and biases 118, set of reward scores 120a-120d, set of reward models 122a-122d, image data 124, etc.
[0017] The operation involves training the prompt augmentation model 106 (prompt augmentation model) and the image generation model 110 in parallel using the multi-reward reinforcement learning model 114. Training the prompt augmentation model 106 and the image generation model 110 in parallel means that both models are optimized or fine-tuned simultaneously. Training the models in parallel helps both models to work together without over-optimization and degradation of metrics. This is done to ensure that the text generated by the prompt augmentation model 106 directly matches how the image generation model 110 interprets the text. By training the prompt augmentation model 106 and the image generation model 110 in parallel, both models receive feedback from the multi-reward reinforcement learning model 114 regarding their respective actions in the environment.
[0018] Joint training of the prompt augmentation model 106 and the image generation model 110 may include iteratively updating parameters such as weights, biases, and coefficients based on feedback from multiple reward criteria 116a-116d. Both the prompt augmentation model 106 and the image generation model 110 can be optimized simultaneously to ensure they function consistently.
[0019] In one embodiment, updating parameters may involve using techniques such as gradient descent or least squares, allowing for adjustment of weights and biases based on received feedback (i.e., a set of reward scores 120a-120d). Parameter updates may be performed iteratively to ensure that the parameters of both agents (i.e., the image generation model 110 and the prompt extension model 106) are improved in response to rewards and errors that occurred during training.
[0020] The joint training of the prompt extension model 106 and the image generation model 110 may be repeated iteratively until the available training data is exhausted or a convergence test is achieved. When used in this disclosure, a convergence test can be used to evaluate whether the model has reached an acceptable level of accuracy. This can be achieved by evaluating consecutive error values. If the difference between these error values falls below a defined threshold, it indicates that the model has stabilized. Alternatively, the error values can be compared to a predetermined threshold to determine whether further training is appropriate.
[0021] Joint training of the prompt extension model 106 and the image generation model 110 includes obtaining input data including user queries 102a-102b. When used in this disclosure, user query 102a refers to a request for information related to image generation. User query 102a may include a description of the desired image data. User query 102a may include information related to theme, aesthetic elements, setting, background, lighting, color tone, emotion, art style, etc. User queries 102a can be received from users through various channels such as online forms, customer service emails, or live chat systems. Users can submit queries by typing text or by selecting an option to describe their request or problem. The descriptions within user queries 102a can range from general inquiries to very specific requests. For example, a general user query might simply be expressed as "car" or "landscape," while a more detailed user query 102a might include a description such as "a pink sports car parked in front of a botanical garden with beautiful pink flowers."
[0022] In one embodiment, user query 102a may include training query 102b. When used in this disclosure, training query 102b is an exemplary user query designed to provide input data for training a model. Training query 102b may be used as training data to help train or fine-tune the prompt augmentation model 106, the image generation model 110, or any other model discussed herein. Training query 102b may be used to help an agent understand the relationship between text input and visual output, or between text input and augmented text input. In some embodiments, training query 102b may be identical or substantially similar to user query 102a. This may mean that training query 102b includes an exemplary description of a desired image.
[0023] System 100 can receive query context data 104a related to user queries 102a-102b. Query context data 104a refers to contextual information related to the user or user queries 102a-102b. Query context data 104a provides additional insights into the circumstances surrounding user queries 102a-102b. Query context data 104a can be used to interpret user queries 102a-102b by considering the broader context in which they arose. Query context data 104a may include a variety of information derived from the user's past interactions with System 100.
[0024] Query context data 104a may include information about the user's previous interactions with system 100. This may include previous queries submitted by the user, links clicked, pages visited, user feedback related to previous images, previous user queries 102a-d, and other relevant actions or behaviors shown within system 100. For example, suppose a user submits a query about "animated raccoon wearing military equipment." Query context data 104a may include details of past feedback the user provided about previously generated images. By analyzing these previous interactions, the system can optimize the image data it generates to more accurately reflect the user's current context.
[0025] In one embodiment, query context data 104a includes training context data 104b. The training context data 104b used in this disclosure is exemplary context data used to train a model. The model may include, but is not limited to, a prompt extension model 106 and an image generation model 110. The training context data 104b is designed to provide the model with exemplary contextual insights associated with a user interaction. The training context data 104b may include historical versions of query context data 104a associated with a user interaction. The training context data 104b may include pairs of submitted queries and image data, feedback on generated images, and patterns in user behavior. The training context data 104b may be generated and optimized for use as training data.
[0026] Continuing to refer to Figures 1A and 1B, System 100 includes a prompt extension model (prompt extension model) 106. When used in this disclosure, the prompt extension model 106 is a model configured to generate extended user queries 108a and 108b. The prompt extension model 106 corresponds to the machine learning models in Figures 4 to 13 described below. The prompt extension model 106 operates by predicting and generating extended versions of user queries 102a and 102b based on query context data 104a and 104b. The prompt extension model 106 evaluates the query context data 104a and 104b to predict and generate extended versions of user queries that better fit the user's context data. For example, if the query context data 104a and 104b indicates that the user previously expressed a preference for lighter colors in a query, the prompt extension model 106 can extend user queries 102a and 102b to include descriptors that reflect this preference.
[0027] In one embodiment, the prompt augmentation model 106 may include a neural network architecture. The prompt augmentation model 106 may include multiple layers of interconnected nodes or neurons, which are configured to process data in a hierarchical manner. Each layer of the neural network may be responsible for a different aspect of the input, enabling the prompt augmentation model 106 to learn complex patterns and relationships within the data. User queries 102a-102b may be processed using these layers, in which case the neural network analyzes the text and identifies key components that can be augmented.
[0028] The nodes within the prompt extension model 106 may be organized within a structured network, such as a convolutional neural network, which includes an input layer, one or more hidden layers, and an output layer for each node. During training of the prompt extension model 106, connections between these nodes may be established by applying elements from the training dataset to the nodes. This may involve using a multi-reward reinforcement learning model 114 to adjust the connections and weights between nodes in adjacent layers based on one or more reward criteria 116a-116d. The adjustment between connections and weights between nodes may be done with the aim of optimizing the prompt extension model 106 to produce a desired output.
[0029] Continuing to refer to Figures 1A and 1B, jointly training the prompt augmentation model 106 and the image generation model 110 involves processing the training query 102b and training context data 104b to generate the augmented training query 108b. The augmented user queries 108a and 108b are user queries 102a and 102b augmented to include more detail and contextual information. These augmented user queries 108a and 108b can be used as input for the image generation model 110. By incorporating additional descriptive elements, the augmented user queries 108a and 108b more effectively capture the user's intent. This includes adding adjectives to specify colors, incorporating details about settings, or inferring emotional context based on the user's history.
[0030] The prompt augmentation model 106 can be configured to augment the text of user queries 102a-102b by providing a more detailed and contextually relevant description of the subject of the query. The prompt augmentation model 106 can be configured to elaborate on user queries 102a-102b by incorporating additional descriptive elements such as color, background details, emotional context, or details relevant to the subject of user queries 102a-102b. Furthermore, the prompt augmentation model 106 can reinforce user queries 102a-102b by incorporating multiple reward criteria 116a-116d into the augmented user queries 108a-108b. For example, if the original user queries 102a-102b emphasize a particular mood, the prompt augmentation model 106 can reinforce it by including descriptive adjectives and contextual elements that fit the image emotional reward criterion 116b.
[0031] The enhanced user queries 108a-108b can elaborate on user queries 102a-102b by adding additional details. In a non-restrictive example, if user queries 102a-102b state "Please provide images of fast cars," the enhanced user queries 108a-108b could include a supplementary query specifying "a stylish lime green sports car parked near a beach at sunset." This supplementary query could include additional details such as color, style, settings, and sentiment context.
[0032] The supplementary text for the enhanced user queries 108a-108b can be generated based on an analysis of user queries 102a-102b and query context data 104a-104b. The prompt enhancement model 106 can evaluate the query context data 104a-104b, such as insights from previous interactions, user preferences, and other contextual factors, to determine which details should be added to the query. This means that the enhanced user queries 108a-108b can vary significantly based on the user's history and interests. For example, a user who frequently requests natural landscapes might receive enhanced user queries 108a-108b that emphasize natural elements, such as "a quiet forest with a gentle river," instead of a more general description.
[0033] In one embodiment, the augmented user query 108a may include the augmented training query 108. The augmented training query 108b is the output of the prompt augmentation model 106 during the training phase. The augmented training query 108b is configured to serve as a reference point for training both the prompt augmentation model 106 and the image generation model 110. The augmented training query 108b is generated from the training query 102b and the training context data 104b during the joint training process of the prompt augmentation model 106 and the image generation model 110. The augmented training query 108b can serve multiple purposes in the training process. The augmented training query 108b can provide a clear target to the image generation model 110 during the learning phase, helping the model understand the types of visual elements it includes in its output. Additionally, the augmented training query 108b can enhance the importance of context and detail in query augmentation, allowing the model to learn from training examples when tuning its parameters. The enhanced training query 108b allows for the evaluation of both the prompt enhancement model 106 and the image generation model 110 through error analysis.
[0034] Continuing with reference to Figures 1A and 1B, the system 100 includes an image generation model 110. As used in this disclosure, an image generation model is a model designed to generate images based on text input and contextual data. The image generation model 110 is configured to interpret text data and generate image data. The generated image data may include, but is not limited to, a training set 112 of image data, a subset 202 of image data, a Pareto-optimal set 204, image data 124, and any other image data discussed herein. In some embodiments, the image generation model 110 may include a large-scale language model optimized for image generation. The LLM can use several natural language processing techniques to convert text prompts into image data. The LLM can be used to generate image data that reflects text input. The image generation model may include any of the following: machine learning models, natural language processing models, image processing models, large-scale language models, etc., which are described below herein in Figures 4 to 13.
[0035] Continuing to refer to Figures 1A and 1B, jointly training the prompt augmentation model 106 and the image generation model 110 includes generating a training set 112 of image data based on the augmented training query 108b. As used in this disclosure, the training set 112 of image data is a collection of images used to jointly train or optimize the prompt augmentation model 106 and the image generation model 110. The training set 112 of image data includes multiple images generated by the image generation model 110 based on the augmented training query 108b. This may include generating multiple iterations of image data using a single augmented training query 108b. Alternatively or additionally, the image generation model 110 may generate multiple iterations of image data using multiple augmented training queries 108b.
[0036] The generation of a training set 112 of image data can be achieved by employing techniques such as probabilistic sampling or controlled randomness. Probabilistic sampling introduces a certain level of variability within the image generation process. This variability can arise in the form of adjustments to parameters such as color saturation, lighting, viewpoint, and composition. This variability helps the image generation model 110 generate different iterations of image data based on identical or similar augmented training queries 108b.
[0037] In one embodiment, each image data in the image data training set 112 may be optimized to emphasize one or more reward criteria 116a-116d. Each image data produced by the image generation model 110 may be conceptualized as a result of a decision-making process that compares these multiple reward criteria 116a-116d. For example, if the augmented training query 108b includes details about images of "cityscapes," the image generation model 110 might, in one iteration, prioritize aesthetic quality associated with reward criterion 116c, resulting in vibrant colors and dramatic lighting. In another iteration, the image generation model 110 might be configured to focus on text-to-image relevance criterion 116a. By optimizing the parameters associated with reward criteria 116a-116d, the image generation model 110 can produce images that emphasize or optimize different reward criteria 116a-116d.
[0038] Optimizing the parameters associated with reward criteria 116a-116d allows the image generation model 110 to optimize the image generation process for a specific outcome. The ability to emphasize one or more reward criteria 116a-116d enables the image generation model 110 to create images that produce balanced representations that simultaneously satisfy multiple reward criteria 116a-116d. This may involve identifying the optimal trade-offs between competing reward criteria 116a-116d. Optimizing the image generation model 110 for one reward criterion may negatively impact a second reward criterion. In such cases, optimizing the image generation model 110 may involve employing strategies to find the most effective balance. By employing techniques such as multi-objective optimization, the image generation model 110 can iteratively adjust the parameters associated with each reward criterion 116a-116d. Multi-objective optimization techniques are described in more detail below.
[0039] Continuing to refer to Figures 1A and 1B, the operation may further include evaluating multiple image data according to multiple reward criteria 116a to 116d. Where used in this disclosure, the multiple reward criteria 116a to 116d are specific metrics or targets used in reinforcement learning to evaluate the performance of an agent, such as an image generation model 110 or a prompt extension model 106. These reward criteria 116a to 116d may be used to guide the training process by providing feedback on the quality of the generated outputs. The reward criteria 116a to 116d may be used to provide measurable targets that guide the agent's behavior. This feedback is used to help the image generation model 110 or the prompt extension model 106 learn which actions in the environment should be prioritized. The multiple reward criteria 116a to 116d may be used to evaluate the quality and effectiveness of the generated outputs (i.e., image data and extended user prompts). This evaluation is then used to guide the agent's behavior toward a desired outcome.
[0040] Each of the multiple reward criteria 116a–116d can function as a separate evaluation axis that can assess the generated image or augmented prompt. The agent receives a reward or penalty based on its actions. These rewards or penalties are used to influence the agent regarding the success or failure of those actions. By defining clear reward criteria 116a–116d for the image generation model 110 or prompt augmentation model 106, the model can quantitatively evaluate "good" images or augmented prompts and those that are "not very effective." This evaluation process helps the model identify which features to prioritize or modify in future iterations.
[0041] Exemplary embodiments of the multiple reward criteria 116a-116d may include the text-image alignment reward criterion 116a. When used in this disclosure, the text-image alignment reward criterion 116a refers to a set of evaluation metrics used to assess how well a generated image corresponds to a given text description or prompt. The text-image alignment reward criterion 116a may be used to evaluate the relationship between generated images and their associated text prompts. These text prompts may include user queries 102a-102b or extended user queries 108a-108b. The text-image alignment reward criterion 116a may include elements that take into account the overall context and tone of the text. The text-image alignment reward criterion 116a may evaluate text input and generated images in the context of query context data 104a-104b. For example, if the text prompt describes an animated frog dancing on a lotus leaf, the image should include identifiable features such as a frog, a dancing posture, and a lotus leaf. The text-image alignment reward standard 116a can be used to ensure that elements and actions described in text are visually present.
[0042] In one embodiment, the text-image alignment reward criterion 116a may include a first text-image alignment reward criterion 116a and a second text-image alignment reward criterion 116a. The first text-image alignment reward criterion 116a may be associated with the relationship between the generated image and the original user queries 102a-102b. In contrast, the second text-image alignment reward criterion 116a may be associated with the extended user queries 108a-108b. The second text-image alignment reward criterion 116a may be used to evaluate how effectively the generated image fits with the broader context of the extended user queries 108a-108b. The first and second text-image alignment reward criterion 116a may be used to ensure that the generated image captures the important elements and subjects explicitly mentioned in the descriptive text of the user queries 102a-102b and the extended user queries 108a-108b.
[0043] Exemplary embodiments of the multiple reward criteria 116a-116d may include one or more image sentiment reward criteria 116b. When used in this disclosure, image sentiment reward criteria 116b are used to evaluate the emotional tone or atmosphere that a generated image conveys in relation to a corresponding text description or prompt. Image sentiment reward criteria 116b may be used to ensure that the generated image data reflects the emotional context described in the text prompt. In some cases, image sentiment reward criteria 116b may be an evaluation of a specific visual role that contributes to the overall sentiment of the image. These cues may include elements such as color palette, facial expression, and composition. For example, bright colors and cheerful facial expressions are generally thought to be associated with feelings of joy, while muted tones and somber postures may suggest sadness or contemplation.
[0044] Exemplary embodiments of the multiple reward criteria 116a-116d may include one or more aesthetic reward criteria 116c. When used in this disclosure, aesthetic reward criterion 116c is a reward criterion used to evaluate the visual quality and artistic appeal of the generated image. Aesthetic reward criterion 116c may focus on various elements that contribute to the overall appearance and appeal of the image. These may include considerations such as composition, color harmony, lighting, detail and texture, and art style. Aesthetic reward criterion 116c may include an evaluation of the arrangement and balance of elements within the image, or the use of color within the image.
[0045] In one embodiment, the aesthetic reward criterion 116c can be derived from historical evaluations of the aesthetic quality of real-world images. System 100 can generate the aesthetic reward criterion 116c from multiple subjective evaluations of the aesthetic appeal of image data. These evaluations can be analyzed to identify patterns in how aesthetic quality is perceived. These annotated datasets can be used to train a machine learning model, such as a third reward model 122c or any other model disclosed herein. By applying techniques such as supervised learning, the machine learning model can gain a deeper understanding of elements that resonate with viewers, such as composition, color harmony, and texture. This trained model can then be used to generate the aesthetic reward criterion 116c.
[0046] Exemplary embodiments of the multiple reward criteria 116a-116d may include one or more human preference reward criteria 116d. Human preference reward criteria 116d may be used to evaluate generated images based on a dataset comprised of human feedback. This may be done to ensure that the output closely matches what individuals find appealing or desirable. Generating human preference reward criteria 116d may involve collecting user feedback through surveys, ratings, or other interactive methods. This feedback can then be aggregated to create a dataset that reflects the diverse preferences of the audience. This dataset may include multiple text-image pairs that reflect user preferences. Once this annotated dataset is established, a machine learning model, such as a fourth reward model 122d, can be applied to analyze the data and extract meaningful insights into user preferences. Using the fourth reward model 122d, key features (i.e., color scheme, composition, and subject matter) can be identified and scored based on the annotated dataset. In some embodiments, this annotated dataset can be used as training data by a fourth reward model 122d.
[0047] Continuing to refer to Figures 1A and 1B, the operation may further include using a set of reward models 122a to 122d to generate a set of reward scores 120a to 120d for each image data in the training set of image data 112. When used in this disclosure, a reward score is a quantitative measure that reflects the effectiveness of an action or decision made by an agent in achieving a specific goal, such as a set of reward criteria 116a to 116d. The agent is configured to optimize its behavior by receiving feedback based on the set of reward criteria 116a to 116d. This feedback is quantified by the set of reward scores 120a to 120d. Each reward score in the set of reward scores 120a to 120d represents a quantification of a distinct aspect of the agent's behavior in the environment. For example, a reward score may reflect how well a generated image fits a text prompt.
[0048] The set of reward scores 120a–120d can facilitate improvements in the agent's decision-making process within the environment. By receiving quantitative feedback on multiple reward criteria 116a–116d, the agent can learn to identify patterns and make informed choices that enhance their overall effectiveness. This encourages the agent to optimize its strategy based on its impact on the set of reward scores 120a–120d. This improvement encourages the agent to leverage successful strategies that yield higher scores.
[0049] In one embodiment, generating a set of reward scores 120a to 120d involves generating at least one distinct reward score 120 for each reward criterion within a plurality of reward criteria 116a to 116d. The first reward score 120a is associated with the text-image alignment reward criterion 116a. The first reward score 120a quantifies how effectively the generated image responds to the text prompt (i.e., user queries 102a to 102b or extended user queries 108a to 108b). The second reward score 120b corresponds to the image emotion reward criterion 116b. The second reward score 120b quantifies the emotional tone conveyed by the generated image and its relevance to the text input. The third reward score 120c is associated with the aesthetic reward criterion 116c. The third reward score 120c quantifies the visual quality and artistic appeal of the generated image. The fourth reward score, 120d, is associated with the human preference reward criterion, 116d. The fourth reward score, 120d, quantifies how well the generated image matches the user's preferences and sensibilities.
[0050] Each reward score within the set of reward scores 120a to 120d may be generated using the reward models of the set of reward models 122a to 122d. Each reward model in the set of reward models 122a to 122d may be optimized for a specific reward criterion 116a to 116d within a group of reward criteria 116a to 116d. This enables the models to effectively analyze and score images. Each reward model within the set of reward models 122a to 122d may be identical or substantially similar to the machine learning models, natural language processing models, image processing models, large-scale language models, etc., discussed below in this specification, as shown in Figures 4 to 13. For example, the first reward model 122a is associated with the text-image alignment reward criterion 116a. Thus, the first reward model 122a can use one or more natural language processing and image processing techniques to evaluate how well the generated image corresponds to the associated text prompt. The first reward model 122a can generate a first reward score 120a by analyzing semantic relationships and visual features.
[0051] Similarly, a second reward model 122b may be dedicated to an image emotion reward criterion 116b. The second reward model 122b can measure the emotional tone of an image using a machine learning algorithm trained on emotion analysis. The second reward model 122b can evaluate image features such as color, facial expression, and overall composition to quantify how well an image matches the desired emotion of a text prompt. A third reward model 122c may be associated with an aesthetic reward criterion 116c. The third reward model 122c can evaluate the aesthetic quality of image data using an image processing model. A fourth reward model 122d may be associated with a human preference reward criterion 116d. The fourth reward model 122d may be trained on a dataset consisting of user feedback and preference data to signal its scoring.
[0052] Continuing to refer to Figures 1A to 1B, the operation may further include using a multi-reward reinforcement learning model 114 to adjust weights and biases 118 associated with multiple reward criteria 116a to 116d based on a set of reward scores 122a to 122d. When used in this disclosure, the multi-reward reinforcement learning model 114 is a model designed to optimize the performance of an agent when producing an output based on multiple reward criteria. The multi-reward reinforcement learning model 114 is configured to adjust biases and weights 118 associated with multiple reward criteria 116a to 116d based on a set of reward scores 122a to 122d. The multi-reward reinforcement learning model 114 is configured to consider the performance of multiple agents across multiple reward criteria 116a to 116d. The multi-reward reinforcement learning model 114 may correspond to machine learning models described later in this specification in Figures 4 to 13. In one embodiment, the multi-reward reinforcement learning model 114 may be configured to identify the optimal trade-off between multiple reward criteria 116a to 116d. Based on these optimal trade-offs, the multi-reward reinforcement learning model 114 can optimize the biases and weights 118 associated with multiple reward criteria 116a to 116d.
[0053] The multi-reward reinforcement learning model 114 can be used to fine-tune the biases and weights 118 that determine the relative importance of each reward criterion. The weights can be used to indicate the degree of influence each reward criterion has in the final decision-making process. Bias is a term related to adjustments made to steer the output results in a desired direction. These adjustments to the weights and biases 118 can be indicated by a set of reward scores 120a-120d generated from the reward models 122a-122d. By analyzing the agent's performance based on these scores, the multi-reward reinforcement learning model 114 can identify which reward criteria 116a-116d are actively contributing to achieving the desired outcome and which reward criteria may require recalibration. For example, if the text-image alignment reward score 120a consistently yields high scores while the aesthetic reward score 120c remains low, the multi-reward reinforcement learning model 114 can maximize the weights assigned to the text-image alignment reward criterion 116a while minimizing the weights of the aesthetic reward criterion 116c. Adjusting the weights and biases 118 can involve an iterative learning process in which the multi-reward reinforcement learning model 114 evaluates the effect of these adjustments on subsequent sets of reward scores 120a to 120d. By using techniques such as gradient descent or other optimization algorithms, the multi-reward reinforcement learning model 114 can systematically improve its parameters to enhance its performance.
[0054] Continuing to refer to Figure 1B, once the prompt augmentation model 106 and the image generation model 110 have been trained, the system 100 can be used to generate image data 124 based on the augmented user query 108a. Generating image data 124 may include using the trained prompt augmentation model 106 to convert the user query 108a into an augmented user query 108a, as described above in this specification. The augmented user query 108a can then be used as input to the trained image generation model 110. The image generation model 110 is then configured to generate image data 124 along reward criteria 116a-116d. This can be done using any of the image generation processes discussed herein.
[0055] In some embodiments, the prompt extension model 106 and the image generation model 110 can be frozen once the training process is complete. During the training phase, both agents can continuously adjust their behavior based on feedback received from the reward score set 122a-122d. After the agents have sufficiently learned the optimal strategy, a convergence test can be used to evaluate whether the agents' performance has stabilized and the consistency of their outputs has reached an acceptable level. Once the agents pass this convergence test and it is indicated that learning has plateaued and further adjustments are unlikely to yield significant improvement, they can be frozen. Freezing an agent may involve fixing its weights and biases and effectively stopping further parameter updates.
[0056] Referring to Figure 2, a block diagram of an exemplary system for generating image data using an image generation model according to an exemplary embodiment of the present disclosure is shown. Figure 2 includes a subset of image data 202, a Pareto optimal set 204, a non-dominant sorting algorithm 206, a policy gradient update 208, and the like.
[0057] In one embodiment, training the prompt extension model 106 and the image generation model 110 involves using a non-dominant sorting algorithm 206 to select a subset 202 of image data from a training set 112 of image data as a function of a set of reward scores 120a to 120d. As used in this disclosure, the subset 202 of image data is a collection of images selected from the broader training set 112 of image data. The subset 202 of image data may be selected based on the performance of the image generation model 110, quantified by the set of reward scores 120a to 120d. The subset 202 of image data may be characterized by a representation of the best trade-off among the various reward scores 120a to 120d.
[0058] In one embodiment, a subset of image data 202 can be represented as a Pareto optimal set 204. As used in this disclosure, a Pareto optimal set 204 refers to a collection of solutions in a multi-objective optimization problem where there is no single solution that improves performance on one criterion while degrading performance on another. A Pareto optimal set 204 can represent a subset of image data 202 that displays the optimal trade-off between various reward scores 120a–120d. The image data within a Pareto optimal set 204 can be considered optimal because it reflects balanced performance across two or more reward criteria 116a–116d. In a non-limiting example, suppose a first image is good on a first reward criterion but weak on a second reward criterion. The second image exhibits the opposite strength, and both the first and second images can be selected to be part of a Pareto optimal set 204.
[0059] The optimal trade-off between various reward scores 120a-120d can refer to identifying the best appropriate outcome among multiple competing reward criteria 116a-116d. Each reward score from the set of reward scores 120a-120d can quantitatively represent the performance of the prompt extension model 106 and the image generation model 110 while generating image data. One of the key challenges in identifying and generating the optimal trade-off is managing the competing reward scores 120. If improving one reward score often leads to a decrease in another, the optimal trade-off between multiple competing reward criteria 116a-116d can be identified.
[0060] The non-superiority sorting algorithm 206 can be used to identify the optimal trade-off between image data based on their respective reward scores 120a to 120d. The non-superiority sorting algorithm 206 can be used to analyze how a change in one reward score affects the remaining reward scores. When used in this disclosure, the non-superiority sorting algorithm 206 is a technique used in multi-objective optimization methods to classify and rank solutions based on their performance across several reward criteria. The non-superiority sorting algorithm 206 is configured to identify one or more non-superior solutions within a training set 112 of image data. An image may be considered non-superior if there are no other images that perform better across two or more reward criteria. In a non-limiting example, if image A performs well on a first reward criterion and image B performs well on a second reward criterion, neither image outperforms the other because both have strengths in different areas.
[0061] The non-superiority sorting algorithm 206 can be configured to evaluate the relationships between images within a training set 112 of image data. This may include classifying images based on their overall performance, as reflected by a set of reward scores 120a–120d. Exceptional categories may include, but are not limited to, non-superiority categories, text-image matching categories, image sentiment categories, aesthetic categories, and human preference categories. Non-superiority categories may be used to represent the best-performing solutions across defined groups of reward criteria 116a–116d. Subsequent categories may be reserved for images that perform well across one or more reward criteria 116a–116d.
[0062] Once images are ranked or classified using the non-dominant sorting algorithm 206, the non-dominant sorting algorithm 206 can select a subset 202 of the image data based on these identified trade-offs. The subset 202 of the image data may be identified from non-dominant categories, which may represent the best trade-offs between the image data.
[0063] Continuing to refer to Figure 2, training the prompt augmentation model 106 and the image generation model 110 may include using a multi-reward reinforcement learning model 114 to adjust weights and biases 118 associated with multiple reward criteria 116a-116d based on a subset 202 of image data. The multi-reward reinforcement learning model 114 may be configured to identify specific weights and biases 118 that need to be adjusted to fine-tune the prompt augmentation model 106 and the image generation model 110 based on a subset 202 of image data or a Pareto optimal set 204. The weights and biases 118 may be fine-tuned based on identified optimal trade-offs between various reward scores. The multi-reward reinforcement learning model 114 may be configured to examine features of these images within the subset 202 of image data or the Pareto optimal set 204. Based on this examination, the multi-reward reinforcement learning model 114 may determine which features contribute to the most favorable outcome. In a non-restrictive example, if the images represented within the Pareto optimal set 204 exhibit a strong fit to the prompt and effective sentiment expression, the multi-reward reinforcement learning model 114 can adjust the weights associated with these aspects. This may be done with the aim of improving the image data generated over subsequent iterations.
[0064] Continuing to refer to Figure 2, training the prompt extension model 106 and the image generation model 110 may include determining a policy gradient update 208 as a function of a subset of image data 202. When used in this disclosure, the policy gradient update 208 is used to optimize the agent's policy. The policy gradient update 208 can be used to help the agent optimize its interactions with each other in the environment to maximize / minimize a set of reward scores 120a-120d. The policy gradient update 208 may be constructed based on a series of evaluations of the model's performance against a subset of image data 202 or a Pareto optimization set 204. The subset of image data 202 may be used as a reference point when determining how well the agent's actions fit the desired outcome. By analyzing the reward scores 116a-116d associated with the subset of image data 202, the multi-reward reinforcement learning model 114 can make gradient determinations that show how policy changes will affect future performance. In one embodiment, a multi-reward reinforcement learning model 114 can evaluate reward scores 116a to 116d of actions performed based on a subset 202 of image data. The gradient can be generated based on the contribution of each image data to the set of reward scores 116a to 116d.
[0065] Once the gradient is calculated, policy gradient update 208 is applied to adjust the policy parameters in a direction that maximizes / minimizes the identified reward criteria 116a-116d. These adjustments can be made to encourage the agent to explore and select actions that lead to higher reward outcomes. Furthermore, by optimizing the biases and weights associated with the reward criteria, the agent can learn to make optimal trade-offs between competing reward criteria. This process can be performed iteratively; that is, the policy gradient can be updated until the agent can pass the convergence test without issue.
[0066] In one embodiment, determining the policy gradient update 208 may involve minimizing reward scores 116a–116d associated with each reward criterion that is not adequately represented within the subset of image data 202 or the Pareto-optimized set 204. By identifying these undervalued reward scores, the multi-reward reinforcement learning model 114 can develop a targeted policy focused on improving performance across the unrepresented reward criteria. To achieve this minimization, the multi-reward reinforcement learning model 114 may use gradient descent to incorporate the gradients of the undervalued reward scores 120a–120d into the overall policy gradient update. By actively minimizing these scores, the agent learns to address performance gaps.
[0067] In an additional embodiment, determining the policy gradient update 208 may also include maximizing the reward scores 120a-120d associated with each reward criterion 116a-116d present in the subset 202 of image data. This can be done to encourage the agent to take advantage of its own merits by reinforcing positive behaviors that lead to favorable reward scores 120a-120d. To achieve this maximization, the multi-reward reinforcement learning model 114 can identify the reward scores 120a-120d corresponding to the reward criteria reflected in the subset 202 of image data. By analyzing successful image outputs, the multi-reward reinforcement learning model 114 can determine which features and actions yield higher rewards. This understanding allows the multi-reward reinforcement learning model 114 to adjust its policy parameters in a manner that prioritizes actions that lead to similar successful outcomes in future iterations. In some cases, the process of maximizing these reward scores 120a-120d may include applying techniques such as stochastic gradient ascent. By calculating the gradients of the derived reward scores 120a to 120d and integrating them into the policy gradient update, the multi-reward reinforcement learning model 114 can perform information-based adjustments that amplify the probability of generating a desired output.
[0068] Referring here to Figure 3, a flowchart of Method 300 for implementing multi-reward reinforcement learning for a text-to-image generative model is shown according to an exemplary embodiment of the present disclosure. Method 300 may be implemented by processing logic, which may include hardware (e.g., processing devices, circuits, dedicated logic, programmable logic, microcode, device hardware, integrated circuits, etc.), software (instructions for invoking or executing processing devices), or a combination thereof. In some embodiments, the method is implemented by a server computing system (e.g., server computing system 60) or a client computing system (e.g., client computing device 50). Although shown in a specific sequence or order, the order of processes may be changed unless otherwise specified. Thus, the illustrated embodiments should be understood as examples only, and the processes shown may be implemented in a different order, and some processes may be performed in parallel. Furthermore, in various embodiments, one or more processors may be omitted. Thus, not all processes are required in all embodiments. Other process flows are possible.
[0069] In operation 302, the processing logic can train a prompt augmentation model and an image generation model in parallel using a multi-reward reinforcement learning model. The joint training of the prompt augmentation model and the image generation model involves iteratively updating the model parameters based on feedback from multiple reward criteria. This feedback helps the image generation model and the prompt augmentation model to learn which actions in the environment to prioritize based on the quality and effectiveness of the generated outputs (i.e., image data and augmented user prompts).
[0070] In operation 304, training the prompt augmentation model and the image generation model involves processing the training query and training context data so that the prompt augmentation model can generate an augmented training query. The prompt augmentation model is configured to generate an augmented training query by augmenting the training query with additional contextual and explanatory elements. Both the training query and the training context data are exemplary representations of the user query and query context data used as training data for the prompt augmentation model. Similar to the training query and training context data, the augmented training query is an exemplary representation of an augmented user query that can be used to train the image generation model.
[0071] In operation 306, training the prompt augmentation model and the image generation model involves the image generation model generating a training set of image data based on the augmented training query. The image generation model is configured to generate a training set of image data from the text input of the augmented training query. The image data training set is a collection of generated images that are exemplary representations of the potential output of the image generation model. The image data training set may represent an optimal trade-off between two or more reward criteria.
[0072] In operation 308, training the prompt augmentation model and the image generation model includes generating a set of reward scores for each image in the training set of image data using a set of reward models, wherein generating a set of reward scores includes generating at least one reward score for each reward criterion in a plurality of reward criteria. In one embodiment, the plurality of reward criteria include one or more text-image alignment reward criteria, one or more image sentiment reward criteria, one or more aesthetic reward criteria, or one or more human preference reward criteria. In some cases, one or more text-image alignment reward criteria may include a first text-image alignment reward criterion associated with a training query and a second text-image alignment reward criterion associated with an augmented training query.
[0073] The quality of the training set of augmented training queries and image data can be represented by a set of reward scores. This set of reward scores quantifies how well the augmented training queries and image data training set conform to a set of reward criteria. These reward scores can be used as feedback to both the prompt augmentation model and the image generation model, respectively.
[0074] In operation 310, training the prompt augmentation model and the image generation model involves using a reward-reinforcement learning model to adjust the weights and biases associated with multiple reward criteria based on a set of reward scores. Both the prompt augmentation model and the image generation model can be fine-tuned based on the optimal trade-offs between the various reward scores present in the training set of image data. The multi-reward reinforcement learning model is used to examine the properties of the training set of image data to determine which features or actions contribute most to the favorable outcome. Based on these optimal trade-offs, the multi-reward reinforcement learning model can simultaneously adjust the weights and biases associated with the reward criteria of both the prompt augmentation model and the image generation model.
[0075] In some embodiments, the processing logic can acquire input data including user queries and query context data. In some embodiments, the processing logic can generate image data based on the user queries and query context data using a trained prompt augmentation model and a trained image generation model. For example, the processing logic can generate image data by processing the user queries and query context data with a trained prompt augmentation model to generate an augmented user query. For example, the processing logic can generate image data by generating image data based on an augmented user query using a trained image generation model. In some examples, the augmented user query may be generated by incorporating multiple reward criteria into the augmented user query.
[0076] In some embodiments, the processing logic may train a prompt extension model, and the image generation model may include selecting a subset of image data from a training set of image data as a function of a set of reward scores using a non-dominant sorting algorithm. In some embodiments, the processing logic may use a multi-reward reinforcement learning model to adjust weights and biases associated with multiple reward criteria based on the subset of image data. In some embodiments, the subset of image data may include a Pareto optimal set associated with a set of reward scores.
[0077] In some embodiments, training prompt augmentation models and image generation models may involve the computing system determining policy gradient updates as a function of a subset of image data. Training prompt augmentation models and image generation models may also involve using a multi-reward reinforcement learning model to adjust weights and biases associated with multiple reward criteria as a function of policy gradient updates. In some cases, determining policy gradient updates may involve minimizing the reward score associated with each reward criterion not represented within the subset of image data. Additionally, determining policy gradient updates may involve maximizing the reward score associated with each reward criterion represented within the subset of image data. Furthermore, policy gradient updates may involve maximizing one or more text-to-image alignment reward criteria.
[0078] Figure 4 is a flowchart illustrating a method 400 for training one or more machine learning models according to an aspect of the present disclosure. For example, exemplary machine learning models may include the prompt extension model 106, the image generation model 110, the multi-reward reinforcement learning model 114, and the set of reward models 122a-122d, which may be used in combination with any other models or algorithms referred to herein.
[0079] Method 400 can be implemented by processing logic, which may include hardware (e.g., processing devices, circuits, dedicated logic, programmable logic, microcode, device hardware, integrated circuits, etc.), software (instructions invoked or executed on the processing device), or a combination thereof. In some embodiments, the method is implemented by a server computing system (e.g., server computing system 60) or a client computing system (e.g., client computing device 50). Although shown in a specific sequence or order, the order of processes can be changed unless otherwise specified. Thus, the illustrated embodiments should be understood as examples only, and the processes shown may be implemented in a different order, and some processes may be performed in parallel. Furthermore, in various embodiments, one or more processors may be omitted. Therefore, not all processes are required in all embodiments. Other process flows are possible.
[0080] In operation 402, the processing logic can obtain a training instance. The training dataset may contain multiple training instances, which are divided among multiple datasets (e.g., a training dataset, a validation dataset, or a test dataset). The training instances may or may not be labeled. Although referred to as “training” instances in exemplary method 400, it should be understood that a training instance can be formed if runtime estimation trains the model using an evaluation of the model’s performance against that runtime instance (e.g., online training / learning). Exemplary data types for training instances and various tasks associated with them are described throughout this disclosure.
[0081] In operation 404, the processing logic may use one or more machine learning models to process the training instance and generate an output. The output may be obtained directly from one or more machine learning models, or it may be the final result of a chain of processing operations that include the outputs of one or more machine learning models.
[0082] In operation 406, the processing logic may receive an evaluation signal associated with an output. The evaluation signal may be obtained using a loss function. Various methods for determining the loss can be used, such as mean squared error, likelihood loss, cross-entropy loss, hinge loss, contrastive loss, or various other loss functions. The evaluation signal may be computed using known ground truth labels (e.g., supervised learning), predicted or estimated labels (e.g., semi-supervised or self-supervised learning), or unlabeled (e.g., unsupervised learning). The evaluation signal may be a reward (e.g., for reinforcement learning). The reward may be computed using a machine learning-prepared reward model configured to generate a reward based on the received output(s). The reward may be computed using feedback data describing human feedback on the output(s).
[0083] In operation 408, the processing logic may update the trained model using the evaluation signal. For example, the parameter values of the trained model(s) may be learned using various training or learning techniques, such as backpropagation, in some embodiments. For example, the evaluation signal may be backpropagated from the output (or another source of the evaluation signal) through the trained model(s) and used to update one or more parameters of the model(s) (for example, based on the gradient of the evaluation signal's parameter(s)). For example, a system(s) containing one or more trained models may be trained in an end-to-end manner. Gradient descent may be used to iteratively update parameters over several training iterations. In some embodiments, performing backpropagation may include performing time-truncate backpropagation. Exemplary method 400 includes implementing several generalization techniques (e.g., weight decay, dropout, etc.) to improve the generalization ability of the trained model.
[0084] In some embodiments, exemplary method 400 may be implemented to train a machine learning model from an initialized state to a fully trained state (for example, when the model exhibits a desired performance profile based on accuracy, precision, recall, etc.).
[0085] In some embodiments, exemplary method 400 may be implemented for specific stages of a training procedure. For example, in some embodiments, exemplary method 400 may be implemented to pre-train a machine learning model. Pre-training may include extensive training on potentially noisy data to achieve a broad base of performance levels across various tasks / data types.
[0086] In some embodiments, exemplary method 400 may be implemented to fine-tune a machine learning model. Fine-tuning may include, for example, smaller-scale training on higher-quality data (e.g., labeled, curated, etc.). Fine-tuning may affect all or some of the parameters of the machine learning model. For example, different parts of the machine learning model may be “frozen” during certain training stages. For example, parameters associated with embedding spaces may be “frozen” during fine-tuning (e.g., to retain information learned from a broader domain than present in the fine-tuning dataset). In some embodiments, exemplary method 400 uses an adapter module. The adapter may be a small trainable layer inserted between existing layers of the pre-trained model. During the fine-tuning process, the original parameters of the pre-trained model are typically frozen, and only the parameters of the adapter are updated.
[0087] In some embodiments, exemplary method 400 may be implemented to perform parameter-efficient fine-tuning methods such as layer-by-layer optimization of residuals (LoRA). LoRA can improve a pre-trained model with only minimal adjustments to the original parameters. This can be achieved by introducing trainable low-rank matrices that modify the behavior of the pre-trained weights without directly altering them. In some embodiments, only these auxiliary matrices are updated during fine-tuning, which significantly reduces the number of parameters to be trained.
[0088] Exemplary fine-tuning techniques include reinforcement learning, which can be based on user feedback regarding the model's performance during use.
[0089] Figure 5 is a block diagram of an exemplary processing flow for processing input 2(multiple) using machine learning model 1(multiple) to generate output 3(multiple).
[0090] The machine-trained model 1(v) may be one or more machine-trained models or model components, or may include them. An exemplary machine-trained model may include a neural network (e.g., a deep neural network). An exemplary machine-trained model may include a nonlinear model or a linear model. An exemplary machine-trained model may use other architectures instead of, or in addition to, a neural network. An exemplary machine-trained model may include decision tree-based models, support vector machines, hidden Markov models, Bayesian networks, linear regression models, k-means clustering models, etc.
[0091] The machine learning model 1(or more) may be any one or more of the machine learning models described above with respect to the aforementioned diagram, or may include them, or may be represented in any other way. For example, machine learning model 1(or more) may be any one or more of the prompt extension model 106, the image generation model 110, the multi-reward reinforcement learning model 114, and the set of reward models 122a-122d, or may be represented in any other way, and may be used in combination with any other models or algorithms referred to herein. Various features, variations, and embodiments of machine learning model 1(or more) are described below, and should be understood as such features, variations, and embodiments are understood in the same way as described with respect to the use of the prompt extension model 106, the image generation model 110, the multi-reward reinforcement learning model 114, the set of reward models 122a-122d, and any other models, algorithms, or any of the machine learning components described herein.
[0092] Exemplary neural networks may include feedforward neural networks, recurrent neural networks (RNNs) including long-short-term memory (LSTM) based recurrent neural networks, convolutional neural networks (CNNs), distributed models, generative adversarial networks, or other forms of neural networks. Exemplary neural networks may also be deep neural networks. Some exemplary machine-learned models may utilize attention mechanisms such as self-attention. For example, some exemplary machine-learned models may include multi-head self-attention models.
[0093] A machine learning model 1(or more) can include one or more instances of the same model configured to work on data from input 2(or more). A machine learning model 1(or more) can include multiple different models, or parts of multiple different models, configured to work on data from input 2(or more).
[0094] A machine learning model 1(or more) can include an ensemble of different models that can interact cooperatively to process data from input 2(or more). For example, a model ensemble may include multiple models with different attributes (e.g., different architectures, trained with different recipes, etc.). The ensemble can produce an overall output based on the individual outputs of the constituent models. In this way, for example, diverse constituent models can work together to provide system-level robustness by effectively aggregating the individual strengths and weaknesses of any given model. Each individual output can be combined into a weighted combination using a voting or routing mechanism, or a trained output layer (e.g., one or more feedforward layers or fully connected layers).
[0095] A machine learning model (one or more) can employ a Mixture-of-Experts structure. For example, Zhou et al., Mixture-of-Experts with Expert Choice Routing. AR X IV See 2202.09368v2 (October 14, 2022). For example, different parts of a model can learn different domains of expertise (explicitly or implicitly), and the paths through the model are selected by a learned routing mechanism that involves the appropriate experts for a given input (e.g., a specific part of the input, such as a token unit). For example, a feedforward network can be slightly activated for a given part of the input based on the output of a routing mechanism that processes that part of the input. In this way, for example, a group of activated weights can form an “expert” selected by the router. In each forward pass, only a subset of the total model weights may be involved, thereby reducing the amount of work performed to process a given input compared to a densely activated model. In this way, for example, the expressiveness and interpretability of a high-parameter model can be achieved with more computationally efficient forward passes.
[0096] Input 2(or more) can generally contain or represent various types of data. Input 2(or more) can contain one type or many different types of data. Output 3(or more) can be the same type(or more) of data as Input 2(or more) or different types of data. Output 3(or more) can contain one type or many different types of data.
[0097] Examples of data types for Input 2(or more) or Output 3(or more) include natural language text data, software code data (e.g., source code, object code, machine code, or other forms of computer-readable instructions or programming languages), machine code data (e.g., binary code, assembly code, or other forms of machine-readable instructions that can be directly executed by a computer's central processing unit), assembly code data (e.g., low-level programming languages that program processing units using symbolic representations of machine code instructions), genetic data or other chemical or biochemical data, image data, audio data, audiovisual data, haptic data, biometric data, medical data, financial data, statistical data, geographic data, astronomical data, historical data, sensor data in general (e.g., digital or analog values, such as voltage or other absolute or relative value measurements obtained from actual or artificial inputs such as sound sensors, light sensors, or displacement sensors), and similar data. The data may be raw or processed, and may be in any format or schema.
[0098] In multimodal input 2 or output 3, exemplary combinations of data types include image data and audio data, image data and natural language data, natural language data and software code data, image data and biometric data, sensor data and medical data, etc. Please understand that any combination of data types is possible in input 2 or output 3.
[0099] An exemplary input 2 may include one or more data types, such as the exemplary data types described above. An exemplary output 3 may include one or more data types, such as the exemplary data types described above. The data types of input 2 may be the same as or different from the data types of output 3. It should be understood that the exemplary data types described above are provided for illustrative purposes only. The data types contemplated within the scope of this disclosure are not limited to the examples given above.
[0100] Figure 6 is a block diagram of an exemplary implementation of an exemplary machine learning model configured to process a sequence of information. For example, an exemplary embodiment of machine learning model 1(or more) may include machine learning sequence processing model 4(or more). The exemplary system can pass input 2(or more) to sequence processing model 4(or more). Sequence processing model 4(or more) may include one or more machine learning components. Sequence processing model 4(or more) can process data from input 2(or more) to obtain input sequence 5. Input sequence 5 may include one or more input elements 5-1, 5-2, ..., 5-M, etc., obtained from input 2(or more). Sequence processing model 4 can process input sequence 5 using prediction layer 6(or more) to generate output sequence 7. Output sequence 7 may include one or more output elements 7-1, 7-2, ..., 7-N, etc., generated based on input sequence 5. The system can generate output 3(or more) based on output sequence 7.
[0101] A sequence processing model 4(or more) may include one or more machine learning model components configured to ingest, generate, or otherwise infer sequences of information. For example, some exemplary sequence processing models in the text domain are referred to as “Large-Scale Language Models,” or LLMs. See, for example, PaLM 2 Technical Report, GOOGLE, https: / / ai.google / static / documents / palm2techreport.pdf (nd). Other exemplary sequence processing models can operate in other domains, such as the image domain. For example, Dosovitskiy et al., An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale. AR X IV :2010.11929v2 (June 3, 2021), for example, regarding the audio domain, see Agostinelli et al., MusicLM: Generating Music From Text, AR X IV :2301.11325v1 (January 26, 2023). For the field of biochemistry, see, for example, Jumper et al., Highly accurate protein structure prediction with AlphaFold, 596 Nature 583 (August 26, 2021). Sequence processing model 4(or more) may process one or more types of data simultaneously. Sequence processing model 4(or more) may include relatively large models (e.g., many parameters, high computational cost), relatively small models (e.g., few parameters, low computational load), or both.
[0102] Generally, a sequence processing model 4(or more) can use data from input 2(or more) to obtain an input sequence 5. For example, input sequence 5 may include a representation of the data from input 2(or more) in a format understood by the sequence processing model 4(or more). One or more machine learning-trained components of the sequence processing model 4(or more) can take data from input 2(or more), parse that data into pieces compatible with the processing architecture of the sequence processing model 4(or more) (e.g., via "tokenization"), and project those pieces into the input space associated with the prediction layer 6(or more) (e.g., via "embedding").
[0103] Sequence processing models 4(or more) can take data from input 2(or more), parse the data into a sequence of elements, and obtain input sequence 5. For example, a portion of the input data from input 2(or more) can be broken down into pieces that collectively represent the content of that portion of the input data. These pieces can provide elements for a sequence.
[0104] Elements 5-1, 5-2, ..., 5-M can, in some cases, represent constituent units for capturing or representing meaningful information in a particular data domain. For example, an element can describe an "atomic" that spans one or more domains. For instance, in the case of a text input source, an element might correspond to a group of one or more word or subword components, such as one or more sets of characters.
[0105] For example, elements 5-1, 5-2, ..., 5-M may represent tokens obtained using a tokenizer. For example, a tokenizer may process a given portion of an input source and output a set of tokens representing that portion of the input source (e.g., corresponding to input elements 5-1, 5-2, ..., 5-M). Various tokenization techniques can be used. For example, a text input source(s) may be tokenized using byte-pair encoding (BPE) techniques. See, for example, Kudo et al., SentencePiece: A simple and language-independent subword tokenizer and detokenizer for Neural Text Processing, PROCEEDINGS OF THE 2018 CONFERENCE ON EMPIRICAL METHODS IN NATURAL LANGUAGE PROCESSING (System Demonstrations), pp. 66-71 (October 31 - November 4, 2018), https: / / aclanthology.org / D18-2012.pdf. Image-based input sources can be tokenized by extracting and serializing patches from the images.
[0106] In general, any data type can be serialized and processed to form input sequence 5. The elements 5-1, 5-2, ..., 5-M shown in Figure 6 may be tokens or their embedded representations.
[0107] The prediction layer 6(or more) can predict one or more output elements 7-1, 7-2, ..., 7-N based on the input elements. The prediction layer 6(or more) can include one or more trained machine learning model architectures, such as one or more trained parameter layers that manipulate and transform the input(or more) to extract higher-dimensional meaning and relationships between elements from the input elements 5-1, 5-2, ..., 5-M. In this way, for example, the exemplary prediction layer 6(or more) can predict new output(or more) taking into account the context provided by the input sequence 5.
[0108] Predictive layers 6(or more) can evaluate associations between parts of input sequence 5 and specific output elements. These associations can provide information for predicting the likelihood that a particular output will follow the input context. For example, consider the text fragment, "The carpenter's toolbox was small and heavy. It was full of ___." An example predictive layer 6(or more) can determine that "It" refers back to "tool box" by determining the relationships between each embedding. The example predictive layer 6(or more) can also link "It" to attributes of the toolbox such as "small" and "heavy". Based on these associations, predictive layer 6(or more) can assign, for example, the word "nails" to the word "sawdust" with a higher probability.
[0109] The converter is an exemplary architecture that can be used in the prediction layer 4(or more). For example, Vaswani et al., Attention Is All You Need. AR X IVSee 1706.03762v7 (August 2, 2023). A transformer is an example of a machine learning model architecture that uses an attention mechanism to compute associations between items in a context window. A context window can include an input sequence 5 and a sequence containing, optionally, one or more output elements 7-1, 7-2, ..., 7-N. A transformer block can include one or more attention layers and one or more post-attention layers (e.g., feedforward layers such as multilayer perceptrons).
[0110] The prediction layer 6(or more) can include, in addition to or instead of, a converter-based architecture, other machine learning model architectures. For example, not only convolutional neural networks (CNNs), but also recurrent neural networks (RNNs) and long-term short-term memory (LSTM) models can be used. In general, the prediction layer 6(or more) can leverage various types of artificial neural networks that can understand or generate sequences of information.
[0111] Output sequence 7 may contain the same or different data types as input sequence 5, or may represent the same or different data types as input sequence 5 in other ways. For example, input sequence 5 may represent text data, and output sequence 7 may represent text data. Input sequence 5 may represent image, audio, or audiovisual data, and output sequence 7 may represent text data (e.g., describing image, audio, or audiovisual data). It will be understood that any other intervening model component of the prediction layer 6(or more) and the sequence processing model 4(or more) may be configured to receive various data types of input sequence 5(or more) and output various data types of output sequence 7(or more).
[0112] Output sequence 7 can have various relationships with input sequence 5. Output sequence 7 can be a continuation of input sequence 5. Output sequence 7 can be a complementary relationship to input sequence 5. Output sequence 7 can translate, transform, extend, or otherwise modify input sequence 5. Output sequence 7 can respond to input sequence 5 by answering, evaluating, confirming, or otherwise. Output sequence 7 can implement (or describe instructions for implementing) instructions provided via input sequence 5.
[0113] Output sequence 7 can be generated autoregressively. For example, in some applications, the output of one or more prediction layers 6 (if any) is passed through one or more output layers (e.g., a softmax layer) to obtain a probability distribution over an output vocabulary (e.g., a text or symbolic vocabulary) conditioned on a set of input elements within a context window. In this way, for example, output sequence 7 can be generated autoregressively by sampling the most likely next element, adding that element to the context window, regenerating the probability distribution based on the updated context window, and then sampling the most likely next element, and so on.
[0114] Output sequence 7 can be generated non - autoregressively. For example, multiple output elements of output sequence 7 can be predicted together without explicit sequential conditioning on each other. For example, see Saharia et al., Non - Autoregressive Machine Translation with Latent Alignments, AR X IV :2004.07437v3 (November 16, 2020).
[0115] Output sequence 7 can contain one or more parts or elements. In an exemplary content generation configuration, output sequence 7 can contain multiple elements (e.g., sentences of text, discretized waveform values, computer code, etc.) corresponding to multiple parts of the generated output sequence. In an exemplary classification configuration, output sequence 7 can contain a single element associated with the classification output. For example, the output "Vocabulary" can contain a set of classes to which the input sequence is classified. For example, a visual transducer block can pass latent state information to a multilayer perceptron that outputs the most likely class values associated with the input image.
[0116] Figure 7 is a block diagram of an exemplary technique for filling an exemplary input sequence 8. The input sequence 8 can include various functional elements that constitute part of the model infrastructure, such as element 8-0, obtained from a task indicator 9 that signals that a particular task is being performed (for example, to help fit the performance of the model(s) to that particular task) to any model(s) that processes the input sequence 8. The input sequence 8 can include various data elements from different data modalities. For example, input modality 10-1 can include one modality of data. The data-to-sequence model 11-1 processes the data from input modality 10-1 and projects that data into a format compatible with input sequence 8 (for example, one or more vectors with dimensions set according to the dimensions of input sequence 8) to obtain elements 8-1, 8-2, and 8-3. Other input modalities 10-2 can include different data modalities. The data-to-sequence model 11-2 can project data from input modality 10-2 into a format compatible with input sequence 8, thereby obtaining elements 8-4, 8-5, and 8-6. Other input modalities 10-3 may contain different data modalities. The data-to-sequence model 11-3 can project data from input modality 10-3 into a format compatible with input sequence 8, thereby obtaining elements 8-7, 8-8, and 8-9.
[0117] Input sequence 8 may be the same as or different from input sequence 5. Input sequence 8 can be a multimodal input sequence containing elements that represent data from different modalities using a common dimensional representation. For example, the embedding space may have P dimensions. Input sequence 8 can be configured to contain multiple elements having P dimensions. In this way, for example, exemplary embodiments can facilitate information extraction and inference across diverse data modalities by projecting data onto elements in the same embedding space for comparison, combination, or other computations between them.
[0118] For example, elements 8-0, ..., 8-9 can represent specific locations within a multidimensional embedding space. Some elements can be mapped to a discrete set of locations within the embedding space. For instance, elements corresponding to individual members of a given vocabulary of tokens can be mapped to discrete locations within the embedding space associated with those tokens. Other elements can be contiguously distributed throughout the embedding space. For example, some data types can be decomposed into contiguously defined parts (e.g., image patches) that can be described using contiguously distributed locations within the embedding space.
[0119] In some embodiments, the expressiveness of the embedding space is not limited to the meaning associated with any particular set of tokens or other constituent units. For example, a continuous embedding space can encode a spectrum of higher-order information. Individual pieces of information (e.g., tokens) can be mapped to specific points within that space; for example, the token for the word "dog" can be projected onto an embedding value that points to a specific location within the embedding space associated with dog-related information. Similarly, an image patch of a dog on grass may also be projected onto the embedding space. In some embodiments, the projection of the dog image may be similar to the projection of the word "dog," while also having similarity to the projection of the word "grass," and at the same time, both may be different. In some embodiments, the projection of the image patch cannot exactly match any single projection of either of the single words. In some embodiments, the projection of the image patch can match a combination of projections of the words "dog" and "grass." In this way, for example, a higher-order embedding space can encode information that is independent of the data modality in which the information is represented.
[0120] The task indicator 9 may include a model or model component configured to identify the task being performed and inject input values represented by elements 8-0, which signal which task is being performed, into the input sequence 8. For example, the input values may be provided as a data type associated with an input modality (e.g., the input values may be text task labels embedded in the input along with other text data, or the input values may be pixel-based representations of tasks embedded in the input along with other image data) and may be projected along with that input modality. The input values may be provided as a data type that is different from, or at least independent of, other inputs(s). For example, the input values represented by elements 8-0 may be learned in a contiguous embedding space.
[0121] Input modalities 10⁻¹, 10⁻², and 10⁻³ can be associated with various different data types (for example, as described above with respect to input 2(or more) and output 3(or more)).
[0122] The data-to-sequence models 11-1, 11-2, and 11-3 may be the same or different from each other. The data-to-sequence models 11-1, 11-2, and 11-3 can be adapted to their respective input modalities 10-1, 10-2, and 10-3. For example, a text data-to-sequence model can subdivide a portion of the input text and project the subdivided portion onto an element(s) in the input sequence 8 (e.g., elements 8-1, 8-2, 8-3, etc.). An image data-to-sequence model can subdivide an input image and project the subdivided portion onto an element(s) in the input sequence 8 (e.g., elements 8-4, 8-5, 8-6, etc.). A data-to-sequence model for any data type can subdivide the input of that data type and project the subdivided portion onto an element(s) in the input sequence 8 (e.g., elements 8-7, 8-8, 8-9, etc.).
[0123] The data-to-sequence models 11-1, 11-2, and 11-3 can form part of a machine learning-based sequence processing model 4(or more). The data-to-sequence models 11-1, 11-2, and 11-3 can be trained in conjunction with the machine learning-based sequence processing model 4(or more) or independently of it. The data-to-sequence models 11-1, 11-2, and 11-3 can be trained end-to-end using the machine learning-based sequence processing model 4(or more).
[0124] Figure 8 is a block diagram of an exemplary model development platform 12 that can facilitate the creation, adaptation, and improvement of exemplary machine learning models (e.g., machine learning model 1(or more), sequence processing model 4(or more)). The model development platform 12 can provide several different toolkits that developer systems can employ in developing new or adapted machine learning models.
[0125] The model development platform 12 may provide one or more model libraries 13 containing the constituent units of a new model. The model library 13 may include one or more pre-trained foundational models 13-1 that can provide a backbone of processing power across various tasks. The model library 13 may also include one or more pre-trained expert models 13-2 that can focus on performance in a specific area of expertise. The model library 13 may include various model primitives 13-3 that can provide a low-level architecture or (optionally pre-trained) components that can be assembled into various configurations as needed. The model primitives 13-3 may include a library of pre-trained adapters or LoRA modules that can adapt the baseline foundational model so that its output matches a desired performance profile, or enhances the model's functionality (e.g., adapting to different input modalities).
[0126] The model development platform 12 can receive a selection of various model components 14. The model development platform 12 can pass the selected model components 14 to the workbench 15, which then incorporates the selected model components 14 into the development model 16.
[0127] The workbench 15 can facilitate further refinement and adaptation of the development model 16 by leveraging several different toolkits integrated with the model development platform 12. For example, the workbench 15 can facilitate the alignment of the development model 16 to a desired performance profile for various tasks by using the model alignment toolkit 17.
[0128] The model alignment toolkit 17 can provide several tools for generating outputs that match the desired operating characteristics of the development model 16. Alignment can include improving the accuracy, precision, recall, etc., of the model output. Alignment can include ensuring that the model output adheres to a desired output style, schema, or other desirable characteristics. Alignment can be general or domain-specific. For example, a pre-trained foundational model 13-1 can start with an initial level of performance across multiple domains. Alignment of the pre-trained foundational model 13-1 can include improving performance in information or tasks in a specific domain (for example, at the expense of performance in other domains of information or tasks).
[0129] The model alignment toolkit 17 can integrate one or more datasets 17-1(or more) to align a development model 16. A curated dataset 17-1(or more) may include labeled or unlabeled training data. A dataset 17-1(or more) can be obtained from public domain datasets. Alternatively, a dataset 17-1(or more) may be obtained from private datasets associated with one or more developer systems(or more) for aligning a bespoke machine learning model(or more) customized for a private use case.
[0130] The pre-training pipeline 17-2 may include a machine learning model training workflow configured to update the development model 16 across a large and potentially noisy dataset. For example, pre-training may utilize unsupervised learning techniques (e.g., denoising) to process a large number of training instances to update model parameters from an initialized state and achieve desired baseline performance. The pre-training pipeline 17-2 may perform pre-training by leveraging an unlabeled dataset within dataset 17-1(or more). The workbench 15 may implement the pre-training pipeline 17-2 to pre-train the development model 16.
[0131] The fine-tuning pipeline 17-3 may include a machine learning model training workflow configured to improve the model parameters of the development model 16 using higher quality data. The fine-tuning pipeline 17-3 can update the development model 16 by performing supervised training using labeled datasets within dataset 17-1(or more). The fine-tuning pipeline 17-3 can update the development model 16 by performing reinforcement learning using reward signals from user feedback signals. The workbench 15 can implement the fine-tuning pipeline 17-3 for fine-tuning the development model 16.
[0132] The prompt library 17-4 may include a set of inputs configured to elicit behavior that meets a desired performance criterion. The prompt library 17-4 may include fusion shot prompts (e.g., inputs that provide an example of a desired model output to prepend to a desired runtime query), thought chain prompts (e.g., inputs that provide step-by-step reasoning in an example to facilitate thorough reasoning by the model), and so on.
[0133] Exemplary prompts can be obtained from the available repositories of the prompt library 17-4. Exemplary prompts can be provided by one or more developer systems using the workbench 15.
[0134] In some embodiments, pre-trained or fine-tuned models can perform adequately even without exemplary inputs. For example, zero-shot prompts may include inputs that lack examples. Zero-shot prompts may be located within or outside the training region(s).
[0135] The prompt library 17-4 may include one or more prompt engineering tools. The prompt engineering tools can provide workflows for obtaining or learning optimized prompt values. The prompt engineering tools can facilitate direct learning of prompt values (e.g., input element values) based on one or more training iterations. The workbench 15 can implement the prompt engineering tools in the development model 16.
[0136] The prompt library 17-4 can include a pipeline for prompt generation. For example, inputs can be generated using the development model 16 itself or other machine learning models. In this way, for example, the first model can process information about a task and output an input that the second model processes to carry out the steps of the task. The second model may be the same as or different from the first model. The workbench 15 can implement the prompt generation pipeline in the development model 16.
[0137] The prompt library 17-4 may include a pipeline for context injection. For example, the performance of the development model 16 for a particular task may improve when additional context is provided to perform that task. The prompt library 17-4 may include software components configured to identify a desired context, retrieve the context from an external source (e.g., a database, a sensor, etc.), and add the context to the input prompt. The workbench 15 can implement a pipeline for context injection in the development model 16.
[0138] While the various training examples described herein with respect to the model development platform 12 refer to "pre-training" and "fine-tuning," it should be understood that the model alignment toolkit 17 can generally support a variety of training techniques adapted for training various machine learning models. Exemplary training techniques can correspond to the aforementioned exemplary training methods 400.
[0139] The model development platform 12 may include a model plugin toolkit 18. The model plugin toolkit 18 may include a variety of tools configured to extend the functionality of machine learning models by integrating them with other systems, devices, and software components. For example, machine learning models can use tools as needed to improve the quality of their performance. For instance, deterministic tasks can be offloaded to dedicated tools instead of probabilistically performing tasks with an increased risk of error. For example, instead of autoregressively predicting the solution to a system of linear equations, a machine learning model can recognize which tool to call to obtain the solution and pass the system of equations to the appropriate tool. The tool can be a conventional system of equations solver that can act deterministically to solve the system of equations. The output of the tool may be returned in response to the original query. In this way, by using tools, some exemplary models can focus on the advantages of machine learning models, such as understanding the intent of unstructured requests for tasks, while simultaneously extending the model's performance by offloading specific tasks to more focused tools that mechanically apply deterministic algorithms to more specific problems.
[0140] The Model Plugin Toolkit 18 may include a validation tool 18-1. The validation tool 18-1 may include a tool that can analyze and verify the output(s) of a machine learning model. The validation tool 18-1 may include a designed heuristic that establishes specific thresholds to apply to the model output. For example, the validation tool 18-1 may justify the output of a machine learning model to a structured data source (e.g., to mitigate hallucination).
[0141] The model plugin toolkit 18 may include a tool package 18-2 for implementing one or more tools, which may include scripts or other executable code that can be run with the development model 16. The tool package 18-2 may include one or more inputs configured to cause a machine learning model(s) to implement a tool (e.g., a fusion shot prompt that guides the model to output a tool call in the appropriate syntax). The tool package 18-2 may include, for example, fine-tuning training data for training the model to use the tool.
[0142] The model plug-in toolkit 18 may include an interface for calling external application programming interfaces (APIs) 18-3. For example, in addition to directly implementing tool calls and tool code in the development model 16, or instead, the development model 16 may be adapted to output instructions that initiate API calls to send or retrieve data via an external system.
[0143] The model plugin toolkit 18 can integrate with the prompt library 17-4 to build a catalog of available tools for use in the development model 16. For example, the model can receive a catalog of available tools as input, and the model can select a tool from the available tools and generate an output that initiates a tool call to use that tool.
[0144] The model development platform 12 may include a computational optimization toolkit 19 for optimizing the computational performance of the development model 16. For example, tools for model compression 19-1 may allow the size of the development model 16 to be reduced while maintaining a desired level of performance. For example, model compression 19-1 may include quantization workflows, weight pruning, and sparsification techniques. Tools for hardware acceleration 19-2 may facilitate the configuration of model storage and executables to work optimally with different hardware resources. For example, hardware acceleration 19-2 may include tools for optimally sharing models for distributed processing across multiple processing units due to increased bandwidth, reduced integrated memory requirements, etc. Tools for knowledge distillation 19-3 may provide training for lighter models based on the knowledge encoded in the development model 16. For example, the development model 16 can be a high-performance, large-scale machine learning model optimized using the model development platform 12. To obtain a lightweight model for execution in resource-constrained environments, the smaller model may be a "student model," which learns by mimicking the development model 16 as a "teacher model." In this way, for example, the investment made in training the parameters and configuration of the development model 16 can be efficiently transferred to a smaller model for more efficient inference.
[0145] Workbench 15 may implement one or more of the toolkits implemented in the model development platform 12, or it may not implement any of them. Workbench 15 can output an output model 20 based on the development model 16. The output model 20 may be a deployed version of the development model 16. The output model 20 may be a development or training checkpoint of the development model 16. The output model 20 may be a knowledge distillation, compression, or other optimized version of the development model 16.
[0146] Figure 9 is a block diagram of an exemplary training flow for training a machine learning-prepared development model 16. One or more parts of the exemplary training flow may be implemented by a computing system including one or more computing devices, such as the computing system described with reference to other figures. Each part of the exemplary training flow may be carried out by any one (or any combination) of one or more computing devices. Furthermore, one or more parts of the exemplary training flow may be implemented on the hardware components of the device described herein, for example, to train one or more systems or models. Figure 9 shows the elements carried out in a particular order for illustrative and explanatory purposes. Those skilled in the art will understand, by using the disclosures provided herein, that any element of the methods described herein can be adapted, rearranged, extended, omitted, combined, or modified in various ways without departing from the scope of this disclosure. Figure 9 is described for illustrative purposes with reference to elements / terms described in relation to other systems and drawings, and is not intended to limit it. One or more parts of the exemplary training flow may be carried out additionally or alternatively by other systems.
[0147] First, the development model 16 can maintain its initial state as an initialized model 21. The development model 16 can be initialized with weight values. The initial weight values can be random or based on an initialization schema. The initial weight values can be based in advance on pre-training on the same or different models.
[0148] The initialized model 21 can undergo pre-training in the pre-training stage 22. The pre-training stage 22 can be implemented using one or more pre-training pipelines 17-2 on data from dataset 17-1(or more). For example, if the initialized model 21 is already pre-trained (e.g., the development model 16 is a pre-trained base model or expert model, or is based on a pre-trained base model or expert model), then pre-training can be omitted.
[0149] The pre-trained model 23 can then become a new version of the development model 16, which can persist as development model 16 or as a new development model. The pre-trained model 23 can be the initial state if the development model 16 is already pre-trained. The pre-trained model 23 can be fine-tuned in the fine-tuning stage 24. The fine-tuning stage 24 can be implemented using one or more fine-tuning pipelines 17-3 on data from dataset 17-1(or more). Fine-tuning can be omitted, for example, if the performance of the pre-trained model is sufficient, if the model is already fine-tuned, or if other tuning methods are preferred.
[0150] The fine-tuned model 29 can then become a new version of the development model 16, which can persist as development model 16 or as a new development model. If the development model 16 is already fine-tuned, the fine-tuned model 29 can be returned to its initial state. The fine-tuned model 29 can be improved by user feedback 26. For example, improvements based on user feedback 26 may include reinforcement learning based on human feedback from human users of the fine-tuned model 25. Since reinforcement learning can be a form of fine-tuning, it should be understood that the fine-tuning stage 24 can include a stage for improvement using user feedback 26. Improvements based on user feedback 26 may result in the creation of an improved model 27. The improved model 27 can be output to a downstream system 28(or more) for deployment or further development.
[0151] In some embodiments, computational optimization operations may be applied before, during, or after each stage. For example, an initialized model 21 may undergo computational optimization 29-1 (e.g., using the computational optimization toolkit 19) before the pre-training stage 22. A pre-trained model 23 may undergo computational optimization 29-2 (e.g., using the computational optimization toolkit 19) before the fine-tuning stage 24. A fine-tuned model 25 may undergo computational optimization 29-3 (e.g., using the computational optimization toolkit 19) before being improved by user feedback 26. An improved model 27 may undergo computational optimization 29-4 (e.g., using the computational optimization toolkit 19) before being output to a downstream system 28(or more). Computational optimizations 29-1, ..., 29-4(or more) may all be the same, all be different, or include at least some different optimization techniques.
[0152] Figure 10 is a block diagram of an inference system for operating one or more machine learning models 1(or more) to perform inference (e.g., training and deployment). A model host 31 can receive one or more machine learning models 1(or more). A model host 31 can host one or more model instances 31-1(or more), which may be one or more instances of one or more models. A model host 31 can host model instances 31-1(or more) using the available computing resources 31-2 associated with the model host 31.
[0153] The model host 31 can perform inference on behalf of one or more clients 32. Clients 32 can send input requests 33 to the model host 31. Using the input requests 33, the model host 31 can obtain inputs 2 for input to a machine-trained model 1. The machine-trained model 1 can process inputs 2 to produce outputs 3. Using outputs 3, the model host 31 can return an output payload 34 to respond to the input requests 33 from the clients 32. The output payload 34 may contain or be based on outputs 3.
[0154] The model host 31 can extend its inference tasks by utilizing various other resources and tools. For example, the model host 31 can communicate with a tool interface 35 to facilitate the use of tools by model instances 31-1(or more). The tool interface 35 can include local or remote APIs. The tool interface 35 can include integrated scripts or other software functions. The model host 31 can involve online learning interfaces 36(or more) to facilitate the continuous improvement of machine-learned models 1(or more). For example, the online learning interfaces 36(or more) can be used within a reinforcement learning loop to obtain user feedback on the inferences provided by the model host 31. The model host 31 can access runtime data sources 37(or more) to extend input 2(or more) with additional contextual information. For example, the runtime data sources 37(or more) can include a knowledge graph 37-1 that facilitates the retrieval of structured information for information associated with input requests 33(or more) (e.g., search engine services). The runtime data source 37(or more) may include a public or private, external, or local database 37-2(or more) that can store information associated with input requests 33(or more) to extend input 2(or more). The runtime data source 37(or more) may include account data 37-3, which is retrieved in association with a user account corresponding to a client 32, and can be used to customize the behavior of the model host 31 accordingly.
[0155] The model host 31 may be implemented by one or more computing devices or systems. The client 2(s) may be implemented by one or more computing devices or systems, which may include computing devices or systems shared with the model host 31.
[0156] For example, the model host 31 can operate on a server system that provides machine learning services to client devices (multiple) running client 32 (or more) (for example, via a local or wide area network). Client devices (multiple) can be end-user devices used by individuals. Client devices (multiple) can be server systems that run client 32 (or more) and provide various functions as services to downstream end-user devices.
[0157] In some embodiments, the model host 31 can operate on the same device or system as the client 32(or more). The model host 31 can be a machine learning service that runs on a device and provides machine learning capabilities to one or more applications running on a client device, which may include application implementation clients 32(or more). The model host 31 can be part of the same application as the client 32(or more). For example, the model host 31 may be a subroutine or method implemented by one part of the application, and the client 32(or more) may be other subroutines or methods that engage with the model host 31 to perform inference functions within the application. It should be understood that the model host 31 and the client 32(or more) can have a variety of different configurations.
[0158] A model instance 31-1(or more) can contain one or more machine-learned models available for performing inference. A model instance 31-1(or more) can contain weights or other model components that are stored in persistent storage, temporarily cached, or loaded into fast memory. A model instance 31-1(or more) can contain multiple instances(or more) of the same model (for example, to run more requests in parallel on the same model). A model instance 31-1(or more) can contain instances(or more) of different models(or more). A model instance 31-1(or more) can contain cached intermediate states of active or inactive models(or more) used to speed up the estimation of those models. For example, an inference session with a particular model can generate a significant amount of computational results that can be reused for future inference runs (for example, using a KV cache for a transformer-based model). These computational results can be stored associated with that inference session so that the session can run more efficiently when it is resumed.
[0159] A computing resource 31-2(or more) may include one or more processors (such as a central processing unit, graphical processing unit, tensor processing unit, or machine learning accelerator) connected to one or more memory devices. A computing resource 31-2(or more) may include a dynamic pool of available resources shared with other processes. A computing resource 31-2(or more) may include a memory device large enough to accommodate an entire model instance in a single memory instance. A computing resource 31-2(or more) may also share a model instance(or more) across multiple memory devices (for example, using data parallelism or tensor parallelism). This may be done to enhance parallelism or to run large models using multiple memory devices where the entire model cannot fit in memory on individual memory devices.
[0160] Input request 33 may contain data for input 2(or more). Model host 31 can process input request 33 to retrieve input 2(or more). Input 2(or more) can be retrieved directly from input request 33 or by using input request 33. Input request 33 may be submitted to model host 31 via API.
[0161] The model host 31 can perform inference in parallel across batches of input requests 33. For example, a model instance 31-1 can consist of an input structure having a batch dimension. Separate inputs 2(or more) can be distributed across the batch dimension (e.g., rows of an array). Separate inputs 2(or more) can contain completely different contexts. Separate inputs 2(or more) can be multiple inference steps of the same task. Separate inputs 2(or more) can be arranged incrementally within the input structure, so that any given inference cycle can act on different parts of each input 2(or more). In this way, for example, the model host 31 can perform inference in parallel for batches, and as a result, the output 3(or more) can also include a batch dimension and return inference results for the batched inputs 2(or more) in parallel. In this way, for example, batches of input requests 33(or more) can be processed in parallel for higher throughput of the output payloads 34(or more).
[0162] The output payload 34 may contain or be based on the outputs 3(or more) from the machine learning model 1(or more). The model host 31 can process the outputs 3(or more) to obtain the output payload 34. This may involve chaining multiple inferences (e.g., iteratively, recursively, across the same model(or more) or different models(or more)) to arrive at the final output of the task returned in the output payload 34. The output payload 34 may be sent to the client 32(or more) via the API.
[0163] The online learning interface 36(multiple) can facilitate reinforcement learning of the machine learning model 1(multiple). The online learning interface 36(multiple) can facilitate reinforcement learning with human feedback (RLHF). The online learning interface 36(multiple) can facilitate federated learning of the machine learning model 1(multiple).
[0164] The model host 31 may have access to a library of pre-trained adapters or LoRA modules that can be used to adapt the baseline model so that its output matches a desired performance profile, model enhancements (e.g., adapting to different input modalities), etc. For example, the model host 31 may receive an input request to load a customized model, and the model host 31 may retrieve one or more components to fit the baseline model to a custom profile. The model host 31 may determine that certain functionality is required for a particular task (e.g., based on the output of a model that preprocesses the input) and retrieve pre-trained components accordingly.
[0165] The model host 31 can run a machine-trained model 1(or more) to perform inference on various tasks using various types of data. For example, various different inputs 2(or more) and outputs 3(or more) can be used for various different tasks. In some embodiments, input 2(or more) may be image data or represent image data in other ways. The machine-trained model 1(or more) can process the image data to generate outputs. For example, the machine-trained model 1(or more) can process the image data to generate image recognition outputs (e.g., recognition of image data, latent embedding of image data, coded representation of image data, hash of image data, etc.). As another example, the machine-trained model 1(or more) can process the image data to generate image segmentation outputs. As yet another example, the machine-trained model 1(or more) can process the image data to generate image classification outputs. As yet another example, the machine-trained model 1(or more) can process the image data to generate image data modification outputs (e.g., modification of image data, etc.). As another example, a machine learning model 1(or more) can process image data to generate encoded image data output (e.g., encoded and / or compressed representations of the image data). As yet another example, a machine learning model 1(or more) can process image data to generate upscaled image data output. As yet another example, a machine learning model 1(or more) can process image data to generate predictive output.
[0166] In some embodiments, the task is a computer vision task. In some cases, input 2(or more) includes pixel data from one or more images, and the task is an image processing task. For example, the image processing task could be image classification, and the output would be a set of scores, each corresponding to a different object class, representing the likelihood that one or more images depict an object belonging to that object class. The image processing task could be object detection, and the image processing output would identify one or more regions within one or more images, and for each region, the likelihood that the region depicts an object of interest. As another example, the image processing task could be image segmentation, and the image processing output would define, for each pixel in one or more images, the likelihood for each category within a given set of categories. For example, the set of categories could be foreground and background. As yet another example, the set of categories could be object classes. As yet another example, the image processing task could be depth estimation, and the image processing output would define, for each pixel in one or more images, the respective depth values. As another example, the image processing task could be motion estimation, where the network input includes multiple images, and the image processing output defines the motion of the scene depicted between the images in the network input, for each pixel in the input images.
[0167] In some embodiments, input 2(or more) may be natural language data or represent natural language data in other ways. A machine learning model 1(or more) can process the natural language data to produce an output. For example, a machine learning model 1(or more) can process natural language data to produce a language coding output. Another example is a machine learning model 1(or more) can process natural language data to produce a latent text embedding output. Another example is a machine learning model 1(or more) can process natural language data to produce a translation output. Another example is a machine learning model 1(or more) can process natural language data to produce a classification output. Another example is a machine learning model 1(or more) can process natural language data to produce a text segmentation output. Another example is a machine learning model 1(or more) can process natural language data to produce a semantic intent output. As another example, a machine learning model 1(or more) can process natural language data to generate upscaled text or natural language output (e.g., text or natural language data of higher quality than the input text or natural language). As yet another example, a machine learning model 1(or more) can process natural language data to generate predictive output (e.g., one or more predicted next parts of natural language content).
[0168] In some embodiments, input 2(or more) can be speech data (e.g., data describing spoken natural language, such as audio data or text data) or can represent the speech data in other ways. A machine learning model 1(or more) can process the speech data to produce an output. For example, a machine learning model 1(or more) can process the speech data to produce a speech recognition output. Another example is a machine learning model 1(or more) can process the speech data to produce a speech translation output. Another example is a machine learning model 1(or more) can process the speech data to produce a latent embedding output. Another example is a machine learning model 1(or more) can process the speech data to produce an encoded speech output (e.g., an encoded or compressed representation of the speech data). Another example is a machine learning model 1(or more) can process the speech data to produce an upscaled speech output (e.g., speech data of higher quality than the input speech data). As another example, a machine learning model 1(or more) can process speech data to generate a textual representation output (e.g., a textual representation of the input speech data). As yet another example, a machine learning model 1(or more) can process speech data to generate a predictive output.
[0169] In some embodiments, input 2 can be latent coded data (e.g., a latent spatial representation of the input) or can be represented in other ways. A machine learning model 1(or more) can process the latent coded data to produce an output. For example, a machine learning model 1(or more) can process the latent coded data to produce a recognition output. As another example, a machine learning model 1(or more) can process the latent coded data to produce a reconstruction output. As yet another example, a machine learning model 1(or more) can process the latent coded data to produce a search output. As yet another example, a machine learning model 1(or more) can process the latent coded data to produce a reclustering output. As yet another example, a machine learning model 1(or more) can process the latent coded data to produce a prediction output.
[0170] In some embodiments, input 2(or more) can be statistical data or represent statistical data in other ways. Statistical data can be computer-processed or computed data from some other data source, or represent it, or include it in other ways. A machine learning model 1(or more) can process the statistical data to produce an output. For example, a machine learning model 1(or more) can process statistical data to produce a recognition output. Another example is a machine learning model 1(or more) can process statistical data to produce a prediction output. Another example is a machine learning model 1(or more) can process statistical data to produce a classification output. Another example is a machine learning model 1(or more) can process statistical data to produce a segmentation output. Another example is a machine learning model 1(or more) can process statistical data to produce a visualization output. Another example is a machine learning model 1(or more) can process statistical data to produce a diagnostic output.
[0171] In some embodiments, input 2(or more) can be sensor data or represent sensor data in other ways. A machine learning model 1(or more) can process the sensor data to generate an output. For example, a machine learning model 1(or more) can process sensor data to generate a recognition output. As another example, a machine learning model 1(or more) can process sensor data to generate a prediction output. As yet another example, a machine learning model 1(or more) can process sensor data to generate a classification output. As yet another example, a machine learning model 1(or more) can process sensor data to generate a segmentation output. As yet another example, a machine learning model 1(or more) can process sensor data to generate a visualization output. As yet another example, a machine learning model 1(or more) can process sensor data to generate a diagnostic output. As yet another example, a machine learning model 1(or more) can process sensor data to generate a detection output.
[0172] In some embodiments, a machine-trained model(s) may be configured to perform tasks that involve encoding input data for reliable or efficient transmission or storage (or corresponding decryption). For example, the task could be an audio compression task. The input may include audio data, and the output may include compressed audio data. In another example, the input may include visual data (e.g., one or more images or videos), and the output may include compressed visual data, and the task is a visual data compression task. In yet another example, the task may involve generating embeddings for the input data (e.g., input audio or visual data). In some cases, the input may include audio data representing spoken utterances, and the task is a speech recognition task. The output may include text output mapped to the spoken utterances. In some cases, the task may involve encrypting or decrypting the input data. In some cases, the task may involve microprocessor performance tasks, such as branch prediction or memory address translation.
[0173] In some embodiments, the task is a generative task, and a machine-trained model 1(or more) can be configured to output content generated considering an input 2(or more). For example, input 2(or more) can represent data from one or more modalities that encode a context for generating additional content, or data from one or more modalities that encode a context for generating additional content in other ways.
[0174] In some embodiments, the task can be a text completion task. A machine-trained model 1(or more) can be configured to process an input 2(or more) representing text data and to produce an output 3(or more) representing additional text data that completes a text sequence containing the input 2(or more). For example, a machine-trained model 1(or more) can be configured to produce an output 3(or more) to complete a sentence, paragraph, or portion of text that follows a portion of text represented by the input 2(or more).
[0175] In some embodiments, a task can be a command that follows the task. A machine learning model 1(or more) can be configured to process an input 2(or more) representing a command to perform a certain function and to produce an output 3(or more) that achieves the goal of satisfying the function of that command (e.g., at least one step of a multi-step procedure to perform that function). The output 3(or more) can represent data of the same or different modality as the input 2(or more). For example, the input 2(or more) can represent text data (e.g., a natural language command for a task to be performed), and the machine learning model 1(or more) can process the input 2(or more) to produce an output 3(or more) representing text data in response to the command (e.g., a natural language response, a programming language response, a machine language response, etc.). Input 2(or more) can represent image data (e.g., image-based instructions for a task to be performed, optionally accompanied by text instructions), and a machine learning model 1(or more) can process input 2(or more) to generate outputs 3(or more) representing text data in response to the instructions (e.g., natural language responses, programming language responses, machine language responses, etc.). One or more outputs 3(or more) can be generated iteratively or recursively to sequentially process and achieve the steps necessary to accomplish the requested function. For example, an initial output can be executed by an external system or processed by a machine learning model 1(or more) to complete the initial steps for performing the function. Multiple steps may be performed, and a final output in response to the first instruction is obtained.
[0176] In some embodiments, the task can be a question-answering task. A machine-trained model 1(or more) can be configured to process an input 2(or more) representing a question to be answered and produce an output 3(or more) that achieves the goal of returning an answer to that question (e.g., at least one step of a multi-step procedure for performing its function). The output 3(or more) can represent data of the same or different modality as the input 2(or more). For example, the input 2(or more) can represent text data (e.g., natural language instructions for the task to be performed), and the machine-trained model 1(or more) can process the input 2(or more) to produce an output 3(or more) that represents text data responding to the question (e.g., a natural language response, a programming language response, a machine language response, etc.). Input 2(or more) can represent image data (e.g., image-based instructions for a task to be performed, optionally accompanied by text instructions), and a machine learning model 1(or more) can process input 2(or more) to generate output 3(or more) representing text data in response to a question (e.g., natural language response, programming language response, machine language response, etc.). One or more outputs 3(or more) can be generated iteratively or recursively to sequentially process and achieve steps to achieve an answer to a question. For example, an initial output can be executed by an external system or processed by a machine learning model 1(or more) to complete initial steps to obtain an answer to a question (e.g., querying a database, performing calculations, executing a script, etc.). Multiple steps may be performed to obtain a final output that responds to the question.
[0177] In some embodiments, the task can be an image generation task. A machine-trained model 1(or more) can be configured to process an input 2(or more) representing context about a desired portion of image content. The context can include text data, image data, audio data, etc. The machine-trained model 1(or more) can be configured to produce an output 3(or more) representing image data that depicts an image relevant to the context. For example, the machine-trained model 1(or more) can be configured to generate pixel data for an image. The values of the channels(or more) associated with the pixels in the pixel data can be selected based on the context (for example, based on probabilities determined based on the context).
[0178] In some embodiments, the task can be an audio generation task. A machine-trained model 1(or more) can be configured to process an input 2(or more) representing context about a desired portion of audio content. The context can include text data, image data, audio data, etc. The machine-trained model 1(or more) can be configured to produce an output 3(or more) representing audio data related to the context. For example, the machine-trained model 1(or more) can be configured to generate waveform data in image format (e.g., spectrogram). The channel(or more) values associated with pixels in the image may be selected based on the context. The machine-trained model 1(or more) can be configured to generate waveform data in the form of a sequence of discrete samples of a continuous waveform. The sequence values may be selected based on the context (e.g., based on probabilities determined based on the context).
[0179] In some embodiments, the task can be a data generation task. A machine-trained model 1(or more) can be configured to process an input 2(or more) representing a context about a desired portion of data (e.g., data from various data domains such as sensor data, image data, multimodal data, statistical data, etc.). The desired data can be, for example, synthetic data for training other machine-trained models. The context can include any data type(or more). The machine-trained model 1(or more) can be configured to produce an output 3(or more) representing data that fits the desired data. For example, the machine-trained model 1(or more) can be configured to generate data values to fill a dataset. The values of the data objects(or more) can be selected based on the context (e.g., based on probabilities determined based on the context).
[0180] Figure 11 is a block diagram of an exemplary network computing system capable of implementing an exemplary embodiment of the present disclosure. The system may include a plurality of computing devices and systems that are communicably coupled over a network 49. An exemplary computing device 50 is described to provide an example of a computing device capable of implementing any embodiment of the present disclosure (e.g., implementing a model host 31, a client 32(or more), or both). An exemplary server computing system 60 is described as an example of a server computing system capable of implementing any embodiment of the present disclosure (e.g., implementing a model host 31, a client 32(or more), or both). The computing devices 50 and the server computing system 60(or more) can interact in cooperation (e.g., over the network 49) to implement any embodiment of the present disclosure (e.g., implementing a model host 31, a client 32(or more), or both). A model development platform system 70 is an exemplary system that can host or provide a model development platform 12(or more) for developing machine learning models. The third-party system 80(or more) is an exemplary system(or more) in which any of the computing device 50, server computing system 60(or more), or model development platform system 70(or more) can interact in the implementation of various aspects of the present invention (e.g., the use of third-party tools, access to third-party databases or other resources).
[0181] Network 49 can be any type of communication network, such as a local area network (e.g., an intranet), a wide area network (e.g., the Internet), or any combination thereof, and can include any number of wired or wireless links. Generally, communication over Network 49 can be conducted over any type of wired or wireless connection using a wide variety of communication protocols (e.g., TCP / IP, HTTP, SMTP, FTP), encoding or formatting (e.g., HTML, XML), or protection schemes (e.g., VPN, Secure HTTP, SSL). Network 49 can also be implemented via a system bus. For example, one or more devices or systems in Figure 11 may be located in the same place as one or more other devices or systems, may be housed in one or more other devices or systems, or may be integrated in any other way.
[0182] The computing device 50 can be any type of computing device, such as a personal computing device (e.g., a laptop or desktop), a mobile computing device (e.g., a smartphone or tablet), a game console or controller, a wearable computing device, an embedded computing device, a server computing device, a virtual machine running on a host device, or any other type of computing device. The computing device 50 can be a client computing device. The computing device 50 can be an end-user computing device. The computing device 50 can be a computing device for the service itself, providing services to the end user (the end user can interact with the computing device 50 using other computing devices).
[0183] The computing device 50 may include one or more processors 51 and memory 52. The processor 51(s) may be any suitable processing device (e.g., a processor core, microprocessor, ASIC, FPGA, controller, microcontroller, etc.) and may be one processor or multiple operably connected processors. The memory 52 may include one or more non-temporary computer-readable storage media such as HBM, RAM, ROM, EEPROM, EPROM, flash memory devices, magnetic disks, etc., and combinations thereof. The memory 52 may store data 53 and instructions 54 executed by the processor 51(s) to cause the computing device 50 to perform operations. The operation may implement any one or more features described herein. The operation may implement exemplary methods and techniques described herein.
[0184] Furthermore, the computing device 50 may include one or more input components that receive user input. For example, a user input component may be a touch-sensitive component (e.g., a touch-sensitive display screen or touchpad) that is sensitive to the touch of a user input object (e.g., a finger or stylus). The touch-sensitive component may serve to implement a virtual keyboard. Other exemplary user input components include a microphone, a camera, LiDAR, a physical keyboard or other buttons, or other means by which the user can provide user input.
[0185] The computing device 50 can store or contain one or more machine learning models 55. The machine learning models 55 may contain one or more machine learning models 1(or more), such as a sequence processing model 4. The machine learning models 55 may contain one or more model instances 31(or more). The machine learning models 55(or more) may be received from a server computing system 60(or more), a model development platform system 70, a third-party system 80(or more) (e.g., an application distribution platform), or developed locally on the computing device 50. The machine learning models 55(or more) may be loaded into memory 52 and used by a processor 51(or more), or implemented in other ways. The computing device 50 can implement multiple parallel instances of the machine learning models 55(or more).
[0186] The server computing system 60 may include one or more processors 61 and memory 62. The processor 61(or more) may be any suitable processing device (e.g., a processor core, microprocessor, ASIC, FPGA, controller, microcontroller, etc.) and may be one processor or multiple operably connected processors. The memory 62 may include one or more non-temporary computer-readable storage media such as HBM, RAM, ROM, EEPROM, EPROM, flash memory devices, magnetic disks, etc., and combinations thereof. The memory 62 may store data 63 and instructions 64 executed by the processor 61(or more) to cause the server computing system 60(or more) to perform operations. The operation may implement any one or more features described herein. The operation may implement exemplary methods and techniques described herein.
[0187] In some embodiments, the server computing system 60 includes one or more server computing devices, or is otherwise implemented by one or more server computing devices. If the server computing system 60 includes multiple server computing devices, such server computing devices can operate in a sequential computing architecture, a parallel computing architecture, or any combination thereof.
[0188] The server computing system 60 may store one or more machine learning models 65, or may otherwise include one or more machine learning models 65. The machine learning models 65 may be the same as or different from the machine learning models 55. The machine learning models 65 may include one or more machine learning models 1, such as sequence processing model 4. The machine learning models 65 may include one or more model instances 31-1. The machine learning models 65 may be received from a computing device 50, a model development platform system 70, a third-party system 80, or developed locally in the server computing system 60. The machine learning models 65 may be loaded into memory 62 and used by a processor 61, or otherwise implemented. The server computing system 60 may implement multiple parallel instances of the machine learning models 65.
[0189] In an exemplary configuration, the machine learning model 65 can be contained in or otherwise stored and implemented by the server computing system 60 to establish a client-server relationship with the computing device 50 for providing model inference. For example, the server computing system 60(s) can implement a model host 31 on the computing device 50 on behalf of the client 32(s). For example, the machine learning model 65 may be implemented by the server computing system 60 as part of a web service (e.g., a remote machine learning model host service such as an online interface for performing machine learning model calculations over a network on the server computing system 60(s)). For example, the server computing system 60(s) can communicate with the computing device 50 via a local intranet or internet connection. For example, computing device 50 may be a workstation or endpoint communicating with a server computing system 60(or more), in which case the implementation of the machine learning model 65 is managed by the server computing system 60(or more) to perform inference remotely (e.g., at runtime or for training operations), and its output(or more) is returned to computing device 50 (e.g., cast, streamed, etc.). The machine learning model 65 can work collaboratively or interactively with the machine learning model 55 on computing device 50 to perform various tasks.
[0190] A model development platform system 70(or more) may include one or more processors 71 and memory 72. The processor 71(or more) may be any suitable processing device (e.g., a processor core, microprocessor, ASIC, FPGA, controller, microcontroller, etc.) and may be a single processor or multiple operably connected processors. The memory 72 may include one or more non-temporary computer-readable storage media, such as HBM, RAM, ROM, EEPROM, EPROM, flash memory devices, magnetic disks, and combinations thereof. The memory 72 may store data 73 and instructions 74 that can be executed by the processor 71(or more) to cause the model development platform system 70(or more) to perform operations. The operations may implement any one or more features described herein. The operations may implement exemplary methods and techniques described herein. Exemplary operations include the functions described herein with respect to the model development platform 12. These functions and other functions may be implemented by developer tools 75(or more).
[0191] A third-party system 80(or more) may include one or more processors 81 and memory 82. The processor 81(or more) may be any suitable processing device (e.g., a processor core, microprocessor, ASIC, FPGA, controller, microcontroller, etc.) and may be one processor or multiple operably connected processors. The memory 82 may include one or more non-temporary computer-readable storage media such as HBM, RAM, ROM, EEPROM, EPROM, flash memory devices, magnetic disks, etc., and combinations thereof. The memory 82 may store data 83 and instructions 84 that can be executed by the processor 81(or more), causing the third-party system 80(or more) to perform operations. The operations may implement any one or more features described herein. The operations may implement exemplary methods and techniques described herein. Exemplary operations include the functionality of tools and other external resources (e.g., third-party resources) that are invoked when performing training or inference using machine learning models 1, 4, 16, 20, 55, 65, etc., as described herein.
[0192] Figure 11 shows one exemplary configuration of a computing system that may be used to implement this disclosure. Other computing system configurations may be used similarly. For example, in some embodiments, one or both of the computing system 50 or server computing system 60(or more) may implement all or part of the operation of the model development platform system 70. For example, the computing system 50 or server computing system 60(or more) may implement developer tools 75(or more) (or extensions thereof) to develop, update / train or improve machine learning models 1, 4, 16, 20, 55, 65, etc., using one or more techniques described herein with respect to the model alignment toolkit 17. In this way, for example, the computing system 50 or server computing system 60(or more) may develop, update / train or improve machine learning models based on local datasets (for example, for personalization / customization of the model, to the extent permitted by user data preference selection).
[0193] Figure 12 shows a block diagram of an exemplary computing device 98 implemented according to an exemplary embodiment of the present disclosure. The computing device 98 may be a user computing device or a server computing device (e.g., computing device 50, server computing system 60, etc.). The computing device 98 may implement a model host 31. For example, the computing device 98 may include multiple applications (e.g., applications 1 to N). Each application may include its own machine learning library and machine-trained model(s). For example, each application may include a machine-trained model. Exemplary applications include a text messaging application, an email application, a dictation application, a virtual keyboard application, a browser application, etc. As illustrated in Figure 12, each application may communicate with multiple other components of the computing device, such as one or more sensors, a context manager, a device state component, or additional components. In some embodiments, each application may communicate with each device component using an API (e.g., a public API). In some embodiments, the API used by each application is specific to that application.
[0194] Figure 13 shows a block diagram of an exemplary computing device 99 implemented according to an exemplary embodiment of the present disclosure. Computing device 99 may be the same as or different from computing device 98. Computing device 99 may be a user computing device or a server computing device (e.g., computing device 50, server computing system 60(or more)). Computing device 98 may implement a model host 31. For example, computing device 99 may contain multiple applications (e.g., applications 1 to N). Each application may communicate with the central intelligence layer. Exemplary applications include text messaging applications, email applications, dictation applications, virtual keyboard applications, browser applications, etc. In some embodiments, each application may communicate with the central intelligence layer (and the models(or more) stored therein) using an API (e.g., a common API across all applications).
[0195] The central intelligence layer can include multiple machine learning models. For example, as illustrated in Figure 13, each machine learning model may be provided for each application and managed by the central intelligence layer. In other embodiments, two or more applications may share a single machine learning model. For example, in some embodiments, the central intelligence layer may provide a single model to all applications. In some embodiments, the central intelligence layer may be contained within the operating system of the computing device 99 or otherwise implemented by the operating system of the computing device 99.
[0196] The central intelligence layer can communicate with the central device data layer. The central device data layer can be a centralized repository of data from the computing device 99. As shown in Figure 13, the central device data layer can communicate with several other components of the computing device, such as one or more sensors, a context manager, a device state component, or additional components. In some embodiments, the central device data layer can communicate with each device component using an API (e.g., a private API).
[0197] The technologies described herein refer to servers, databases, software applications, and other computer-based systems, as well as the actions performed and the information transmitted to and from such systems. The inherent flexibility of computer-based systems allows for a wide variety of feasible configurations, combinations, and divisions of tasks and functions between their components. For example, the processes described herein can be implemented using a single device or component, or multiple devices or components working together. Databases and applications can be implemented on a single system or distributed across multiple systems. Distributed components can operate sequentially or in parallel.
[0198] While the subject matter has been described in detail with respect to various specific exemplary embodiments, each example is provided for illustrative purposes only and does not limit the disclosure. Those skilled in the art, having attained the foregoing understanding, will readily be able to produce modifications, variations, and equivalents to such embodiments. Therefore, the disclosure of the subject matter does not preclude the inclusion of such modifications, variations, or additions to the subject matter that would be readily apparent to those skilled in the art. For example, features illustrated or described as part of one embodiment may be used in another embodiment to produce yet another embodiment. Thus, the disclosure is intended to cover such modifications, variations, and equivalents.
[0199] The aspects of this disclosure have been described in relation to their exemplary embodiments. Any and all of the following claims and features should not be construed as limiting the scope of possible combinations of features disclosed herein by the dependency of the exemplary claims enumerated herein, so that combinations can be combined or rearranged in any way, including combinations of claims not expressly enumerated. Accordingly, the scope of this disclosure is illustrative and not limiting, and the disclosure of the subject matter does not preclude such modifications, variations or additions to the subject matter, as would be readily apparent to those skilled in the art. Furthermore, terms are described herein using lists of exemplary elements joined by conjunctions such as “and,” “or,” and “but.” It should be understood that such conjunctions are provided for illustrative purposes only. For example, clauses and other sets of items joined by certain conjunctions such as “or” may refer to “and / or,” “at least one of the exemplary elements enumerated therein,” “any combination of,” etc. Terms such as “based on” should be understood as “at least partially based on.”
[0200] The term "can" should not be understood as referring to a capability that necessarily exists in all embodiments, but rather as referring to the potential of features in various embodiments. For example, the phrase "X can perform Y" should be understood as indicating that in various embodiments, X may be configured to perform Y, and not as indicating that X can always perform Y in all examples. In various embodiments, X may not be able to perform Y and may remain within the scope of this disclosure.
[0201] The term "can" should not be understood as referring to a capability that necessarily exists in all embodiments, but rather as referring to the potential of features in various embodiments. For example, the phrase "X can perform Y" should be understood as indicating that in various embodiments, X may be configured to perform Y, and not as indicating that X can always perform Y in all examples. In various embodiments, X may not be able to perform Y and may remain within the scope of this disclosure.
[0202] This specification uses examples to disclose the invention in the best form and to enable the implementation of the invention, including the fabrication and use of any device or system by any person skilled in the art, and the implementation of any method of incorporating the invention. The claims of the invention are defined by the claims and may include other embodiments that a person skilled in the art could conceive. Such other embodiments are intended to be within the claims if they include structural elements that do not differ from the language of the claims, or if they include equivalent structural elements that do not substantially differ from the language of the claims.
Claims
1. A method implemented in a computer, This involves training a prompt extension model and an image generation model in parallel using a multi-reward reinforcement learning model. To generate an enhanced training query, the prompt enhancement model processes the training query and training context data, The image generation model generates a training set of image data based on the extended training query, For each image data in the training set of the image data, a set of reward scores is generated using a set of reward models, wherein generating the set of reward scores includes generating at least one reward score for each reward criterion in a plurality of reward criteria. A computer-implemented method comprising training the multi-reward reinforcement learning model by adjusting the weights and biases associated with the multiple reward criteria based on the set of reward scores.
2. The computing system retrieves input data, including user queries and query context data. Using the trained prompt extension model and the trained image generation model, generate image data based on the user query and the query context data, To generate an extended user query, the trained prompt extension model processes the user query and the query context data, A computer-implemented method according to claim 1, comprising generating the image data based on the extended user query using the trained image generation model.
3. The method implemented on a computer according to claim 2, wherein generating the extended user query includes incorporating the plurality of reward criteria into the extended user query.
4. A method for implementing the plurality of reward criteria in a computer according to claim 1, wherein the plurality of reward criteria include one or more text-image alignment reward criteria.
5. One or more text-to-image alignment reward criteria, A first text-image alignment reward criterion associated with the training query, A computer-implemented method according to claim 2 or 4, comprising a second text-image alignment reward criterion associated with the extended training query.
6. The method for implementing the plurality of reward criteria in a computer according to claim 1, wherein the plurality of reward criteria include one or more image emotion reward criteria.
7. The method of implementing the plurality of reward criteria in a computer according to claim 1, wherein the plurality of reward criteria include one or more aesthetic reward criteria.
8. The method for implementing the plurality of reward criteria in a computer according to claim 1, wherein the plurality of reward criteria include one or more human preference reward criteria.
9. Training the prompt extension model and the image generation model is From the training set of the aforementioned image data, a subset of the image data is selected using a non-dominant sort algorithm as a function of the set of reward scores, A computer-implemented method according to claim 1, comprising using the multi-reward reinforcement learning model to adjust the weights and biases associated with the plurality of reward criteria based on a subset of the image data.
10. The computer-implemented method according to claim 1, wherein the subset of the image data includes a Pareto optimal set associated with the set of reward scores.
11. Training the prompt extension model and the image generation model is The computing system determines the policy gradient update as a function of the subset of the image data, A computer-implemented method according to claim 9 or 10, comprising using the multi-reward reinforcement learning model as a function of the policy gradient update to adjust the weights and biases associated with the multiple reward criteria.
12. A computer-implemented method according to claim 11, further comprising determining the policy gradient update by minimizing the reward score associated with each reward criterion not represented in the subset of image data.
13. A computer-implemented method according to claim 11, further comprising determining the policy gradient update to maximize the reward score associated with each reward criterion represented within the subset of image data.
14. A computer-implemented method according to claim 11, wherein determining the policy gradient update includes maximizing one or more text-image alignment reward criteria.
15. A computing system, One or more processors, The system includes one or more temporary or non-temporary computer-readable media that store instructions executable for causing one or more processors to perform an operation, and the operation is The one or more processors acquire input data including user queries and query context data, This includes training a prompt extension model and an image generation model in parallel using a multi-reward reinforcement learning model, Training the prompt extension model and the image generation model is To generate an enhanced training query, the prompt enhancement model processes the training query and training context data, Using the aforementioned image generation model, a training set of image data is generated based on the aforementioned augmented training query. For each image data in the training set of the image data, a set of reward scores is generated using a set of reward models, wherein generating the set of reward scores includes generating at least one reward score for each reward criterion in a plurality of reward criteria. From the training set of the aforementioned image data, a subset of the image data is selected using a non-dominant sort algorithm as a function of the set of reward scores, Using the multi-reward reinforcement learning model, the weights and biases associated with each reward criterion not represented within the subset of image data are minimized. A computing system comprising generating image data based on user queries and query context data using the trained prompt extension model and the trained image generation model.
16. The generation of the aforementioned image data is To generate an extended user query, the trained prompt extension model processes the user query and the query context data, The computing system according to claim 15, further comprising generating the image data based on the extended user query using the trained image generation model.
17. The computing system according to claim 16, wherein generating the extended user query includes incorporating the plurality of reward criteria into the extended user query.
18. Training the prompt extension model and the image generation model is The computing system determines the policy gradient update as a function of the subset of the image data, The computing system according to any one of claims 15 to 17, further comprising using the multi-reward reinforcement learning model as a function of the policy gradient update to adjust the weights and biases associated with the multiple reward criteria.
19. The computing system according to claim 18, wherein determining the policy gradient update further includes maximizing the reward score associated with each reward criterion represented within the subset of image data.
20. A method implemented in a computer, This involves training a prompt extension model and an image generation model in parallel using a multi-reward reinforcement learning model. To generate an enhanced training query, the prompt enhancement model processes the training query and training context data, The image generation model generates a training set of image data based on the extended training query, A computer-implemented method comprising training the prompt extension model and the image generation model in parallel using the multi-reward reinforcement learning model based on the training set of the image data.
Citation Information
Patent Citations
Method, model and device for training text graph model, and electronic equipment
CN116894880A
Generation apparatus, generation method, and generation program
JP2021149716A
JPP7404596B
Generating images using sequences of generative neural networks
US20230377226A1
Systems and methods for generation of machine-learned multitask models
WO2022019913A1