SYSTEMS AND METHODS FOR A MULTIPLE-REWARD REINFORCEMENT LEARNING FRAMEWORK FOR TEXT-TO-IMAGE GENERATION

The multi-reward reinforcement learning framework trains a prompt expansion model and image generation model jointly to enhance image quality and alignment with textual prompts, addressing the challenge of contextualization in existing image generation models.

BR112025009523A2Pending Publication Date: 2026-07-28GOOGLE LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
BR112025009523
Authority / Receiving Office
BR · BR
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-01-10
Filing Date
2024-11-13
Publication Date
2026-07-28

AI Technical Summary

Technical Problem

Existing image generation models struggle to generate images that meet performance metrics due to insufficient contextualization of textual prompts, leading to a loss of focus on the original prompt and degradation of image quality.

Method used

A multi-reward reinforcement learning framework is employed to train a prompt expansion model and an image generation model in tandem, using a joint optimization approach to adjust weights and trends based on multiple reward criteria, ensuring the generated images align with both the original and expanded textual prompts.

Benefits of technology

This approach enhances image generation by maintaining focus on the original prompt while improving image quality, clarity, resolution, and alignment with user intent, resulting in images that outperform baseline methods across various quality criteria.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

Systems and methods for multi-reward reinforcement learning framework for text-to-image generation are disclosed. The method includes training, in tandem, a prompt expansion model and an image generation model using a multi-reward reinforcement learning model by: processing, by the prompt expansion model, a training query and training context data to generate an expanded training query; generating, by the image generation model, a training set of image data based on the expanded training query; generating a set of reward scores for each image datum within the training set of image data using a set of reward models, wherein generating the set of reward scores comprises generating at least one reward score for each reward criterion within a plurality of reward criteria; and adjusting weights and biases associated with the plurality of reward criteria based on the set of reward scores using the multi-reward reinforcement learning model.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application claims priority over U.S. Provisional Application No. 63 / 619,632, entitled “MULTI-REWARD REINFORCEMENT LEARNING FRAMEWORK FOR TEXT-TO-IMAGE GENERATION”, filed January 10, 2024, disclosure of which is incorporated herein by reference in its entirety. FIELD

[0002] The present invention relates generally to machine learning processes and machine learning devices and systems. More particularly, the present disclosure relates to a multi-reward reinforcement learning framework for text-to-image generation. BACKGROUND

[0003] A computer can receive input (or inputs). The computer can execute instructions to process the input (or inputs) to generate output (or outputs) using a parameterized model. The computer can obtain feedback on its performance in generating the outputs with the model. The computer can generate feedback by evaluating its performance. The computer can receive feedback from an external source. The computer can update model parameters based on the feedback to improve its performance. In this way, the computer can iteratively “learn” to generate the desired outputs. The resulting model is often called a machine learning model. BRIEF DESCRIPTION

[0004] Aspects and advantages of the invention according to the present disclosure will be presented in part in the following description, Petition 870250038841, dated 05 / 13 / 2025, page 15 / 142 2 / 96 either may be obvious from the description, or they may be learned through practice with the technology.

[0005] According to one embodiment, a method for a multi-reward reinforcement learning framework for text-to-image generation is provided. The method includes training, in tandem, a prompt expansion model (prompt expansion model) and an image generation model (image generation models) using a multi-reward reinforcement learning model by: processing, by the prompt expansion model, a training query and training context data to generate an expanded training query; generating, by the image generation model, a training image dataset based on the expanded training query;Generate a set of reward scores for each image datum within the training image dataset using a set of reward models, wherein generating the set of reward scores comprises generating at least one reward score for each reward criterion within a plurality of reward criteria; and adjusting weights and trends associated with the plurality of reward criteria based on the set of reward scores using the multiple reward reinforcement learning model.

[0006] According to another embodiment, a system for a multi-reward reinforcement learning framework for text-to-image generation is provided. The system includes one or more processors; and one or more computer-readable transient or non-transient media that store executable instructions to cause the one or more processors to perform operations, the operations comprising: obtaining, by the one or more processors, input data comprising a user query and data Petition 870250038841, dated 05 / 13 / 2025, page 16 / 142 3 / 96 of query context; train, in tandem, a prompt expansion model (prompt expansion model) and an image generation model using a multi-reward reinforcement learning model, wherein training the prompt expansion model and the image generation model comprises: processing, by the prompt expansion model, a training query and training context data to generate an expanded training query; generating, using the image generation model, a training image dataset based on the expanded training query; generating a set of reward scores for each image data within the training image dataset using a set of reward models, wherein generating the set of reward scores comprises generating at least one reward score for each reward criterion within a plurality of reward criteria;Select a subset of image data from the training image dataset as a function of the reward score set using a non-dominated classification algorithm; and minimize weightings and biases associated with each reward criterion that is not represented within the image data subset using the multi-reward reinforcement learning model; and generate, using the trained prompt expansion model and the trained image generation model, image data based on the user query and query context data.

[0007] According to one embodiment, a method for a multi-reward reinforcement learning framework for text-to-image generation is provided. The method includes training, in tandem, a prompt expansion model (prompt expansion model) and an image generation model using a multi-reward reinforcement learning model by: processing, by Petition 870250038841, dated 05 / 13 / 2025, page 17 / 142 4 / 96 prompt expansion model, a training query, and training context data to generate an expanded training query; generate, by the image generation model, a training image dataset based on the expanded training query; and train, in tandem, the prompt expansion model and the image generation model using the multi-reward reinforcement learning model based on the training image dataset.

[0008] These and other features, aspects and advantages of the present invention will become better understood with reference to the following description and appended claims. The attached drawings, which are incorporated herein and form a part of this descriptive report, illustrate embodiments of the technology and, together with the description, serve to explain the principles of the technology. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] A complete and empowering disclosure of the present invention, including the best mode for producing and using the present systems and methods, directed to one of ordinary skill in the art, is presented in the descriptive report, which refers to the accompanying Figures, in which:

[0010] Figures 1A-B represent a diagram of an exemplary framework for multi-reward reinforcement learning for generative text-to-image models according to exemplary modalities of the present disclosure;

[0011] Figure 2 represents a block diagram of an exemplary system for generating image data using an image generation model according to exemplary embodiments of the present disclosure;

[0012] Figure 3 represents a flowchart of an exemplary modality of a method for reinforcement learning of multiples Petition 870250038841, dated 05 / 13 / 2025, page 18 / 142 5 / 96 rewards for generative text-to-image models according to the exemplified modalities of this disclosure;

[0013] Figure 4 is a flowchart diagram illustrating an exemplary method for training a machine learning model according to exemplary implementations of aspects of the present disclosure;

[0014] Figure 5 is a block diagram of an exemplary processing flow for using a machine learning model (or models) to process input (or inputs) to generate output (or outputs) in accordance with exemplary implementations of aspects of this disclosure;

[0015] Figure 6 is a block diagram of an exemplary sequence processing model according to exemplary implementations of aspects of the present disclosure;

[0016] Figure 7 is a block diagram of an exemplary technique for populating an exemplary input sequence for processing by a sequence processing model according to exemplary implementations of aspects of the present disclosure;

[0017] Figure 8 is a block diagram of an exemplary model development platform according to exemplary implementations of aspects of the present disclosure;

[0018] Figure 9 is a block diagram of an exemplary training workflow for training a machine learning model according to exemplary implementations of aspects of this disclosure;

[0019] Figure 10 is a block diagram of an inference system for operating one or more machine learning models to perform inference according to exemplary implementations of aspects of this disclosure; Petition 870250038841, dated 05 / 13 / 2025, page 19 / 142 6 / 96

[0020] Figure 11 is a block diagram of an exemplary network computing system according to exemplary implementations of aspects of the present disclosure;

[0021] Figure 12 is a block diagram of an exemplary computing device according to exemplary implementations of aspects of the present disclosure; and

[0022] Figure 13 is a block diagram of an exemplary computing device according to exemplary implementations of aspects of the present disclosure. DETAILED DESCRIPTION

[0023] Generally, the present disclosure is directed to a multi-reward reinforcement learning framework for image generation models. More particularly, the present disclosure provides a means to train image generation models and prompt expansion models in tandem using a multi-reward reinforcement learning model. Training the image generation models and the prompt expansion model in tandem can provide improved fine-tuning of both the image generation models and the prompt expansion model.

[0024] This disclosure provides the use of a Pareto optimal selection in batches. The system can identify optimal trade-offs between various rewards during the training phase. The system can utilize a joint optimization approach for the image generation models and the prompt expansion model to facilitate the generation of expanded text prompts. The joint optimization approach can provide improved image generation while refocusing the image generation models on the original text prompt. This can be achieved by assigning and optimizing rewards for both the original text prompt and the expanded text prompts. Petition 870250038841, dated 05 / 13 / 2025, page 20 / 142 7 / 96 dios. This allows image generation models to generate images that utilize the expanded prompt while maintaining the original text prompt within the context window for image generation.

[0025] This disclosure provides several technical benefits to address technical problems. For example, existing image generation models have encountered several challenges in generating images that meet performance metrics. This is true, at least in part, because the textual prompts provided are often devoid of sufficient context for image generation models to generate images that meet performance metrics such as image quality, resolution, clarity, or other performance metrics. To overcome this problem, prompt expansion models have been employed to enrich the textual prompts with additional details for use by image generation models for image generation. Existing methods using prompt expansion models present several technical challenges.Existing systems simply provide an expanded prompt to the image generation models; however, as discussed, providing only an expanded prompt can lead to a loss of focus on the original prompt by the image generation models. This results in the generation of images that fail to meet performance metrics.

[0026] This disclosure can prevent overoptimization and degradation of metrics resulting from training with simple aggregation of multiple rewards by jointly optimizing multiple rewards in the use of the multiple reward reinforcement learning model. By adjusting the image generation models and the prompt expansion model simultaneously, the system can provide improved adjustment of both the prompt expansion model and the image generation models. This can provide a model of ge Petition 870250038841, dated 05 / 13 / 2025, page 21 / 142 8 / 96 Enhanced image conversion that outperforms baseline text-to-image methods across a variety of quality criteria metrics. Quality criteria metrics may include, for example, visual clarity, contrast, pixelation, resolution, aesthetics, human preferences, image sentiment, or text-to-image alignment.

[0027] The enhancements associated with the systems and methods discussed in this document can be better understood with reference to the Figures. Reference is now made to the Figures, which provide exemplary arrangements of computing systems, model structures, and data flows for illustrative purposes only.

[0028] With reference now to the drawings, Figure 1 illustrates an exemplary block diagram of a system for multi-reward reinforcement learning for image generation models. The system 100 may include a user query 102a-b, query context data 104a-b, prompt expansion model (prompt expansion model) 106, expanded user query 108a-b, image generation model (image generation models) 110, training image dataset 112, multi-reward reinforcement learning model 114, plurality of reward criteria 116a-d, weightings and trends 118, reward score set 120a-d, reward model set 122a-d, image data 124 and the like.

[0029] The operations involve training, in tandem, a prompt expansion model 106 (prompt expansion model) and an image generation model 110 using a multi-reward reinforcement learning model 114. Training the prompt expansion model 106 and the image generation model 110 in tandem means that both models are optimized or fine-tuned simultaneously. Training the models in tandem helps both models to Petition 870250038841, dated 05 / 13 / 2025, page 22 / 142 9 / 96 work in harmony without over-optimization and degradation of metrics. This is done to ensure that the text generated by the prompt expansion model 106 is directly aligned with how the image generation model 110 will interpret the text. By training the prompt expansion model 106 and the image generation model 110 in tandem, both models receive feedback from the multi-reward reinforcement learning model 114 regarding their actions within the environment.

[0030] Joint training of the prompt expansion model (prompt expansion model) 106 and the image generation model (image generation models) 110 may include iteratively updating parameters such as weights, trends, and coefficients based on feedback from the plurality of reward criteria 116a-d. Both the prompt expansion model 106 and the image generation model 110 may be simultaneously optimized to ensure that they work cohesively.

[0031] In one embodiment, updating parameters may include employing techniques such as gradient descent or least squares processes, allowing for the adjustment of weightings and trends based on received feedback (i.e., reward score set 120a-d). Parameter updates may be performed iteratively to ensure that the parameters of both agents (i.e., image generation model 110 and prompt expansion model 106) are refined in response to rewards and errors encountered during training.

[0032] Joint training of the prompt expansion model 106 and the image generation model 110 can be iteratively repeated until the available training data are exhausted or a convergence test is achieved. As used in this disclosure, a convergence test can be used to evaluate Petition 870250038841, dated 05 / 13 / 2025, page 23 / 142 10 / 96 if the models have reached an acceptable level of accuracy. This can be achieved by evaluating successive error values. If these differences in error values ​​fall below a defined threshold, this indicates that the models have stabilized. Alternatively, the error values ​​can be compared to a predetermined threshold to determine if further training is warranted.

[0033] Joint training of the prompt expansion model Figure 106 and the image generation model 110 include obtaining input data comprising a user query 102a-b. As used in this disclosure, a user query 102a refers to a request for information related to generating an image. The user query 102a may include a description of the desired image data. The user query 102a may include information related to subject matter, aesthetics, location, background, lighting, tone, emotions, art style, and the like. The user query 102a may be received from a user through various channels such as online forms, customer service emails, or live chat systems. Users may submit their queries by entering text or selecting options that describe their request or problem. Descriptions within the user query 102a may range from generalized questions to highly specific requests.For example, a generalized user query might simply state "a car" or "landscape," while a more detailed 102a user query might contain a description such as "a pink sports car parked in front of a beautiful botanical garden of pink flowers."

[0034] In one embodiment, a user query 102a may include a training query 102b. As used in this disclosure, training query 102b is an exemplary user query designed to provide input data for training. Petition 870250038841, dated 05 / 13 / 2025, page 24 / 142 11 / 96 a model. Training query 102b can be used as training data to help train or fine-tune prompt expansion model 106, image generation model 110, or any other model discussed in this document. Training query 102b can be used to help agents understand the relationship between textual input and visual output or a textual input and an expanded textual input. In some embodiments, training query 102b may be the same as or substantially similar to user query 102a. This may mean that training query 102b includes example descriptions of a desired image.

[0035] System 100 can receive query context data 104a associated with user query 102a-b. Query context data 104a refers to contextual information associated with the user or user query 102a-b. Query context data 104a provides additional knowledge of the circumstances surrounding user query 102a-b. Query context data 104a can be used to interpret user queries 102a-b by considering the broader context in which they arise. Query context data 104a may include a range of information derived from the user's previous interactions with system 100.

[0036] Query context data 104a may include information regarding the user's previous interactions with system 100. This may include previous queries submitted by the user, links that were clicked, pages that were visited, user feedback related to previous images, previous user queries 102a—or other relevant actions or behaviors displayed within system 100. For example, suppose a user submitted a question about an “animated raccoon ves Petition 870250038841, dated 05 / 13 / 2025, page 25 / 142 12 / 96 wearing military uniform”. The context data for query 104a may contain details about the user's historical feedback on previously generated images. By analyzing these past interactions, the system can adapt the generated image data to more accurately reflect the user's current context.

[0037] In one embodiment, query context data 104a includes training context data 104b. As used in this disclosure, training context data 104b are exemplary contextual data that are used to train a model. The model may include, but is not limited to, the prompt expansion model 106 and the image generation model 110. Training context data 104b are designed to provide the models with exemplary contextual knowledge associated with user interactions. Training context data 104b may include historical versions of query context data 104a associated with user interactions. Training context data 104b may include pairs of forwarded queries and image data, feedback on generated images, and patterns in user behavior. Training context data 104b may be generated and optimized for use as training data.

[0038] With continued reference to Figures 1A-B, system 100 includes prompt expansion model (prompt expansion model) 106. As used in this disclosure, prompt expansion model 106 is a model that is configured to generate an expanded user query 108a-b. Prompt expansion model 106 may be consistent with the machine learning models described herein below in Figures 4-13. Prompt expansion model 106 operates by predicting and generating expanded versions of user query 102a-b based on query context data 104a-b. Prompt expansion model 106 evaluates the data Petition 870250038841, dated 05 / 13 / 2025, page 26 / 142 13 / 96 query context 104a-b to predict and generate expanded versions of the user query that more closely align with the user's contextual data. For example, if query context data 104a-b indicates that the user has previously expressed a preference for bright colors in their queries, prompt expansion model 106 can expand user query 102a-b to include descriptors that reflect this preference.

[0039] In one embodiment, the 106 prompt expansion model may include a neural network architecture. The 106 prompt expansion model may include multiple layers of interconnected nodes or neurons, which are configured to process data in a hierarchical manner. Each layer of the neural network may be responsible for different aspects of the input, enabling the 106 prompt expansion model to learn complex patterns and relationships within the data. User query 102a-b may be processed using these layers where the neural network analyzes the text and identifies key components upon which it can expand.

[0040] The nodes in the prompt expansion model 106 can be arranged in a structured network, such as a convolutional neural network, which includes an input layer of nodes, one or more intermediate layers, and an output layer of nodes. During the training of the prompt expansion model 106, connections between these nodes can be established by applying elements of the training dataset to the nodes. This can include using the multi-reward reinforcement learning model 114 to adjust the connections and weightings between nodes in adjacent layers based on one or more reward criteria 116a-d. Adjustments between the connections and weightings between nodes can be made with the goal of optimizing the prompt expansion model 106 to produce the desired outputs.

[0041] With continued reference to Figures 1A-B, the training Petition 870250038841, dated 05 / 13 / 2025, page 27 / 142 The 14 / 96 assembly of the prompt expansion model 106 and the image generation model 110 includes processing the training query 102b and the training context data 104b to generate an expanded training query 108b. The expanded user query 108a-b is a user query 102a-b that has been augmented to include more details and contextual information. This expanded user query 108a-b can be used as an input to the image generation model 110. By incorporating additional descriptive elements, the expanded user query 108a-b captures user intent more efficiently. This might include adding adjectives to specify colors, incorporating location details, or inferring emotional context based on the user's history.

[0042] The prompt expansion template 106 can be configured to expand the text of user query 102a-b by providing a more detailed and contextually relevant description of the query's subject matter. The prompt expansion template 106 is configured to elaborate on user query 102a-b by incorporating additional descriptive elements such as colors, background details, emotional context, or details related to the subject matter of user query 102ab. Additionally, the prompt expansion template 106 can augment user query 102a-b by incorporating the plurality of reward criteria 116a-d into the expanded user query 108a-b. For example, if the original user query 102a-b emphasizes a specific mood, the prompt expansion template 106 can augment it to include descriptive adjectives and contextual elements that align with the image sentiment reward criteria 116b.

[0043] The expanded user query 108a-b can elaborate on user query 102a-b by adding additional details. In a non-limiting example, if user query 102a-b states "provide an image of a fast car", the user query Petition 870250038841, dated 05 / 13 / 2025, page 28 / 142 The expanded 15 / 96 query 108a-b may include an augmented query that specifies a stylish, lime-green sports car parked near a beach during sunset. This augmented query includes additional details such as color, style, location, and emotional context.

[0044] The augmented text of the expanded user query 108ab can be generated based on an analysis of user query 102a-b context query data 104a-b. The prompt expansion model 106 can evaluate context query data 104a-b such as knowledge from previous interactions, user preferences, and other contextual factors to determine what details should be added to the query. This means that the expanded user query 108a-b can vary significantly based on the user's interests and history. For example, a user who frequently requests nature scenes might receive an expanded user query 108a-b that emphasizes natural elements, such as a tranquil forest with a flowing river, instead of a more generic description.

[0045] In one embodiment, the expanded user query 108a can include an expanded training query 108.The expanded training query 108b is an output of the prompt expansion model 106 during the training phase. The expanded training query 108b is configured to serve as a reference point for training both the prompt expansion model 106 and the image generation model 110. The expanded training query 108b is generated from the training query 102b and the training context data 104b during the joint training process of the prompt expansion model 106 and the image generation model 110. The expanded training query 108b can serve multiple purposes in the training process. The expanded training query 108b can provide a clear target for the image generation model 110 during its learning phase, helping the... Petition 870250038841, dated 05 / 13 / 2025, page 29 / 142 16 / 96 model to understand the types of visual elements to include in its outputs. Additionally, the expanded training query 108b can reinforce the importance of context and detail in query expansion, allowing the model to learn from training examples as it adjusts its parameters. The expanded training query 108b can enable the evaluation of both the prompt expansion model 106 and the image generation model 110 through error analysis.

[0046] With continued reference to Figures 1A-B, system 100 includes image generation model 110. As used in this disclosure, the image generation model is a model that is designed to generate images based on textual inputs and contextual data. Image generation model 110 is configured to interpret textual data to generate image data. The generated image data may include, but are not limited to, the training image dataset 112, the image data subset 202, the Pareto optimal dataset 204, the image data 124, and any other image data discussed in this document. In some embodiments, image generation model 110 may include a large language model that is tuned for image generation. The LLM may use various natural language processing techniques to convert textual prompts into image data. The LLM may be used to generate image data that reflect the textual input.The image generation model can include any of the machine learning models, natural language processing models, image processing models, big language models, and similar models that are discussed in this document below in Figures 4-13.

[0047] With continued reference to Figures 1A-B, jointly train the prompt 106 expansion model and the generation model. Petition 870250038841, dated 05 / 13 / 2025, page 30 / 142 Image generation model 110 includes generating a training image dataset 112 based on the expanded training query 108b. As used in this disclosure, the training image dataset 112 is a collection of images that are used to jointly train or optimize the prompt expansion model 106 and the image generation model 110. The training image dataset 112 includes a plurality of images generated by the image generation model 110 based on the expanded training query 108b. This may include generating multiple iterations of image data using a single expanded training query 108b. Alternatively or additionally, the image generation model 110 may generate multiple iterations of image data using multiple expanded training queries 108b.

[0048] The generation of image data training sets 112 can be achieved by employing techniques such as stochastic sampling or controlled randomness. Stochastic sampling introduces a level of variance within the image generation process. This variance can come in the form of parameter adjustments such as color saturation, lighting, perspective, composition, and the like. This variance can help the image generation model 110 produce different iterations of image data based on an equal or similar expanded training query 108b.

[0049] In one embodiment, each image datum within the training image dataset 112 can be tuned to emphasize one or more reward criteria 116a-d. Each image datum produced by the image generation model 110 can be conceptualised as the result of a decision-making process that weighs these multiple reward criteria 116a-d. For example, when the expanded training query 108b includes of Petition 870250038841, dated 05 / 13 / 2025, page 31 / 142 18 / 96 details on an image of an “urban landscape”, the image generation model 110 can prioritize the reward criterion 116c associated with aesthetic quality in one iteration, leading to bright colors and dramatic lighting. In another iteration, the image generation model 110 can be configured to focus on the text alignment criterion for image 116a. By optimizing parameters associated with reward criteria 116a-d, the image generation model 110 can create images that emphasize or optimize different reward criteria 116a-d.

[0050] Optimizing the parameters associated with the 116a-d reward criteria can enable the 110 image generation model to optimize the image generation process for specific outcomes. The ability to emphasize one or more 116a-d reward criteria can allow the 110 image generation model to produce images that create a balanced representation that satisfies multiple 116a-d reward criteria simultaneously. This may include identifying optimal trade-offs between conflicting 116a-d reward criteria. Optimizing the 110 image generation model for one reward criterion may negatively impact a second reward criterion. In such cases, optimizing the 110 image generation model may involve employing strategies to find the most effective balance.By employing techniques such as multi-objective optimization, the image generation model 110 can iteratively adjust the parameters associated with each reward criterion 116ad. Multi-objective optimization techniques are discussed in greater detail below.

[0051] With continued reference to Figures 1A-B, operations may further include evaluating the plurality of image data according to the plurality of reward criteria 116a-d. As used in this disclosure, the plurality of reward criteria 116a Petition 870250038841, dated 05 / 13 / 2025, page 32 / 142 19 / 96 d are specific metrics or objectives used in reinforcement learning to evaluate the performance of an agent, such as the image generation model 110 or the prompt expansion model 106. These reward criteria 116a-d can be used to guide the training process by providing feedback on the quality of generated outputs. The reward criteria 116a-d can be used to provide measurable objectives that guide agent behavior. This feedback is used to help the image generation model 110 or the prompt expansion model 106 learn which actions to prioritize within the environment. The plurality of reward criteria 116a-d can be used to help evaluate the quality and effectiveness of the generated outputs (i.e., image data and expanded user prompt). This evaluation is subsequently used to direct agent behaviors toward desired outcomes.

[0052] Each criterion of the plurality of reward criteria 116a-d can serve as a distinct lens through which generated images or expanded prompts can be evaluated. Agents receive rewards or penalties based on their actions. These rewards or penalties are used to influence agents regarding the success or failure of such actions. For the 110 image generation model or the 106 prompt expansion model, defining clear 116a-d reward criteria enables the models to quantify a “good” image or expanded prompt versus a “less effective” one. This evaluation process can fine-tune the models to identify which features to prioritize or modify in future iterations.

[0053] Exemplary forms of the plurality of reward criteria 116a-d may include text-image alignment reward criteria 116a. As used in this disclosure, text-image alignment reward criteria 116a are Petition 870250038841, dated 05 / 13 / 2025, page 33 / 142 20 / 96 refers to a set of evaluation metrics used to assess how well a generated image matches a given textual description or prompt. The text-image alignment reward criteria 116a can be used to evaluate the relationship between generated images and their associated textual prompts. These textual prompts may include the user query 102a-b or the expanded user query 108a-b. The text-image alignment reward criteria 116a may include considerations of the overall context and tone of the text. The text-image alignment reward criteria 116a may evaluate the textual input and the generated image within the context of the query context data 104a-b. For example, if the textual prompt describes an animated frog dancing on a water lily, the image should include identifiable features such as the frog, the dance pose, and the water lily.The text-image alignment reward criteria 116a can be used to ensure that the elements and actions described textually are visually present.

[0054] In one embodiment, the text-image alignment reward criteria 116a may include first text-image alignment reward criteria 116a and second text-image alignment reward criteria 116a. The first text-image alignment reward criteria 116a may be associated with the relationship between the generated image and the original user query 102a-b. In contrast, second text-image alignment reward criteria 116a may be associated with the expanded user query 108a-b. The second text-image alignment reward criteria 116a may be used to evaluate how efficiently the generated image aligns with the broader context of the expanded user query 108a-b. The first and second text-image alignment reward criteria 116a may be used to ensure that the generated image captures the elements Petition 870250038841, dated 05 / 13 / 2025, page 34 / 142 21 / 96 key and themes explicitly mentioned in the written text of user query 102a-b and expanded user query 108a-b.

[0055] Exemplary forms of the plurality of reward criteria 116a-d may include one or more image sentiment reward criteria 116b. As used in this disclosure, image sentiment reward criteria 116b are used to evaluate the emotional tone and mood conveyed by generated images in relation to their corresponding textual descriptions or prompts. Image sentiment reward criteria 116b may be used to ensure that the generated image data reflect the emotional context that was described in the textual prompts. In some cases, image sentiment reward criteria 116b may be an evaluation of specific visual cues that contribute to the overall sentiment of an image. These cues may include factors such as color palettes, facial expressions, and composition.For example, bright colors and cheerful facial expressions are generally considered to be associated with feelings of joy, while muted tones and solemn postures may suggest sadness or contemplation.

[0056] Exemplary forms of the plurality of reward criteria 116a-d may include one or more aesthetic reward criteria 116c. As used in this disclosure, aesthetic reward criteria 116c are reward criteria that are used to evaluate the visual quality and artistic appeal of generated images. Aesthetic reward criteria 116c may focus on various elements that contribute to the overall appearance and attractiveness of an image. This may include considerations such as composition, color harmony, lighting, detail and texture, art style, and the like. Aesthetic reward criteria 116c may include an assessment of how well the elements within the image are arranged. Petition 870250038841, dated 05 / 13 / 2025, page 35 / 142 22 / 96 and balanced or the use of color within the image.

[0057] In one modality, the aesthetic reward criteria 116c can be derived from historical classifications of the aesthetic qualities of real images. System 100 can generate the 116c aesthetic reward criteria from various subjective evaluations of the aesthetic appeal of image data. These evaluations can be analyzed to identify patterns in how aesthetic qualities are perceived. These annotated datasets can be used to train a machine learning model such as the third 122c reward model or any other model disclosed in this paper. By applying techniques such as supervised learning, the machine learning model can develop an understanding of factors such as composition, color harmony, and texture that resonate with viewers. This trained model can then be used to generate the 116c aesthetic reward criteria.

[0058] Exemplary modalities of the plurality of 116a-d reward criteria may include one or more 116d human preference reward criteria. 116d human preference reward criteria can be used to evaluate images generated based on a dataset composed of human feedback. This can be done to ensure that the output closely aligns with what individuals find attractive or desirable. Generating 116d human preference reward criteria may involve collecting user feedback through surveys, ratings, or other interactive methods. This feedback can then be aggregated to create a dataset that reflects the diverse preferences of the audience. This dataset will include a plurality of text-image pairs that reflect user preferences. When this annotated dataset is established, Petition 870250038841, dated 05 / 13 / 2025, page 36 / 142 23 / 96 machine learning models, such as the 122d fourth reward model, can be applied to analyze data and extract meaningful insights into user preferences. The 122d fourth reward model can be used to identify and score key features (i.e., color schemes, compositions, and matter) based on the annotated dataset. In some modalities, this annotated dataset can be used as training data by the 122d fourth reward model.

[0059] With continued reference to Figures 1A-B, operations may further include generating a set of reward scores 120a-d for each image datum within the training image dataset 112 using a set of reward models 122a-d. As used in this disclosure, a reward score is a quantitative measure that reflects the effectiveness of an action or decision taken by an agent in achieving specific goals, such as the plurality of reward criteria 116a-d. Agents are configured to optimize their behavior by receiving feedback based on the plurality of reward criteria 116a-d. This feedback is quantified by the set of reward scores 120a-d. Each reward score in the set of reward scores 120a-d represents a quantification of a distinct aspect of the agents' behavior within the environment.For example, a reward score might reflect how well a generated image aligns with a textual prompt.

[0060] The 120a-d reward scoring system can facilitate the refinement of agents' decision-making processes within the environment. By receiving quantitative feedback on multiple 116a-d reward criteria, agents can learn to identify patterns and make informed choices that enhance their overall effectiveness. This encourages agents to optimize their strategies with Petition 870250038841, dated 05 / 13 / 2025, page 37 / 142 24 / 96 based on its impact on the 120ad reward score set. Refinement encourages agents to capitalize on successful strategies that deliver higher scores.

[0061] In one embodiment, generating the set of reward scores 120a-d includes producing at least one distinct reward score 120 for each reward criterion within the plurality of reward criteria 116a-d. The first reward score 120a is associated with the text-image alignment reward criteria 116a. The first reward score 120a is a quantification of how efficiently the generated image matches the textual prompt (i.e., user query 102a-b or the expanded user query 108a-b). The second reward score 120b corresponds to the image sentiment reward criteria 116b. The second reward score 120b represents a quantification of the alignment of the emotional tone conveyed by the generated images with the textual input. The third reward score 120c is associated with the aesthetic reward criteria 116c.The third reward score, 120c, quantifies the visual quality and artistic appeal of the generated images. The fourth reward score, 120d, is associated with the human preference reward criteria, 116d. The fourth reward score, 120d, quantifies how well the generated images align with user preferences and tastes.

[0062] Each reward score within the 120a-d reward score set can be generated using a reward model from the 122a-d reward model set. Each reward model from the 122a-d reward model set can be fitted to a specific 116a-d reward criterion within the plurality of 116a-d reward criteria. This can enable the model to analyze and score images effectively. Each model Petition 870250038841, dated 05 / 13 / 2025, page 38 / 142 25 / 96 of the reward within the 122a-d reward model set may be equal to or substantially similar to the machine learning models, natural language processing models, image processing models, big language models, and the like that are discussed in this document below in Figures 413. For example, the first 122a reward model is associated with the text-image alignment reward criteria 116a. Therefore, the first 122a reward model may employ one or more natural language processing techniques and image processing techniques to evaluate how well a generated image matches its attached textual prompt. The first 122a reward model may analyze semantic relationships and visual features to produce the first 120a reward score.

[0063] Similarly, a second reward model 122b can be dedicated to the image sentiment reward criteria 116b. The second reward model 122b can employ machine learning algorithms trained in sentiment analysis to measure the emotional tone of the images. The second reward model 122b can evaluate image features such as color, expression, and overall composition to quantify how well the images align with the desired sentiment of the textual prompt. The third reward model 122c can be associated with the aesthetic reward criteria 116c. The third reward model 122c can use an image processing model to evaluate the artistic quality of the image data. The fourth reward model 122d is associated with the human preference reward criteria 116d.The fourth 122d reward model can be trained on a dataset consisting of user preference data and feedback to inform its score.

[0064] With continued reference to Figures 1A-B, the operations Petition 870250038841, dated 05 / 13 / 2025, pp. 39 / 142 26 / 96 may also include adjusting weights and trends 118 associated with the plurality of reward criteria 116a-d based on the set of reward scores 122a-d using the multi-reward reinforcement learning model 114. As used in this disclosure, the multi-reward reinforcement learning model 114 is a model that is designed to optimize the performance of agents in generating outputs based on multiple reward criteria. The multi-reward reinforcement learning model 114 is configured to adjust the trends and weights 118 associated with the plurality of reward criteria 116a-d based on the set of reward scores 122a-d. The multi-reward reinforcement learning model 114 is configured to consider the performance of multiple agents across multiple reward criteria 116a-d.The multi-reward reinforcement learning model 114 can be consistent with the machine learning models described in this document below in Figures 4-13. In one embodiment, the multi-reward reinforcement learning model 114 can be configured to identify optimal trade-offs among the plurality of reward criteria 116a-d. Based on these optimal trade-offs, the multi-reward reinforcement learning model 114 can optimize the trends and weightings 118 associated with the plurality of reward criteria 116a-d.

[0065] The multiple reward reinforcement learning model 114 can be used to fine-tune the biases and weights 118 that determine the relative importance of each reward criterion. Weights can be used to determine the degree of influence each reward criterion has on the final decision-making process. Bias is a term associated with adjustments that can shift the output in a desired direction. Adjusting these weights Petition 870250038841, dated 05 / 13 / 2025, pages 40 / 142 27 / 96 and trends 118 can be informed by the set of reward scores 120a-d generated from the reward models 122ad. By analyzing agent performance based on these scores, the multi-reward reinforcement learning model 114 can identify which reward criteria 116a-d are contributing positively to achieving the desired outcomes, and which may require recalibration. For example, if the text-image alignment reward score 120a consistently provides high scores, while the aesthetic reward score 120c is insufficient, the multi-reward reinforcement learning model 114 can maximize the weighting assigned to the text-image alignment reward criterion 116a while minimizing the weighting for the aesthetic reward criteria 116c.Adjusting weights and trends 118 may involve an iterative learning process in which the multi-reward reinforcement learning model 114 evaluates the effects of these adjustments on subsequent sets of reward scores 120a-d. By employing techniques such as gradient descent or other optimization algorithms, the multi-reward reinforcement learning model 114 can systematically refine its parameters to improve performance.

[0066] With continued reference to Figure 1B, once the training of the prompt expansion model 106 and the image generation model 110 has been completed, the system 100 can be used to generate image data 124 based on the expanded user query 108a. Generating the image data 124 may include converting the user query 108a into the expanded user query 108a using the trained prompt expansion model 106, as described in this document above. The expanded user query 108a can be used as an input in the trained image generation model. Petition 870250038841, dated 05 / 13 / 2025, pages 41 / 142 28 / 96 110. The image generation model 110 is then configured to generate image data 124 that are aligned with the reward criteria 116a-d. This can be done using any of the image generation processes discussed in this document.

[0067] In some modes, the prompt expansion model 106 and the image generation model 110 can be frozen when the training process is complete. During the training phase, both agents can continuously adjust their behavior based on feedback received from the reward score set 122a-d. When the agents have sufficiently learned the optimal strategies, a convergence test can be employed to assess whether the agents' performance has stabilized and reached an acceptable level of consistency in their outputs. When the agents pass this convergence test, indicating that their learning is stabilized and that further adjustments are unlikely to provide significant improvements, they can be frozen. Freezing the agents may include locking their weightings and trends in place, effectively halting any further updates to their parameters.

[0068] With reference now to Figure 2, an exemplary representation of a block diagram of an exemplary system for generating image data using an image generation model according to exemplary embodiments of the present disclosure. Figure 2 includes a subset of image data 202, Pareto optimal set 204, non-dominated classification algorithm 206, policy gradient update 208 and the like.

[0069] In one embodiment, training the prompt expansion model 106 and the image generation model 110 includes selecting a subset of image data 202 from the dataset of Petition 870250038841, dated 05 / 13 / 2025, pp. 42 / 142 29 / 96 training image 112 as a function of the reward score set 120a-d using a non-dominated classification algorithm 206. As used in this disclosure, the image data subset 202 is a collection of images selected from the larger training image dataset 112. The image data subset 202 can be chosen based on the performance of the image generation model 110 as quantified by the reward score set 120a-d. The image data subset 202 can be characterized by its representation of the optimal trade-offs among the various reward scores 120a-d.

[0070] In one embodiment, the image data subset 202 can be represented as a Pareto optimal set 204. As used in this disclosure, a Pareto optimal set 204 refers to a collection of solutions in a multi-objective optimization problem in which no single solution can be improved in one criterion without degrading performance in another. The Pareto optimal set 204 can represent a subset of image data 202 that exhibits optimal trade-offs among various reward scores 120a-d. The image data within the Pareto optimal set 204 can be considered optimal because they reflect balanced performance across two or more reward criteria 116a-d. In a non-limiting example, it is assumed that a first image excels in a first reward criterion but is weaker in a second reward criterion.A second image displays the opposing strengths; both the first and second images can be selected to be part of the Pareto optimal set 204.

[0071] Optimal trade-offs between various reward scores 120a-d may refer to identifying the best possible outcomes. Petition 870250038841, dated 05 / 13 / 2025, pages 43 / 142 30 / 96 through multiple competing reward criteria 116a-d. Each reward score with the set of reward scores 120a-d can represent a quantification of the performance of the prompt expansion model 106 and the image generation model 110 while generating image data. One of the main challenges in identifying and producing optimal trade-offs is managing the competing reward scores 120. Optimal trade-offs between multiple competing reward criteria 116a-d can be identified when improving one reward score frequently leads to a reduction in another reward score.

[0072] The non-dominated classification algorithm 206 can be used to identify optimal trade-offs between image data based on their respective reward scores 120a-d. The non-dominated classification algorithm 206 can be used to analyze how changes in one reward score impact the remaining reward scores. As used in this disclosure, the non-dominated classification algorithm 206 is a technique used in multi-objective optimization to categorize and classify solutions based on their performance across multiple reward criteria. The non-dominated classification algorithm 206 is configured to identify one or more unnamed solutions within the training image dataset 112. An image can be considered non-dominated if there is no other image with better performance across two or more reward criteria.In a non-limiting example, if Image A excels in a first reward criterion while Image B is superior in a second reward criterion, neither image dominates the other, since both have strengths in different areas.

[0073] The non-dominated classification algorithm 206 can be configured to evaluate the relationship between the images within the set. Petition 870250038841, dated 05 / 13 / 2025, pages 44 / 142 31 / 96 of training image data 112. This may include classifying images based on their overall performance as reflected by the set of reward scores 120a-d. Exemplary categories may include, but are not limited to, non-dominated categories, text-image alignment categories, image sentiment categories, aesthetic categories, human preference categories, and the like. A non-dominated category may be a category that is used to represent the best-performing solutions across a defined set of reward criteria 116a-d. Subsequent categories may be reserved for images that perform well across one or more reward criteria 116a-d.

[0074] When images have been ordered or classified using the non-dominated classification algorithm 206, the non-dominated classification algorithm 206 can select the image data subset 202 based on these identified exchanges. The image data subset 202 can be identified from the non-dominated category. This category can represent the most optimal exchanges between the image data.

[0075] With continued reference to Figure 2, training the prompt expansion model 106 and the image generation model 110 may include adjusting the weightings and trends 118 associated with the plurality of reward criteria 116a-d based on the image data subset 202 using the multi-reward reinforcement learning model 114. The multi-reward reinforcement learning model 114 can be configured to identify specific weightings and trends 118 that need to be adjusted to fine-tune the prompt expansion model 106 and the image generation model 110 based on the image data subset 202 or the Pareto optimal set 204. The weightings and trends Petition 870250038841, dated 05 / 13 / 2025, pages 45 / 142 32 / 96 118 can be fine-tuned based on the optimal trade-offs identified between various reward scores. The multi-reward reinforcement learning model 114 can be configured to examine the characteristics of these images within the image data subset 202 or the Pareto optimal set 204. Based on this examination, the multi-reward reinforcement learning model 114 can determine which features contribute to the most favorable outcomes. In a non-limiting example, if the images represented within the Pareto optimal set 204 demonstrate strong alignment with prompts and effective emotional expression, the multi-reward reinforcement learning model 114 can adjust the weightings associated with these dimensions. This can be done with the goal of improving the generated image data through subsequent iterations.

[0076] With continued reference to Figure 2, training the prompt expansion model 106 and the image generation model 110 can include determining a policy gradient update 208 as a function of the image data subset 202. As used in this disclosure, the policy gradient update 208 is used to optimize agent policies. The policy gradient update 208 can be used to help agents optimize their interaction within the environment to maximize / minimize the reward score set 120a-d. The policy gradient update 208 can be created based on a series of evaluations of the model's performance against the image data subset 202 or the Pareto optimal set 204. The image data subset 202 can be used as a benchmark when making a determination of how well the agents' actions align with desired outcomes.By analyzing the reward scores 116a-d associated with the image data subset 202, the learning model... Petition 870250038841, dated 05 / 13 / 2025, pp. 46 / 142 33 / 96 by reinforcement of multiple rewards 114 can make a determination of gradients that indicate how policy changes would impact future performance. In one embodiment, the reinforcement learning model of multiple rewards 114 can evaluate reward scores 116a-d for actions performed based on the image data subset 202. Gradients can be generated based on the contribution of each image data to the set of reward scores 116a-d.

[0077] When gradients are computed, policy gradient update 208 is applied to adjust policy parameters in the direction that maximizes / minimizes the identified reward criteria 116a-d. These adjustments can be made to encourage agents to explore and select actions that lead to high-reward outcomes. Additionally, by optimizing the trends and weightings associated with the reward criteria, agents can learn to make optimal trade-offs between competing reward criteria. This process can be performed iteratively, meaning that the policy gradient can be updated until agents can successfully pass a convergence test.

[0078] In one embodiment, determining the policy gradient update 208 may involve minimizing the reward scores 116a-d associated with each reward criterion that is not adequately represented within the image data subset 202 or the Pareto optimal set 204. By identifying these underrepresented reward scores, the multi-reward reinforcement learning model 114 can develop a targeted policy focused on improving performance through the unrepresented reward criteria. To achieve this minimization, the multi-reward reinforcement learning model 114 can employ Petition 870250038841, dated 05 / 13 / 2025, pages 47 / 142 34 / 96 using a gradient descent approach, incorporate the gradients of the underrepresented reward scores 120a-d into the overall policy gradient update. By actively minimizing these scores, agents learn to address performance gaps.

[0079] In an additional embodiment, determining the policy gradient update 208 may also include maximizing the reward scores 120a-d associated with each reward criterion 116ad that is present within the image data subset 202. This can be done to encourage agents to capitalize on their strengths by reinforcing positive behaviors that lead to favorable reward scores 120a-d. To achieve this maximization, the multi-reward reinforcement learning model 114 can identify the reward scores 120a-d that correspond to the reward criteria reflected in the image data subset 202. By analyzing the successful image outputs, the multi-reward reinforcement learning model 114 can ascertain which features and actions provide superior rewards.This understanding allows the multi-reward reinforcement learning model 114 to adjust its policy parameters in a way that favors actions that lead to similar successful outcomes in future iterations. In some cases, the process of maximizing these reward scores 120a-d may include applying techniques such as stochastic gradient ascent. By calculating the gradients of the represented reward scores 120a-d and integrating them into the policy gradient update, the multi-reward reinforcement learning model 114 can make informed adjustments that amplify the probability of producing desirable outputs.

[0080] With reference now to Figure 3, it represents a flowchart Petition 870250038841, dated 05 / 13 / 2025, pages 48 / 142 35 / 96 of a method 300 for performing multiple reward reinforcement learning for generative text-to-image models according to exemplary embodiments of the present disclosure. The method 300 can be performed by processing logic that may include hardware (e.g., processing device, circuit set, dedicated logic, programmable logic, microcode, device hardware, integrated circuit, etc.), software (e.g., instructions run or executed on a processing device), or a combination thereof. In some embodiments, the method is performed by a server computing system (e.g., server computing system 60) or a client computing system (e.g., client computing device 50). Although shown in a particular sequence or order, unless otherwise specified, the order of the processes may be modified.Thus, the modalities illustrated should be understood only as examples, and the processes illustrated may be executed in a different order, and some processes may be executed in parallel. Furthermore, one or more processors may be omitted in various modalities. Therefore, not all processes are necessary in all modalities. Other process flows are possible.

[0081] In operation 302, the processing logic can train, in tandem, a prompt expansion model (prompt expansion model) and an image generation model (image generation models) using a multi-reward reinforcement learning model. The joint training of the prompt expansion model and the image generation models includes iteratively updating the model parameters based on feedback from the plurality of reward criteria. This feedback is used to help the image generation models and the prompt expansion model learn which actions to prioritize within the environment. Petition 870250038841, dated 05 / 13 / 2025, pp. 49 / 142 36 / 96 based on the quality and effectiveness of the generated outputs (i.e., image data and expanded user prompt).

[0082] In operation 304, training the prompt expansion model and the image generation model involves processing, by the prompt expansion model, a training query and training context data to generate an expanded training query. The prompt expansion model is configured to augment the training query with additional contextual and descriptive elements to generate the expanded training query. Both the training query and the training context data are exemplary representations of the user query and the query context data that are used as training data for the prompt expansion model. Similar to the training query and training context data, the expanded training query is an exemplary representation of the expanded user query that can be used to train the image generation models.

[0083] In operation 306, training the prompt expansion model and the image generation model includes generating, by the image generation model, a training image dataset based on the expanded training query. The image generation models are configured to generate the training image dataset from the textual input of the expanded training query. The training image dataset is a collection of generated images that are exemplary representations of potential outputs from the image generation models. The training image dataset may represent optimal trade-offs between two or more reward criteria.

[0084] In operation 308, training the prompt expansion model and the image generation model includes generating a set of reward scores for each image datum within the set of Petition 870250038841, dated 05 / 13 / 2025, page 50 / 142 37 / 96 training image data using a set of reward models, wherein generating the set of reward scores includes generating at least one reward score for each reward criterion within a plurality of reward criteria. In one embodiment, the plurality of reward criteria includes one or more text-image alignment reward criteria, one or more image sentiment reward criteria, one or more aesthetic reward criteria, or one or more human preference reward criteria. In some cases, the one or more text-image alignment reward criteria may include: first text-image alignment reward criteria associated with the training query; and second text-image alignment reward criteria associated with the expanded training query.

[0085] The quality of the expanded training query and the training image dataset can be represented by the reward score set. The reward score set are quantifications of how well the expanded training query and the training image dataset align with the reward criteria set. These reward scores can be used as feedback to both the prompt expansion model and the image generation models, respectively.

[0086] In operation 310, training the prompt expansion model and the image generation model includes adjusting weights and trends associated with the plurality of reward criteria based on the set of reward scores using the multi-reward reinforcement learning model. Both the prompt expansion model and the image generation models can be Petition 870250038841, dated 05 / 13 / 2025, page 51 / 142 38 / 96 finely tuned based on optimal trade-offs between various reward scores present within the training image dataset. The multi-reward reinforcement learning model is used to examine the characteristics of the training image dataset to determine which features or actions contribute to the most favorable outcomes. Based on these optimal trade-offs, the multi-reward reinforcement learning model can simultaneously adjust the weightings and trends associated with the reward criteria for both the prompt expansion model and the image generation models.

[0087] In some implementations, the processing logic may obtain input data that includes a user query and query context data. In some implementations, the processing logic may generate, using the trained prompt expansion model and the trained image generation model, image data based on the user query and query context data. For example, the processing logic may generate the image data by processing, using the trained prompt expansion model, the user query and query context data to generate an expanded user query. For example, the processing logic may generate the image data by generating, using the trained image generation model, the image data based on the expanded user query. In some cases, the expanded user query may be generated by incorporating the plurality of reward criteria in the expanded user query.

[0088] In some implementations, the processing logic may train the prompt expansion model and the image generation model, including selecting a subset of image data. Petition 870250038841, dated 05 / 13 / 2025, page 52 / 142 39 / 96 from the training image dataset as a function of the reward score set using a non-dominated classification algorithm. In some implementations, the processing logic may adjust the weightings and trends associated with the plurality of reward criteria based on the image dataset subset using the multi-reward reinforcement learning model. In some implementations, the image dataset subset may include a Pareto optimum associated with the reward score set.

[0089] In some embodiments, training the prompt expansion model and the image generation model may include determining, by the computing system, a policy gradient update as a function of the image data subset. Training the prompt expansion model and the image generation model may include adjusting the weightings and trends associated with the plurality of reward criteria using the multi-reward reinforcement learning model as a function of the policy gradient update. In some cases, determining the policy gradient update may include minimizing the reward scores associated with each reward criterion that is not represented within the image data subset. Additionally, determining the policy gradient update may include maximizing the reward scores associated with each reward criterion that is represented within the image data subset.Furthermore, the policy gradient update may include maximizing one or more text-image alignment reward criteria.

[0090] Figure 4 represents a flowchart of a 400 method for training one or more machine learning models in accordance with aspects of the present disclosure. For example, an example model Petition 870250038841, dated 05 / 13 / 2025, page 53 / 142 40 / 96 mplificative machine learning may include the prompt expansion model 106, image generation model 110, multi-reward reinforcement learning model 114, Reward Model Set 122a-d, along with any other model or algorithm mentioned herein.

[0091] Method 400 can be performed by processing logic that may include hardware (e.g., processing device, circuit set, dedicated logic, programmable logic, microcode, device hardware, integrated circuit, etc.), software (e.g., instructions run or executed on a processing device), or a combination thereof. In some embodiments, the method is performed by a server computing system (e.g., server computing system 60) or a client computing system (e.g., client computing device 50). Although shown in a particular sequence or order, unless otherwise specified, the order of the processes may be modified. Thus, the embodiments illustrated should be understood as examples only, and the processes illustrated may be executed in a different order, and some processes may be executed in parallel.Furthermore, one or more processors can be omitted in various modes. Therefore, not all processes are required in all modes. Other process flows are possible.

[0092] In operation 402, the processing logic can obtain a training case. A training dataset can include a plurality of training cases spread across multiple datasets (e.g., a training dataset, a validation dataset, or a test dataset). A training case can be labeled or unlabeled. Although deno Petition 870250038841, dated 05 / 13 / 2025, pp. 54 / 142 41 / 96 mined in example method 400 as a “training” instance, it should be understood that runtime inferences can form training instances when a model is trained using an evaluation of the model’s performance in that runtime instance (e.g., online training / learning). Exemplary data types for the training case and various tasks associated with it are described throughout this disclosure.

[0093] In the 404 operation, the processing logic can process, using one or more machine learning models, the training case to generate an output. The output can be directly obtained from one or more machine learning models or can be a downstream result of a chain of processing operations that includes an output from one or more machine learning models.

[0094] In operation 406, the processing logic can receive an evaluation signal associated with the output. The evaluation signal can be obtained using a loss function. Several loss determinations can be used, such as mean squared error, probability loss, cross-entropy loss, articulation loss, or various other loss functions. The evaluation signal can be computed using known basic truth labels (e.g., supervised learning), predicted or estimated labels (e.g., semi- or self-supervised learning), or no labels (e.g., unsupervised learning). The evaluation signal can be a reward (e.g., for reinforcement learning). The reward can be computed using a machine learning reward model configured to generate rewards based on the output(s) received. The reward can be computed using feedback data that describes human feedback on the output (or Petition 870250038841, dated 05 / 13 / 2025, page 55 / 142 42 / 96 exits).

[0095] In operation 408, the processing logic can update the machine learning model using the evaluation signal. For example, values ​​for parameters of the machine learning model (or models) can be learned, in some modes, using various training or learning techniques, such as backpropagation. For example, the evaluation signal can be backpropagated from the output (or other source of the evaluation signal) through the machine learning model (or models) to update one or more parameters of the model (or models) (e.g., based on a gradient of the evaluation signal relative to the parameter value (or values)). For example, system(s) containing one or more machine learning models can be trained end-to-end. Gradient descent techniques can be used to iteratively update the parameters over several training iterations.In some implementations, performing backward error propagation may include performing truncated backpropagation through time. The exemplary 400 method may include implementing various generalization techniques (e.g., weight decays, dropouts, etc.) to enhance the generalization capability of the models being trained.

[0096] In some implementations, the exemplary method 400 can be implemented to train a machine learning model from an initialized state to a fully trained state (for example, when the model exhibits a desired performance profile, such as based on accuracy, precision, recall, etc.).

[0097] In some implementations, the exemplary method 400 can be implemented for particular stages of a training procedure. For example, in some implementations, the exemplary method 400 can be implemented for pre-training a Petition 870250038841, dated 05 / 13 / 2025, page 56 / 142 43 / 96 machine learning model. Pre-training may include, for example, large-scale training using potentially noisy data to achieve a broad range of performance levels across a variety of tasks / data types.

[0098] In some implementations, the exemplary method 400 can be implemented to fine-tune a machine learning model. Fine-tuning may include, for example, smaller-scale training on higher-quality data (e.g., labeled, curated, etc.). Fine-tuning can affect all or a portion of the parameters of a machine learning model. For example, various portions of the machine learning model can be “frozen” to certain training stages. For example, parameters associated with an embedding space can be “frozen” during fine-tuning (e.g., to retain information learned from a broader domain (or domains) than that present in the fine-tuning dataset(s)). In some implementations, the exemplary method 400 uses adapter modules. Adapters can be small trainable layers that are inserted between pre-existing layers of a pre-trained model.During the fine-tuning process, the original parameters of the pre-trained model are typically frozen, and only the parameters of the adapters are updated.

[0099] In some implementations, the exemplary method 400 can be implemented to perform parameter-efficient fine-tuning methods, such as Residual Layer Optimization (LoRA). LoRA can refine pre-trained models with minimal adjustments to the original parameters. This can be achieved by introducing trainable low-degree matrices that modify the behavior of the pre-trained weights without directly altering them. In some Petition 870250038841, dated 05 / 13 / 2025, page 57 / 142 In 44 / 96 implementations, during fine-tuning, only these auxiliary matrices are updated, which significantly reduces the number of parameters that are trained.

[0100] An exemplary fine-tuning approach includes reinforcement learning. Reinforcement learning can be based on user feedback on model performance during use.

[0101] Figure 5 is a block diagram of an exemplary processing flow for using machine learning model (or models) 1 to process input (or inputs) 2 to generate output (or outputs) 3.

[0102] The machine learning model (or models) 1 may be or include one or multiple machine learning models or model components. Exemplary machine learning models may include neural networks (e.g., deep neural networks). Examples of machine learning models may include nonlinear models or linear models. Exemplary machine learning models may use other architectures instead of or in addition to neural networks. Exemplary machine learning models may include decision tree-based models, support vector machines, hidden Markov models, Bayesian networks, linear regression models, k-means clustering models, etc.

[0103] The machine learning model (or models) 1 may be or include or otherwise be representative of any one or more of the machine learning models described above in relation to the preceding Figures. For example, the machine learning model (or models) 1 may be or include or otherwise be representative of any one or more of the prompt expansion model 106, the image generation model 110, the model of Petition 870250038841, dated 05 / 13 / 2025, pages 58 / 142 45 / 96 multi-reward reinforcement learning 114, the Reward Model Set 122a-d, together with any other model or algorithm mentioned herein. Although various features, variations, and implementations described below are described in relation to machine learning model (or models) 1, it should be understood that such features, variations, and implementations should be understood as described in relation to each of the prompt expansion model 106, the image generation model 110, the multi-reward reinforcement learning model 114, the Reward Model Set 122a-d, together with any other machine learning model, algorithm, or component described in this document.

[0104] Exemplary neural networks may include feedforward neural networks, recurrent neural networks (RNNs), including long-term and short-term memory (LSTM) based recurrent neural networks, convolutional neural networks (CNNs), diffusion models, generative adversarial networks, or other forms of neural networks. Examples of neural networks may be deep neural networks. Some examples of machine learning models may leverage an attention mechanism, such as self-attention. For example, some exemplary machine learning models may include multi-head self-attention models.

[0105] Machine learning model (or models) 1 may include a single or multiple instances of the same model configured to operate on input data (or inputs) 2. Machine learning model (or models) 1 may include multiple different models or multiple different model portions configured to operate on input data (or inputs) 2.

[0106] The machine learning model (or models) 1 may include a set of different models that can interact in a way Petition 870250038841, dated 05 / 13 / 2025, pp. 59 / 142 46 / 96 Cooperative for processing input data (or inputs) 2. For example, a model set may include multiple models that have different attributes (e.g., different architectures, trained with different recipes, etc.). The set may output an overall output based on the individual outputs of the constituent models. In this way, for example, the various constituent models may work together to provide system-level robustness by effectively aggregating individual strengths and weaknesses of any given model. The respective individual outputs may be combined into a weighted combination, using a voting or routing mechanism, or a learning output layer (e.g., one or more feedforward or fully connected layers).

[0107] The machine learning model (or models) 1 may employ a mixture-of-experts structure. See, for example, Zhou et al., Mixture-of-Experts with Expert Choice Routing, arXiv:2202.09368v2 (October 14, 2022). For example, different portions of a model may learn (explicitly or implicitly) different areas of expertise, with paths through the model being selected by a learning routing mechanism that engages the appropriate expert for a given input (e.g., a given portion of an input, such as token-based). For example, a feedforward network may be sparsely activated for a given portion of an input based on an output from a routing mechanism that processes the portion of the input. In this way, for example, the group of activated weightings may form an “expert” that is selected by the router.In each pass forward, only a subset of the total model weights can be engaged, thus reducing the number of operations performed to process a given input compared to a densely activated model. In this way, for example, the power. Petition 870250038841, dated 05 / 13 / 2025, pages 60 / 142 47 / 96 expressive and interpretative of a high-parameter counting model can be achieved with more computationally efficient forward passes.

[0108] Input (or inputs) 2 may generally include, or otherwise represent, various data types. Input (or inputs) 2 may include one type or many different types of data. Output (or outputs) 3 may consist of data of the same type (or types) or different types of data compared to input (or inputs) 2. Output (or outputs) 3 may include one type or many different types of data.

[0109] Exemplary data types for input(s) 2 or output(s) 3 include natural language text data, software code data (e.g., source code, object code, machine code, or any other form of computer-readable instructions or programming languages), machine code data (e.g., binary code, assembly code, or other forms of machine-readable instructions that can be executed directly by a computer's central processing unit), assembly code data (e.g., low-level programming languages ​​that use symbolic representations of machine code instructions to program a processing unit), genetic data or other chemical or biochemical data, image data, audio data, audiovisual data, haptic data, biometric data, medical data, financial data, statistical data, geographic data, astronomical data, historical data,Sensor data generally (e.g., digital or analog values, such as voltage or other absolute or relative level measurement values ​​from a real or artificial input, such as from an audio sensor, light sensor, displacement sensor, etc.) and similar. The data... Petition 870250038841, dated 05 / 13 / 2025, pp. 61 / 142 48 / 96 of the data can be raw or processed and can be in any format or schema.

[0110] In multimodal inputs 2 and outputs 3, exemplary combinations of data types include image data and audio data, image data and natural language data, natural language data and software code data, image data and biometric data, sensor data and medical data, etc. It should be understood that any combination of data types in an input 2 or an output 3 may be present.

[0111] An example input 2 may include one or more data types, such as the example data types noted above. An example output 3 may include one or more data types, such as the example data types noted above. The data type(s) of input 2 may be the same as or different from the data type(s) of output 3. It should be understood that the example data types noted above are provided for illustrative purposes only. The data types contemplated within the scope of this disclosure are not limited to those examples noted above.

[0112] Figure 6 is a block diagram of an exemplary implementation of an exemplary machine learning model configured to process sequences of information. For example, an exemplary implementation of machine learning model (or models) 1 might include machine learning sequence processing model (or models) 4. An exemplary system might pass input (or inputs) 2 to sequence processing model (or models). Sequence processing model (or models) 4 might include one or more machine learning components. The sequence processing model (or models) Petition 870250038841, dated 05 / 13 / 2025, pages 62 / 142 49 / 96 Sequence processing model 4 can process the input data (or inputs) 2 to obtain an input sequence 5. The input sequence 5 can include one or more input elements 5-1, 5-2, ..., 5M, etc. obtained from the input (or inputs) 2. The sequence processing model 4 can process the input sequence 5 using prediction layer (or layers) 6 to generate an output sequence 7. The output sequence 7 can include one or more output elements 7-1, 7-2, ..., 7-N, etc. generated based on the input sequence 5. The system can generate output (or outputs) 3 based on the output sequence 7.

[0113] The sequence processing model (or models) 4 may include one or multiple machine learning model components configured to ingest, generate, or otherwise rationalize with respect to information sequences. For example, some example sequence processing models in the text domain are called “Large Language Models” or LLMs. See, for example, PaLM 2 Technical Report, Google, https: / / ai.google / static / documents / palm2techreport.pdf (nd). Other example sequence processing models can operate in other domains, such as image domains, see, for example, Dosovitskiy et al., An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale, arXiv:2010.11929v2 (June 3, 2021), audio domains, see, for example, Agostinelli et al., MusicLM: Generating Music From Text, ARXIV:2301.11325v1 (January 26, 2023), biochemical domains, see, for example, Jumper et al., Highly accurate protein structure prediction with AlphaFold, 596 Nature 583 (August 26, 2021), as an example. The sequence processing model(s) 4 can process one or more data types simultaneously. The sequence processing model(s) 4 can include relatively large models (e.g., Petition 870250038841, dated 05 / 13 / 2025, pages 63 / 142 50 / 96 more parameters, computationally expensive, etc.), relatively small models (e.g., fewer parameters, computationally lightweight, etc.), or both.

[0114] In general, the sequence processing model (or models) 4 can obtain the input sequence 5 using data from the input (or inputs) 2. For example, the input sequence 5 can include a representation of data from the input (or inputs) 2 in a format understood by the sequence processing model (or models) 4. One or more machine learning components of the sequence processing model (or models) 4 can ingest the data from the input (or inputs) 2, analyze the data into chunks compatible with the processing architectures of the sequence processing model (or models) 4 (e.g., through “tokenization”), and project the chunks into an input space associated with the prediction layer (or layers) 6 (e.g., through “embedding”).

[0115] The sequence processing model (or models) 4 can ingest the input data (or inputs) 2 and parse the data into a sequence of elements to obtain the input sequence 5. For example, a portion of input data from input (or inputs) 2 can be decomposed into chunks that collectively represent the content of the portion of input data. The chunks can provide the elements of the sequence.

[0116] The 5-1, 5-2, . . . , 5-M elements can, in some cases, represent building blocks for capturing or expressing meaningful information in a specific data domain. For example, the elements can describe “atomic units” in one or more domains. For example, for textual input sources, the elements can correspond to groups of one or more words or components of subwords, such as sets of one or more characters. Petition 870250038841, dated 05 / 13 / 2025, pages 64 / 142 51 / 96

[0117] For example, the elements 5-1, 5-2, . . . , 5-M can represent tokens obtained using a tokenizer. For example, a tokenizer can process a given portion of an input source and generate a series of tokens (e.g., corresponding to the input elements 5-1, 5-2, . . . , 5-M) that represent that portion of the input source. Several approaches to tokenization can be used. For example, textual input sources can be tokenized using a byte pair encoding (BPE) technique. See, for example, Kudo et al., SentencePiece: A simple and language-independent subword tokenizer and detokenizer for Neural Text Processing, Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing (system demonstrations), pages 66-71 (October 31 to November 4, 2018), https: / / aclanthology.org / D18-2012.pdf.The image-based input source (or sources) can be tokenized by extracting and serializing patches from an image.

[0118] In general, arbitrary data types can be serialized and processed in a 5-input sequence. It should be understood that the 5-1, 5-2, . . . , 5-M elements represented in Figure 6 may be the tokens or may be embedded representations thereof.

[0119] Prediction layer (or layers) 6 can predict one or more output elements 7-1, 7-2, . . . , 7-N based on the input elements. Prediction layer (or layers) 6 can include one or more machine learning model architectures, such as one or more learned parameter layers that manipulate and transform the input(s) to extract higher-order meaning and relationships between the input elements 5-1, 5-2, . . . , 5-M. In this way, for example, example prediction layer (or layers) 6 can predict new output element(s) in view of the context provided by the input sequence 5. Petition 870250038841, dated 05 / 13 / 2025, pages 65 / 142 52 / 96

[0120] Prediction layer 6 can evaluate associations between portions of the input sequence 5 and a specific output element. These associations can inform a prediction of the probability that a specific output will follow the input context. For example, consider the text excerpt: “The carpenter’s toolbox was small and heavy. It was full of.” Example prediction layer 6 can identify that “it” refers to “toolbox” by determining a relationship between the respective embeddings. Example prediction layer 6 can also link “it” to attributes of the toolbox, such as “small” and “heavy.” Based on these associations, prediction layer 6 can, for example, assign a higher probability to the word “nails” than to the word “sawdust.”

[0121] A transformer is an example of an architecture that can be used in prediction layers 4. See, for example, Vaswani et al., Attention Is All You Need, ARXIV:1706.03762v7 (August 2, 2023). A transformer is an example of a machine learning model architecture that uses an attention mechanism to compute associations between items within a context window. The context window may include a sequence containing the input sequence 5 and potentially one or more output elements 7-1, 7-2, . . . , 7-N. A transformer block may include one or more attention layers and one or more post-attention layers (e.g., feedforward layers, such as a multilayer perceptron).

[0122] Prediction layer(s) 6 may include other machine learning model architectures, in addition to or instead of transformer-based architectures. For example, recurrent neural networks (RNNs) and long-term memory models (LSTMs) may also be used, as well as convolutional neural networks. Petition 870250038841, dated 05 / 13 / 2025, pages 66 / 142 53 / 96 (CNNs). In general, the prediction layer(s) 6 can leverage various types of artificial neural networks that can understand or generate sequences of information.

[0123] Output sequence 7 may include or represent the same or different data types as input sequence 5. For example, input sequence 5 may represent textual data, and output sequence 7 may represent textual data. Input sequence 5 may represent image, audio, or audiovisual data, and output sequence 7 may represent textual data (e.g., describing the image, audio, or audiovisual data). It should be understood that the prediction layer (or layers) 6 and any other components of the interstitial model of the sequence processing model(s) 4 may be configured to receive a variety of data types in input sequence(s) 5 and output a variety of data types in output sequence(s) 7.

[0124] Output sequence 7 may have several relationships with input sequence 5. Output sequence 7 may be a continuation of input sequence 5. Output sequence 7 may be complementary to input sequence 5. Output sequence 7 may translate, transform, augment, or otherwise modify input sequence 5. Output sequence 7 may respond to, evaluate, confirm, or otherwise respond to input sequence 5. Output sequence 7 may implement (or describe instructions to implement) an instruction provided by input sequence 5.

[0125] Output sequence 7 can be generated in an autoregressive manner. For example, for some applications, an output from one or more prediction layers 6 can be passed through one or more output layers (e.g., softmax layer) to obtain a probability distribution over an output vocabulary (e.g., a Petition 870250038841, dated 05 / 13 / 2025, pp. 67 / 142 54 / 96 textual or symbolic vocabulary) conditioned to a set of input elements in a context window. In this way, for example, the output sequence 7 can be generated autoregressively by sampling a likely next output element, adding that element to the context window and regenerating the probability distribution based on the updated context window, and sampling a likely next output element, and so on.

[0126] Output sequence 7 can also be generated in a non-autoregressive manner. For example, multiple output elements of output sequence 7 can be predicted together without explicit sequential conditioning between them. See, for example, Saharia et al., Non-Autoregressive Machine Translation with Latent Alignments, arXIV:2004.07437v3 (November 16, 2020).

[0127] Output sequence 7 may include one or multiple portions or elements. In a content generation configuration example, output sequence 7 may include multiple elements corresponding to multiple parts of a generated output sequence (e.g., a textual phrase, values ​​of a discretized waveform, computer code, etc.). In a classification configuration example, output sequence 7 may include a single element associated with a classification output. For example, an output “vocabulary” may include a set of classes into which an input sequence should be classified. For example, a vision transformer block may pass latent state information to a multilayer perceptron that outputs a likely class value associated with an input image.

[0128] Figure 7 is a block diagram of an exemplary technique for populating an exemplary input sequence 8. The input sequence 8 may include several functional elements that Petition 870250038841, dated 05 / 13 / 2025, pp. 68 / 142 55 / 96 form part of the model infrastructure, as an 8-0 element obtained from a task indicator 9 that signals to any model (or models) processing the input sequence 8 that a particular task is being performed (e.g., to help tailor the performance of the model (or models) to that particular task). The input sequence 8 may include multiple data elements of different data modalities. For example, an input modality 10-1 may include a data modality. A data model for sequence 11-1 may process data from input modality 10-1 to project the data into a format compatible with input sequence 8 (e.g., one or more vectors scaled according to the dimensions of input sequence 8) to obtain the elements 8-1, 8-2, 8-3. Another input modality 10-2 may include a different data modality.A data model for sequence 11-2 can project data from input mode 10-2 into a format compatible with input sequence 8 to obtain elements 8-4, 8-5, 8-6. Another input mode 10-3 can include yet another different data mode. A data model for sequence 11-3 can project data from input mode 10-3 into a format compatible with input sequence 8 to obtain elements 8-7, 8-8, 8-9.

[0129] Input sequence 8 may be the same as or different from input sequence 5. Input sequence 8 may be a multimodal input sequence containing elements that represent data of different modalities using a common dimensional representation. For example, an embedding space may have dimensions P. Input sequence 8 may be configured to contain a plurality of elements that have dimensions P. In this way, for example, example implementations may facilitate information extraction and reasoning across various data modalities by designing Petition 870250038841, dated 05 / 13 / 2025, pp. 69 / 142 56 / 96 data in elements within the same embedding space for comparison, combination, or other calculations between them.

[0130] For example, the elements 8-0, . . . , 8-9 may indicate specific locations within a multidimensional embedding space. Some elements may be mapped to a set of discrete locations in the embedding space. For example, elements that correspond to discrete members of a predetermined vocabulary of tokens may be mapped to discrete locations in the embedding space that are associated with those tokens. Other elements may be continuously distributed throughout the embedding space. For example, some data types may be decomposed into continuously defined portions (e.g., image patches) that can be described using continuously distributed locations within the embedding space.

[0131] In some implementations, the expressive power of the embedding space cannot be limited to meanings associated with any particular set of tokens or other building blocks. For example, a continuous embedding space can encode a spectrum of high-order information. An individual piece of information (e.g., a token) can be mapped to a specific point in that space: for example, a token for the word “dog” can be projected onto an embedded value that points to a specific location in the embedding space associated with information related to dogs. Similarly, an image patch of a dog in the grass can also be projected onto the embedding space. In some implementations, the projection of the dog image may be similar to the projection of the word “dog,” while also having similarity to a projection of the word “grass,” while potentially being different from both. In some implementations, the projection of the Petition 870250038841, dated 05 / 13 / 2025, pp. 70 / 142 57 / 96 image patch cannot align exactly with any single projection of a single word. In some implementations, the image patch projection may align with a combination of the projections of the words “dog” and “grass”. In this way, for example, a higher-order embedding space can encode information that may be independent of the data modalities in which the information is expressed.

[0132] Task indicator 9 may include a model or model component configured to identify a task being performed and inject, into input sequence 8, an input value represented by element 8-0 that signals which task is being performed. For example, the input value may be provided as a data type associated with an input mode and designed along with that input mode (e.g., the input value may be a textual task label that is embedded along with other textual data in the input; the input value may be a pixel-based representation of a task that is embedded along with other image data in the input; etc.). The input value may be provided as a data type that differs from, or is at least independent of, another input (or inputs). For example, the input value represented by element 8-0 may be learned within a continuous embedding space.

[0133] The 10-1, 10-2 and 10-3 input modes can be associated with several different data types (for example, as described above in relation to input (or inputs) 2 and output (or outputs) 3).

[0134] The data models for sequences 11-1, 11-2, and 11-3 may be the same or different from each other. The data models for sequences 11-1, 11-2, and 11-3 may be adapted to each respective input mode 10-1, 10-2, and 10-3. For example, a data model Petition 870250038841, dated 05 / 13 / 2025, pp. 71 / 142 Textual sequence models can subdivide a portion of the input text and project the subdivisions onto element(s) in the input sequence (e.g., elements 8-1, 8-2, 8-3, etc.). An image sequence data model can subdivide an input image and project the subdivisions onto element(s) in the input sequence (e.g., elements 8-4, 8-5, 8-6, etc.). A data model for sequences of arbitrary data types can subdivide an input of that arbitrary data type and project the subdivisions onto element(s) in the input sequence (e.g., elements 8-7, 8-8, 8-9, etc.).

[0135] The data models for sequences 11-1, 11-2, and 11-3 can form part of the machine learning sequence processing model (or models) 4. The data models for sequences 11-1, 11-2, and 11-3 can be trained together or trained independently of the machine learning sequence processing model (or models) 4. The data models for sequences 11-1, 11-2, and 11-3 can be trained end-to-end with the machine learning sequence processing model (or models) 4.

[0136] Figure 8 is a block diagram of an exemplary model development platform 12 that can facilitate the creation, adaptation, and refinement of exemplary machine learning models (e.g., machine learning model (or models) 1, sequence processing model (or models) 4, etc.). The model development platform 12 can provide several different toolkits that developer systems can employ in developing new or adapted machine learning models.

[0137] The model development platform 12 may provide one or more model libraries 13 containing building blocks for new models. The model libraries 13 may include Petition 870250038841, dated 05 / 13 / 2025, pp. 72 / 142 59 / 96 one or more pre-trained fundamental models 13-1, which can provide a foundation of processing power in various tasks. Model libraries 13 may include one or more pre-trained expert models 13-2, which may be focused on performance in specific domains of expertise. Model libraries 13 may include several model primitives 13-3, which may provide low-level (optionally pre-trained) architectures or components that can be assembled in various arrangements as desired. Model primitives 13-3 may include a library of pre-trained LoRA adapters or modules that can adapt a baseline fundamental model to align its outputs with a desired performance profile, augment the model's capabilities (e.g., to adapt to a different input modality, etc.), and the like.

[0138] The model development platform 12 can receive selections of various model components 14. The model development platform 12 can pass selected model components 14 to a Workbench 15 that combines selected model components 14 into a development model 16.

[0139] Workbench 15 can facilitate further refinement and adaptation of the development model 16 by leveraging several different toolkits integrated into the model development platform 12. For example, Workbench 15 can facilitate aligning the development model 16 with a desired performance profile across various tasks using a model alignment toolkit 17.

[0140] The model alignment toolkit 17 can provide various tools to make the development model 16 generate outputs aligned with the desired behavioral characteristics. Alignment can include increased accuracy, Petition 870250038841, dated 05 / 13 / 2025, pp. 73 / 142 60 / 96 accuracy, recall, etc. of model outputs. Alignment may include applying output styles, schemas, or other preferred model output characteristics. Alignment can be general or domain-specific. For example, a pre-trained 13-1 fundamental model may start with an initial level of performance in multiple domains. Alignment of the pre-trained 13-1 fundamental model may include improving performance in a specific information or task domain (e.g., even at the expense of performance in another information or task domain).

[0141] The model alignment toolkit 17 can integrate one or more datasets 17-1 to align the development model 16. The selected datasets 17-1 can include labeled or unlabeled training data. The dataset(s) 17-1 can be obtained from publicly available datasets. The dataset(s) 17-1 can be obtained from private datasets associated with one or more developer systems for aligning customized machine learning models for private use cases.

[0142] Pretraining pipelines 17-2 may include a machine learning model training workflow configured to update the development model 16 on large-scale and potentially noisy datasets. For example, pretraining may leverage unsupervised learning techniques (e.g., noise reduction, etc.) to process a large number of training instances to update model parameters from an initialized state and achieve a desired baseline performance. Pretraining pipelines 17-2 may leverage unlabeled datasets in dataset(s) 17-1 to perform pretraining. Workbench 15 may implement Petition 870250038841, dated 05 / 13 / 2025, pp. 74 / 142 61 / 96 a pre-training pipeline 17-2 to pre-train the development model 16.

[0143] Fine-tuning pipelines 17-3 may include a machine learning model training workflow configured to refine the parameters of the development model 16 with higher quality data. Fine-tuning pipelines 17-3 may update the development model 16 by conducting supervised training with datasets labeled in datasets 17-1. Fine-tuning pipelines 17-3 may update the development model 16 by conducting reinforcement learning using reward signals from user feedback signals. Workbench 15 may implement a fine-tuning pipeline 17-3 to fine-tune the development model 16.

[0144] 17-4 prompt libraries may include sets of inputs configured to induce behavior aligned with desired performance criteria. 17-4 prompt libraries may include “few-shot” prompts (e.g., inputs that provide examples of desired model outputs to be added to a desired runtime query), chain-of-thought prompts (e.g., inputs that provide step-by-step reasoning within the examples to facilitate complete model reasoning), and the like.

[0145] Prompt examples can be retrieved from an available repository of prompt libraries 17-4. Prompt examples can be contributed by one or more developer systems using Workbench 15.

[0146] In some implementations, pre-trained or tuned models can achieve satisfactory performance without exemplars in the inputs. For example, “zero-shot” prompts can include inputs that have no exemplars. The “zero-shot” prompts can be within Petition 870250038841, dated 05 / 13 / 2025, pp. 75 / 142 62 / 96 of a domain within a training dataset or outside the training domain(s).

[0147] Prompt libraries 17-4 may include one or more prompt engineering tools. Prompt engineering tools may provide workflows for retrieving or learning optimized prompt values. Prompt engineering tools may facilitate the direct learning of prompt values ​​(e.g., input element values) based on one or more training iterations. Workbench 15 may implement prompt engineering tools in the development model 16.

[0148] The prompt libraries 17-4 can include pipelines for generating prompts. For example, inputs can be generated using the development model itself 16 or other machine learning models. In this way, for example, a first model can process information about a task and generate an input for a second model to process in order to perform a step of the task. The second model can be the same as or different from the first model. Workbench 15 can implement prompt generation pipelines in the development model 16.

[0149] Prompt libraries 17-4 may include pipelines for context injection. For example, the performance of development model 16 on a specific task may improve if additional context is provided to perform the task. Prompt libraries 17-4 may include software components configured to identify the desired context, retrieve the context from an external source (e.g., a database, a sensor, etc.), and add the context to the input prompt. Workbench 15 may implement context injection pipelines in development model 16.

[0150] Although several training examples are described in the pre Petition 870250038841, dated 05 / 13 / 2025, pp. 76 / 142 63 / 96 document relating to the model development platform 12 refers to “pre-training” and “fine-tuning”, it should be understood that the model alignment toolkit 17 can generally support a wide variety of training techniques adapted to train a wide variety of machine learning models. Example training techniques may correspond to the example training method 400 described above.

[0151] The model development platform 12 may include a model plug-in toolkit 18. The model plug-in toolkit 18 may include a variety of tools configured to enhance the functionality of a machine learning model by integrating the machine learning model with other systems, devices, and software components. For example, a machine learning model may use tools to enhance performance quality when appropriate. For example, deterministic tasks may be offloaded to dedicated tools instead of performing the task probabilistically, with a higher risk of error. For example, instead of autoregressively predicting the solution to a system of equations, a machine learning model may recognize a tool to obtain the solution and pass the system of equations to the appropriate tool.The tool can be a traditional equation system solver that can operate deterministically to solve the equation system. The tool's output can be returned in response to the original query. In this way, the use of tools can allow some example models to focus on the strengths of machine learning models—for example, understanding an intent in an unstructured request for a task—while simultaneously increasing model performance by offloading certain tasks to a tool. Petition 870250038841, dated 05 / 13 / 2025, pp. 77 / 142 64 / 96 is a more focused approach on the mechanical application of deterministic algorithms to a well-defined problem.

[0152] The 18-model plug-in toolkit may include 18-1 validation tools. 18-1 validation tools may include tools that can analyze and confirm output(s) of a machine learning model. 18-1 validation tools may include engineered heuristics that establish certain limits applied to the model outputs. For example, 18-1 validation tools may base the outputs of machine learning models on structured data sources (e.g., to mitigate “hallucinations”).

[0153] The model plug-in toolkit 18 may include tool packages 18-2 to implement one or more tools that may include scripts or other executable code that can be run along with the development model 16. The tool packages 18-2 may include one or more configured inputs to cause machine learning model(s) to implement the tools (e.g., “few-shot” prompts that induce a model to issue tool calls in the appropriate syntax, etc.). The tool packages 18-2 may include, for example, fine-tuning of training data to train a model to use a tool.

[0154] The 18-model plug-in toolkit may include interfaces for calling external application programming interfaces (APIs) 18-3. For example, in addition to or instead of implementing tool calls or tool code directly with the 16 development model, the 16 development model may be aligned to generate statements that initiate API calls to send or obtain data through external systems. Petition 870250038841, dated 05 / 13 / 2025, pp. 78 / 142 65 / 96

[0155] The template plug-in toolkit 18 can be integrated with the prompt libraries 17-4 to create a catalog of tools available for use with the development template 16. For example, a template can receive, as an input, a catalog of available tools, and the template can generate an output that selects a tool from among the available tools and initiates a tool call to use it.

[0156] The model development platform 12 may include a computational optimization toolkit 19 to optimize the computational performance of the development model 16. For example, model compression tools 19-1 may allow the development model 16 to be reduced in size while maintaining a desired level of performance. For example, model compression 19-1 may include quantization workflows, weight reduction and sparsification techniques, etc. Hardware acceleration tools 19-2 may facilitate the configuration of model storage and execution formats to operate optimally on different hardware resources. For example, hardware acceleration 19-2 may include tools for optimized model fragmentation for distributed processing across multiple processing units for higher bandwidth, lower unified memory requirements, etc.Tools for distillation 19-3 can provide training for lighter models based on the knowledge encoded in the development model 16. For example, the development model 16 could be a high-performance machine learning model optimized using the model development platform 12. To obtain a lightweight model for execution in resource-constrained environments, a smaller model could be a “student model” that learns to mimic the development model 16 as a “teacher model”. In this way, for example, Petition 870250038841, dated 05 / 13 / 2025, pp. 79 / 142 66 / 96 the investment in learning the parameters and settings of the development model 16 can be efficiently transferred to a smaller model for more efficient inference.

[0157] Workbench 15 can implement one, multiple, or none of the toolkits implemented in the model development platform 12. Workbench 15 can output a model 20 based on the development model 16. The output model 20 can be a deployment version of the development model 16. The output model 20 can be a development or training checkpoint of the development model 16. The output model 20 can be a distilled, compressed, or otherwise optimized version of the development model 16.

[0158] Figure 9 is a block diagram of an exemplary training flow for training a machine learning development model 16. One or more portions of the exemplary training flow may be implemented by a computing system that includes one or more computing devices, such as, for example, computing systems described with reference to the other Figures. Each respective portion of the example training flow may be performed by any (or any combination of) one or more computing devices. Furthermore, one or more portions of the exemplary training flow may be implemented in the hardware components of the device (or devices) described in this document, for example, to train one or more systems or models. Figure 9 represents elements performed in a particular order for purposes of illustration and discussion.Those of ordinary skill in the art, using the disclosures provided in this document, will understand that elements of any of the methods discussed in this document may be adapted, rearranged, expanded, omitted, combined, or modified in various ways without prejudice. Petition 870250038841, dated 05 / 13 / 2025, pages 80 / 142 67 / 96 that falls outside the scope of this disclosure. Figure 9 is described with reference to elements / terms described in relation to other systems and figures for illustrative purposes only and is not intended to be limiting. One or more portions of the illustrative training flow may be performed additionally or alternatively by other systems.

[0159] Initially, development model 16 can persist in an initial state as an initialized model 21. Development model 16 can be initialized with weight values. The initial weight values ​​can be random or based on an initialization scheme. The initial weight values ​​can be based on previous pretraining for the same model or for a different model.

[0160] The initialized model 21 can undergo pretraining in a pretraining stage 22. The pretraining stage 22 can be implemented using one or more pretraining pipelines 17-2 on data from dataset(s) 17-1. Pretraining can be omitted, for example, if the initialized model 21 is already pretrained (for example, the development model 16 contains, is, or is based on a pretrained fundamental model or an expert model).

[0161] The pre-trained model 23 can then be a new version of the development model 16, which can persist as development model 16 or as a new development model. The pre-trained model 23 can be the initial state if the development model 16 has already been pre-trained. The pre-trained model 23 can undergo fine-tuning in a fine-tuning stage 24. The fine-tuning stage 24 can be implemented using one or more fine-tuning pipelines 17-3 with respect to data from the dataset (or datasets) 17-1. Fine-tuning can be omitted, for example, if a pre-trained model Petition 870250038841, dated 05 / 13 / 2025, pages 81 / 142 If the trained 68 / 96 model performs satisfactorily, the model may already be fine-tuned, or other tuning approaches may be preferred.

[0162] The fine-tuned model 29 can then be a new version of the development model 16, which can persist as development model 16 or as a new development model. The fine-tuned model 29 can be the initial state if the development model 16 has already been fine-tuned. The fine-tuned model 29 can undergo refinement with user feedback 26. For example, refinement with user feedback 26 can include reinforcement learning, optionally based on human feedback from human users of the fine-tuned model 25. Since reinforcement learning can be a form of fine-tuning, it should be understood that the fine-tuning stage 24 can replace the refinement stage with user feedback 26. Refinement with user feedback 26 can produce a refined model 27. The refined model 27 can be sent to downstream system(s) 28 for deployment or further development.

[0163] In some implementations, computational optimization operations may be applied before, during, or after each stage. For example, the initialized model 21 may undergo computational optimization 29-1 (e.g., using the computational optimization toolkit 19) before the pretraining stage 22. The pretrained model 23 may undergo computational optimization 29-2 (e.g., using the computational optimization toolkit 19) before the fine-tuning stage 24. The tuned model 25 may undergo computational optimization 29-3 (e.g., using the computational optimization toolkit 19) before refinement with user feedback 26. The refined model 27 may undergo computational optimization 29-4 (e.g., using the computational optimization toolkit 19) before output to Petition 870250038841, dated 05 / 13 / 2025, pages 82 / 142 69 / 96 the downstream system(s) 28. The computational optimizations 29-1, . . . , 29-4 may all be the same, all different, or include at least some different optimization techniques.

[0164] Figure 10 is a block diagram of an inference system for operating one or more machine learning model(s) 1 to perform inference (e.g., for training, for deployment, etc.). A model host 31 can receive the machine learning model(s) 1. The model host 31 can host one or more model instances 31-1, which can be one or multiple instances of one or multiple models. The model host 31 can host the model instance(s) 31-1 using available computer resources 31-2 associated with the model host 31.

[0165] The model host 31 can perform inference on behalf of one or more clients 32. The client (or clients) 32 can transmit an input request 33 to the model host 31. Using the input request 33, the model host 31 can obtain the input (or inputs) 2 for input into the machine learning model (or models) 1. The machine learning model (or models) 1 can process the input (or inputs) 2 to generate the output (or outputs) 3. Using the output (or outputs) 3, the model host 31 can return an output payload 34 to respond to the input request 33 from the client (or clients) 32. The output payload 34 can include or be based on the output (or outputs) 3.

[0166] The model host 31 can leverage several other resources and tools to augment the inference task. For example, the model host 31 can communicate with the tool interfaces 35 to facilitate the use of the tool by the model instance(s) 31-1. The tool interfaces 35 may include local APIs. Petition 870250038841, dated 05 / 13 / 2025, pages 83 / 142 70 / 96 or remote. Tool interfaces 35 may include embedded scripts or other software functionality. The model host 31 may engage online learning interface(s) 36 to facilitate continuous enhancements to the machine learning model(s) 1. For example, the online learning interface(s) 36 may be used within reinforcement learning loops to retrieve user feedback on inferences served by the model host 31. The model host 31 may access runtime data source(s) 37 to augment the input(s) 2 with additional contextual information. For example, the runtime data source(s) 37 may include a knowledge graph 37-1 that facilitates the retrieval of structured information for information associated with the input request(s) 33 (e.g., a search engine service).The runtime data source(s) 37 may include public or private, external or local database(s) 37-2 that may store information associated with the input request(s) 33 to augment the input(s) 2. The runtime data source(s) 37 may include account data 37-3, which may be retrieved in association with a user account corresponding to a client 32 to customize the behavior of the template host 31 accordingly.

[0167] The model 31 host can be implemented by one or multiple computing devices or systems. Client(s) 2 can be implemented by one or multiple computing devices or systems, which may include computing devices or systems shared with the model 31 host.

[0168] For example, the model 31 host can operate on a server system that provides a machine learning service. Petition 870250038841, dated 05 / 13 / 2025, pages 84 / 142 71 / 96 for client device(s) operating client(s) 32 (for example, on a local or wide area network). Client device(s) may be end-user devices used by individuals. Client device(s) may be server systems operating client(s) 32 to provide various functionalities as a service to downstream end-user devices.

[0169] In some implementations, model host 31 may operate on the same device or system as client(s) 32. Model host 31 may be a machine learning service running on the device to provide machine learning functionality to one or more applications operating on a client device, which may include an application that implements client(s) 32. Model host 31 may be part of the same application as client(s) 32. For example, model host 31 may be a subroutine or method implemented by a part of an application, and client(s) 32 may be another subroutine or method that engages model host 31 to perform inference functions within the application. It should be understood that model host 31 and client(s) 32 may have several different configurations.

[0170] The 31-1 model instance(s) may include one or more machine learning models that are available to perform inference. The 31-1 model instance(s) may include weights or other model components that are stored in persistent storage, with temporary caching, or loaded into high-speed memory. The 31-1 model instance(s) may include multiple instances of the same model (for example, for parallel execution of more requests on the same model). The 31-1 model instance(s) may include different model instance(s). The 31-1 model instance(s) may include intermediate states of active model cache(s) or. Petition 870250038841, dated 05 / 13 / 2025, pages 85 / 142 72 / 96 inactive caches are used to accelerate the inference of such models. For example, an inference session with a particular model may generate significant amounts of computational results that can be reused for future inference runs (e.g., using a KV cache for transformer-based models). These computational results can be saved in association with such an inference session, so that the session can be run more efficiently when resumed.

[0171] Computational resource(s) 31-2 may include one or more processors (central processing units, graphics processing units, tensor processing units, machine learning accelerators, etc.) connected to one or more memory devices. Computational resource(s) 31-2 may include a dynamic set of available resources shared with other processes. Computational resource(s) 31-2 may include memory devices large enough to fit an entire model instance into a single memory instance. Computational resource(s) 31-2 may also fractionate the model instance(s) across multiple memory devices (e.g., using data parallelization or tensor parallelization, etc.).This can be done to increase parallelization or to run a large model using multiple memory devices that may not individually be able to fit the entire model into memory.

[0172] Input request 33 may include data for input (or inputs) 2. Template host 31 may process input request 33 to obtain input (or inputs) 2. Input (or inputs) 2 may be obtained directly from input request 33 or may be retrieved using input request 33. Input request 33 may be forwarded to the host of Petition 870250038841, dated 05 / 13 / 2025, pages 86 / 142 73 / 96 model 31 via an API.

[0173] The model host 31 can perform inference on batches of input requests 33 in parallel. For example, a model instance 31-1 can be configured with an input structure that has a batch dimension. The separate input (or inputs) 2 can be distributed across the batch dimension (e.g., rows of an array). The separate input (or inputs) 2 can include completely different contexts. The separate input (or inputs) 2 can be multiple inference steps of the same task. The separate input (or inputs) 2 can be staggered within an input structure, so that any given inference cycle can be operating on different portions of the respective input (or inputs) 2.In this way, for example, the model host 31 can perform batch inference in parallel, so that output (or outputs) 3 can also contain the batch dimension and return the inference results to the input (or inputs) in batches 2 in parallel. In this way, for example, batches of input request (or requests) 33 can be processed in parallel for superior output payload (or payloads) productivity 34.

[0174] Output payload 34 may include or be based on output(s) 3 from machine learning model(s) 1. Model host 31 may process output(s) 3 to obtain output payload 34. This may include chaining multiple rounds of inference (e.g., iteratively, recursively, through the same model(s) or different model(s)) to arrive at a final output for which a task is returned in output payload 34. Output payload 34 may be transmitted to client(s) 32 via an API.

[0175] The online learning interface (or interfaces) 36 can Petition 870250038841, dated 05 / 13 / 2025, page 87 / 142 74 / 96 facilitate reinforcement learning of the machine learning model (or models) 1. The online learning interface (or interfaces) 36 can facilitate reinforcement learning with human feedback (RLHF). The online learning interface (or interfaces) 36 can facilitate federated learning of the machine learning model (or models) 1.

[0176] Model host 31 can access a library of pre-trained LoRA adapters or modules that can adapt a baseline model to align its outputs with a desired performance profile, augment model capabilities (e.g., to adapt to a different input modality, etc.), and the like. For example, model host 31 might receive an input request to load a custom model, and model host 31 might retrieve one or more components to adapt a baseline model to the custom profile. Model host 31 might determine that a particular functionality is needed for a particular task (e.g., based on the output of a model that pre-processes an input) and retrieve a pre-trained component accordingly.

[0177] The model host 31 can run the machine learning model (or models) 1 to perform inference for various tasks using various data types. For example, several different inputs 2 and outputs 3 can be used for several different tasks. In some implementations, the input (or inputs) 2 can be, or otherwise represent, image data. The machine learning model (or models) 1 can process the image data to generate an output. As an example, the machine learning model (or models) 1 can process the image data to generate an image recognition output (e.g., recognition of the image data, latent embedding of the image data, Petition 870250038841, dated 05 / 13 / 2025, pages 88 / 142 75 / 96 an encoded representation of the image data, a hash of the image data, etc.). As another example, machine learning model 1 can process the image data to generate an image segmentation output. As another example, machine learning model 1 can process the image data to generate an image classification output. As another example, machine learning model 1 can process the image data to generate an image data modification output (e.g., an alteration of the image data, etc.). As another example, machine learning model 1 can process the image data to generate an encoded image data output (e.g., an encoded or compressed representation of the image data, etc.).As another example, machine learning model 1 can process image data to generate a sophisticated image data output. As another example, machine learning model 1 can process image data to generate a predictive output.

[0178] In some implementations, the task is a computer vision task. In some cases, the input(s) 2 includes pixel data for one or more images and the task is an image processing task. For example, the image processing task might be image classification, where the output is a set of scores, where each score corresponds to a different object class and represents the probability that one or more images represent an object belonging to that object class. The image processing task might be object detection, where the image processing output identifies one or more regions in one or more images and, for each region, a probability that such a region represents an object of interest. As another example, the Petition 870250038841, dated 05 / 13 / 2025, pp. 89 / 142 76 / 96 Image processing tasks can be image segmentation, where the image processing output defines, for each pixel in one or more images, a respective probability for each category in a predetermined set of categories. For example, the set of categories could be foreground and background. As another example, the set of categories could be object classes. As another example, the image processing task can be depth estimation, where the image processing output defines, for each pixel in one or more images, a respective depth value. As another example, the image processing task can be motion estimation, where the network input includes multiple images, and the image processing output defines, for each pixel of one of the input images, a motion of the scene represented in the pixel among the images in the network input.

[0179] In some implementations, input(s) 2 may be or represent natural language data. The machine learning model (or models) 1 may process the natural language data to generate an output. As an example, the machine learning model (or models) 1 may process the natural language data to generate a language encoding output. As another example, the machine learning model (or models) 1 may process the natural language data to generate a latent text embedding output. As another example, the machine learning model (or models) 1 may process the natural language data to generate a translation output. As another example, the machine learning model (or models) 1 may process the natural language data to generate a classification output. As another example, the machine learning model (or models) 1 may process the natural language data to generate a Petition 870250038841, dated 05 / 13 / 2025, pages 90 / 142 77 / 96 Text segmentation output. As another example, machine learning model 1 can process natural language data to generate semantic intent output. As another example, machine learning model 1 can process natural language data to generate sophisticated natural language or text output (e.g., text or natural language data that is of higher quality than the input text or natural language, etc.). As another example, machine learning model 1 can process natural language data to generate predictive output (e.g., one or more predicted near portions of natural language content).

[0180] In some implementations, input (or inputs) 2 may be, or otherwise represent, speech data (e.g., data describing spoken natural language, such as audio data, textual data, etc.). Machine learning model (or models) 1 may process the speech data to generate an output. As an example, machine learning model (or models) 1 may process the speech data to generate a speech recognition output. As another example, machine learning model (or models) 1 may process the speech data to generate a speech translation output. As another example, machine learning model (or models) 1 may process the speech data to generate a latent embedding output. As another example, machine learning model (or models) 1 may process the speech data to generate an encoded speech output (e.g., an encoded or compressed representation of the speech data, etc.).As another example, machine learning model 1 might process speech data to generate sophisticated speech output (e.g., speech data that is of higher quality than the input speech data, etc.). As another example, the model (or models) of... Petition 870250038841, dated 05 / 13 / 2025, pages 91 / 142 78 / 96 Machine learning 1 can process speech data to generate a textual representation output (e.g., a textual representation of the input speech data, etc.). As another example, the machine learning model (or models) 1 can process speech data to generate a predictive output.

[0181] In some implementations, input (or inputs) 2 may be, or otherwise represent, latent encoding data (e.g., a latent space representation of an input, etc.). Machine learning model (or models) 1 may process the latent encoding data to generate an output. As an example, machine learning model (or models) 1 may process the latent encoding data to generate a recognition output. As another example, machine learning model (or models) 1 may process the latent encoding data to generate a reconstruction output. As another example, machine learning model (or models) 1 may process the latent encoding data to generate a search output. As another example, machine learning model (or models) 1 may process the latent encoding data to generate a reassembly output.As another example, machine learning model 1 can process latent encoding data to generate a predictive output.

[0182] In some implementations, input (or inputs) 2 may be, or otherwise represent, statistical data. The statistical data may be, represent, or otherwise include data computed or calculated from some other data source. The machine learning model (or models) 1 may process the statistical data to generate an output. As an example, the machine learning model (or models) 1 may process the statistical data to generate a recognition output. As another example, Petition 870250038841, dated 05 / 13 / 2025, pages 92 / 142 79 / 96 The machine learning model (or models) 1 can process statistical data to generate a prediction output. As another example, the machine learning model (or models) 1 can process statistical data to generate a classification output. As another example, the machine learning model (or models) 1 can process statistical data to generate a segmentation output. As another example, the machine learning model (or models) 1 can process statistical data to generate a visualization output. As another example, the machine learning model (or models) 1 can process statistical data to generate a diagnostic output.

[0183] In some implementations, input (or inputs) 2 may be, or otherwise represent, sensor data. Machine learning model (or models) 1 may process the sensor data to generate an output. For example, machine learning model (or models) 1 may process the sensor data to generate a recognition output. For another example, machine learning model (or models) 1 may process the sensor data to generate a prediction output. For another example, machine learning model (or models) 1 may process the sensor data to generate a classification output. For another example, machine learning model (or models) 1 may process the sensor data to generate a segmentation output. For another example, machine learning model (or models) 1 may process the sensor data to generate a visualization output.As another example, machine learning model 1 can process sensor data to generate a diagnostic output. As another example, machine learning model 1 can process sensor data to generate a detection output. Petition 870250038841, dated 05 / 13 / 2025, pages 93 / 142 80 / 96

[0184] In some implementations, the machine learning model (or models) 1 may be configured to perform a task that includes encoding input data for reliable or efficient transmission or storage (or corresponding decoding). For example, the task might be an audio compression task. The input might include audio data, and the output might comprise compressed audio data. In another example, the input includes visual data (e.g., one or more images or videos), the output comprises compressed visual data, and the task is a visual data compression task. In another example, the task might comprise generating an embedding for input data (e.g., input audio or visual data). In some cases, the input includes audio data representing spoken utterance, and the task is a speech recognition task. The output might comprise text output that is mapped to the spoken utterance.In some cases, the task involves encrypting or decrypting input data. In some cases, the task involves a microprocessor performance task, such as branch prediction or memory address translation.

[0185] In some implementations, the task is a generative task, and the machine learning model (or models) 1 can be configured to output generated content in view of the input (or inputs) 2. For example, the input (or inputs) 2 can be, or otherwise represent, data from one or more modalities that encode the context to generate additional content.

[0186] In some implementations, the task may be a text completion task. The machine learning model (or models) 1 may be configured to process the input (or inputs) 2 representing textual data and to generate the output (or outputs) 3 representing additional textual data that complete a sequence. Petition 870250038841, dated 05 / 13 / 2025, pp. 94 / 142 81 / 96 textual that includes input (or inputs) 2. For example, machine learning model (or models) 1 can be configured to generate output (or outputs) 3 to complete a sentence, paragraph, or portion of text that follows from a portion of text represented by input (or inputs) 2.

[0187] In some implementations, the task may be an instruction-following task. The machine learning model (or models) 1 may be configured to process the input (or inputs) 2 which represents instructions to perform a function and to generate the output (or outputs) 3 which advances a goal of satisfying the instruction function (e.g., at least one step of a multi-step procedure to perform the function). The output (or outputs) 3 may represent modality data that is the same as or different from the input (or inputs) 2.For example, input(s) 2 could represent textual data (e.g., natural language instructions for a task to be performed), and machine learning model(s) 1 could process input(s) 2 to generate output(s) 3 which represents textual data responsive to the instructions (e.g., natural language responses, programming language responses, machine language responses, etc.). Input(s) 2 could represent image data (e.g., image-based instructions for a task to be performed, optionally accompanied by textual instructions), and machine learning model(s) 1 could process input(s) 2 to generate output(s) 3 which represents textual data responsive to the instructions (e.g., natural language responses, programming language responses, machine language responses, etc.).One or more outputs (3) can be generated iteratively or recursively to sequentially process and fulfill the steps to achieve the requested functionality. For example, an initial output. Petition 870250038841, dated 05 / 13 / 2025, pages 95 / 142 82 / 96 can be executed by an external system or processed by the machine learning model (or models) 1 to complete an initial step in performing a function. Multiple steps can be performed, with a final output being obtained that is responsive to the initial instructions.

[0188] In some implementations, the task may be a question-answering task. The machine learning model (or models) 1 may be configured to process the input (or inputs) 2 representing a question to be answered and to generate the output (or outputs) 3 that advances a goal of returning an answer to the question (e.g., at least one step of a multi-step procedure to perform the function). The output (or outputs) 3 may represent modality data that is the same as or different from the input (or inputs) 2. For example, the input (or inputs) 2 may represent textual data (e.g., natural language instructions for a task to be performed), and the machine learning model (or models) 1 may process the input (or inputs) 2 to generate the output (or outputs) 3 representing question-responsive textual data (e.g., natural language answers, programming language answers, machine language answers, etc.).Input 2 may represent image data (e.g., image-based instructions for a task to be performed, optionally accompanied by textual instructions), and machine learning model 1 may process input 2 to generate output 3 representing textual data responsive to questions (e.g., natural language responses, programming language responses, machine language responses, etc.). One or more outputs 3 may be generated iteratively or recursively to process sequentially and fulfill steps toward answering the question. For example, an initial output might be... Petition 870250038841, dated 05 / 13 / 2025, pp. 96 / 142 83 / 96 executed by an external system or processed by the machine learning model (or models) 1 to complete an initial step of obtaining an answer to the question (e.g., querying a database, performing a computation, executing a script, etc.). Multiple steps may be performed, with a final output being obtained that is responsive to the question.

[0189] In some implementations, the task may be an image generation task. The machine learning model (or models) 1 may be configured to process the input (or inputs) 2 that represents the context in relation to a desired portion of image content. The context may include text data, image data, audio data, etc. The machine learning model (or models) 1 may be configured to generate the output (or outputs) 3 that represents the image data that represent images related to the context. For example, the machine learning model (or models) 1 may be configured to generate pixel data from an image. The values ​​for the channel (or channels) associated with the pixels in the pixel data may be selected based on the context (e.g., based on a probability determined based on the context).

[0190] In some implementations, the task may be an audio generation task. The machine learning model (or models) 1 may be configured to process the input (or inputs) 2 that represents the context in relation to a desired portion of audio content. The context may include text data, image data, audio data, etc. The machine learning model (or models) 1 may be configured to generate the output (or outputs) 3 that represents the audio data related to the context. For example, the machine learning model (or models) 1 may be configured to generate waveform data in the form of an image (by Petition 870250038841, dated 05 / 13 / 2025, pages 97 / 142 84 / 96 example, a spectrogram). The values ​​for the channel (or channels) associated with image pixels can be selected based on context. The machine learning model(s) 1 can be configured to generate waveform data in the form of a sequence of discrete samples of a continuous waveform. The sequence values ​​can be selected based on context (e.g., based on a context-determined probability).

[0191] In some implementations, the task may be a data generation task. The machine learning model (or models) 1 may be configured to process the input (or inputs) 2 that represents the context with respect to a desired portion of data (e.g., data from various data domains such as sensor data, image data, multimodal data, statistical data, etc.). The desired data may be, for example, synthetic data for training other machine learning models. The context may include arbitrary data type (or types). The machine learning model (or models) 1 may be configured to generate the output (or outputs) 3 that represents the data that aligns with the desired data. For example, the machine learning model (or models) 1 may be configured to generate data values ​​to populate a dataset.The values ​​for the data object (or objects) can be selected based on context (for example, based on a probability determined by the context).

[0192] Figure 11 is a block diagram of an example networked computing system that can accomplish aspects of example implementations of the present disclosure. The system may include various computing devices and systems that are communicatively coupled through a network 49. An example of a device Petition 870250038841, dated 05 / 13 / 2025, pages 98 / 142 Computing device 50 is described to provide an example of a computing device that can perform any aspect of this disclosure (e.g., implement model host 31, client(s) 32, or both). An example server computing system 60 is described as an example of a server computing system that can perform any aspect of this disclosure (e.g., implement model host 31, client(s) 32, or both). Computing device 50 and server computing system(s) 60 can cooperatively interact (e.g., via network 49) to perform any aspect of this disclosure (e.g., implement model host 31, client(s) 32, or both). Model development platform system 70 is an example system that can host or serve model development platform(s) 12 for machine learning model development.The third-party system(s) 80 are example systems with which any computing device 50, server computing system(s) 60 or model development platform system(s) 70 may interact in performing various aspects of this disclosure (for example, engaging third-party tools, accessing third-party databases or other resources, etc.).

[0193] Network 49 can be any type of communications network, such as a local area network (e.g., intranet), wide area network (e.g., Internet), or some combination thereof, and can include any number of wired or wireless links. In general, communication over network 49 can be carried over any type of wired or wireless connection, using a wide variety of communication protocols (e.g., TCP / IP, HTTP, SMTP, FTP), encodings or formats (e.g., HTML, XML), or security schemes (e.g., VPN, secure HTTP, SSL). Network 49 Petition 870250038841, dated 05 / 13 / 2025, pages 99 / 142 86 / 96 can also be implemented via a system bus. For example, one or more devices or systems from Figure 11 can be co-located, contained, or otherwise integrated within one or more other devices or systems.

[0194] Computing device 50 can be any type of computing device, such as, for example, a personal computing device (e.g., laptop or desktop computer), a mobile computing device (e.g., smartphone or tablet computer), a game console or controller, a wearable computing device, an embedded computing device, a server computing device, a virtual machine operating on a host device, or any other type of computing device. Computing device 50 can be a client computing device. Computing device 50 can be an end-user computing device. Computing device 50 can be a service-providing computing device that provides a service to an end-user (who may use another computing device to interact with computing device 50).

[0195] The computing device 50 may include one or more processors 51 and a memory 52. ​​The processor(s) 51 may be any suitable processing device (e.g., a processor core, a microprocessor, an ASIC, an FPGA, a controller, a microcontroller, etc.) and may be a single processor or a plurality of processors that are operationally connected. The memory 52 may include one or more non-transient computer-readable storage media, such as HBM, RAM, ROM, EEPROM, EPROM, flash memory devices, magnetic disks, etc., and combinations thereof. The memory 52 may store... [Petition 870250038841, dated 13 / 05 / 2025, page 100 / 142] 87 / 96 narrate data 53 and instructions 54 that can be executed by the processor(s) 51 to cause the computing device 50 to perform operations. The operations may implement any one or several features described in this document. The operations may implement example methods and techniques described in this document.

[0196] The computing device 50 may also include one or more input components that receive user input. For example, a user input component may be a touch-sensitive component (e.g., a touch-sensitive display screen or a touchpad) that is sensitive to the touch of a user input object (e.g., a finger or a pen). The touch-sensitive component may serve to implement a virtual keyboard. Other examples of user input components include a microphone, a camera, a LIDAR, a physical keyboard or other buttons, or other means by which a user may provide user input.

[0197] The computing device 50 may store or include one or more machine learning models 55. The machine learning models 55 may include one or more machine learning models 1, such as a sequence processing model 4.Machine learning models 55 may include one or multiple instances of model 31-1. The machine learning model (or models) 55 may be received from a server computing system (or systems) 60, model development platform system 70, third-party system (or systems) 80 (e.g., an application distribution platform), or developed locally on the computing device 50. The machine learning model(s) 55 may be loaded into memory 52 and used or otherwise implemented by the processor(s) 51. The computing device 50 may implement multiple instances. Petition 870250038841, dated 05 / 13 / 2025, pages 101 / 142 88 / 96 parallel machine learning model(s) 55.

[0198] The server computing system(s) 60 may include one or more processors 61 and memory 62. The processor(s) 61 may be any suitable processing device (e.g., a processor core, a microprocessor, an ASIC, an FPGA, a controller, a microcontroller, etc.) and may be a single processor or a plurality of processors that are operationally connected. The memory 62 may include one or more non-transient computer-readable storage media, such as HBM, RAM, ROM, EEPROM, EPROM, flash memory devices, magnetic disks, etc., and combinations thereof. The memory 62 may store data 63 and instructions 64 that may be executed by the processor(s) 61 to cause the server computing system(s) 60 to perform operations. The operations may implement any one or several features described herein.Operations can implement example methods and techniques described in this document.

[0199] In some implementations, the server computing system 60 includes, or is otherwise implemented by, one or more server computing devices. In cases where the server computing system 60 includes multiple server computing devices, such server computing devices may operate according to sequential computing architectures, parallel computing architectures, or some combination thereof.

[0200] The server computing system 60 may store or otherwise include one or more machine learning models 65. The machine learning model (or models) 65 may be the same as or different from the machine learning model (or models) 55. The machine learning models 65 may include one or more machine learning models 1, such as a model of Petition 870250038841, dated 05 / 13 / 2025, pages 102 / 142 89 / 96 Sequence processing 4. Machine learning models 65 may include one or multiple model instances 31-1. The machine learning model (or models) 65 may be received from the computing device 50, the model development platform system 70, the third-party system (or systems) 80, or developed locally on the server computing system (or systems) 60. The machine learning model(s) 65 may be loaded into memory 62 and used or otherwise implemented by the processor(s) 61. The server computing system(s) 60 may implement multiple parallel instances of machine learning model(s) 65.

[0201] In an example configuration, machine learning models 65 can be included or stored and implemented by the server computing system 60 to establish a client-server relationship with the computing device 50 to serve model inferences. For example, the server computing system(s) 60 can implement the model host 31 on behalf of the client(s) 32 on the computing device 50. For example, the machine learning models 65 can be implemented by the server computing system 60 as a portion of a web service (e.g., remote machine learning model hosting service, as an online interface to perform machine learning model operations over a network on server computing system(s) 60).For example, server computing system(s) 60 can communicate with computing device 50 via a local intranet or internet connection. For example, computing device 50 can be a workstation or endpoint communicating with server computing system(s) 60, with mode implementation. Petition 870250038841, dated 05 / 13 / 2025, pages 103 / 142 90 / 96 machine learning models 65 that are managed by the server computing system(s) 60 to perform inference remotely (e.g., for runtime or training operations), with output(s) returned (e.g., converted, transmitted, etc.) to the computing device 50. The machine learning models 65 can work cooperatively or interoperably with machine learning models 55 on the computing device 50 to perform various tasks.

[0202] The model 70 development platform system(s) may include one or more processors 71 and a memory 72. The processor(s) 71 may be any suitable processing device (e.g., a processor core, a microprocessor, an ASIC, an FPGA, a controller, a microcontroller, etc.) and may be a single processor or a plurality of processors that are operationally connected. The memory 72 may include one or more non-transient computer-readable storage media, such as HBM, RAM, ROM, EEPROM, EPROM, flash memory devices, magnetic disks, etc., and combinations thereof. The memory 72 may store data 73 and instructions 74 that may be executed by the processor(s) 71 to cause the model 70 development platform system(s) to perform operations. The operations may implement any one or several features described herein.Operations can implement example methods and techniques described in this document. Example operations include the functionality described in this document with respect to the model development platform 12. This and other functionalities can be implemented by the developer tool(s) 75.

[0203] Third-party systems 80 may include one or more processors 81 and memory 82. The processor(s) 81 may be Petition 870250038841, dated 05 / 13 / 2025, pages 104 / 142 91 / 96 any suitable processing device (e.g., a processor core, a microprocessor, an ASIC, an FPGA, a controller, a microcontroller, etc.) and may be a processor or a plurality of processors that are operationally connected. Memory 82 may include one or more non-transient computer-readable storage media such as HBM, RAM, ROM, EEPROM, EPROM, flash memory devices, magnetic disks, etc., and combinations thereof. Memory 82 may store data 83 and instructions 84 that may be executed by the processor(s) 81 to cause third-party systems 80 to perform operations. The operations may implement any or several features described herein. The operations may implement example methods and techniques described herein.The illustrative operations include the functionality described in this document in relation to external tools and other resources called during training or performing inference with machine learning model(s) 1,4, 16, 20, 55, 65, etc. (e.g., third-party resource(s) 85).

[0204] Figure 11 illustrates an exemplary arrangement of computing systems that can be used to implement the present disclosure. Other computing system configurations may also be used. For example, in some implementations, one or both computing systems 50 or server computing systems 60 may implement all or part of the operations of the model development platform system 70. For example, computing system 50 or server computing system(s) 60 may implement developer tool(s) 75 (or extensions thereof) to develop, update / train, or refine machine learning models 1, 4, 16, 20, 55, 65, etc. using one or more techniques described in this document with respect to Petition 870250038841, dated 05 / 13 / 2025, pages 105 / 142 92 / 96 model alignment toolkit 17. In this way, for example, computing system 50 or server computing system (or systems) 60 can develop, update / train, or refine machine learning models based on local datasets (e.g., for mode personalization / customization, as permitted by user data preference selections).

[0205] Figure 12 is a block diagram of an exemplary computing device 98 that performs according to exemplary embodiments of the present disclosure. Computing device 98 can be a user computing device or a server computing device (e.g., computing device 50, server computing system (or systems) 60, etc.). Computing device 98 can implement model host 31.For example, computing device 98 might include several applications (e.g., applications 1 through N). Each application might contain its own machine learning library and machine learning models. For example, each application might include one machine learning model. Examples of applications include a text messaging application, an email application, a dictation application, a virtual keyboard application, a browser application, etc. As illustrated in Figure 12, each application might communicate with several other components of the computing device, such as one or more sensors, a context manager, a device state component, or additional components. In some implementations, each application might communicate with each device component using an API (e.g., a public API). In some implementations, the API used by each application is specific to that application.

[0206] Figure 13 is a block diagram of an exemplary computing device 99 that has performance according to Petition 870250038841, dated 05 / 13 / 2025, pages 106 / 142 93 / 96 are exemplary embodiments of the present disclosure. Computing device 99 may be the same as or different from computing device 98. Computing device 99 may be a user computing device or a server computing device (e.g., computing device 50, server computing system(s) 60, etc.). Computing device 98 may implement model host 31. For example, computing device 99 may include several applications (e.g., applications 1 to N). Each application is in communication with a central intelligence layer. Examples of applications include a text messaging application, an email application, a dictation application, a virtual keyboard application, a browser application, etc.In some implementations, each application can communicate with the central intelligence layer (and the models stored within it) using an API (for example, a common API across all applications).

[0207] The core intelligence layer may include several machine learning models. For example, as illustrated in Figure 13, a respective machine learning model may be provided for each application and managed by the core intelligence layer. In other implementations, two or more applications share a single machine learning model. For example, in some implementations, the core intelligence layer may provide a single model for all applications. In some implementations, the core intelligence layer is included within or otherwise implemented by a computing device operating system 99.

[0208] The central intelligence layer can communicate with a central device data layer. The central device data layer can be a centralized data repository for the computing device 99. As illustrated in Figure 13, the layer Petition 870250038841, dated 05 / 13 / 2025, pages 107 / 142 94 / 96 central device data can communicate with various other computing device components, such as one or more sensors, a context manager, a device state component, or additional components. In some implementations, the central device data layer can communicate with each device component using an API (e.g., a private API).

[0209] The technology discussed in this document refers to servers, databases, software applications, and other computer-based systems, as well as actions taken and information sent to and from such systems. The inherent flexibility of computer-based systems allows for a wide variety of possible configurations, combinations, and divisions of tasks and functionalities among and within components. For example, the processes discussed in this document can be implemented using a single device or component or multiple devices or components working in combination. Databases and applications can be implemented on a single system or distributed across multiple systems. Distributed components can operate sequentially or in parallel.

[0210] Although the present matter has been described in detail with respect to several specific embodiments thereof, each example is provided by way of explanation, not limitation of the disclosure. Those skilled in the art, upon an understanding of what has been set forth, may easily produce alterations, variations and equivalents to such embodiments. Consequently, the disclosure in question does not preclude the inclusion of such modifications, variations or additions to the present matter as would be readily apparent to a person of ordinary skill in the art. For example, the features Petition 870250038841, dated 05 / 13 / 2025, pp. 108 / 142 95 / 96 illustrated or described as part of one modality may be used with another modality to produce yet another modality. Thus, it is intended that this disclosure covers such alterations, variations and equivalents.

[0211] Aspects of the disclosure have been described in terms of illustrative embodiments thereof. Any and all features in the following claims may be combined or rearranged in any possible manner, including combinations of claims not explicitly enumerated in combination, as the example claim dependencies listed herein should not be read as limiting the scope of possible combinations of features disclosed herein. Consequently, the scope of this disclosure is by way of example, not limitation, and this disclosure does not preclude the inclusion of such modifications, variations, or additions to the present matter, as would be readily apparent to one of ordinary skill in the art. Furthermore, terms are described herein using lists of illustrative elements joined by conjunctions such as “and,” “or,” “but,” etc.It should be understood that such conjunctions are provided for explanatory purposes only. Clauses and other sequences of items joined by a particular conjunction such as "or," for example, may refer to "and / or," "at least one of," "any combination of" illustrative elements listed therein, etc. Terms such as "based on" should be understood as "based on at least partly on."

[0212] The term “can” should be understood as referring to a possibility of a feature in various implementations, and not as prescribing a capability that is necessarily present in every implementation. For example, the sentence “X can perform Y” should be understood as indicating that, in various implementations, X Petition 870250038841, dated 05 / 13 / 2025, pp. 109 / 142 96 / 96 has the potential to be configured to perform Y, and not as indicating that, in every instance, X must always be able to perform Y. It should be understood that, in various implementations, X may be unable to perform Y and remain within the scope of this disclosure.

[0213] The term “can” should be understood as referring to a possibility of a feature in various implementations, and not as prescribing a capability that is necessarily present in every implementation. For example, the phrase “X can perform Y” should be understood as indicating that, in various implementations, X has the potential to be configured to perform Y, and not as indicating that, in every instance, X must always be able to perform Y. It should be understood that, in various implementations, X may be unable to perform Y and remain within the scope of this disclosure.

[0214] This technical description uses examples to disclose the invention, including the improved manner, and also to enable anyone skilled in the art to practice the invention, including the production and use of any devices or systems and the performance of any methods incorporated herein. The patentable scope of the invention is defined by the claims, and may include other examples that occur to those skilled in the art. Such other examples are intended to be covered by the scope of the claims if they include structural elements that do not differ from the literal language of the claims, or if they include equivalent structural elements with insubstantial differences from the literal language of the claims.

Claims

1. A computer-implemented method, characterized in that it comprises: training, in tandem, a prompt expansion model and an image generation model using a multi-reward reinforcement learning model by: processing, by the prompt expansion model, a training query and training context data to generate an expanded training query; generating, by the image generation model, a training image dataset based on the expanded training query; generating a set of reward scores for each image data within the training image dataset using a set of reward models, wherein generating the set of reward scores comprises generating at least one reward score for each reward criterion within a plurality of reward criteria;and adjust weightings and trends associated with the plurality of reward criteria based on the set of reward scores using the multiple reward reinforcement learning model.

2. A computer-implemented method according to claim 1, characterized in that it comprises: obtaining, through a computer system, input data comprising a user query and query context data; generating, using the trained prompt expansion model and the trained image generation model, image data based on the user query and the query context data by: Petition 870250038841, dated 05 / 13 / 2025, page 111 / 142 2 / 7 processing, by the trained prompt expansion model, the user query and the query context data to generate an expanded user query; and generating, using the trained image generation model, the image data based on the expanded user query.

3. A computer-implemented method according to claim 2, characterized in that generating the expanded user query comprises incorporating the plurality of reward criteria into the expanded user query.

4. A computer-implemented method, according to any of the preceding claims, characterized in that the plurality of reward criteria comprises one or more text-image alignment reward criteria.

5. A computer-implemented method according to claim 2 or 4, characterized in that one or more text-image alignment reward criteria comprise: a first text-image alignment reward criterion associated with the training query; and a second text-image alignment reward criterion associated with the expanded training query.

6. A computer-implemented method, according to any of the preceding claims, characterized in that the plurality of reward criteria comprises one or more image sentiment reward criteria.

7. A computer-implemented method, according to any of the preceding claims, characterized in that the plurality of reward criteria comprises one or more aesthetic reward criteria.

8. A computer-implemented method, according to any of the preceding claims, characterized by the fact that the plurality of reward criteria comprises one or more human preference reward criteria.

9. A computer-implemented method, according to any of the preceding claims, characterized in that training the prompt expansion model and the image generation model comprises: selecting a subset of image data from the training image dataset as a function of the reward score set using a non-dominated classification algorithm; and adjusting the weightings and trends associated with the plurality of reward criteria based on the image data subset using the multi-reward reinforcement learning model.

10. A computer-implemented method, according to any of the preceding claims, characterized in that the image data subset comprises a Pareto optimal set associated with the reward score set.

11. A computer-implemented method according to claim 9 or 10, characterized in that training the prompt expansion model and the image generation model comprises: determining, by a computer system, a policy gradient update as a function of the image data subset; and adjusting the weightings and trends associated with the plurality of reward criteria using the multi-reward reinforcement learning model as a function of the policy gradient update.

12. Computer-implemented method according to claim 11, characterized in that determining the policy gradient update further comprises minimizing the reward scores associated with each reward criterion that is not represented within the image data subset.

13. A computer-implemented method according to claim 11 or 12, characterized in that determining the policy gradient update further comprises maximizing the reward scores associated with each reward criterion that is represented within the image data subset.

14. A computer-implemented method, according to any one of claims 11 to 13, characterized in that determining the policy gradient update comprises maximizing one or more text-image alignment reward criteria.

15. A computing system, characterized in that it comprises: one or more processors; and one or more computer-readable transient or non-transient media that store executable instructions to cause the one or more processors to perform operations, wherein the operations comprise: obtaining, by the one or more processors, input data comprising a user query and query context data; training, in tandem, a prompt expansion model and an image generation model using a multi-reward reinforcement learning model, wherein training the prompt expansion model and the image generation model comprises: processing, by the prompt expansion model, a training query and training context data to generate an expanded training query;Generate, using the image generation model, a set of training image data based on the expanded training query; generate a set of reward scores for each image data within the training image data set using a set of reward models, wherein generating the set of reward scores comprises generating at least one reward score for each reward criterion within a plurality of reward criteria; select a subset of image data from the training image data set as a function of the set of reward scores using a non-dominated classification algorithm; and minimize weightings and biases associated with each reward criterion that is not represented within the subset of image data using the multiple reward reinforcement learning model;and generate, using the trained prompt expansion model and the trained image generation model, image data based on the user query and the query context data.

16. A computing system according to claim 15, characterized in that generating the image data further comprises: processing, by the trained prompt expansion model, the user query and the query context data to generate an expanded user query; and generating, by the trained image generation model, the image data based on the expanded user query.

17. Computer system, according to claim 16, characterized in that generating the expanded user query comprises incorporating the plurality of reward criteria Petition 870250038841, dated 05 / 13 / 2025, pp. 115 / 142 6 / 7 in the expanded user query.

18. A computing system according to any one of claims 15 to 17, characterized in that training the prompt expansion model and the image generation model further comprises: determining, by the computing system, a policy gradient update as a function of the image data subset; and adjusting the weightings and trends associated with the plurality of reward criteria using the multi-reward reinforcement learning model as a function of the policy gradient update.

19. A computing system according to claim 18, characterized in that determining the policy gradient update further comprises maximizing the reward scores associated with each reward criterion that is represented within the image data subset.

20. A computer-implemented method characterized in that it comprises: training, in tandem, a prompt expansion model (prompt expansion model) and an image generation model using a multi-reward reinforcement learning model by: processing, by the prompt expansion model, a training query and training context data to generate an expanded training query; generating, by the image generation model, a training image dataset based on the expanded training query; and training, in tandem, the prompt expansion model and the image generation model using the multi-reward reinforcement learning model based on the training image dataset.