Computational budget for generative models
Patent Information
- Application Number
- KR1020267025106
- Authority / Receiving Office
- KR · KR
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-01-31
- Publication Date
- 2026-09-01
Smart Images

Figure P1020267025106_ABST
Abstract
Description
Background Technology
[0001] Generative models have demonstrated state-of-the-art performance across a wide range of tasks, such as text generation (e.g., writing, summarizing, translation, coding), image generation, and audio generation. Generative models utilize very large neural network models trained on vast amounts of data. For example, Large Language Models (LLMs) may include Transformer-based neural network models with self-attention capabilities, enabling general-purpose language understanding and generation in response to queries. Consequently, generative models are being deployed in various applications, for instance, as coding assistants, email composition assistants, and to generate images in presentations.
[0002] However, generative models are often computationally intensive due to their large size, and a given model may use approximately the same amount of computation for any type of task. For example, in a decoder-only transformer-based generative model, the length of the prompt and the length of the output sequence can determine the amount of computation performed by the generative model. Therefore, the amount of computational resources consumed by a generative model when performing different types of tasks can be very large and inflexible. means of solving the problem
[0003] This specification describes a system and a technique for automatically determining and adjusting the amount of computational resources consumed by a generative model when processing user input for a desired task. In particular, the system and the technique determine feature data of the user input (e.g., context data) and process the feature data using a machine learning (ML) model to generate a compute budget corresponding to the settings of the generative model (e.g., the number of rounds of magnetic substrates, the level of planning, image quality, or the number of samples).
[0004] Generally, one innovative aspect of the subject matter of the invention described herein may be implemented by a method comprising acquiring at least two sound signals captured by at least two audio sensors at two different locations at a first time point, generating an input characterizing at least two sound signals, and using a neural network to process the input characterizing at least two sound signals to generate a prediction at the first time point, wherein the prediction comprises (a) individual scores for each of a plurality of sound event classes representing a predicted class for a first sound event detected in at least two sound signals, and (b) a predicted location of a first sound source that emitted the first sound event. Another embodiment of the present aspect comprises a corresponding computer system, device, and a computer program recorded on one or more computer storage devices, each configured to perform an action of the method. One or more computer systems configured to perform a specific operation or action means that software, firmware, hardware, or a combination thereof is installed in the system to cause the system to perform the operation or action when the system is in operation. The statement that one or more computer programs are configured to perform a specific operation or action means that one or more programs include instructions that cause the device to perform the operation or action when executed by a data processing device.
[0005] The above-described embodiments and other embodiments may each optionally include one or more of the following features, either alone or in combination. In particular, one embodiment includes all of the following features in combination. The output includes a predicted level of the computational budget of the generative model for user input, and different levels of the computational budget correspond to different settings of the generative model; the action further includes mapping the predicted level of the computational budget of the generative model for user input to the corresponding settings of the generative model; and processing the user input by the generative model includes processing the user input by the generative model using the corresponding settings of the generative model. The output includes a setting of the generative model that defines the computational budget of the generative model for user input, and processing the user input by the generative model according to the computational budget defined by the output includes processing the user input by the generative model using the settings of the generative model. The computational budget defines a number of rounds of self-critique that the generative model uses to process the user input. The computational budget defines a level of planning that a generative model uses to process user input. The computational budget defines the number of samples that a generative model generates from user input. The computational budget defines the quality level of the generated output that a generative model generates from user input. The generated output is an image, and the quality level of the generated output includes the resolution of the image. The feature data of the user input includes the context data of the user input, and processing the feature data involves processing the context data of the user input using a machine learning model to generate an output that defines the computational budget of the generative model for the user input.User input includes a request to generate output text, the generative model includes a large-scale language model (LLM) trained to generate text data, and feature data is context data in which the output text will be used. User input includes a request to generate an output image, the generative model includes a generative image model trained to generate image data, and feature data includes features of one or more images within the context data in which the output image will be used. Feature data of the user input includes the importance level of the response to the user input. Feature data of the user input includes the difficulty level of the task defined by the user input. An action includes updating an output that defines the computational budget of the generative model for the user input using the metadata of the generative model; and processing the user input by the generative model based on the computational budget defined by the updated output.
[0006] Specific embodiments of the subject matter of the invention described in this specification may be implemented to realize one or more of the following advantages.
[0007] The system and method described herein increase the computational efficiency of a generative model because the amount of computational resources consumed by the generative model can be adjusted for the desired quality expected by the user according to the characteristics of the task, e.g., task difficulty, task importance, or context. For example, the system and method can perform more computations using the generative model for difficult tasks or inputs that seek high-quality output, and can perform fewer computations using the generative model for easy tasks or inputs that do not seek high-quality output.
[0008] Some existing techniques allocate the same amount of computational resources to all requests, and thus, the generated output for difficult or critical tasks may need to be regenerated multiple times until a satisfactory result is obtained. The system and method can avoid situations where the output from a generative model needs to be regenerated by determining the expected quality of the output based on the context of the input query and adjusting the amount of computational resources consumed by the generative model based on the expected quality of the output. Since inference by a generative model can incur very high computational costs considering how large the generative model is, the system and method can save a significant amount of computational costs. In some embodiments, by reducing or eliminating the need to regenerate the output, the system can generate output of better quality within a fixed amount of available computational resources. In some embodiments, the system and method can adjust the quality of the generated output through the computational budgeting mechanism disclosed herein. For example, the system and method can achieve a target quality level of the generated output by adjusting the computational budget by a percentage corresponding to the target level.
[0009] Details regarding one or more embodiments of the subject matter of the invention described herein are provided in the accompanying drawings and the following description. Other features, aspects, and advantages of the subject matter of the invention will become apparent from the detailed description, drawings, and claims. Brief explanation of the drawing
[0010] Figure 1 is a diagram of an exemplary system. Figure 2a is a diagram of an exemplary system for generating output text. Figure 2b is a diagram of an exemplary system for generating an output image. Figure 2c is a diagram of an exemplary system for generating an output dialogue. Figure 3 is a flowchart of an exemplary process for using a generative model according to the computational budget. In various drawings, the same reference number and name represent the same element. Specific details for implementing the invention
[0011] FIG. 1 is a diagram of an exemplary system (100). The system (100) is an example of a system implemented by a computer program on one or more computers located at one or more locations where the system, components, and techniques described herein are implemented. One or more computers may include personal computers, mobile communication devices, servers, and other devices capable of sending and receiving data over a network. A network (not shown), such as a local area network ('LAN'), a wide area network ('WAN'), the Internet, or a combination thereof, connects one or more computers implementing the system. The system may use multiple computers operating together, including a single computer or, for example, a set of remote computers deployed as a cloud computing service.
[0012] The system (100) dynamically determines the amount of computational resources that the generative model consumes for each user input. For example, the system (100) may use the generative model to perform more computations for difficult tasks or inputs that seek high-quality output, and may use the generative model to perform fewer computations for easy tasks or inputs that do not seek high-quality output.
[0013] A generative model is a machine learning (ML) model that generates content including text, images, audio, or other synthetic data based on input. A generative model (114) can be a very large neural network model and can be trained on a vast amount of data. As a result, after training, performing inference using the generative model can be very computationally expensive.
[0014] During inference, the generative model (114) can generate a generative output (116), for example, content of a specific type, in response to a query. A query is a question or a search for specific information. In some implementations, the generative model (114) can generate multimodal output, such as an image and corresponding text describing the image.
[0015] In some embodiments, the generative model (114) may be configured to process an input sequence of tokens to generate an output sequence of tokens. Tokens may represent any suitable type of content, e.g., text, images, videos, audio, or some combination thereof. For example, the generative model may be a large-scale language model (LLM) and may be configured to process an input sequence of tokens from a vocabulary of text tokens to generate an output sequence of tokens from the vocabulary.
[0016] More generally, the generative model (114) may be any suitable neural network that receives an input sequence consisting of text tokens selected from types of content and autoretroactively generates an output sequence consisting of text tokens from types of content. For example, the generative model (114) may be a transformer-based language model neural network or a recurrent neural network-based language model neural network.
[0017] In some situations, when the neural network used to implement the language model autoregressively generates an output sequence of tokens, the generative model (114) may be referred to as an autoregressive neural network. More specifically, the autoregressively generated output is created by generating each specific token of the output sequence conditioned on the current input sequence, which includes a context input that provides context for the output sequence and a token already generated for any previous position before a specific token of the output sequence, that is, a token already generated for any previous position before a specific position of the specific token in the output sequence.
[0018] For example, when generating a token at a given position in an output sequence, the current input sequence may include the input sequence and a token at any previous position preceding the given position in the output sequence. As a specific example, the current input sequence may include the input sequence and a token at any previous position preceding the given position in the output sequence. Optionally, the input sequence and the current output sequence may be separated into one or more predetermined tokens within the current input sequence.
[0019] More specifically, to generate a specific token at a specific location within an output sequence, the generative model (114) can process the current input sequence to generate a score distribution (e.g., a probability distribution) that assigns an individual score, e.g., an individual probability, to each token in the token vocabulary. Then, the language model neural network can use the score distribution to select a token from the vocabulary as a specific token. For example, the language model neural network can carefully select the token with the highest score, or it can sample tokens from the distribution using, for example, nucleus sampling or other sampling techniques.
[0020] As a specific example, the generative model (114) may be an autoregressive transformer-based neural network comprising (i) a plurality of attention blocks each applying a self-attention operation and (ii) an output subnetwork that processes the output of the last attention block to generate a score distribution.
[0021] 생성형 ન(114)은 임의읝 다양한 읰배 신경맸 아키텍처를 수 있다. snowflake snowflakes[J. Hoffmann , S. Borgeaud , A. Mensch , E. Buchatskaya , T. Cai , E. Rutherford , D. d. L. Casas, LA Hendricks, J. Welbl, A. Clark, et al. Training Compute-Optimal Large Language Models, arXiv Preprint arXiv:2203.15556, 2022; JW Rae, S Borgeaud, T Cai, K Millican, J Hoffmann, HF Song, J Aslanides, S Henderson, Ring, S Young, E Rutherford, T Hennigan, J Menick, A Cassirer, R Powell, G van den Driessche, LA Hendricks, M Rauh, P Huang; Glaese , J Welbl , S Dathathri , S Huang , J Uesato , J Mellor , I Higgins , A Creswell , N McAleese , A Wu , E Elsen , SM Jayakumar , E Buchatskaya , D Budden , E Sutherland , K Simonyan , M Paganini , L Sifre , L Martens , XL Li , A . Kuncoro , A. Nematzadeh , E. Gribovskaya , D. Donato , A. Lazaridou , A. Mensch , J. Lespiau , M. Tsimpoukelli , N. Grigorev , D. Fritz , T. Sottiaux , M. Pajarskas , T. Pohlen , Z. Gong , D. Toyama , C. de Masson d'Autume , Y. Li , T. TerziMikulik, I. Babuschkin, A. Clark, D. de Las Casas, A. Guy, C. Jones, J. Bradbury, M. Johnson, B. A. Hechtman, L. Weidinger, I. Gabriel, W. S. Isaac, E. Lockhart, S. Osindero, L. Rimell, C. Dyer, O. Vinyals, K. Ayoub, J. Stanway, L. Bennett, D. Hassabis, K. Kavukcuoglu, and G. Irving. Scaling language models: Methods, analysis & insights from training gopher. CoRR, abs / 2112.11446, 2021; Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. Exploring the limits of transfer learning with a unified text-to-text transformer. arXiv preprint arXiv:1910.10683, 2019; Daniel Adiwardana, Minh-Thang Luong, David R. So, Jamie Hall, Noah Fiedel, Romal Thoppilan, Zi Yang, Apoorv Kulshreshtha, Gaurav Nemade, Yifeng Lu, and Quoc V. Le. Towards a human-like open-domain chatbot. CoRR, abs / 2001.09977, 2020; and Tom B Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. Language models are few-shot learners. Includes those described in arXiv preprint arXiv:2005.14165, 2020].
[0022] In some embodiments, the generative model (114) may use a decoder-only architecture that includes many decoder blocks without using an encoder. Each decoder block may include a self-attention layer and a feed-forward neural network. A transformer-based generative model is an example of a generative model to which the systems and techniques of this specification may be applicable.
[0023] The systems and techniques described in this specification are applicable to other types of generative models, such as diffusion models, Bayesian networks, Generative Adversarial Networks (GANs), and Variable Autoencoders (VAEs).
[0024] An example of a generative model (114) may be a latent diffusion model.
[0025] As another example, the generative model (114) may be a diffusion model that generates a first image using a text-to-image diffusion model and then generates a final image by applying one or more super-resolution diffusion models.
[0026] As another example, the generative model (114) may be an autoregressive generative model that autoregressively generates tokens representing audio, video, images, or other data.
[0027] As another example, the generative model (114) may be a masked token generative model that sequentially unmasks tokens representing text, video, audio, image, or other data during generation.
[0028] The computational budget of a generative model defines the amount of computational resources consumed by the generative model when processing user input. The system (100) can adjust the amount of computational resources consumed by the generative model so that the amount is maintained within the computational budget of the generative model in any various ways.
[0029] In some embodiments, the computational budget (112) may define the number of rounds of self-criticism that the generative model (114) uses to process user input (102). The generative model analyzes and compares its initial outputs through self-criticism and then returns a final response. For example, in one round of self-criticism, the LLM may generate one or more initial outputs, analyze one or more initial outputs to identify errors, and then attempt to correct one or more initial outputs or regenerate a new output based on such errors before providing the final output. More rounds of self-criticism mean performing more inference steps and consuming more computational resources, but it can generate a generative output of higher quality or better than the generative output generated using fewer rounds of self-criticism. For example, the LLM may generate higher or lower quality output text by performing more or fewer rounds of self-criticism.
[0030] In some embodiments, the computational budget (112) may define the level of planning that the generative model (114) uses to process user input (102). The generative model (114) may use hierarchical planning. In hierarchical planning, a task is broken down into smaller tasks, and then broken down into smaller tasks, and so on. After forming a hierarchical plan, the system (100) can find a path through the hierarchical plan to solve the task by determining the level of planning according to the computational budget (112). For example, the system (100) may instruct the generative model (114) to construct a more detailed plan, for example, to divide the task into smaller tasks, which may result in using more steps, such as more detailed inference traces. Compared to a generative model that directly generates output using fewer steps, the generative model (114) is more likely to generate high-quality output through more steps associated with the inference traces of the generative algorithm. In some implementations, using a higher level of planning can increase accuracy for tasks requiring inference at an additional computational cost. In some implementations, using a lower level of planning can improve task resolution speed at the cost of reduced quality, such as by leaving subtasks to the user instead of resolving them.
[0031] In some embodiments, the computational budget (112) may define the number of samples that the generative model (114) generates from user input (102). For example, the system (100) may generate multiple samples using the generative model (114). The system (100) may increase or decrease the number of samples generated by the generative model. In some embodiments, multiple samples may be presented to the user on a device, thus providing the user with more choices. In some embodiments, the system may use multiple samples for other processing, such as self-consistency evaluation, which automatically selects one of the samples to be presented to the user as output. In some embodiments, the system may measure the quality (and / or variety) of multiple samples using one or more metrics. Based on the quality measurement, the system may select one or more samples from the multiple samples and provide one or more samples to the user as output to be presented to the user on a user device, for example.
[0032] In some embodiments, the computation budget (112) may define the quality level of the generated output (116) that the generative model (114) generates from the user input (102), for example, the resolution of the image that the generative model (114) generates from the user input (102). For example, the system (100) may generate high-quality generated output (116). The type of quality may be determined by the type of generative model (114) used by the system (100). In some embodiments, the image or audio generating model may be configured to generate finer tokens that capture finer details of the image or audio, and thus generate a higher resolution image output or audio output with a lower compression rate. For example, the generative model (114) may include an audio generative model such as AudioLM (https: / / google-research.github.io / seanet / audiolm / examples / ) that uses Residual Vector Quantization (RVQ) tokenization. The system (100) can generate audio output with different levels of compression by adjusting the RVQ level of the audio generative model. Thus, the audio output may have finer details or coarser details.
[0033] Some image generation models may include multiple upsampling stages configured to generate images at multiple different resolutions. The system may adjust the number of upsampling stages for running the image generation model based on a computational budget representing the desired resolution of the output image. For example, the image generation model may include three upsampling stages configured to generate images at three different resolutions, for example, 128x128, 256x256, and 512x512. The system (100) may determine, based on the computational budget (112), that an output image of 256x256 is sufficient. The system may process user input using the upsampling stage of the image generation model that generates a 256x256 image, rather than using the upsampling stage of the image generation model that generates a 512x512 image.
[0034] In some embodiments, based on a computational budget (112), the system (100) may determine from a plurality of generative models a generative model (114) to be used to process user input (102). For example, the system (100) may decide to use a larger or smaller generative model depending on the computational budget (112).
[0035] In some embodiments, the generative model (114) may be a diffusion model, and the computational budget (112) may define the number of denoising steps performed by the diffusion model. For example, the diffusion model may perform a large number of denoising steps if the computational budget is high, which indicates that a high-quality output is required. In some cases, the diffusion model may perform a small number of denoising steps if the computational budget is low, which indicates that a high-quality output is not required.
[0036] In some embodiments, based on the computation budget (112), the system (100) may control the amount of tokens used in the generative model, which may affect the computation cost or the quality of the generated output in a predictable manner. For example, some generative models may control the number of tokens for the use of one or more tools for specific features, for example. In some embodiments, the system (100) may control the tokens allowed during the decoding process to enable or disable specific features. Thus, the system (100) may directly control the computation cost of the generative model, the quality of the generated output, or both in a predictable manner.
[0037] In some embodiments, based on the computational budget (112), the system may adjust the input prompt for the generative model, for example, after several steps of generating the output, or at the start time. For example, the prompt may include a detailed description of the desired output, which allows for a higher quality output to be obtained at the cost of higher computation.
[0038] The system (100) receives user input (102) to be processed by a generative model (114). User input (102) may be input for an operation that can be performed by the generative model, such as text generation, image generation, or audio generation. In some embodiments, user input (102) may include text data, image data, video data, audio data, or a combination of some of these. User input (102) may include a query to be processed by the generative model (114). For example, user input (102) may be "Write a thank-you email to everyone who attended the meeting." In some implementations, user input (102) may include a prompt, which may be natural language text, requesting the generative model (114) to perform a specific operation. The prompt may rephrase the query, style it, provide relevant context, or provide examples.
[0039] In some embodiments, the system (100) may receive user input (102) from the user interface (UI) of the application. For example, an email composition system may receive user input indicating an action to create a draft email. A document composition system may receive user input indicating an action to rewrite a paragraph in a document. An image creation system may receive user input indicating an action to create a new image to be used in a presentation. A conversational system using LLM may receive user input on various topics.
[0040] To generate a calculation budget (112), the system (100) may consider a number of different features of user input (102).
[0041] In particular, the system (100) may include a feature determination engine (104) that generates feature data (106) of user input (102).
[0042] In some embodiments, feature data (106) may be context data of user input (102). For example, context data may be context signals within the application being used. For example, context data may be multiple recipients in a distribution list of a draft email, multiple collaborators in a document, the quality of an existing image in the rest of a presentation, or the time spent editing text.
[0043] In some embodiments, feature data (106) may include hint data included in or inferred from user input (102). For example, feature data (106) may be the difficulty level of a task defined by user input, or the importance level of a response to user input, or both. The feature determination engine (104) may process user input (102) to determine an explicit or implicit queue or hint.
[0044] For example, the feature determination engine (104) can process the user input (102) to predict whether the user input (102) provides any hint that the generated output (116) is particularly important. In some embodiments, the feature determination engine (104) can predict the difficulty of the task by processing historical interaction data between the system (100) and the user. For example, the feature determination engine (104) can predict the difficulty of the task based on how often the user accepts outputs for similar queries.
[0045] The system (100) includes a computational budget determination engine (108) that generates a computational budget (112) for generating a generated output (116) using a generative model (114).
[0046] In some embodiments, the computation budget determination engine (108) can automatically determine the computation budget (112) for generating the generated output (116) using a generative model (114) based on the feature data (106) of the user input (102).
[0047] In some embodiments, the computational budget determination engine (108) can use the model (110) to process feature data (106) of the user input (102) to generate an output that defines the computational budget (112) of the generative model (114) for the user input (102).
[0048] In some embodiments, the model (110) may include one or more machine learning (ML) models, such as a classification model or a regression model. In some embodiments, the model (110) may be an LLM trained to understand the contextual features of the user input (102) and to generate a computational budget (112) for processing the user input (102) using a generative model (114).
[0049] The model (110) may be trained using supervised learning on labeled training examples. In some embodiments, the system may train the ML model on user feedback data. The user feedback data may include data indicating whether the generated output was accepted. The system may maintain a list of user feedback data items, and each data item may include a query, a computational budget for the query, and feedback data for the generated output produced according to the computational budget.
[0050] For example, when a user requests that the generated output for a query be improved or regenerated, the system can determine feedback data indicating that the computational budget for generating the output was insufficient. For instance, the user may provide feedback such as "This isn't very good" and / or "Could you regenerate it?". The system can determine feedback data indicating an insufficient computational budget.
[0051] For example, when there is missing information, such as when a user provides feedback like "Thank you, but could you add X as well," the system can determine feedback data indicating that there is missing information and that the computational budget was unnecessarily consumed by the query. The feedback data may suggest that it would have been better to output the answer using a smaller computational budget and fewer computational resources.
[0052] In some embodiments, the system may use a search strategy to generate various feedback to be included in the training data for training the model (110). For example, the system may randomly increase or decrease the computational budget for a query and obtain feedback data indicating how the user responds to the corresponding generated output.
[0053] In some implementations, the model (110) may be trained using online learning. Instead of training the model on the entire training data set, the system may train the model (110) as training data becomes available in a sequential order. Thus, the system may update the model (110) for future user input at each step. For example, the system may train the model (110) when user feedback data becomes available and may use a generative model (114) to provide a better estimate of the computational budget for processing future user input.
[0054] In some embodiments, the output defining the computational budget (112) may be a predicted level of the computational budget (112) of the generative model (114) for the user input (102). The model (110) may be a classification model trained to generate likelihood scores for multiple levels of the computational budget. For example, the output may be one of high, medium, or low levels.
[0055] In some embodiments, the output defining the computation budget (112) may be one or more settings of the generative model (114). In some embodiments, the output may be a value for a set of parameters regarding how inference should be performed by the generative model (114). The model (110) may be a regression model trained to generate a predicted value for each of one or more parameters regarding how inference should be performed by the generative model (114).
[0056] The system (100) processes user input (102) by a generative model (114) according to a computational budget (112) to generate a generated output (116). The system (100) may use the computational budget (112) to parameterize the inference of the generative model (114) on the user input (102). For example, the system (100) may apply more or fewer rounds of self-criticism to the output of the generative model (114). The system (100) may apply more or fewer planning. The system (100) may use the generative model (114) to sample more or fewer times. The system (100) may generate outputs of higher or lower quality.
[0057] After performing inference of the generative model (114) according to the computational budget (112), the system (100) can obtain a generated output (116). The system can present the generated output (116) on the device to the user.
[0058] FIG. 2a is a diagram of an exemplary system (200A) for generating output text. A user is composing an email to be sent to a wide distribution list. For example, the number of recipients of the email may be greater than a threshold. The system (200A) receives a request (202) from the user's device to generate output text, for example, to rephrase a paragraph in the email.
[0059] The feature determination engine (204) can determine the context data (206) of the request (202). For example, the system (200A) can access the 'To:' field of the email and determine that the email's distribution list is extensive, for example, that the number of recipients of the email is greater than a threshold.
[0060] Context data (206) may include various features of the request (202), such as the email distribution list, the email subject line, the email body, and the email signature. For example, the email subject line may include words or phrases describing the email, such as 'IMPORTANT' or 'URGENT,' indicating the importance or urgency of the email. The email body may indicate the style of the email, such as using formal business language or casual language. Whether the email includes a signature may indicate whether the email is a formal business email or a casual personal email.
[0061] Based on context data (206), the computational budget decision engine (208) may use LLM (214) to decide to use more computational budget (212) to process the request (202). The computational budget decision engine (208) may process the context data (206) using an ML model (210) to produce an output indicating a higher level of computational budget (212). In some implementations, the computational budget decision engine (208) may process the context data (206) using an ML model (210) to generate a value for a setting to process the request (202) using LLM (214), for example, using more rounds of self-criticism.
[0062] In some embodiments, an ML model (210), such as an LLM, may be trained to achieve linguistic understanding of contextual data. For example, the LLM may be trained to predict the importance or desired quality of an email by processing text data such as the email's distribution list, the email's subject line, the email's body, and the email's signature. Additionally, the ML model (210) may be trained to generate a computational budget based on the prediction of the email's importance or desired quality. For example, if the email's desired quality is high, the ML model (210) may generate a high level of computational budget. If the email's desired quality is average, the ML model (210) may generate a normal level of computational budget.
[0063] The system (200A) processes the request (202) by the LLM (214) according to the computational budget (212) to generate output text (216). For example, the system can generate additional suggestions for a method to improve a paragraph, a more improved version of a paragraph after multiple rounds of self-criticism, or both.
[0064] FIG. 2b is a diagram of an exemplary system (200B) for generating an output image. A designer is creating a presentation for a client. The presentation already includes some existing images. The system (200B) receives a request (222) from the user's device to generate an output image, for example, to generate a picture of a cat on the moon to be used in the presentation.
[0065] The feature determination engine (224) can determine context data (226) of the request (222). In some embodiments, the context data (226) may include an existing image in the presentation. The quality or complexity of the existing image may indicate the importance, complexity, or desired quality of the output image. In some embodiments, the context data (226) may include other text data before and after the location of the request (222) and the requested image. The text data may indicate the importance, complexity, or desired quality of the output image.
[0066] Based on context data (226), the computational budget determination engine (228) may determine the computational budget (232) to process the request (222) using an ML model (230) and an image generation model (234). In some embodiments, the ML model (230) may be a computer vision ML model trained to achieve image understanding of an existing image in a presentation. For example, a convolutional neural network model may be trained to predict the importance, complexity, or desired quality of an output image (236) by processing an existing image. In some embodiments, the ML model (230) may include an LLM trained to predict the importance, complexity, or desired quality of a presentation by processing text data, such as the request (222), and other text data before and after the location of the requested image. Additionally, the ML model (230) may be trained to generate a computational budget based on the prediction of the importance, complexity, or desired quality of the presentation.
[0067] For example, because the context data (226) indicates that an existing image within the presentation has high quality, the system (200B) may determine that a high-quality image would be desirable to match with the rest of the presentation. The computational budget determination engine (228) may process the context data (226) using an ML model (230) to generate an output that indicates a high level of computational budget (232). In some embodiments, the computational budget determination engine (228) may process the context data (226) using an ML model (230) to generate a value for the setting that processes the request (222) using an image generation model (234), for example, a resolution of the output image (236) that is higher than a threshold.
[0068] The system (200B) generates an output image (236) by processing a request (222) by an image generation model (234) according to a computational budget (232). For example, the system can generate a higher resolution image that corresponds to a longer token sequence for the image generation model (234). The system (200B) can insert the high-quality output image (236) into a presentation.
[0069] FIG. 2c is a diagram of an exemplary system (200C) for generating an output conversation. A user is conversing with an LLM (254), for example, in a chatbot application. The system (200C) receives a conversation (242) with the chatbot from the user's device. For example, the user initiates a conversation with the LLM (254) about a first topic, for example, planning a vacation trip. As the conversation continues, the conversation may indicate the importance or desired quality of the response to the first topic. For example, as the conversation continues, the user conveys some preferences regarding hotels and locations for the vacation trip. However, as the conversation progresses, the user is unable to make a decision, and the chatbot highlights various aspects of the vacation choice. For example, the location may be crowded, or tourists may complain about a specific aspect of the trip.
[0070] The feature determination engine (244) can determine the context data (246) of the conversation (242). The system (200C) can predict the importance or desired quality of the response to the first topic based on the historical conversation. For example, the system can process the historical conversation using, for example, LLM and predict a reduced level of importance for the first topic, for example, planning a vacation trip. For example, the system can determine that helping the user decide what they want themselves, rather than creating a perfectly planned trip, provides greater benefit to the user.
[0071] Based on context data (246), the computational budget determination engine (248) can process the conversation (242) using an LLM (254) by determining the computational budget (252) using an ML model (259). In some embodiments, the ML model (250) may include an LLM trained to predict the importance or desired quality of a response to a first topic by processing a historical conversation, for example, a historical conversation indicating a reduced level of importance in creating a plan for a vacation trip. Additionally, the ML model (250) may be trained to generate a computational budget based on the prediction of the importance or desired quality of the response to the first topic.
[0072] For example, because the context data (246) indicates that the level of importance of planning for the first topic, e.g., a vacation trip, has decreased, the system (200C) may determine that a perfectly planned trip may not be necessary. The computational budget determination engine (248) may process the context data (246) using an ML model (250) to generate an output indicating a low level of computational budget (252). In some embodiments, the computational budget determination engine (248) may process the context data (246) using an LLM (254) to generate a value for a setting to process the conversation (242) using, for example, less planning or fewer samplings.
[0073] The system (200C) processes the conversation (242) by the LLM (254) according to the computational budget (252) to generate an output conversation (256). For example, the system can generate a less detailed plan for a vacation trip. In some embodiments, the system can generate suggestions for other alternative activities for spending the vacation.
[0074] FIG. 3 is a flowchart of an exemplary process (300) for using a generative model according to a computational budget. The process will be described as being performed by a properly programmed computer system such as the system (100).
[0075] The system receives user input to be processed by a generative model (302). The system determines feature data of the user input (304).
[0076] In some implementations, user input may include a request to generate output text, and the generative model may include a large-scale language model (LLM) trained to generate text data. Feature data may be context data in which the output text will be used.
[0077] In some implementations, user input may include a request to generate an output image, and the generative model may include a generative image model trained to generate image data. Feature data may include features of one or more images within context data in which the output image will be used.
[0078] In some embodiments, the feature data of the user input may include a level of importance of the response to the user input. In some embodiments, the feature data of the user input may include a level of difficulty of the task defined by the user input.
[0079] The system uses a machine learning model to process feature data of user input to generate an output that defines the computational budget of the generative model for the user input (306). The computational budget of the generative model defines the amount of computational resources consumed when the generative model processes the user input.
[0080] In some implementations, the computational budget may define the number of rounds of self-criticism that the generative model uses to process user input. In some implementations, the computational budget may define the level of planning that the generative model uses to process user input. In some implementations, the computational budget may define the number of samples that the generative model generates from user input. In some implementations, the computational budget may define the quality level of the generated output that the generative model generates from user input. In some implementations, the generated output may be an image, and the quality level of the generated output may include the resolution of the image.
[0081] In some embodiments, feature data of user input may include context data of user input, and processing the feature data may include processing the context data of user input using a machine learning model to generate an output that defines the computational budget of a generative model for user input.
[0082] The system processes user input by a generative model according to the computational budget defined by the output to generate a generated output (308).
[0083] In some embodiments, the output may include a predicted level of the computational budget of the generative model for user input, and different levels of the computational budget may correspond to different settings of the generative model. The system may map the predicted level of the computational budget of the generative model for user input to the corresponding settings of the generative model. The system may process the user input by the generative model using the corresponding settings of the generative model.
[0084] In some implementations, the output may include settings of the generative model that define the computational budget of the generative model for user input. The system can process user input by the generative model using the settings of the generative model.
[0085] In some implementations, the system can use metadata from the generative model to update an output that defines the computational budget of the generative model for user input. The system can process the user input by the generative model based on the computational budget defined by the updated output. The initial output may define an estimated computational budget to generate a generated output of target quality, taking into account the characteristics of the user input. However, some generative models can generate this target quality more easily than others.
[0086] Metadata of a generative model can include available task libraries (e.g., a collection of prompts), training data used to train the generative model, or a combination of both. Metadata of a generative model can indicate how well a task is represented in the model's training data. For example, solving the (N+1)th mathematical benchmark may be easier if the model already has N mathematical benchmarks in its training data or task library than if the same model has N summary benchmarks. Based on the generative model's metadata, the system can perform improvements to the output that define the computational budget. Metadata of the generative model can be used to increase, decrease, or maintain the computational budget unchanged.
[0087] In some implementations, the system may use user feedback data to improve an output that defines a computational budget for subsequent processing using a generative model in response to user feedback. After generating a generative output based on the computational budget, the system may receive user feedback regarding the generated output. Based on the user feedback, the system may determine a computational budget for generating the next generative output.
[0088] For example, based on user feedback, the system may determine that the previous computational budget was insufficient and decide to increase the computational budget to produce higher-quality output. As another example, based on user feedback, the system may determine that the previous computational budget was more than sufficient and decide to decrease the computational budget to produce lower-quality output.
[0089] This specification uses the term "configured" in relation to systems and computer program components. "One or more computer systems are configured to perform specific operations or actions" means that software, firmware, hardware, or a combination thereof is installed in the system to cause the system to perform the operations or actions when the system is in operation. "One or more computer programs are configured to perform specific operations or functions" means that one or more programs include instructions that cause the device to perform the operations or functions when executed by a data processing device.
[0090] Embodiments of the subject matter of the invention and functional operations described herein may be implemented in digital electronic circuits, in tangibly implemented computer software or firmware, in computer hardware comprising structures and their structural equivalents disclosed herein, or in a combination of one or more of these. Embodiments of the subject matter of the invention described herein may be implemented as one or more computer programs, namely, as one or more modules of computer program instructions encoded on a tangible non-transient storage medium for execution by a data processing device or for controlling the operation of a data processing device. The computer storage medium may be a machine-readable storage device, a machine-readable storage substrate, a random or serial access memory device, or a combination of one or more of these. Alternatively or additionally, program instructions may be encoded on an artificially generated radio signal, for example, a machine-generated electrical, optical, or electromagnetic signal generated to encode information for transmission to a receiver device suitable for execution by a data processing device.
[0091] The term 'data processing device' refers to data processing hardware and encompasses all types of devices, devices, and machines for processing data, including, for example, programmable processors, computers, or multiple processors or computers. The device may also be an off-the-shelf or custom-made parallel processing subsystem, for example, a GPU or other types of special-purpose processing subsystem, or may additionally include such. The device may also be a special-purpose logic circuit, for example, a field programmable gate array (FPGA) or an application-specific integrated circuit (ASIC), or may additionally include such. Optionally, in addition to the hardware, the device may include code that creates an execution environment for computer programs, for example, processor firmware, a protocol stack, a database management system, an operating system, or a combination of one or more of these.
[0092] A computer program, which may be referred to or described as a program, software, software application, app, module, software module, script, or code, may be written in any form of programming language, including compiled or interpreted languages, or declarative or procedural languages, and may be distributed in any form, including a standalone program, a module, a component, a subroutine, or other units suitable for use in a computing environment. A program may correspond to a file within a file system, but is not required to do so. A program may be stored in a part of a file containing one or more scripts stored in a markup language document, in a single file dedicated to the program, or in multiple coordinated files, for example, files storing one or more modules, subprograms, or parts of code. A computer program may be distributed to be executed on a single computer or a single site, or on multiple computers distributed across multiple sites and interconnected by a data communication network.
[0093] As used herein, “engine” or “software engine” refers to a software-implemented input / output system that provides an input and a different output. An engine may be an encoded block of function, such as a library, platform, software development kit (“SDK”), or object. Each engine may be implemented on any suitable type of computing device, e.g., servers, mobile phones, tablet computers, notebook computers, music players, e-book readers, laptop or desktop computers, PDAs, smartphones, or other stationary or portable devices comprising one or more processors and computer-readable media. Additionally, two or more of the engines may be implemented on the same computing device or on different computing devices.
[0094] The processes and logic flows described herein may be performed by one or more programmable computers that execute one or more computer programs to perform functions by operating on input data and generating outputs. The processes and logic flows may also be performed by special purpose logic circuits, e.g., FPGAs or ASICs, or by a combination of special purpose logic circuits and one or more programmed computers.
[0095] Computers suitable for executing computer programs may be based on general-purpose or special-purpose microprocessors, or both, or any other type of central processing unit. Generally, the central processing unit will receive instructions and data from read-only memory or random access memory, or both. The essential components of a computer are the central processing unit for executing or performing instructions and one or more memory devices for storing instructions and data. The central processing unit and memory may be supplemented by or integrated into special-purpose logic circuits. Generally, the computer will also include one or more mass storage devices for storing data, such as magnetic, magneto-optical disks, or optical disks, or will be operably coupled to receive data from them, transmit data to them, or perform both. However, the computer does not need to have such devices. In addition, computers can be embedded in other devices, for example, to name just a few, mobile phones, personal digital assistants (PDAs), mobile audio or video players, game consoles, Global Positioning System (GPS) receivers, or portable storage devices, for example, Universal Serial Bus (USB) flash drives.
[0096] Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and memory devices, including, for example, semiconductor memory devices, e.g., EPROM, EEPROM, and flash memory devices; magnetic disks, e.g., internal hard disks or removable disks; magneto-optical disks; and CD-ROM and DVD-ROM disks.
[0097] To provide interaction with a user, embodiments of the subject matter of the invention described herein may be implemented on a computer having a presence-sensing display or other surface, such as a display device for displaying information to a user, e.g., a cathode ray tube (CRT) or liquid crystal display (LCD) monitor, and a keyboard and pointing device, e.g., a mouse, a trackball, or a keyboard and pointing device for which a user can provide input to the computer. Other types of devices may also be used to provide interaction with a user, for example, the feedback provided to the user may be any form of sensory feedback, such as visual feedback, auditory feedback, or tactile feedback, and input from the user may be received in any form including sound, voice, or tactile input. Additionally, the computer may interact with the user by transmitting documents to a device used by the user and receiving documents from the device, for example, by transmitting a web page to a web browser on the user's device in response to a request received from a web browser. Additionally, the computer may interact with the user by transmitting text messages or other forms of messages to a personal device, e.g., a smartphone, running a messaging application, and receiving response messages from the user in return.
[0098] Although this specification contains many specific implementation details, they should not be interpreted as limitations on the scope of any invention or claimable scope, but rather as descriptions of features that may be specific to specific embodiments of specific inventions. Specific features described in this specification in the context of separate embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented individually or in any suitable sub-combination in multiple embodiments. Furthermore, while features may be described above as operating in specific combinations and may even be initially claimed as such, one or more features from the claimed combination may be omitted from the combination in some cases, and the claimed combination may relate to a sub-combination or a variation of a sub-combination.
[0099] Similarly, although operations are depicted in a specific order in the drawings, this should not be understood as requiring that such operations be performed in the specific order depicted or in a sequential order, or that all illustrated operations be performed, in order to achieve desirable results. In certain situations, multitasking and parallel processing may be advantageous. Furthermore, the separation of various system modules and components in the aforementioned embodiments should not be understood as requiring such separation in all embodiments, and it should be understood that the described program components and systems can generally be integrated together into a single software product or packaged into multiple software products.
[0100] Specific embodiments of the subject matter of the invention have been described. Other embodiments fall within the scope of the following claims. For example, the actions described in the claims may be performed in different orders and still achieve desirable results. As an example, the processes illustrated in the accompanying drawings do not necessarily require the specific order or sequential order illustrated to achieve desirable results. In some cases, multitasking and parallel processing may be advantageous.
Claims
Claim 1 A method performed by one or more computers, comprising: receiving user input to be processed by a generative model; determining feature data of said user input; processing said feature data of said user input using a machine learning model to generate an output that defines a computational budget of said generative model for said user input, wherein said computational budget of said generative model defines the amount of computational resources consumed by said generative model when processing said user input; and processing said user input by said generative model according to said computational budget defined by said output to generate a generated output. Claim 2 In claim 1, the output includes a predicted level of the computational budget of the generative model for the user input, and different levels of the computational budget correspond to different settings of the generative model, and the method further includes the step of mapping the predicted level of the computational budget of the generative model for the user input to a corresponding setting of the generative model, and the step of processing the user input by the generative model includes the step of processing the user input by the generative model using the corresponding setting of the generative model. Claim 3 In claim 1, the output includes a setting of the generative model that defines the computational budget of the generative model for the user input, and the step of processing the user input by the generative model according to the computational budget defined by the output is A method comprising the step of processing the user input by the generative model using the settings of the generative model. Claim 4 A method according to claim 1, wherein the computational budget defines the number of rounds of self-criticism that the generative model uses to process the user input. Claim 5 In claim 1, the method wherein the computational budget defines the level of planning used by the generative model to process the user input. Claim 6 A method according to claim 1, wherein the computational budget defines the number of samples generated by the generative model from the user input. Claim 7 A method according to claim 1, wherein the computational budget defines the quality level of the generated output that the generative model generates from the user input. Claim 8 A method according to claim 7, wherein the generated output is an image, and the quality level of the generated output includes the resolution of the image. Claim 9 A method according to any one of claims 1 to 8, wherein the feature data of the user input includes context data of the user input, and the step of processing the feature data includes the step of processing the context data of the user input using the machine learning model to generate the output that defines the computational budget of the generative model for the user input. Claim 10 A method according to claim 9, wherein the user input includes a request to generate output text, the generative model includes a large-scale language model (LLM) trained to generate text data, and the feature data is context data in which the output text will be used. Claim 11 A method according to claim 9, wherein the user input includes a request to generate an output image, the generative model includes a generative image model trained to generate image data, and the feature data includes features of one or more images within context data in which the output image is to be used. Claim 12 A method according to any one of claims 1 to 8, wherein the feature data of the user input includes the level of importance of the response to the user input. Claim 13 A method according to any one of claims 1 to 8, wherein the feature data of the user input includes a difficulty level of a task defined by the user input. Claim 14 A method comprising, in any one of claims 1 to 8, the step of updating the output that defines the computational budget of the generative model for the user input using the metadata of the generative model; and the step of processing the user input by the generative model based on the computational budget defined by the updated output. Claim 15 A system comprising one or more computers and one or more storage devices for storing instructions, wherein the instructions are operable to cause the one or more computers to perform operations when executed by the one or more computers, and the operations include: an operation of receiving user input to be processed by a generative model; an operation of determining feature data of the user input; an operation of processing the feature data of the user input using a machine learning model to generate an output that defines a computational budget of the generative model for the user input, wherein the computational budget of the generative model defines the amount of computational resources consumed by the generative model when processing the user input; and an operation of processing the user input by the generative model according to the computational budget defined by the output to generate a generated output. Claim 16 In paragraph 15, the output includes a predicted level of the computational budget of the generative model for the user input, and different levels of the computational budget correspond to different settings of the generative model, and the operations further include an operation of mapping the predicted level of the computational budget of the generative model for the user input to a corresponding setting of the generative model, and the operation of processing the user input by the generative model includes an operation of processing the user input by the generative model using the corresponding setting of the generative model. Claim 17 In paragraph 15, the output includes a setting of the generative model that defines the computational budget of the generative model for the user input, and the operation of processing the user input by the generative model according to the computational budget defined by the output includes an operation of processing the user input by the generative model using the setting of the generative model. Claim 18 In paragraph 15, the above computational budget defines the number of rounds of self-criticism that the generative model uses to process the user input, in a system. Claim 19 In paragraph 15, the above computational budget is a system that defines the level of planning used by the generative model to process the user input. Claim 20 One or more non-transient storage media encoded in instructions, wherein the instructions cause the computing device to perform operations when executed by the computing device, the operations include: receiving user input to be processed by a generative model; determining feature data of the user input; processing the feature data of the user input using a machine learning model to generate an output that defines a computational budget of the generative model for the user input, wherein the computational budget of the generative model defines the amount of computational resources consumed by the generative model when processing the user input; and processing the user input by the generative model according to the computational budget defined by the output to generate a generated output.