Aligning sequence processing models with recommended knowledge
By generating auxiliary prompts to train the sequence processing model, the challenges of traditional recommender systems in data storage and user interaction data processing are solved, and a more efficient, secure and reliable recommendation system is achieved.
Patent Information
- Application Number
- CN202411849285.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-12-15
- Filing Date
- 2024-12-16
- Publication Date
- 2025-05-06
AI Technical Summary
Traditional recommender systems have challenges in data storage and handling of user interaction data, including storage requirements for large numbers of users, projects and interactive data, data privacy and security issues, as well as prone to failure and maintenance difficulties caused by system complexity.
Sequence processing models are trained by generating auxiliary prompts, and recommendation-related knowledge is introduced to improve the performance of the model on the recommended tasks. The method includes obtaining project datasets, generating auxiliary prompts, and using these prompts to train sequences to process the model to replace some of the functions of traditional recommendation systems.
Reduces dependence on large databases, reduces the storage needs of long-term user historical data, improves system reliability and maintenance simplicity, and enhances data privacy and security.
Smart Images

Figure CN119939018A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates generally to machine learning and more particularly to systems and methods for enhancing the performance of sequence processing models on recommendation tasks. Background Art
[0002] The traditional software stack used in a typical recommender system includes various components, each serving a specialized function in the overall recommendation process. The architecture usually starts with the database layer, which is the base component where all user, item, and interaction data is stored. This layer is usually built using traditional relational database systems such as MySQL or more advanced NoSQL databases such as MongoDB based on the size of the data. Typically, the next component in the stack is the data processing layer, which generally consists of big data processing systems such as MapReduce, Spark, or Flink. These tools help transform, filter, and prepare data for recommendation algorithms. The recommendation engine itself forms the core of the software stack. This component uses various algorithms such as collaborative filtering, content-based filtering, or hybrid models to generate recommendations. The application layer forms the topmost component of the stack, interfacing with users and triggering the recommender system when users interact. This layer is usually built using various programming languages and web technologies, including Python, Java, JavaScript, HTML, and CSS. Last but not least, the recommender system can also include a feedback loop for continuous learning and improvement. The feedback loop collects user feedback on the recommendations, incorporates the user feedback into the system, and refines future recommendations accordingly.
[0003] However, these traditional recommender systems are accompanied by certain disadvantages, particularly those related to their data storage requirements and handling of user interaction data. The need for a database layer capable of storing a large amount of user, item, and interaction data means that these systems rely on a wide range of databases. Meeting this requirement can be technically challenging, especially for systems that handle an increasing number of users and items. Additionally, continuously updating these databases with new user interaction data can also be technically demanding. In addition, traditional recommendation systems typically require storage of long-term user interaction data to provide accurate and personalized recommendations. This requirement poses challenges in terms of data privacy and security because it involves storing sensitive user data over an extended period of time. In addition to these technical challenges, traditional recommender systems also tend to be complex, consisting of multiple components, each serving a specialized function. This complexity can make these systems prone to failure and difficult to maintain and update. Summary of the invention
[0004] Aspects and advantages of embodiments of the present disclosure will be set forth in part in the following description, or may be learned from the description, or may be learned through practice of the embodiments.
[0005] An example aspect of the present disclosure provides a computer-implemented method for training a sequence processing model for use in a recommendation system. The method includes obtaining, by a computing system including one or more computing devices, a project data set describing a plurality of projects included in a candidate pool. The method includes generating, by the computing system, a plurality of auxiliary prompts for use in training the sequence processing model, wherein each auxiliary prompt includes a prompt input and a prompt output, and wherein the plurality of auxiliary prompts encode recommendation-related knowledge about the plurality of projects. The method includes training the sequence processing model using the plurality of auxiliary prompts by the computing system. The method includes providing, by the computing system, a trained sequence processing model for use in a recommendation system.
[0006] Other exemplary aspects of the present disclosure relate to other systems, methods, devices, tangible non-transitory computer-readable media and apparatus for performing the functions described herein. These and other features, aspects and advantages of various implementations will be better understood with reference to the following description and the appended claims. The accompanying drawings, which are incorporated into and constitute a part of this specification, illustrate implementations of the present disclosure and, together with the description, help explain the relevant principles. BRIEF DESCRIPTION OF THE DRAWINGS
[0007] A detailed discussion of embodiments for those of ordinary skill in the art is set forth in this specification with reference to the accompanying drawings, in which:
[0008] Figure 1 is a flowchart illustrating an example method for training a model for machine learning according to an example implementation of aspects of the present disclosure;
[0009] Figure 2 is a block diagram of an example process flow for processing an input to generate an output using a machine learning model according to an example implementation of aspects of the present disclosure;
[0010] Figure 3 is a block diagram of an example sequence processing model according to an example implementation of aspects of the present disclosure;
[0011] Figure 4 is a block diagram of an example technique for populating an example output sequence for processing by a sequence processing model according to an example implementation of aspects of the present disclosure;
[0012] Figure 5 is a block diagram of an example model development platform according to an example implementation of aspects of the present disclosure;
[0013] Figure 6 is a block diagram of an example training workflow for training a model for machine learning according to an example implementation of aspects of the present disclosure;
[0014] Figure 7is a block diagram of an inference system for operating one or more machine learning models to perform inference according to an example implementation of aspects of the present disclosure;
[0015] Figure 8 is a block diagram of an example networked computing system according to an example implementation of aspects of the present disclosure;
[0016] Fig. 9 is a block diagram of an example computing device according to an example implementation of aspects of the present disclosure; and
[0017] Fig.10 is a block diagram of an example computing device according to an example implementation of aspects of the present disclosure. DETAILED DESCRIPTION
[0018] Traditional recommender systems, while effective, are also accompanied by certain disadvantages, particularly those associated with their data storage requirements and handling of user interaction data. On the other hand, sequence processing models, such as, for example, so-called "large language models", have strong generalization capabilities but lack specific knowledge related to recommendation tasks. In view of this dichotomy, the present disclosure provides systems and methods for enhancing the performance of sequence processing models with respect to recommendation tasks. The improved recommendation performance provided by the proposed technology enables conventional recommender systems to be replaced with sequence processing models for recommendation generation. Therefore, the proposed technology aims to bridge the knowledge gap between sequence processing models and conventional recommender systems, while resulting in many technical benefits, such as reduced demand for large databases of large user-item interaction datasets required to store multiple recommendation system components in a deployed system, reduced storage of long-term user history data, potential on-device implementations, and other benefits.
[0019] Specifically, example implementations of the present disclosure mitigate this knowledge gap by aligning the sequence processing model with the recommendation knowledge. The example training system can generate natural language prompts that can be referred to as "auxiliary prompts" that encode different types of recommendation-related knowledge such as item attributes and user preferences. These auxiliary prompts encode various operations and losses into a natural language format that can be used to impart recommendation knowledge to the sequence processing model, including item embedding, Bayesian personalized ranking (BPR), and masked item modeling.
[0020] The proposed techniques improve the performance of sequence processing models on basic recommendation tasks such as retrieval, ranking, and rating prediction. The auxiliary hint structure effectively introduces recommendation-related knowledge even for domains that are relatively unfamiliar or "unobserved" to the sequence processing model. The resulting recommendation-tuned sequence processing model provides a number of technical benefits, as outlined below.
[0021] More specifically, one example aspect of the present disclosure relates to systems and methods for training sequence processing models for use in recommendation systems, with a focus on improving performance of retrieval, ranking, and rating prediction tasks. An example method may include obtaining a project dataset describing a variety of projects included in a candidate pool. The project dataset may describe, for each project in the project, one or more attribute values of one or more project attributes such as a title, category, brand, price, attributes, description, and comments.
[0022] Then, the example method may include generating a series of auxiliary prompts for use in training the sequence processing model. Each auxiliary prompt may include a prompt input and a prompt output, which together encode recommendation-related knowledge about the item. For example, the auxiliary prompt may ask questions about the characteristics of the item in the input and provide answers in the output. In some cases, the auxiliary prompts encoding different types of recommendation-related knowledge may correspond to natural language representations of operations and losses that show good performance when used by conventional recommendation systems.
[0023] An example auxiliary prompt structure, which may be referred to as an item embedding prompt, encodes an item embedding method into a natural language representation. Item embedding is a common practice employed by representation-based recommenders, in which items are represented in a common vector space. The present disclosure uses natural language expressions to simulate the item embedding process. Item embedding prompts are generated by asking questions about the characteristics of items in the input and answering the questions in the output.
[0024] Another example auxiliary prompt structure (which may be referred to as a BPR loss reduction prompt) represents the reduction of the Bayesian personalized ranking (BPR) loss into a natural language prompt format. The BPR loss is a loss function that has been used in conventional recommenders to learn to optimize model parameters for ranking. The BPR loss reduction prompt converts this process into a natural language prompt format. Specifically, by asking and answering questions about the user's choice between positive items and negative items, a BPR loss reduction prompt can be generated.
[0025] Another example auxiliary prompt structure (which may be referred to as a masked item modeling prompt) represents the conversion of the masked item modeling training framework to a natural language prompt format. Specifically, a typical masked modeling application Cloze objective, where random entries in an input sequence are replaced with a special token "[mask]", and the model learns to predict masked entries based on their left and right contexts. The present disclosure converts this process into a natural language prompt format by masking and identifying random items within a user's purchase sequence. This helps capture different historical interactions and their contexts.
[0026] Other aspects of the present disclosure relate to fine-tuning and evaluating a recommender system based on a sequence processing model. This can be done by first generating a recommendation task and an auxiliary task prompt. Next, the sequence processing model backbone is fine-tuned using the recommendation task prompt along with the auxiliary task prompt. For example, the sequence processing model backbone can be first trained using the auxiliary task prompt, and then subsequent training is performed using the recommendation task prompt. Finally, the fine-tuned model is evaluated on and / or deployed for basic recommendation tasks (e.g., retrieval, ranking, and rating prediction).
[0027] Another aspect of the present disclosure relates to a method for simplifying the representation of items to reduce the complexity of the input / output space. For example, in some or all of the auxiliary prompts in the auxiliary prompts, the item identifier of the item can be replaced with a shortened identifier. This makes it easier for the sequence processing model to process and understand the information.
[0028] Another aspect of the disclosure relates to techniques that operate to relieve a sequence processing model from memorizing user identifiers (IDs), which can be challenging due to the significant number of user IDs. For example, in some or all of the training prompts, a user identifier of a user can be replaced with a sequence of items that the user has interacted with. This can enable the sequence processing model to model the item space more directly, rather than trying to memorize a large number of user IDs.
[0029] Once fine-tuned or otherwise trained, the sequence processing model can be deployed for use in a recommendation system. For example, the trained sequence processing model can be instructed or otherwise used to perform a retrieval task, a ranking task, or a rating prediction task. In some cases, the sequence processing model deployed to perform these tasks may not have been previously trained on data specific to the candidate pool. Therefore, the proposed technology leads to a sequence processing model as a flexible and adaptable solution for a variety of recommendation systems.
[0030] The systems and methods of the present disclosure provide many technical effects and benefits. As an example, since the trained sequence processing model is the primary component to be deployed, the need for one or more large databases in the deployed recommendation system is eliminated. This represents a significant technical improvement over traditional recommender systems, which require extensive databases to store and process user, item, and interaction data. The proposed method can leverage the inherent ability of the sequence processing model to understand and generate recommendations based on encoded knowledge, thereby reducing the reliance on large databases.
[0031] As another related technical benefit, the proposed technique provides more privacy and security because the entire long-term user data set does not need to be stored. In conventional systems, user data about all past user interactions are stored and processed in a database that can potentially be exploited or compromised. However, using a sequence processing model-based approach, only short-term user interactions can be used to generate recommendations based on the knowledge encoded in the model, thereby significantly reducing the need for extensive data storage and thereby enhancing data privacy and security.
[0032] As another example technical benefit, the proposed technology is designed so that some implementations of the method can be implemented on a device. For example, because the sequence processing model has learned representations of items included in the item data set, access to the item data set (e.g., as stored in a database) is not required, and therefore, some implementations of the model can be run on the device. This is a significant technical advancement because it enables the recommendation system to operate locally on the user's device, thereby reducing reliance on constant network connectivity and server-side processing, thereby saving network bandwidth and reducing data transmission costs. This on-device implementation also enhances the speed and responsiveness of the system, thereby providing a robust solution for real-time recommendations. In addition, the ability to implement a recommendation system on a device represents a further privacy benefit because data about user interactions does not necessarily need to leave the user's device. Similarly, the use of an on-device sequence processing model provides the opportunity to create a personalized recommendation model in a privacy-preserving manner.
[0033] As another example technical benefit, the refined approach provides fewer points of failure compared to traditional systems. By simplifying the architecture and transitioning from a multi-component system to a single sequence processing model, the risk of system crashes and failures is significantly reduced. This results in a more reliable and robust recommendation system, leading to enhanced system performance.
[0034] As another example technical benefit, the use of sequence processing models opens up the possibility of leveraging specialized hardware and investments in sequence processing model performance. Since sequence processing models are a major area of research and development in the field of machine learning, there are numerous ongoing efforts to optimize the performance of sequence processing models through specialized hardware and software solutions, including advances in specialized hardware accelerators such as application specific integrated circuits (ASICs) and graphics processing units (GPUs). By adopting the sequence processing models in the proposed technology, recommendation systems can benefit from these advances, providing a technically superior and future-proof solution.
[0035] The proposed sequence processing model based recommendation system can be applied to many different applications or use cases. As an example, the sequence processing model based recommendation system can be used to provide personalized content recommendations. For example, the trained sequence processing model can be used to create personalized content recommendations for users on various platforms such as video streaming services, e-commerce websites, or social media platforms. The system can analyze the user's past behavior, preferences, and interactions to suggest relevant content, products, or posts.
[0036] As another example, a recommendation system based on a sequence processing model can be incorporated into a search engine or other information retrieval system. For example, a search engine can use a sequence processing model to improve search results by better understanding the intent behind a user's search query. The model and the search engine can work together to provide more relevant search results, thereby enhancing the user's experience and satisfaction. As yet another example, a recommendation system based on a sequence processing model can be incorporated into an e-learning platform. In an e-learning platform, a sequence processing model can provide personalized learning recommendations to students based on their learning style, progress, and preferences, thereby enhancing the learning experience.
[0037] Referring now to the drawings, example embodiments of the present disclosure will be discussed in further detail.
[0038] Example Method
[0039] Figure 1 A flow chart of a method 100 for training one or more machine-learned models according to aspects of the present disclosure is depicted. For example, an example machine-learned model may include a sequence processing model.
[0040] One or more portions of the example method 100 may be implemented by a computing system including one or more computing devices, such as, for example, the computing systems described with reference to other figures. Each respective portion of the example method 100 may be performed by any one (or any combination) of the one or more computing devices. In addition, one or more portions of the example method 100 may be implemented on hardware components of the devices described herein, for example, to train one or more systems or models. Figure 1 Elements performed in a specific order are depicted for purposes of illustration and discussion. One of ordinary skill in the art will understand, using the disclosure provided herein, that the elements of any method discussed herein may be adjusted, rearranged, expanded, omitted, combined or modified in various ways without departing from the scope of the present disclosure. Figure 1 Elements / terms described with reference to other systems and figures are described for illustrative purposes and are not intended to be limiting. One or more portions of the example method 100 may additionally or alternatively be performed by other systems.
[0041] At 102, the example method 100 may include obtaining training instances. The set of training data may include multiple training instances divided between multiple data sets (e.g., a training data set, a validation data set, or a test data set). The training instances may be labeled or unlabeled. Although referred to as "training" instances in the example method 100, it should be understood that when the model is trained (e.g., online training / learning) using an evaluation of the performance of the model on the runtime instance, the runtime inference may form the training instance. Example data types for training instances and various tasks associated therewith are described throughout this disclosure. As an example, the training instance may be or include an auxiliary prompt.
[0042] At 104, example method 100 may include processing the training instance using one or more machine-learned models to generate an output. The output may be obtained directly from the one or more machine-learned models, or may be a downstream result of a chain of processing operations that includes the output of the one or more machine-learned models. For example, processing the training instance may include processing an auxiliary prompt (e.g., a prompt input portion of the auxiliary prompt).
[0043] At 106, the example method 100 may include receiving an evaluation signal associated with the output. The evaluation signal may be obtained using a loss function. For example, the loss function may compare the output of the model with the prompt output portion of the auxiliary prompt. Various losses may be determined, such as mean square error, likelihood loss, cross entropy loss, hinge loss, contrast loss, or various other loss functions. The evaluation signal may be calculated using a known benchmark true label (e.g., supervised learning), a predicted or estimated label (e.g., semi-supervised learning or self-supervised learning), or without a label (e.g., unsupervised learning). The evaluation signal may be a reward (e.g., for reinforcement learning). The reward may be calculated using a reward model of a machine learning configured to generate a reward based on the received output. The reward may be calculated using feedback data describing human feedback to the output.
[0044] At 108, the example method 100 may include updating the machine-learned model using the evaluation signal. For example, in some embodiments, various training or learning techniques (such as, for example, back propagation) may be used to learn the values of the parameters of the machine-learned model. For example, the evaluation signal may be back-propagated from the output (or another source of the evaluation signal) through the machine-learned model to update one or more parameters of the model (e.g., based on the gradient of the evaluation signal relative to the parameter value). For example, a system including one or more machine-learned models may be trained in an end-to-end manner. Gradient descent techniques may be used to iteratively update parameters in multiple training iterations. In some implementations, performing error back propagation may include performing truncated back propagation through time. The example method 100 may include implementing a variety of generalization techniques (e.g., weight decay, dropout, etc.) to improve the generalization ability of the model being trained.
[0045] In some implementations, example method 100 may be implemented to train a machine learning model from an initialized state to a fully trained state (e.g., when the model exhibits a desired performance profile, such as based on accuracy, precision, recall, etc.).
[0046] In some implementations, the example method 100 may be implemented for a specific stage of the training process. For example, in some implementations, the example method 100 may be implemented for pre-training a machine learning model. Pre-training may include, for example, large-scale training on potentially noisy data to achieve a wide range of performance level bases across a variety of tasks / data types. In some implementations, the example method 100 may be implemented for fine-tuning a machine learning model. Fine-tuning may include, for example, smaller-scale training on higher quality (e.g., labeled, selected, etc.) data. Fine-tuning may affect all or part of the parameters of the machine learning model. For example, various parts of the machine learning model may be "frozen" for certain training stages. For example, parameters associated with the embedding space may be "frozen" during fine-tuning (e.g., to retain information learned from a wider domain than that present in the fine-tuning data set). Example fine-tuning methods include reinforcement learning. Reinforcement learning may be based on user feedback on model performance during use.
[0047] Example machine learning model
[0048] Figure 2 is a block diagram of an example processing flow for using a machine learning model 1 to process input 2 to generate output 3.
[0049] The machine learning model 1 may be or include one or more machine learning models or model components. An example machine learning model may include a neural network (e.g., a deep neural network). An example machine learning model may include a nonlinear model or a linear model. An example machine learning model may use other architectures to replace or supplement a neural network. An example machine learning model may include a decision tree-based model, a support vector machine, a hidden Markov model, a Bayesian network, a linear regression model, a k-means clustering model, and the like.
[0050] Example neural networks may include feedforward neural networks, recurrent neural networks (RNNs) (including recurrent neural networks based on long short-term memory (LSTM)), convolutional neural networks (CNNs), diffusion models, generative adversarial networks, or other forms of neural networks. Example neural networks may be deep neural networks. Some example machine learning models may utilize attention mechanisms, such as self-attention. For example, some example machine learning models may include multi-head self-attention models.
[0051] The machine learning model 1 may include a single or multiple instances of the same model configured to operate on data from the input 2. The machine learning model 1 may include an ensemble of different models that can interact collaboratively to process data from the input 2. For example, the machine learning model 1 may adopt a mixed expert structure. See, for example, Zhou et al. Mixture-of-Experts with Expert Choice Routing, arXiv:2202.09368v2 (October 14, 2022).
[0052] Input 2 may generally include or otherwise represent various types of data. Input 2 may include one type of data or many different types of data. Output 3 may be the same type of data or a different type of data compared to input 2. Output 3 may include one type of data or many different types of data.
[0053] Example data types for input 2 or output 3 include natural language text data, software code data (e.g., source code, object code, machine code, or any other form of computer-readable instructions or programming language), machine code data (e.g., binary code, assembly code, or other form of machine-readable instructions that can be directly executed by a central processing unit of a computer), assembly code data (e.g., a low-level programming language that uses symbolic representations of machine code instructions to program a processing unit), genetic data or other chemical or biochemical data, image data, audio data, audio-visual data, tactile data, biometric data, medical data, financial data, statistical data, geographic data, astronomical data, historical data, sensor data (e.g., digital or analog values, such as voltage or other absolute or relative level measurements from real or artificial inputs (such as from audio sensors, light sensors, displacement sensors, etc.)), etc. The data can be raw or processed and can be in any format or mode.
[0054] In the multimodal input 2 or output 3, example combinations of data types include image data and audio data, image data and natural language data, natural language data and software code data, image data and biometric data, sensor data and medical data, etc. It should be understood that any combination of data types in the input 2 or output 3 may exist.
[0055] Example input 2 may include one or more data types, such as the example data types noted above. Example output 3 may include one or more data types, such as the example data types noted above. The data type of input 2 may be the same as or different from the data type of output 3. It should be understood that the example data types noted above are provided for illustrative purposes only. The data types contemplated within the scope of the present disclosure are not limited to those examples noted above.
[0056] Example machine learning sequence processing model
[0057] Figure 3is a block diagram of an example implementation of an example machine learning model configured to process a sequence of information. For example, an example implementation of a machine learning model 1 may include a machine learning sequence processing model 4. An example system may pass an input 2 to the sequence processing model 4. The sequence processing model 4 may include one or more machine learning components. The sequence processing model 4 may process data from the input 2 to obtain an input sequence 5. The input sequence 5 may include one or more input elements 5-1, 5-2, ..., 5-M, etc. obtained from the input 2. The sequence processing model 4 may process the input sequence 5 using a prediction layer 6 to generate an output sequence 7. The output sequence 7 may include one or more output elements 7-1, 7-2, ..., 7-N, etc. generated based on the input sequence 5. The system may generate an output 3 based on the output sequence 7.
[0058] The sequence processing model 4 may include one or more machine learning model components configured to ingest, generate, or otherwise reason about sequences of information. For example, some example sequence processing models in the text domain are referred to as "large language models" or LLMs. See, for example, Palm 2 Technical Report, Google, https: / / ai.google / static / documents / palm2techre port.pdf (nd). Other example sequence processing models can operate in other domains, for example, such as image domains, for example, see Dosovitskiy et al. An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale, arXiv:2010.11929v2 (June 3, 2021); audio domains, for example, see Agostinelli et al. MusicLM: Generating Music From Text, arXiv:2301.11325v1 (January 26, 2023); biochemical domains, for example, see Jumper et al. Highly accurate protein structure prediction with AlphaFold, 596 Nature 583 (August 26, 2021). The sequence processing model 4 can process one or more types of data simultaneously. The sequence processing model 4 can include a relatively large model (e.g., more parameters, computationally expensive, etc.), a relatively small model (e.g., fewer parameters, computationally lightweight, etc.), or both.
[0059] In general, the sequence processing model 4 can use the data from the input 2 to obtain the input sequence 5. For example, the input sequence 5 can include a representation of the data from the input 2 in a format understood by the sequence processing model 4. One or more machine learning components of the sequence processing model 4 can ingest data from the input 2, parse the data into segments compatible with the processing architecture of the sequence processing model 4 (e.g., via "lemmaization"), and project the segments into an input space associated with the prediction layer 6 (e.g., via "embedding").
[0060] The sequence processing model 4 may ingest data from the input 2 and parse the data into a sequence of elements to obtain an input sequence 5. For example, a portion of the input data from the input 2 may be decomposed into fragments that collectively represent the content of the portion of the input data. The fragments may provide elements of the sequence.
[0061] In some cases, elements 5-1, 5-2, ..., 5-M may represent building blocks for capturing or expressing meaningful information in a particular data domain. For example, an element may describe an "atomic unit" across one or more domains. For example, for a text input source, an element may correspond to a group of one or more word or sub-word components (such as a set of one or more characters).
[0062] For example, elements 5-1, 5-2, ..., 5-M can represent word units obtained using a word unit analyzer. For example, a word unit analyzer can process a given part of an input source and output a series of word units representing the part of the input source (e.g., corresponding to input elements 5-1, 5-2, ..., 5-M). Various methods can be used to lemmatize. For example, a text input source can be lemmatized using a byte pair encoding (BPE) technique. For example, see Kudo et al., Sentence Piece: A simple and language independent sub word tokenizer and detokenizer for Neural Text Processing, Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing (System Demonstrations), pp. 66 to 71 (October 31 to November 4, 2018), https: / / aclanthology.org / D18-2012.pdf. Image-based input sources can be tokenized by extracting and serializing patches from the image.
[0063] In general, any data type can be serialized and processed into an input sequence 5. It should be understood that Figure 3 The elements 5-1, 5-2, ..., 5-M depicted in FIG. 5 may be word-grams or may be embedded representations thereof.
[0064] The prediction layer 6 can predict one or more output elements 7-1, 7-2, ..., 7-N based on the input elements. The prediction layer 6 can include one or more machine learning model architectures, such as one or more learned parameter layers, which manipulate and transform the input to extract high-order meaning from the input elements 5-1, 5-2, ..., 5-M and the relationships between the input elements. In this way, for example, the example prediction layer 6 can predict new output elements based on the context provided by the input sequence 5.
[0065] The prediction layer 6 can evaluate associations between parts of the input sequence 5 and specific output elements. These associations can inform predictions about the likelihood that a specific output follows the input context. For example, consider the text snippet "The carpenter's tool kit is small and heavy. It is filled with ___." The example prediction layer 6 can recognize that "it" mentions "tool kit" again by determining the relationship between the corresponding embeddings. The example prediction layer 6 can also link "it" to attributes of the tool kit, such as "small" and "heavy." Based on these associations, for example, the prediction layer 6 can assign a higher probability to the word "nails" than to the word "sawdust."
[0066] Transformer is an example architecture that can be used in prediction layer 4. For example, see Vaswani et al., Attention Is All You Need, arXiv:1706.03762v7 (August 2, 2023). Transformer is an example of a model architecture for machine learning that uses an attention mechanism to compute associations between items within a context window. The context window may include a sequence that contains an input sequence 5 and potentially one or more output elements 7-1, 7-2, ..., 7-N. A Transformer block may include one or more attention layers and one or more post-attention layers (e.g., feed-forward layers such as multi-layer perceptrons).
[0067] The prediction layer 6 may include other machine learning model architectures in addition to or instead of the Transformer-based architecture. For example, recurrent neural networks (RNN) and long short-term memory (LSTM) models, as well as convolutional neural networks (CNNs), may also be used. In general, the prediction layer 6 may utilize various artificial neural networks that can understand or generate information sequences.
[0068] The output sequence 7 may include or otherwise represent the same or different data types as the input sequence 5. For example, the input sequence 5 may represent text data, and the output sequence 7 may represent text data. The input sequence 5 may represent image, audio, or audio-visual data, and the output sequence 7 may represent text data (e.g., describing the image, audio, or audio-visual data). It should be understood that the prediction layer 6 and any other gap model components of the sequence processing model 4 can be configured to receive multiple data types in the input sequence 5 and output multiple data types in the output sequence 7.
[0069] The output sequence 7 may have various relationships with the input sequence 5. The output sequence 7 may be a continuation of the input sequence 5. The output sequence 7 may be complementary to the input sequence 5. The output sequence 7 may translate, transform, enhance, or otherwise modify the input sequence 5. The output sequence 7 may answer, evaluate, confirm, or otherwise respond to the input sequence 5. The output sequence 7 may implement instructions provided via the input sequence 5 (or describe instructions for implementing the instructions).
[0070] The output sequence 7 can be generated autoregressively. For example, for some applications, the output of one or more prediction layers 6 can be passed through one or more output layers (e.g., a softmax layer) to obtain a probability distribution of an output vocabulary (e.g., a text or symbol vocabulary) conditioned on the set of input elements in the context window. In this way, for example, the output sequence 7 can be generated autoregressively by sampling a possible next output element, adding the element to the context window, and regenerating the probability distribution based on the updated context window, and sampling a possible next output element, etc.
[0071] The output sequence 7 can also be generated non-autoregressively. For example, multiple output elements of the output sequence 7 can be predicted together without explicit order conditions with respect to each other. See, for example, Saharia et al., Non-Autoregressive Machine Translation with Latent Alignments, arXiv:2004.07437v3 (November 16, 2020).
[0072] The output sequence 7 may include one or more parts or elements. In an example content generation configuration, the output sequence 7 may include multiple elements corresponding to the multiple parts of the generated output sequence (e.g., text sentences, values of discrete waveforms, computer code, etc.). In an example classification configuration, the output sequence 7 may include a single element associated with a classification output. For example, the output "vocabulary" may include the set of classes into which the input sequence will be classified. For example, a visual Transformer block may pass latent state information to a multilayer perceptron that outputs possible class values associated with the input image.
[0073] Figure 4 is a block diagram of an example technique for populating an example input sequence 8. The input sequence 8 may include various functional elements that form part of the model infrastructure, such as an element 8-0 obtained from a task indicator 9, which signals any model that processes the input sequence 8 that a specific task is being performed (e.g., to help adapt the performance of the model to the specific task). The input sequence 8 may include various data elements from different data modalities. For example, the input modality 10-1 may include a data modality. The data-to-sequence model 11-1 may process data from the input modality 10-1 to project the data into a format compatible with the input sequence 8 (e.g., one or more vectors of dimensions determined according to the dimensions of the input sequence 8) to obtain elements 8-1, 8-2, 8-3. Another input modality 10-2 may include different data modalities. The data-to-sequence model 11-2 may project the data from the input modality 10-2 into a format compatible with the input sequence 8 to obtain elements 8-4, 8-5, 8-6. Another input modality 10-3 may include yet another different data modality. The data-to-sequence model 11-3 may project the data from the input modality 10-3 into a format compatible with the input sequence 8 to obtain elements 8-7, 8-8, 8-9.
[0074] Input sequence 8 may be the same as or different from input sequence 5. Input sequence 8 may be a multimodal input sequence that includes elements that represent data from different modalities using a common dimensional representation. For example, the embedding space may have P dimensions. Input sequence 8 may be configured to include multiple elements having P dimensions. In this way, for example, example implementations may facilitate information extraction and reasoning across different data modalities by projecting data into elements in the same embedding space to perform comparisons, combinations, or other calculations between them.
[0075] For example, elements 8-0, ..., 8-9 may indicate specific locations within a multidimensional embedding space. Some elements may be mapped to a discrete set of locations in the embedding space. For example, elements corresponding to discrete members in a predetermined word-gram vocabulary may be mapped to discrete locations associated with those word-grams in the embedding space. Other elements may be distributed continuously across the embedding space. For example, some data types may be decomposed into continuously defined parts (e.g., image patches) that may be described using continuously distributed locations within the embedding space.
[0076] In some implementations, the expressive power of the embedding space may not be limited to the meaning associated with any particular set of word-grams or other building blocks. For example, a continuous embedding space can encode a series of high-order information. Individual fragments of information (e.g., word-grams) can be mapped to specific points in the space: for example, the word-gram of the word "dog" can be projected to an embedding value that points to a specific position in the embedding space associated with canine-related information. Similarly, an image patch of an image of a dog on grass can also be projected into the embedding space. In some implementations, the projection of the image of the dog can be similar to the projection of the word "dog", while also being similar to the projection of the word "grass", but potentially different from both. In some implementations, the projection of the image patch may not be completely aligned with any single projection of a single word. In some implementations, the projection of the image patch may be aligned with a combination of the projections of the words "dog" and "grass". In this way, for example, a high-order embedding space can encode information that may be independent of the data modality of the expressed information.
[0077] The task indicator 9 may include a model or model component configured to identify the task being performed and inject an input value represented by an element 8-0 into the input sequence 8, which element 8-0 signals which task is being performed. For example, the input value may be provided as a data type associated with the input modality and projected along with the input modality (e.g., the input value may be a text task label embedded along with other text data in the input; the input value may be a pixel-based representation of the task embedded along with other image data in the input; etc.). The input value may be provided as a data type that is different from or at least independent of other inputs. For example, the input value represented by the element 8-0 may be learned within a continuous embedding space.
[0078] The input modalities 10 - 1 , 10 - 2 , and 10 - 3 may be associated with a variety of different data types (eg, as described above with respect to input 2 and output 3 ).
[0079] The data-to-sequence models 11-1, 11-2, and 11-3 may be the same or different from each other. The data-to-sequence models 11-1, 11-2, and 11-3 may be adapted to each corresponding input modality 10-1, 10-2, and 10-3. For example, a text data-to-sequence model may subdivide a portion of the input text and project the subdivision into an element in the input sequence 8 (e.g., element 8-1, 8-2, 8-3, etc.). An image data-to-sequence model may subdivide an input image and project the subdivision into an element in the input sequence 8 (e.g., element 8-4, 8-5, 8-6, etc.). A data-to-sequence model of any data type may subdivide the input of the arbitrary data type and project the subdivision into an element in the input sequence 8 (e.g., element 8-7, 8-8, 8-9, etc.).
[0080] The data-to-sequence models 11-1, 11-2, and 11-3 may form part of the machine-learned sequence processing model 4. The data-to-sequence models 11-1, 11-2, and 11-3 may be trained jointly with the machine-learned sequence processing model 4 or independently of the machine-learned sequence processing model 4. The data-to-sequence models 11-1, 11-2, and 11-3 may be trained end-to-end with the machine-learned sequence processing model 4.
[0081] Model development platform for sample machine learning
[0082] Figure 5 1 is a block diagram of an example model development platform 12 that can facilitate the creation, adaptation, and refinement of example machine learning models (e.g., machine learning model 1, sequence processing model 4, etc.). The model development platform 12 can provide a number of different tool kits that a developer system can use to develop new or adapted machine learning models.
[0083] The model development platform 12 may provide one or more model libraries 13 containing building blocks for new models. The model library 13 may include one or more pre-trained base models 13-1, which may provide a backbone of processing capabilities across various tasks. The model library 13 may include one or more pre-trained expert models 13-2, which may focus on performance in a specific professional field. The model library 13 may include various model primitives 13-3, which may provide low-level architectures or components (optionally pre-trained) that may be assembled in various arrangements as needed.
[0084] The model development platform 12 may receive a selection of various model components 14. The model development platform 12 may pass the selected model components 14 to a workbench 15, which combines the selected model components 14 into a development model 16.
[0085] The workbench 15 may further refine and adapt the development model 16 by utilizing a number of different tool suites integrated with the model development platform 12. For example, the workbench 15 may utilize a model alignment tool suite 17 to facilitate alignment of the development model 16 with desired performance profiles on various tasks.
[0086] The model alignment toolkit 17 can provide a variety of tools for causing the development model 16 to generate outputs that are aligned with desired behavioral characteristics. Alignment can include increasing the accuracy, precision, recall, etc. of the model output. Alignment can include enforcing the use of output styles, patterns, or other preferred characteristics of the model output. Alignment can be general or domain-specific. For example, the pre-trained base model 13-1 can start from an initial performance level across multiple domains. Alignment of the pre-trained base model 13-1 can include improving performance in a specific information or task domain (e.g., even at the expense of performance in another information or task domain).
[0087] The model alignment toolkit 17 can integrate one or more datasets 17-1 for aligning the development model 16. The curated dataset 17-1 can include labeled or unlabeled training data. The dataset 17-1 can be obtained from a public domain dataset. The dataset 17-1 can be obtained from a private dataset associated with one or more developer systems for alignment of custom machine learning models customized for private use cases.
[0088] The pre-training pipeline 17-2 may include a model training workflow for machine learning configured to update the development model 16 on a large-scale, potentially noisy dataset. For example, pre-training may utilize unsupervised learning techniques (e.g., denoising, etc.) to process a large number of training instances to update model parameters from an initialized state and achieve a desired baseline performance. The pre-training pipeline 17-2 may utilize an unlabeled dataset in the dataset 17-1 to perform pre-training. The workbench 15 may implement the pre-training pipeline 17-2 to pre-train the development model 16.
[0089] The fine-tuning pipeline 17-3 may include a model training workflow of machine learning configured to use higher quality data to refine model parameters of the development model 16. The fine-tuning pipeline 17-3 may update the development model 16 by performing supervised training using a labeled dataset in the dataset 17-1. The fine-tuning pipeline 17-3 may update the development model 16 by performing reinforcement learning using a reward signal from a user feedback signal. The workbench 15 may implement the fine-tuning pipeline 17-3 to fine-tune the development model 16.
[0090] The prompt library 17-4 may include a set of inputs configured to induce behavior aligned with a desired performance standard. The prompt library 17-4 may include a few-trial prompt (e.g., an input that provides an example of the desired model output so that the head is appended to the desired runtime query), a thought chain prompt (e.g., an input that provides step-by-step reasoning within an example to encourage the model to perform comprehensive reasoning), etc.
[0091] Example prompts may be retrieved from an available repository of prompt libraries 17-4. One or more developer systems may use the workbench 15 to facilitate example prompts.
[0092] In some implementations, a pre-trained or fine-tuned model can achieve satisfactory performance without examples in the input. For example, a zero-trial prompt can include an input that lacks examples. The zero-trial prompt can be within the domain within the training dataset or outside the training domain.
[0093] The prompt library 17-4 may include one or more prompt engineering tools. The prompt engineering tool may provide a workflow for retrieving or learning optimized prompt values. The prompt engineering tool may facilitate directly learning prompt values (e.g., input element values) based on one or more training iterations. The workbench 15 may implement the prompt engineering tool in the development model 16.
[0094] The prompt library 17-4 may include a pipeline for prompt generation. For example, the development model 16 itself or another machine-learned model may be used to generate inputs. In this way, for example, a first model may process information about a task and output inputs for a second model to process in order to perform the steps of the task. The second model may be the same as or different from the first model. The workbench 15 may implement the prompt generation pipeline in the development model 16.
[0095] The prompt library 17-4 may include a pipeline for context injection. For example, the performance of the development model 16 on a particular task may be improved if provided with additional context for performing the task. The prompt library 17-4 may include software components configured to identify the desired context, retrieve the context from an external source (e.g., a database, a sensor, etc.), and add the context to the input prompt. The workbench 15 may implement the context injection pipeline in the development model 16.
[0096] Although various training examples described herein with respect to the model development platform 12 involve "pre-training" and "fine tuning", it should be understood that the model alignment toolkit 17 can generally support a wide variety of training techniques suitable for training a wide variety of machine learning models. The example training techniques can correspond to the example training method 100 described above.
[0097] The model development platform 12 may include a model plug-in tool kit 18. The model plug-in tool kit 18 may include a variety of tools that are configured to enhance the functionality of the machine learning model by integrating the machine learning model with other systems, devices, and software components. For example, the machine learning model can use tools to improve the quality of performance where appropriate. For example, deterministic tasks can be offloaded to a dedicated tool instead of performing tasks with increased risk of errors in a probabilistic manner. For example, the machine learning model does not autoregressively predict the solution of the system of equations, but can identify the tools called to obtain the solution and pass the system of equations to the appropriate tool. The tool can be a traditional system of equations solver that can operate deterministically to solve the system of equations. The output of the tool can be returned in response to the original query. In this way, the use of tools can allow some example models to focus on the advantages of machine learning models-for example, understanding the intent in unstructured requests for tasks-while enhancing the performance of the model by offloading certain tasks to more focused tools to mechanically apply deterministic algorithms to well-defined problems.
[0098] The model plug-in tool suite 18 may include a validation tool 18-1. The validation tool 18-1 may include a tool that may parse and validate the output of a machine-learned model. The validation tool 18-1 may include an engineered heuristic that establishes certain thresholds that are applied to the model output. For example, the validation tool 18-1 may ground the output of a machine-learned model against a structured data source (e.g., to mitigate "hallucinations").
[0099] The model plug-in toolkit 18 may include a toolkit 18-2 for implementing one or more tools, which may include scripts or other executable code that can be executed with the development model 16. The toolkit 18-2 may include one or more inputs configured to cause the machine learning model to implement the tool (e.g., a few-trial prompts that induce the model to output tool calls with the correct syntax, etc.). For example, the toolkit 18-2 may include fine-tuning training data for training the model to use the tool.
[0100] Model plug-in toolkit 18 may include an interface for calling external application programming interface (API) 18-3. For example, in addition to or instead of using development model 16 to directly implement tool calls or tool code, development model 16 may also align with output instructions that initiate API calls to send or obtain data via an external system.
[0101] The model plug-in toolkit 18 may be integrated with the hint library 17-4 to build a catalog of available tools for use with the development model 16. For example, the model may receive in input a catalog of available tools, and the model may generate an output that selects a tool from the available tools and initiates a tool call for using the tool.
[0102] The model development platform 12 may include a computing optimization tool suite 19 for optimizing the computing performance of the development model 16. For example, a model compression 19-1 tool may allow the development model 16 to reduce size while maintaining a desired performance level. For example, the model compression 19-1 may include quantization workflows, weight pruning, and sparsification techniques, etc. The hardware acceleration 19-2 tool may facilitate the configuration of model storage and execution formats to operate optimally on different hardware resources. For example, the hardware acceleration 19-2 may include tools for optimally slicing the model for distributed processing on multiple processing units to increase bandwidth, reduce unified memory requirements, etc. The distillation 19-3 tool may be used to train a more lightweight model based on the knowledge encoded in the development model 16. For example, the development model 16 may be a high-performance large-scale machine learning model optimized using the model development platform 12. In order to obtain a lightweight model that runs in a resource-constrained environment, the smaller model may be a "student model" that learns to imitate the development model 16 as a "teacher model." In this way, for example, the investment in learning the parameters and configuration of the development model 16 can be efficiently transferred to a smaller model for more efficient inference.
[0103] The workbench 15 may implement one, more than one, or no tool suite implemented in the model development platform 12. The workbench 15 may output an output model 20 based on the development model 16. The output model 20 may be a deployed version of the development model 16. The output model 20 may be a development or training checkpoint of the development model 16. The output model 20 may be a distilled, compressed, or otherwise optimized version of the development model 16.
[0104] Figure 6 1 is a block diagram of an example training process for training a development model 16 for machine learning. One or more portions of the example training process may be implemented by a computing system including one or more computing devices, such as, for example, the computing systems described with reference to other figures. Each respective portion of the example training process may be performed by any one (or any combination) of the one or more computing devices. In addition, one or more portions of the example training process may be implemented on hardware components of the devices described herein, for example, to train one or more systems or models. Figure 6Elements performed in a specific order are depicted for purposes of illustration and discussion. One of ordinary skill in the art will understand, using the disclosure provided herein, that the elements of any method discussed herein may be adjusted, rearranged, expanded, omitted, combined or modified in various ways without departing from the scope of the present disclosure. Figure 6 The elements / terms described with reference to other systems and figures are described for illustrative purposes and are not intended to be limiting. Additionally or alternatively, one or more portions of the example training process may be performed by other systems.
[0105] Initially, the development model 16 may be maintained in an initial state as an initialization model 21. The development model 16 may be initialized using weight values. The initial weight values may be random or based on an initialization pattern. The initial weight values may be based on previous pre-training for the same or a different model.
[0106] The initialization model 21 may be pretrained in a pretraining stage 22. The pretraining stage 22 may be implemented using one or more pretraining pipelines 17-2 on data from the dataset 17-1. For example, if the initialization model 21 has been pretrained (e.g., the development model 16 includes, is, or is based on a pretrained base model or expert model), pretraining may be omitted.
[0107] The pre-trained model 23 may then be a new version of the development model 16, which may be maintained as the development model 16 or a new development model. If the development model 16 has been pre-trained, the pre-trained model 23 may be the initial state. The pre-trained model 23 may be fine-tuned in a fine-tuning phase 24. The fine-tuning phase 24 may be implemented using one or more fine-tuning pipelines 17-3 on data from the dataset 17-1. For example, if the pre-trained model has satisfactory performance, if the model has already been fine-tuned, or if other methods of adjustment are preferred, fine-tuning may be omitted.
[0108] The fine-tuned model 29 may then be a new version of the development model 16, which may be maintained as the development model 16 or a new development model. If the development model 16 has already been fine-tuned, the fine-tuned model 29 may be the initial state. The fine-tuned model 29 may undergo refinement 26 with user feedback. For example, refinement 26 with user feedback may include performing reinforcement learning, optionally based on human feedback from human users of the fine-tuned model 25. Since reinforcement learning may be a form of fine-tuning, it should be understood that the fine-tuning stage 24 may include a stage of refinement 26 with user feedback. Refinement 26 with user feedback may produce a refined model 27. The refined model 27 may be output to a downstream system 28 for deployment or further development.
[0109] In some implementations, computational optimization operations may be applied before, during, or after each stage. For example, the initialized model 21 may be computationally optimized 29-1 (e.g., using the computational optimization toolkit 19) prior to the pre-training stage 22. The pre-trained model 23 may be computationally optimized 29-2 (e.g., using the computational optimization toolkit 19) prior to the fine-tuning stage 24. The fine-tuned model 25 may be computationally optimized 29-3 (e.g., using the computational optimization toolkit 19) prior to refinement 26 utilizing user feedback. The refined model 27 may be computationally optimized 29-4 (e.g., using the computational optimization toolkit 19) prior to output to a downstream system 28. The computational optimizations 29-1, ..., 29-4 may all be the same, all different, or include at least some different optimization techniques.
[0110] Example machine learning model inference system
[0111] Figure 7 1 is a block diagram of an inference system for operating one or more machine learning models 1 to perform inference (e.g., for training, for deployment, etc.). A model host 31 may receive a machine learning model 1. The model host 31 may host one or more model instances 31-1, which may be one or more instances of one or more models. The model host 31 may host the model instance 31-1 using available computing resources 31-2 associated with the model host 31.
[0112] Model master 31 can perform inference on behalf of one or more clients 32. Client 32 can transmit input request 33 to model master 31. Using input request 33, model master 31 can obtain input 2 to input into machine learning model 1. Machine learning model 1 can process input 2 to generate output 3. Using output 3, model master 31 can return output payload 34 in response to input request 33 from client 32. Output payload 34 can include or be based on output 3.
[0113] The model host 31 can utilize various other resources and tools to enhance the inference task. For example, the model host 31 can communicate with the tool interface 35 to facilitate the use of tools by the model instance 31-1. The tool interface 35 may include a local or remote API. The tool interface 35 may include integrated scripts or other software functions. The model host 31 may use the online learning interface 36 to promote the continuous improvement of the machine learning model 1. For example, the online learning interface 36 can be used in a reinforcement learning loop to retrieve user feedback on the inference served by the model host 31. The model host 31 can access the runtime data source 37 for enhancing the input 2 with additional context information. For example, the runtime data source 37 may include a knowledge graph 37-1 that facilitates structured information retrieval for information associated with the input request 33 (e.g., a search engine service). The runtime data source 37 may include a public or private, external or local database 37-2 that can store information associated with the input request 33 for enhancing the input 2. The runtime data source 37 may include account data 37 - 3 , which may be retrieved in association with a user account corresponding to the client 32 to customize the behavior of the model host 31 accordingly.
[0114] The model host 31 may be implemented by one or more computing devices or systems. The client 2 may be implemented by one or more computing devices or systems, which may include a computing device or system shared with the model host 31 .
[0115] For example, the model host 31 may be operated on a server system that provides machine learning services (e.g., over a local area network or wide area network) to client devices operating the client 32. The client device may be an end-user device used by an individual. The client device may be a server system that operates the client 32 to provide various functions as services to downstream end-user devices.
[0116] In some implementations, the model host 31 may operate on the same device or system as the client 32. The model host 31 may be a machine learning service that runs on a device to provide machine learning functionality to one or more applications operating on a client device, which may include an application that implements the client 32. The model host 31 and the client 32 may be part of the same application. For example, the model host 31 may be a subroutine or method implemented by a portion of the application, and the client 32 may be another subroutine or method that uses the model host 31 to perform inference functionality within the application. It should be understood that the model host 31 and the client 32 may have a variety of different configurations.
[0117] Model instance 31-1 may include one or more machine learning models that can be used to perform inference. Model instance 31-1 may include weights or other model components stored on / in a persistent storage device, temporarily cached, or loaded into a high-speed memory. Model instance 31-1 may include multiple instances of the same model (e.g., for executing more requests in parallel on the same model). Model instance 31-1 may include instances of different models. Model instance 31-1 may include cached intermediate states of active or inactive models, which are used to accelerate the inference of those models. For example, an inference session with a particular model may generate a significant amount of computational results that can be reused for future inference runs (e.g., using a KV cache for a Transformer-based model). These computational results may be stored in association with the inference session so that the session can be executed more efficiently when it is resumed.
[0118] Computing resources 31-2 may include one or more processors (central processing units, graphics processing units, tensor processing units, machine learning accelerators, etc.) connected to one or more memory devices. Computing resources 31-2 may include a dynamic pool of available resources shared with other processes. Computing resources 31-2 may include a memory device large enough to fit the entire model instance in a single memory instance. Computing resources 31-2 may also share model instances across multiple memory devices (e.g., using data parallelism or tensor parallelism, etc.). Doing so can increase parallelism or execute large models using multiple memory devices that individually may not be able to fit the entire model in memory.
[0119] The input request 33 may include data for the input 2. The model host 31 may process the input request 33 to obtain the input 2. The input 2 may be obtained directly from the input request 33 or may be retrieved using the input request 33. The input request 33 may be submitted to the model host 31 via an API.
[0120] The model host 31 can perform inference on multiple batches of input requests 33 in parallel. For example, the model instance 31-1 can be configured with an input structure having a batch dimension. Individual inputs 2 can be distributed across the batch dimension (e.g., rows of an array). Individual inputs 2 can include completely different contexts. Individual inputs 2 can be multiple inference steps of the same task. Individual inputs 2 can be interleaved in the input structure so that any given inference cycle can operate on different parts of the corresponding input 2. In this way, for example, the model host 31 can perform inference on batches in parallel, so that the output 3 can also contain a batch dimension and return the inference results of the batched inputs 2 in parallel. In this way, for example, multiple batches of input requests 33 can be processed in parallel to achieve a higher throughput of the output payload 34.
[0121] The output payload 34 may include or be based on the output 3 from the machine-learned model 1. The model host 31 may process the output 3 to obtain the output payload 34. This may include chaining multiple rounds of inference (e.g., iteratively, recursively, across the same model or different models) to obtain the final output of the task to be returned in the output payload 34. The output payload 34 may be transmitted to the client 32 via an API.
[0122] The online learning interface 36 can facilitate reinforcement learning of the machine-learned model 1. The online learning interface 36 can facilitate reinforcement learning with human feedback (RLHF). The online learning interface 36 can facilitate federated learning of the machine-learned model 1.
[0123] The model host 31 can execute the machine learning model 1 to perform inference for various tasks using various types of data. For example, various different inputs 2 and outputs 3 can be used for various different tasks. In some implementations, the input 2 can be or otherwise represent image data. The machine learning model 1 can process the image data to generate an output. As an example, the machine learning model 1 can process the image data to generate an image recognition output (e.g., recognition of image data, potential embedding of image data, encoded representation of image data, hash of image data, etc.). As another example, the machine learning model 1 can process the image data to generate an image segmentation output. As another example, the machine learning model 1 can process the image data to generate an image classification output. As another example, the machine learning model 1 can process the image data to generate an image data modification output (e.g., a change in image data, etc.). As another example, the machine learning model 1 can process the image data to generate an encoded image data output (e.g., an encoded and / or compressed representation of image data, etc.). As another example, the machine learning model 1 can process the image data to generate an amplified image data output. As another example, the machine learning model 1 can process the image data to generate a predicted output.
[0124] In some implementations, the task is a computer vision task. In some cases, the input 2 includes pixel data of one or more images, and the task is an image processing task. For example, the image processing task may be image classification, where the output is a set of scores, each score corresponding to a different object class and representing the possibility that one or more images depict an object belonging to the object class. The image processing task may be object detection, where the image processing output identifies one or more regions in one or more images, and for each region, identifies the possibility that the region depicts an object of interest. As another example, the image processing task may be image segmentation, where the image processing output defines the corresponding possibility of each category in a predetermined category set for each pixel in one or more images. For example, the category set may be a foreground and a background. As another example, the category set may be an object class. As another example, the image processing task may be depth estimation, where the image processing output defines a corresponding depth value for each pixel in one or more images. As another example, the image processing task may be motion estimation, where the network input includes multiple images, and the image processing output defines the motion of the scene depicted at the pixel between images in the network input for each pixel of one of the input images.
[0125] In some implementations, input 2 may be or otherwise represent natural language data. The machine-learned model 1 may process natural language data to generate an output. As an example, the machine-learned model 1 may process natural language data to generate a language encoding output. As another example, the machine-learned model 1 may process natural language data to generate a potential text embedding output. As another example, the machine-learned model 1 may process natural language data to generate a translation output. As another example, the machine-learned model 1 may process natural language data to generate a classification output. As another example, the machine-learned model 1 may process natural language data to generate a text segmentation output. As another example, the machine-learned model 1 may process natural language data to generate a semantic intent output. As another example, the machine-learned model 1 may process natural language data to generate an amplified text or natural language output (e.g., text or natural language data of higher quality than the input text or natural language, etc.). As another example, the machine-learned model 1 may process natural language data to generate a prediction output (e.g., one or more predicted next parts of natural language content).
[0126] In some implementations, input 2 may be or otherwise represent speech data (e.g., data describing spoken natural language, such as audio data, text data, etc.). The machine-learned model 1 may process the speech data to generate an output. As an example, the machine-learned model 1 may process the speech data to generate a speech recognition output. As another example, the machine-learned model 1 may process the speech data to generate a speech translation output. As another example, the machine-learned model 1 may process the speech data to generate a potential embedding output. As another example, the machine-learned model 1 may process the speech data to generate an encoded speech output (e.g., an encoded and / or compressed representation of the speech data, etc.). As another example, the machine-learned model 1 may process the speech data to generate an amplified speech output (e.g., speech data of higher quality than the input speech data, etc.). As another example, the machine-learned model 1 may process the speech data to generate a text representation output (e.g., a text representation of the input speech data, etc.). As another example, the machine-learned model 1 may process the speech data to generate a predicted output.
[0127] In some implementations, input 2 may be or otherwise represent potentially coded data (e.g., a latent space representation of an input, etc.). The machine-learned model 1 may process the potentially coded data to generate an output. As an example, the machine-learned model 1 may process the potentially coded data to generate a recognition output. As another example, the machine-learned model 1 may process the potentially coded data to generate a reconstruction output. As another example, the machine-learned model 1 may process the potentially coded data to generate a search output. As another example, the machine-learned model 1 may process the potentially coded data to generate a re-clustering output. As another example, the machine-learned model 1 may process the potentially coded data to generate a prediction output.
[0128] In some implementations, input 2 may be or otherwise represent statistical data. Statistical data may be, represent, or otherwise include data calculated and / or measured from some other data source. Machine-learned model 1 may process statistical data to generate an output. As an example, machine-learned model 1 may process statistical data to generate an identification output. As another example, machine-learned model 1 may process statistical data to generate a prediction output. As another example, machine-learned model 1 may process statistical data to generate a classification output. As another example, machine-learned model 1 may process statistical data to generate a segmentation output. As another example, machine-learned model 1 may process statistical data to generate a visualization output. As another example, machine-learned model 1 may process statistical data to generate a diagnostic output.
[0129] In some implementations, input 2 may be or otherwise represent sensor data. Machine-learned model 1 may process sensor data to generate an output. As an example, machine-learned model 1 may process sensor data to generate a recognition output. As another example, machine-learned model 1 may process sensor data to generate a prediction output. As another example, machine-learned model 1 may process sensor data to generate a classification output. As another example, machine-learned model 1 may process sensor data to generate a segmentation output. As another example, machine-learned model 1 may process sensor data to generate a visualization output. As another example, machine-learned model 1 may process sensor data to generate a diagnostic output. As another example, machine-learned model 1 may process sensor data to generate a detection output.
[0130] In some implementations, the machine learning model 1 can be configured to perform tasks including encoding input data to achieve reliable and / or efficient transmission or storage (and / or corresponding decoding). For example, the task can be an audio compression task. The input can include audio data, and the output can include compressed audio data. In another example, the input includes visual data (e.g., one or more images or videos), the output includes compressed visual data, and the task is a visual data compression task. In another example, the task can include generating embeddings for input data (e.g., input audio or visual data). In some cases, the input includes audio data representing spoken utterances, and the task is a speech recognition task. The output can include text output mapped to spoken utterances. In some cases, the task includes encrypting or decrypting the input data. In some cases, the task includes microprocessor performance tasks, such as branch prediction or memory address conversion.
[0131] In some implementations, the task is a generative task, and the machine-learned model 1 can be configured to output content generated according to input 2. For example, input 2 can be or otherwise represent data of one or more modalities that encodes the context for generating additional content.
[0132] In some implementations, the task may be a text completion task. The machine-learned model 1 may be configured to process an input 2 representing text data and generate an output 3 representing additional text data that completes a text sequence including the input 2. For example, the machine-learned model 1 may be configured to generate an output 3 to complete a sentence, paragraph, or portion of text following a portion of text represented by the input 2.
[0133] In some implementations, the task may be an instruction-following task. The machine-learned model 1 may be configured to process an input 2 representing an instruction for performing a function and generate an output 3 that advances a goal of satisfying the instruction function (e.g., at least one step of a multi-step process for performing the function). The output 3 may represent data of the same or different modality as the input 2. For example, the input 2 may represent text data (e.g., a natural language instruction for a task to be performed), and the machine-learned model 1 may process the input 2 to generate an output 3 that represents text data in response to the instruction (e.g., a natural language response, a programming language response, a machine language response, etc.). The input 2 may represent image data (e.g., an image-based instruction for a task to be performed, optionally with a text instruction), and the machine-learned model 1 may process the input 2 to generate an output 3 that represents text data in response to the instruction (e.g., a natural language response, a programming language response, a machine language response, etc.). One or more outputs 3 may be generated iteratively or recursively to sequentially process and complete steps toward completing the requested function. For example, the initial output may be executed by an external system or processed by the machine-learned model 1 to complete the initial steps of executing the function. Multiple steps may be performed and a final output obtained in response to the initial instructions.
[0134] In some implementations, the task may be a question-answering task. The machine learning model 1 may be configured to process an input 2 representing a question to be answered and generate an output 3 that advances the goal of returning an answer to the question (e.g., at least one step of a multi-step process for performing the function). The output 3 may represent data of the same or different modality as the input 2. For example, the input 2 may represent text data (e.g., natural language instructions for a task to be performed), and the machine learning model 1 may process the input 2 to generate an output 3 that represents text data in response to the question (e.g., a natural language response, a programming language response, a machine language response, etc.). The input 2 may represent image data (e.g., image-based instructions for a task to be performed, optionally with text instructions), and the machine learning model 1 may process the input 2 to generate an output 3 that represents text data in response to the question (e.g., a natural language response, a programming language response, a machine language response, etc.). One or more outputs 3 may be generated iteratively or recursively to sequentially process and complete steps toward answering the question. For example, the initial output can be executed by an external system or processed by the machine learning model 1 to complete the initial steps of obtaining the answer to the question (e.g., querying a database, performing calculations, executing a script, etc.). Multiple steps can be performed and a final output responsive to the question is obtained.
[0135] In some implementations, the task may be an image generation task. The machine-learned model 1 may be configured to process an input 2 representing a context about a desired portion of the image content. The context may include text data, image data, audio data, etc. The machine-learned model 1 may be configured to generate an output 3 representing image data depicting an image associated with the context. For example, the machine-learned model 1 may be configured to generate pixel data for an image. The value of a channel associated with a pixel in the pixel data may be selected based on the context (e.g., based on a probability determined based on the context).
[0136] In some implementations, the task may be an audio generation task. The machine learning model 1 may be configured to process an input 2 representing a context about a desired portion of the audio content. The context may include text data, image data, audio data, etc. The machine learning model 1 may be configured to generate an output 3 representing audio data related to the context. For example, the machine learning model 1 may be configured to generate waveform data in the form of an image (e.g., a spectrogram). The values of the channels associated with the pixels of the image may be selected based on the context. The machine learning model 1 may be configured to generate waveform data in the form of a sequence of discrete samples of a continuous waveform. The value of the sequence may be selected based on the context (e.g., based on a probability determined according to the context).
[0137] In some implementations, the task may be a data generation task. The machine learning model 1 may be configured to process an input 2 representing a context about an expected portion of data (e.g., data from various data domains, such as sensor data, image data, multimodal data, statistics, etc.). For example, the expected data may be synthetic data used to train other machine learning models. The context may include any data type. The machine learning model 1 may be configured to generate an output 3 representing data aligned with the expected data. For example, the machine learning model 1 may be configured to generate data values for populating a data set. The value of a data object may be selected based on the context (e.g., based on a probability determined based on the context).
[0138] Example computing systems and devices
[0139] Figure 84 is a block diagram of an example networked computing system that can perform aspects of an example implementation of the present disclosure. The system can include multiple computing devices and systems that are communicatively coupled via a network 49. An example computing device 50 is described to provide an example of a computing device that can perform any aspect of the present disclosure (e.g., implement a model host 31, a client 32, or both). An example server computing system 60 is described as an example of a server computing system that can perform any aspect of the present disclosure (e.g., implement a model host 31, a client 32, or both). The computing device 50 and the server computing system 60 can interact collaboratively (e.g., via a network 49) to perform any aspect of the present disclosure (e.g., implement a model host 31, a client 32, or both). The model development platform system 70 is an example system that can host or serve a model development platform 12 for developing models for machine learning. The third-party system 80 is an example system that any of the computing device 50, the server computing system 60, or the model development platform system 70 can interact with when performing various aspects of the present disclosure (e.g., using third-party tools, accessing third-party databases or other resources, etc.).
[0140] Network 49 may be any type of communications network, such as a local area network (e.g., an intranet), a wide area network (e.g., the Internet), or some combination thereof, and may include any number of wired or wireless links. In general, communications over network 49 may be carried via any type of wired or wireless connection using a wide variety of communications protocols (e.g., TCP / IP, HTTP, SMTP, FTP), encodings or formats (e.g., HTML, XML), or protection schemes (e.g., VPN, secure HTTP, SSL). Network 49 may also be implemented via a system bus. For example, Figure 8 One or more devices or systems may be co-located with, housed by, or otherwise integrated into one or more other devices or systems.
[0141] Computing device 50 may be any type of computing device, such as, for example, a personal computing device (e.g., a laptop or desktop computer), a mobile computing device (e.g., a smartphone or tablet computer), a game console or controller, a wearable computing device, an embedded computing device, a server computing device, a virtual machine operating on a host device, or any other type of computing device. Computing device 50 may be a client computing device. Computing device 50 may be an end-user computing device. Computing device 50 may be a computing device for a provided service, which provides services to an end user (who may use another computing device to interact with computing device 50).
[0142] The computing device 50 may include one or more processors 51 and a memory 52. The processor 51 may be any suitable processing device (e.g., a processor core, a microprocessor, an ASIC, an FPGA, a controller, a microcontroller, etc.) and may be one processor or a plurality of processors operatively connected. The memory 52 may include one or more non-transitory computer-readable storage media, such as HBM, RAM, ROM, EEPROM, EPROM, flash memory devices, disks, etc., and combinations thereof. The memory 52 may store data 53 and instructions 54 that may be executed by the processor 51 to cause the computing device 50 to perform operations. The operations may implement any one or more features described herein. The operations may implement the example methods and techniques described herein.
[0143] The computing device 50 may also include one or more input components that receive user input. For example, the user input component may be a touch-sensitive component (e.g., a touch-sensitive display screen or touchpad) that is sensitive to the touch of a user input object (e.g., a finger or stylus). The touch-sensitive component may be used to implement a virtual keyboard. Other example user input components include a microphone, a camera, a LIDAR, a physical keyboard or other buttons, or other devices through which a user can provide user input.
[0144] The computing device 50 may store or include one or more machine-learned models 55. The machine-learned models 55 may include one or more machine-learned models 1, such as a sequence processing model 4. The machine-learned models 55 may include one or more model instances 31-1. The machine-learned models 55 may be received from a server computing system 60, a model development platform system 70, a third-party system 80 (e.g., an application distribution platform), or developed locally on the computing device 50. The machine-learned models 55 may be loaded into the memory 52 and used or otherwise implemented by the processor 51. The computing device 50 may implement multiple parallel instances of the machine-learned models 55.
[0145] The server computing system 60 may include one or more processors 61 and memory 62. The processor 61 may be any suitable processing device (e.g., a processor core, a microprocessor, an ASIC, an FPGA, a controller, a microcontroller, etc.) and may be one processor or a plurality of processors operatively connected. The memory 62 may include one or more non-transitory computer-readable storage media, such as HBM, RAM, ROM, EEPROM, EPROM, flash memory devices, disks, etc., and combinations thereof. The memory 62 may store data 63 and instructions 64, which may be executed by the processor 61 to cause the server computing system 60 to perform operations. The operations may implement any one or more of the features described herein. The operations may implement the example methods and techniques described herein.
[0146] In some implementations, the server computing system 60 includes or is otherwise implemented by one or more server computing devices. In instances where the server computing system 60 includes multiple server computing devices, such server computing devices may operate according to a sequential computing architecture, a parallel computing architecture, or some combination thereof.
[0147] The server computing system 60 may store or otherwise include one or more machine-learned models 65. The machine-learned models 65 may be the same as or different from the machine-learned models 55. The machine-learned models 65 may include one or more machine-learned models 1, such as a sequence processing model 4. The machine-learned models 65 may include one or more model instances 31-1. The machine-learned models 65 may be received from the computing device 50, the model development platform system 70, a third-party system 80, or developed locally on the server computing system 60. The machine-learned models 65 may be loaded into the memory 62 and used or otherwise implemented by the processor 61. The server computing system 60 may implement multiple parallel instances of the machine-learned models 65.
[0148] In an example configuration, the machine-learned model 65 may be included in the server computing system 60 or otherwise stored and implemented by the server computing system 60 to establish a client-server relationship with the computing device 50 for providing model inference. For example, the server computing system 60 may implement the model host 31 on behalf of the client 32 on the computing device 50. For example, the machine-learned model 65 may be implemented by the server computing system 60 as part of a network service (e.g., a remote machine-learned model hosting service, such as an online interface for performing machine-learned model operations on the server computing system 60 over a network). For example, the server computing system 60 may communicate with the computing device 50 via a local intranet or Internet connection. For example, the computing device 50 may be a workstation or endpoint in communication with the server computing system 60, wherein the implementation of the machine-learned model 65 is managed by the server computing system 60 to remotely perform inference (e.g., for runtime or training operations), wherein the output is returned (e.g., projected, streamed, etc.) to the computing device 50. The machine-learned model 65 may work cooperatively or interoperably with the machine-learned model 55 on the computing device 50 to perform various tasks.
[0149] The model development platform system 70 may include one or more processors 71 and a memory 72. The processor 71 may be any suitable processing device (e.g., a processor core, a microprocessor, an ASIC, an FPGA, a controller, a microcontroller, etc.) and may be a processor or a plurality of processors operatively connected. The memory 72 may include one or more non-temporary computer-readable storage media, such as HBM, RAM, ROM, EEPROM, EPROM, flash memory devices, disks, etc., and combinations thereof. The memory 72 may store data 73 and instructions 74, which may be executed by the processor 71 to cause the model development platform system 70 to perform operations. The operation may implement any one or more features described herein. The operation may implement the example methods and techniques described herein. The example operation includes the functions described herein with respect to the model development platform 12. This function and other functions may be implemented by a developer tool 75.
[0150] The third-party system 80 may include one or more processors 81 and a memory 82. The processor 81 may be any suitable processing device (e.g., a processor core, a microprocessor, an ASIC, an FPGA, a controller, a microcontroller, etc.) and may be a processor or multiple processors operatively connected. The memory 82 may include one or more non-temporary computer-readable storage media, such as HBM, RAM, ROM, EEPROM, EPROM, flash memory devices, disks, etc., and combinations thereof. The memory 82 may store data 83 and instructions 84, which may be executed by the processor 81 to enable the third-party system 80 to perform operations. The operation may implement any one or more features described herein. The operation may implement the example methods and techniques described herein. The example operations include the functions described herein with respect to tools and other external resources (e.g., third-party resources 85) called when training machine learning models 1, 4, 16, 20, 55, 65, etc. or performing inference using the machine learning model.
[0151] Figure 8An example arrangement of a computing system that can be used to implement the present disclosure is illustrated. Other computing system configurations may also be used. For example, in some implementations, one or both of the computing system 50 or the server computing system 60 may implement all or part of the operations of the model development platform system 70. For example, the computing system 50 or the server computing system 60 may implement the developer tool 75 (or its extension) to develop, update / train, or refine machine learning models 1, 4, 16, 20, 55, 65, etc. using one or more techniques described herein with respect to the model alignment toolkit 17. In this way, for example, the computing system 50 or the server computing system 60 may develop, update / train, or refine machine learning models based on a local data set (e.g., for model personalization / customization, as allowed by user data preference selections).
[0152] Fig. 9 is a block diagram of an example computing device 98 executed according to an example embodiment of the present disclosure. Computing device 98 may be a user computing device or a server computing device (e.g., computing device 50, server computing system 60, etc.). Computing device 98 may implement model host 31. For example, computing device 98 may include multiple applications (e.g., applications 1 to N). Each application may contain its own machine learning library and machine learning model. For example, each application may include a machine learning model. Example applications include text messaging applications, email applications, dictation applications, virtual keyboard applications, browser applications, etc. Fig. 9 As shown, each application can communicate with multiple other components of the computing device (e.g., such as one or more sensors, a context manager, a device state component, or additional components). In some implementations, each application can communicate with each device component using an API (e.g., a public API). In some implementations, the API used by each application is specific to the application.
[0153] Fig.10 is a block diagram of an example computing device 99 executed in accordance with an example embodiment of the present disclosure. Computing device 99 may be the same as or different from computing device 98. Computing device 99 may be a user computing device or a server computing device (e.g., computing device 50, server computing system 60, etc.). Computing device 98 may implement model host 31. For example, computing device 99 may include multiple applications (e.g., applications 1 to N). Each application may communicate with a central intelligence layer. Example applications include text messaging applications, email applications, dictation applications, virtual keyboard applications, browser applications, etc. In some implementations, each application may use an API (e.g., a public API across all applications) to communicate with the central intelligence layer (and the models stored therein).
[0154] The central intelligence layer can include multiple machine learning models. Fig.10 As shown, a corresponding machine learning model can be provided for each application, and the corresponding machine learning model is managed by the central intelligence layer. In other implementations, two or more applications can share a single machine learning model. For example, in some implementations, the central intelligence layer can provide a single model for all applications. In some implementations, the central intelligence layer is included in the operating system of the computing device 99 or is otherwise implemented by the operating system.
[0155] The central intelligence layer can communicate with the central device data layer. The central device data layer can be a centralized data repository for the computing device 99. Fig.10 As shown, the central device data layer can communicate with multiple other components of the computing device (e.g., such as one or more sensors, context managers, device state components, or additional components). In some implementations, the central device data layer can communicate with each device component using an API (e.g., a private API).
[0156] Additional disclosure
[0157] The technology discussed herein relates to servers, databases, software applications, and other computer-based systems, as well as the actions taken and the information sent to and from such systems. The inherent flexibility of computer-based systems allows for a variety of possible configurations, combinations, and partitioning of tasks and functions between and among components. For example, the processes discussed herein can be implemented using a single device or component or multiple devices or components working in combination. Databases and applications can be implemented on a single system or distributed across multiple systems. Distributed components can operate sequentially or in parallel.
[0158] Although the subject matter has been described in detail with respect to various specific example embodiments of the subject matter, each example is provided by way of explanation, rather than limitation of the present disclosure. After obtaining an understanding of the foregoing, those skilled in the art can easily produce changes, modifications, and equivalents to such embodiments. Therefore, the present disclosure does not exclude such modifications, changes, and / or additions to the subject matter that will be readily understood by those of ordinary skill in the art. For example, a feature shown or described as part of one embodiment can be used together with another embodiment to produce yet another embodiment. Therefore, the present disclosure is intended to cover such changes, changes, and equivalents.
[0159] Aspects of the present disclosure have been described with respect to its illustrative embodiments. Any and all features in the appended claims may be combined or rearranged in any possible manner, including combinations of claims that are not explicitly listed in a combined manner, because the example claim dependencies listed herein should not be understood as limiting the scope of the possible combinations of features disclosed herein. Therefore, the scope of the present disclosure is illustrative rather than limiting, and the present disclosure does not exclude such modifications, changes and / or additions to this theme that will be easily understood by those of ordinary skill in the art. In addition, the example element lists connected by conjunctions such as "and", "or", "but" are used herein to describe terms. It should be understood that such conjunctions are provided only for the purpose of explanation. For example, clauses and other project sequences connected by specific conjunctions (such as, for example, "or") can refer to "and / or", "at least one of", "any combination" of the example elements listed therein, etc. Terms such as "based on" should be understood as "based at least in part on".
[0160] The term "may" should be understood to refer to the possibility of a feature in various implementations, rather than to specify a capability that must be present in every implementation. For example, the phrase "X may perform Y" should be understood to indicate that in various implementations, X may be configured to perform Y, rather than to indicate that in every instance X must always be able to perform Y. It should be understood that in various implementations, X may not be able to perform Y and still be within the scope of the present disclosure.
[0161] The term "may" should be understood to refer to the possibility of a feature in various implementations, rather than to specify a capability that must be present in every implementation. For example, the phrase "X may perform Y" should be understood to indicate that in various implementations, X may be configured to perform Y, rather than to indicate that in every instance X must always be able to perform Y. It should be understood that in various implementations, X may not be able to perform Y and still be within the scope of the present disclosure.
Claims
1. A computer-implemented method for training a sequence processing model for use in a recommender system, the method comprising: Obtaining, by a computing system including one or more computing devices, an item data set describing a plurality of items included in a candidate pool; generating, by the computing system, a plurality of auxiliary prompts for use in training the sequence processing model, wherein each auxiliary prompt comprises a prompt input and a prompt output, and wherein the plurality of auxiliary prompts encode recommendation-related knowledge about the plurality of items; training, by the computing system, the sequence processing model using the plurality of auxiliary cues; as well as The trained sequence processing model is provided by the computing system for use in a recommender system.
2. The computer-implemented method of claim 1, wherein the one or more auxiliary prompts include item-embedded prompts that encode knowledge about the plurality of items.
3. The computer-implemented method of claim 2, wherein: The item data set describes, for each item in the plurality of items, one or more attribute values of one or more item attributes; and For at least one of the item embedding prompts: the prompt input identifies an item among the plurality of items, and the prompt output includes an attribute value for at least one of the one or more item attributes for the identified item.
4. A computer-implemented method as claimed in claim 2 or 3, wherein: The item data set describes, for each item in the plurality of items, one or more attribute values of one or more item attributes; and For at least one of the item embedding prompts: the prompt input includes a property value for at least one of the one or more item properties, and the prompt output identifies one or more items in the project having the provided property value.
5. The computer-implemented method of claim 2 or 3, wherein the one or more item attributes include a title, a category, a brand, a description, or a review. 6 . The computer-implemented method of claim 1 , wherein the project data set specifies a plurality of users and historical interactions between each of the plurality of users and one or more of the plurality of projects.
7. The computer-implemented method of claim 6, wherein the one or more auxiliary prompts include a Bayesian Personalized Ranking (BPR) loss reduction prompt that encodes knowledge about the historical interactions between the user and the item.
8. The computer-implemented method of claim 7, wherein for at least one of the BPR loss reduction prompts: The prompt input identifies a user, a positive item for the user, and a negative item for the user; and The prompt output identifies the positive item.
9. A computer-implemented method as claimed in any one of claims 6 to 8, wherein: The one or more auxiliary prompts include a masked item modeling prompt; and For at least one of the masked item modeling hints: The prompt input includes a list of items, the list of items including a masked item; and The hint output identifies an item that is masked to generate the masked item.
10. The computer-implemented method of claim 9, wherein generating the masked item modeling hint comprises: identifying, based on the item dataset, a sequence of items with which one of the plurality of users has interacted; as well as masking an item in the item sequence using a masking item to generate the prompt input; The one item in the sequence of items that is masked is a non-terminal item in the sequence of items.
11. The computer-implemented method of claim 10, wherein identifying, based on the item dataset, the sequence of items with which one of the plurality of users has interacted comprises applying a sliding window to extract the sequence of items.
12. The computer-implemented method of claim 1, wherein in some or all of the plurality of auxiliary prompts, a user identifier of a user of the plurality of users is replaced with a sequence of items with which the user has interacted.
13. The computer-implemented method of claim 1, wherein in some or all of the plurality of auxiliary prompts, item identifiers of items of the plurality of items are replaced with shortened identifiers.
14. The computer-implemented method of claim 1, further comprising, after the computing system uses the plurality of auxiliary prompts to train the sequence processing model, but before the computing system provides the trained sequence processing model for use in the recommendation system: training the sequence processing model using one or more recommendation task prompts by the computing system.
15. The computer-implemented method of claim 1, wherein providing, by the computing system, the trained sequence processing model for use in the recommendation system comprises instructing, by the computing system, the trained sequence processing model to perform a retrieval task.
16. The computer-implemented method of claim 1, wherein providing, by the computing system, the trained sequence processing model for use in the recommender system comprises instructing, by the computing system, the trained sequence processing model to perform a ranking task.
17. A computer-implemented method as described in claim 1, wherein providing the trained sequence processing model by the computing system for use in the recommendation system includes instructing the trained sequence processing model by the computing system to perform a rating prediction task.
18. The computer-implemented method of claim 1, wherein the sequence processing model has not been previously trained on data specific to the candidate pool.
19. A recommender system comprising a sequence processing model that has been trained on the auxiliary prompt as described in any preceding claim.
20. One or more non-transitory computer-readable media collectively storing a sequence processing model that has been trained on the auxiliary prompt described in any one of claims 1 to 18.
21. A computer system comprising: One or more processors and one or more non-transitory computer-readable media that collectively store: A sequence processing model trained using one or more auxiliary cues as described in any one of claims 1 to 18; as well as Computer executable instructions for performing operations comprising: receiving a query associated with a user; generating a recommended task prompt based on the query; processing the recommendation task prompt using the sequence processing model to obtain a recommendation output identifying one or more items; and The one or more items are provided to the user as a recommendation output.
22. The computer system of claim 21, wherein the computer system comprises a user computing device and the sequence processing model is implemented on a device on the user computing device.
23. The computer system of claim 21 or 22, wherein the computer system does not store a data set of items such that the recommendation output identifying the one or more items is generated without accessing the data set of items.