Adapting cues selected from set of cue tasks
By prompting the development system to optimize the NLP ML model, the problem of poor performance in specific applications is solved, and the performance and interaction capabilities of the model are improved.
Patent Information
- Application Number
- CN202380084868.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-12-16
- Filing Date
- 2023-12-12
- Publication Date
- 2025-07-18
AI Technical Summary
The existing technology is difficult to effectively discover and optimize prompts of machine learning models for natural language processing, resulting in poor performance of models in specific applications.
Through the development prompt development system, interfaces and tools are provided to allow users to submit, evaluate and optimize prompts for NLP ML models, and use various technologies such as prompt discovery, adaptation of job management and model deployment to optimize model performance.
Improves the performance and accuracy of the NLP ML model in specific applications, and enhances the interaction ability of systems, services or applications with human users.
Smart Images

Figure CN120344972A_ABST
Abstract
Description
BACKGROUND OF THE INVENTION
[0001] Machine learning models and data-driven systems have increasingly been used to help make decisions in application areas such as financial services, healthcare, education, and human resources. These applications provide benefits such as increased accuracy, increased productivity, and cost savings. This trend is the result of a convergence of factors such as ubiquitous connectivity, the ability to use cloud computing to collect, aggregate, and process large amounts of fine-grained data, and improved access to increasingly sophisticated machine learning models that can analyze this data. BRIEF DESCRIPTION OF THE DRAWINGS
[0002] Figure 1 A logic block diagram of a prompt development system for a natural language processing machine learning model for generating prompt recommendations and prompt development, according to some embodiments, is shown.
[0003] Figure 2 An example provider network for implementing machine learning services that implement prompt discovery and development for natural language processing tasks, according to some embodiments, is shown.
[0004] Figure 3 A logic block diagram of an interaction for submitting prompts for discovery, selection, and development via a machine learning service, according to some embodiments, is shown.
[0005] Figure 4 A logic block diagram showing prompt discovery, according to some embodiments, is shown.
[0006] Figure 5A An example prompt discovery interface for submitting a discovery request, according to some embodiments, is shown.
[0007] Figure 5B An example prompt discovery interface for providing prompt recommendations, according to some embodiments, is shown.
[0008] Figure 6 A logic block diagram showing prompt development for NLP tasks, according to some embodiments, is shown.
[0009] Figure 7A An example prompt development interface for performing prompt and NLP ML model tuning, according to some embodiments, is shown.
[0010] Figure 7B An example prompt development interface for adapting job results, according to some embodiments, is shown.
[0011] Figure 8 A logic block diagram showing model deployment for adjusting an NLP ML model, according to some embodiments, is shown.
[0012] Figure 9is a high - level flowchart showing various methods and techniques for generating prompt recommendations for NLP processing tasks according to some embodiments.
[0013] Figure 10 is a high - level flowchart showing various methods and techniques for performing prompt development and tuning using selected prompts from a set of tasks according to some embodiments.
[0014] Figure 11 shows an example system implementing the various methods, techniques, and systems described herein according to some embodiments.
[0015] Although embodiments are described herein by way of example with respect to several embodiments and illustrative figures, those skilled in the art will recognize that the embodiments are not limited to the described embodiments or figures. It should be understood that the figures and the detailed description thereof are not intended to limit the embodiments to the particular forms disclosed, but rather, are intended to cover all modifications, equivalents, and alternatives falling within the spirit and scope of the appended claims. The headings used herein are for organizational purposes only and are not intended to limit the scope of the specification or the claims. As used throughout this application, the word "may" is used in an allowable sense (i.e., meaning "it is possible") rather than in a mandatory sense (i.e., meaning "must"). Similarly, the words "include, including, and includes" mean including but not limited to.
[0016] It should also be understood that although terms such as first and second may be used herein to describe various elements, these elements should not be limited by these terms. These terms are only used to distinguish one element from another. For example, without departing from the scope of the present invention, a first contact may be referred to as a second contact, and similarly, a second contact may be referred to as a first contact. The first contact and the second contact are both contacts, but they are not the same contact. Detailed Description
[0017] Various techniques for generating prompt recommendations and performing prompt development for natural language processing (NLP) machine learning (ML) models are described herein. When used in different applications, machine learning models may be shaped by many different factors. In various scenarios, understanding why a machine learning model makes a decision (e.g., predicts or infers) and what influences that prediction may be important for determining whether the model is performing a task correctly in an application that uses the machine learning model. For example, NLP may use an ML model to extract deep insights about the content of a document. Such NLP models develop deep insights by identifying entities, key phrases, language, sentiment, events, and other common elements in the document. Both pre - trained NLP ML models and custom NLP ML models can be used for entity recognition, classification, or various other natural language processing tasks.
[0018] Machine learning refers to the discipline of training computer systems to recognize patterns by repeated exposure to training data. In unsupervised learning, self-organizing algorithms learn previously unknown patterns in a dataset without any labels being provided. In supervised learning, this training data contains inputs (automatically or by human annotators) labeled with "ground truth" outputs corresponding to the inputs. To evaluate / validate the performance of the trained model, a portion of the training dataset is typically kept outside of the training process. The use of the trained model in production is commonly referred to as "inference" or "prediction", during which the model receives new data not in its training dataset and provides an output based on its learned parameters. The training and validation process can be repeated periodically or intermittently by optimizing the previously learned parameters of the production model with new training data and deploying a new production model for inference, in order to mitigate the degradation of model accuracy over time. For example, a training dataset of documents or other natural language (e.g., human language) sets can be used to train an NLP ML model to perform various natural language processing tasks, including but not limited to information extraction (e.g., named entity recognition, relation extraction, coreference resolution, event extraction, and joint entity relation extraction), text classification (e.g., classification, sentiment, relation classification, topic classification, paraphrase recognition, word sense disambiguation, and natural language inference), question answering (e.g., extractive QA and closed-book QA), summarization (e.g., extractive summarization and abstractive summarization), generation (e.g., sentence completion and structure-to-text), etc.
[0019] In various embodiments, the NLP ML model can use a prompt as an input to generate an inference. A prompt can be a sentence, phrase, or other sequence of words, symbols, or other parts of text in natural language. For example, an NLP ML model can be trained to answer questions about an input text. A prompt for generating an inference about the input text can be "The names mentioned in this text are". Based on the prompt and the input text, the NLP ML model can generate an inference of the names "John, Jane, Bill, Bob, and Alice" for the input text. The use of an NLP ML model to perform a task (such as the question answering task above) can be integrated into various systems, services, or applications in order to enhance the ability of such systems, services, or applications to interact better with human users. Therefore, there is a great need for technologies that can produce better NLP ML models for integration with these systems, services, or applications.
[0020] Different NLP ML models can achieve different results for NLP tasks. For example, some NLP ML models may produce better results for text summarization because these NLP ML models may have been specifically trained or developed for that task. Other NLP ML models can be generalized to perform multiple NLP tasks. In addition to the training data and design of the NLP ML models, the impact of prompts on NLP ML model performance can vary with the performance of the prompts used for a particular NLP task and across different NLP ML models.
[0021] For example, with respect to a particular NLP task and / or NLP ML model, some prompts may provide better (e.g., more accurate) inferences than other prompts. In some scenarios, developers may not have been involved in the initial training or design of the NLP ML model and thus may not have insights into the performance of specific prompts for the NLP ML model. For example, in the above example, the prompts "What are the names in the text?" or "List the names in the text" may provide better inferences given the same input text. Given the large number of possible variations of NLP task prompts, it may be difficult to discover which prompt is optimal for developer use. Thus, in various embodiments, various techniques for discovering prompts for NLP tasks can be implemented, which can take into account the use of the NLP ML model and prompts for a particular application.
[0022] Additionally, even if an optimal prompt can be identified (or can be known or otherwise recommended, e.g., by the NLP ML model designer), other aspects of the NLP ML model task performance may affect the performance of the prompt and NLP ML model for a particular application. Various techniques can be implemented for tuning both the prompt and the NLP ML model and for controlling the performance of various other features of the prompt and NLP ML model. Thus, there can be various possible ways to develop NLP ML models and prompts for a particular application. In various embodiments, various techniques for prompt-based development of NLP ML models can be implemented to allow for a full set of development tools that treat both the NLP ML model and the prompt as inputs to the NP ML model.
[0023] Figure 1 A logical block diagram of a prompt development system for a natural language processing machine learning model for generating prompt recommendations and prompt development according to some embodiments is shown. The prompt development system 110 for the NLP ML model can be a stand-alone system, application, or service, or can be implemented as a machine learning service of a provider network (e.g., as described below with respect to Figures 2 - 8Part of the machine learning service 210 being discussed. In various embodiments, the prompt development system 110 for the NLP ML model can host, support, and collect the NLP ML model (e.g., NLP ML model 124) and task prompts (e.g., prompt 101), and the task prompts can be grouped into sets for a specific NLP ML model 124 and / or a specific NLP task (e.g., name recognition, location recognition, summarization, etc.), such as task prompt sets 122a, 122b, and 122c. In various embodiments, a prompt can contain an input for performing an NLP task. As Figure 1 shown, the prompt 101 can contain command text 102 (e.g., a question or an instruction) and target text 103 (e.g., an input document, a paragraph, a phrase, a document collection, etc.). In some embodiments, the prompt 101 can contain only the command text 102. In some embodiments, the prompt 101 can contain only the target text 103 (e.g., a document or a sentence to be completed). In some embodiments, there can be multiple features or components of a prompt. For example, the command text 102 can have an instruction feature (e.g., summarize) and a style feature (e.g., in Shakespearean style).
[0024] The NLP ML model 124 can be a pre-trained or customized NLP ML model and can be based on various different ML model architectures (e.g., a generative pre-trained transformer (GPT)-based ML model) and frameworks (e.g., PyTorch, TensorFlow, etc.). These NLP models can be trained to perform various NLP processing tasks, such as document or conversation summarization, paraphrasing, structure-to-text, relation extraction, and / or coreference resolution, etc.
[0025] In some embodiments, the prompt development system 110 for the NLP ML model can provide an interface for submitting or uploading prompts and NLP models (e.g., as discussed below with respect to Figure 3 ), as shown at 110. For example, NLP ML model developers can upload various versions of the NLP ML model 124 for selection and development. Similarly, various task prompts can be provided and associated with the task sets of the NLP ML model and / or the NLP task. In this way, the prompt development system 110 for the NLP ML model can act as a development repository that makes different models and prompts available for analysis, optimization, or other development purposes by model producers and prompt producers for NLP ML model developers who integrate the NLP ML model and the prompt into an application.
[0026] As discussed above, techniques for discovering prompts for different NLP tasks can help optimize the performance of NLP tasks for integration with specific applications. Prompt discovery 130 for NLP tasks can be implemented as part of a prompt development system 110 for an NLP ML model to provide various searches and recommendations for prompts based on various analysis and search techniques, as described below with respect to Figures 4 - 5B and Figure 9 in detail. For example, various discovery interactions 132 that include guided, interactive, or other types of requests through various types of interfaces can allow submission of information to discover prompts for NLP tasks and to optimize, evaluate, and select from among many available task prompts 122.
[0027] For prompt-based NLP ML model development, integrated prompt model development 140 for NLP tasks can be incorporated as part of a prompt development system 110 for an NLP ML model to provide development tools for optimizing or otherwise developing both task prompts and NLP ML models. Task prompts can be selected through prompt discovery 130 or can be directly selected from available task prompts and models 120. Development interactions 142, described below with respect to Figures 6 - 7B and Figure 10 in detail, can allow implementation and evaluation of different configurations, parameters, or other characteristics of both prompts and NLP ML models. In this way, an optimal prompt-NLP ML model combination can be developed and deployed in various applications 150, as indicated at 114, where the applications can invoke prompts 152 as part of performing application operations using NLP ML models 154, as described below with respect to Figure 8 in detail.
[0028] Note that the previous description is a logical illustration of NLP ML models, task prompts, discovery interactions, and development interactions and should not be construed as a limitation on machine learning systems.
[0029] This specification continues with a general description of a provider network that implements multiple different services, including machine learning services that can implement prompt discovery and development for natural language processing tasks. Various examples of components (including different components) or arrangements of components that can implement these are then described. Then, many different methods and techniques for implementing prompt discovery and development for natural language processing tasks are described, some of which are illustrated in the accompanying flowcharts. Finally, a description of an example computing system on which various components, systems, devices, and / or nodes can be implemented is provided. Various examples are provided throughout the specification.
[0030] Figure 2Shows an example provider network that can implement machine learning services, where the machine learning services implement prompt discovery and development for natural language processing tasks. In one embodiment, the provider network 200 can be a dedicated or closed system, or can be established by an entity such as a company or a public sector organization to provide one or more services (such as various types of cloud-based storage) accessible via the Internet and / or other networks to clients 250. In one embodiment, the provider network 200 can be implemented in a single location, or can include a number of data centers hosting various resource pools such as a collection of physical and / or virtualized computer servers, storage devices, networking equipment, etc. (e.g., the computing system 2000 described below with reference to Figure 11 ), and these data centers are required to implement and distribute the infrastructure and services provided by the provider network 200. In some embodiments, the provider network 200 can implement various computing resources or services. In some embodiments, for example, machine learning services 210, storage services 230, and / or any other type of network-based service 240 (which can include virtual computing services and various other types of storage, databases, or data processing, analysis, communication, event processing, visualization, data cataloging, data ingestion (e.g., ETL), and security services).
[0031] In various embodiments, Figure 2 the components shown in can be implemented directly as instructions executable by computer hardware (e.g., a microprocessor or a computer system) or directly within computer hardware using a combination of these techniques. For example, in one embodiment, Figure 2 the components of can be implemented by a system including several computing nodes (or simply referred to as nodes), and each computing node can be similar to Figure 11 the computer system embodiment shown and described below. In various embodiments, the functionality of a given system or service component (e.g., the components of the machine learning service 210) can be implemented by a specific node or can be distributed across several nodes. In some embodiments, a given node can implement the functionality of more than one service system component (e.g., more than one data storage component).
[0032] Machine learning 210 may implement an interface 211 to allow a client (e.g., client 250 or a client implemented internally within provider network 200, e.g., a client application hosted on another provider network service such as an event-driven code execution service or a virtual computing service) to adapt and deploy a machine learning model (e.g., an NLP ML model). For example, machine learning service 210 may implement interface 211 (e.g., a graphical user interface, a programming interface implementing an application programming interface (API) and / or a command line interface) such that a client may submit, edit, or otherwise provide an adaptation job for a machine learning model stored in a storage service. For example, interface 211 may include a Prompt - NLP ML model development and management environment 213, which may provide a visual editor, an adaptation script, or other code editors with various development tools to create, submit, and / or otherwise interact with prompt discovery 224, prompt development 226, and model deployment 228, as discussed below. In some embodiments, development and management environment 213 may be a graphical interface, and in some embodiments, an interface to past results generated for other models may be provided. In various embodiments, interface 211 may allow a client to request the execution of adaptation, deployment, or other machine learning service features. Although not shown, interface 211 may include interfaces for other machine learning service 210 features not shown, such as the adaptation, deployment, or use of non-NLP ML models.
[0033] Machine learning service 210 may implement a control plane 212 to perform various control operations to implement the features of machine learning service 210. For example, the control plane may monitor the health and performance of requests at different components (e.g., which may be managed by prompt ingestion 222, prompt discovery 224, prompt development 225, and model deployment 228), such as model adaptation on adaptation host 214, model deployment on model host 215, and model analysis on analysis host 218. These hosts 214, 215, and 218 may implement various machine learning frameworks (e.g., Tensorflow, Pytorch, MxNet, etc.), and may utilize general-purpose processing (e.g., CPU) and / or machine learning - specific hardware (e.g., GPU or tensor processing unit (TPU)) to perform tasks. If a host fails, a request fails, or other interruptions occur, control plane 212 is able to restart a job or other process to complete the request (e.g., rather than sending a failure response to the client). In some embodiments, control plane 212 may arbitrate, balance, select, or dispatch requests to different hosts in various embodiments. For example, control plane 212 may receive request interface 211, which may be a programming interface, and identify available hosts to begin processing the request.
[0034] The data storage service 230 can implement different types of data storage areas for storing, accessing, and managing data on behalf of a client 250 as a network-based service, and the network-based service enables the client 250 to operate a data storage system in a cloud or network computing environment. In some embodiments, the data storage service 230 can also include various types of relational or non-relational databases. In some embodiments, the data storage service 230 can include an object or file data storage area for placing, updating, and obtaining data objects or files. For example, a data storage service 230 can be an object-based data storage area that allows different data objects of different formats or types to be stored and managed according to a key value or other unique identifier that identifies the object, such as structured data (e.g., database data stored in different database schemas), unstructured data (e.g., different types of documents or media content), or semi-structured data (e.g., different log files, human-readable data in different formats such as JavaScript Object Notation (JSON) or Extensible Markup Language (XML)). In at least some embodiments, the data storage service 230 can be regarded as a data lake. For example, an organization can generate many different kinds of data, which are stored in one or more collections of data objects in the data storage service 230. The data objects in the collection can include related or homogeneous data objects, such as a database partition of sales data, and unrelated or heterogeneous data objects, such as image data files (e.g., digital photos or video files), audio files, and website log files. The data storage service 230 can be accessed via a programming interface (e.g., an API) or a graphical user interface.
[0035] Generally speaking, the client 250 can cover any type of client that can submit network-based requests to the provider network 200 via the network 260, and the requests include requests for the machine learning service 210 (e.g., requests for interacting with the development and management environment 213, etc.). For example, a given client 250 can include a suitable version of a web browser, or can include a plug-in module or other type of code module that can be an extension of the execution environment provided by the web browser or execute within the execution environment. In some embodiments, such applications can include sufficient protocol support (e.g., for a suitable version of the Hypertext Transfer Protocol (HTTP)) for generating and processing network-based service requests without having to implement comprehensive browser support for all types of network-based data. That is to say, the client 250 can be an application that can directly interact with the provider network 200. In some embodiments, the client 250 can generate network-based service requests according to a Representational State Transfer (REST)-style network-based service architecture, a document- or message-based network-based service architecture, or another suitable network-based service architecture.
[0036] In some embodiments, the client 250 may provide access to the provider network 200 to those applications in a manner that is transparent to other applications. In one embodiment, the client 250 may transmit a network-based service request (e.g., an access request for configuring or performing an interpretive job) via the network 260. In various embodiments, the network 260 may encompass any suitable combination of networking hardware and protocols necessary to establish network-based communication between the client 250 and the provider network 200. For example, the network 260 may generally encompass the various telecommunications networks and service providers that jointly implement the Internet. In one embodiment, the network 260 may also include a private network, such as a local area network (LAN) or a wide area network (WAN), as well as a public or private wireless network. For example, a given client 250 and provider network 200 may each be provisioned within an enterprise that has its own internal network. In such embodiments, the network 260 may include the hardware (e.g., modems, routers, switches, load balancers, proxy servers, etc.) and software (e.g., protocol stacks, accounting software, firewall / security software, etc.) necessary to establish networking links between a given customer 250 and the Internet and between the Internet and the provider network 200. It should be noted that in some embodiments, the client 250 may communicate with the service provider network 200 using a private network rather than the public Internet.
[0037] Figure 3 A logical block diagram showing an interaction for submitting a prompt for discovery, selection, and development via a machine learning service is shown. The prompt ingestion 222 may support various features to allow NLP ML models and / or prompt developers (including prompt developers who develop using pre-trained NLP ML models developed by others) to submit prompts to be included in a task set. The prompt submission 310 may include the prompt 311 as a text string. The prompt submission 310 may also include various other information for the prompt 311, such as the NLP processing task 313 (e.g., indicated by the code, label, category, or natural language description of the NLP task). In some embodiments, the NLP processing task 313 may be one of a set of task categories provided by the machine learning service 210. In some embodiments, the description 315 may be included in the prompt submission 310, which may include various other information about the prompt, such as information that may be included in an entry of the prompt 311 when displayed or provided as part of prompt discovery and selection. In some embodiments, the submission 310 may include a sample output 317. For example, if the prompt is "What is the name of the speaker?", the sample output 317 may be "The name of the speaker is [first name][last name]". In some embodiments, the prompt submission 310 may include an identifier of the NLP ML model 319 that the prompt 311 is targeted at.
[0038] In at least some embodiments, the prompt ingestion 222 may implement a prompt submission verification 320. The prompt submission verification 320 may implement one or more prompt verification tests 320. Some verification tests 320 may check for a valid format (e.g., valid characters, grammar, etc.), duplicates (e.g., whether this prompt has been submitted), and inappropriate content (e.g., offensive content, content violating the terms of service, etc.). In some embodiments, the prompt verification test 352 may include a performance evaluation. For example, one (or more) analysis hosts 330 may be used to host the NLP ML model 340 (e.g., the identified model 319 or the model identified by the machine learning service 210). Using test data for NLP processing tasks, such as task 313, may use the prompt 311 for inference and obtain the inference for verification, as indicated at 353. These inferences may then be compared with the sample output 317 and the ground truth labels of the test data to determine whether the required sample output 317 of the prompt has been achieved. Such tests may be performed over the entire test set, and a threshold metric (e.g., accuracy) may be applied to accept or reject the prompt submission 310. The prompt submission confirmation 370 may indicate whether the prompt 311 is accepted or rejected. If rejected, the prompt submission 370 may identify the verification tests that the prompt 311 failed to meet.
[0039] The prompt may be stored in or associated with a prompt task set 360 in the storage service 230. The prompt registration 354 may be implemented as part of the prompt ingestion 222 to maintain and update the prompt task set 360 for accepting prompt submissions. For example, the prompt registration 354 may update various data structures or systems accessing the prompt task set 360. In one such example, the search index of the prompt task set may be updated to identify the prompt 311. The search index may be, for example, an inverted index that lists the unique words occurring in the set. In other embodiments, various other search indexes or other data structures may be implemented to make the prompt available. In some embodiments, an entry for the prompt may be created that provides the information included in the submission, such as the description 315, task 313, sample output 317, and NLP ML model 319, and then the entry may be displayed or otherwise provided via the interface 211 (e.g., as part of a discovery or development interaction via the interface 213). As indicated at 355, the prompt ingestion may add the prompt to one or more task sets 360 (e.g., add to the task set for NLP tasks and the task set for a particular NLP ML model).
[0040] Figure 4A logic block diagram showing prompt discovery according to some embodiments is shown. Prompt discovery 224 may provide an interactive prompt interface for processing discovery requests (e.g., discovery request 400). The discovery request 400 may be received via a graphical interface (e.g., a web console or other user interface implemented as part of the Prompt-NLP ML model development and management 213). In some embodiments, a command line or programming interface may also be supported for discovery requests. The discovery request 400 may include various features, such as a description 411. The description 411 may be keywords, terms, or other search criteria that can be used to discover prompts. In some embodiments, the discovery request 400 may include sample input / output 412. For example, the sample input may be a sample document, file, or other text that can be analyzed using an NLP ML model, and the sample output may be an example of the expected result (e.g., the most frequent sentiment of the reviewers on a post is "[sentiment]"). In some embodiments, the discovery request may include performance criteria 415. The performance criteria 415 may include information such as a time range or limit for returning an inferred response, a resource range (or limit) for hosting the NLP ML model to generate the inference, the expected size of the input, etc. Although in some embodiments, the discovery request 400 may not be NLP model specific, in other embodiments, the NLP ML model 417 may be specified (e.g., by name or identifier). Similarly, although in some embodiments, the discovery request 400 may not include an NLP task, in other embodiments, the NLP task may be explicitly specified as 417 (e.g., by the name or selection of the supported / available tasks).
[0041] In some embodiments, prompt discovery 224 may perform NLP task classification. The prompt task classification 410 may perform rule-based (e.g., using heuristics, rule sets, decision trees, etc.) task classification or other techniques, such as an ML task classification model, to determine the NLP task for the discovery request 400. For example, for rule-based task classification, the NLP task classification may parse the description 411, sample input / output 413, performance criteria 415 (and / or other features of the discovery request 400) to determine one (or more) applicable classification rules to apply. For example, keyword searches or comparisons of the description 411 or sample output 413 may be used to identify a set of task classification rules (e.g., action words such as "list", "summarize", "explain", "identify", etc., entity words such as "person", "company", "country", "organization"). In some embodiments, an ML model may be trained to generate an encoding of the sample output 412 and return an inference identifying the NLP task classification. In some embodiments, the NLP task classification 410 may identify an explicitly specified task 419.
[0042] In various embodiments, the hint discovery 224 may implement hint and NLP ML candidate selection 420. The hint and NLP ML candidate selection 420 may utilize NLP task classification and other information from the discovery request 400 to select candidate hints and (when not specifically requested) candidate NLP models for consideration. For example, different hint task sets 421 may be utilized to organize or identify hints. Each hint task set 421 may correspond to a task classification. In some embodiments, each hint task set 421 may also correspond to one (or more) NLP ML models. For example, a text summarization task hint set may correspond to a text summarization NLP task. Some hints in the text summarization task may correspond to NLP ML model A, and other hints in the same task set may correspond to NLP ML model B. In some embodiments, access control may be implemented for portions of the hints (or for the entire task set). For example, in some embodiments, account or organizational access restrictions to hint sets may be implemented to limit access to developers associated with the account or organization.
[0043] Candidate hint selection may be implemented in various ways. For example, a text search based on keywords identified in the description 411 or the sample output 413 may be used to search a hint index in the task set. In some embodiments, hints in the task set may have sample outputs. A comparison of the sample output 413 with the sample outputs of items in the hint may be used. Sample inputs may be used to perform similar techniques (e.g., comparing the sample input text with the sample input text used for the hint). In some embodiments, a combination of the various comparisons (or other analyses) discussed above may be used. In some embodiments, a similarity score may be generated for hints in the hint task set 421, and the similarity score may be used to determine which (if any) hint candidates to select. For example, if 15 hints are scored, the sorting of the hints by score may be used to select the top 3 hints as candidate hints. In some embodiments, the score may be compared to a minimum score threshold (e.g., a candidate must have a similarity score greater than 0.7 (in the range of 0 to 1)). In some embodiments, candidate hints may be selected and then further filtered by other considerations such as, for example, a candidate NLP ML model (or a specified model) or performance criteria (e.g., if fast hint performance is desired, hints with slow performance may be discarded). In some embodiments, the ML model that generates the similarity score may be used to compare based on the information received in the discovery request and the hints in the hint task set 421.
[0044] In some embodiments, an NLP ML model may be specified (as indicated at 417). In such cases, the candidate NLP ML model can be the specified model. However, in some embodiments, the prompt and NLP ML candidate selection 420 can select one or more NLP ML models from the available NLP ML models 423. For example, metadata or other information describing the NLP ML models can be maintained, which indicates which NLP tasks the NLP ML model supports. Such metadata can be evaluated to determine that the NLP ML model performs the NLP tasks identified for the discovery request 410. In some embodiments, performance criteria 415 can be used to identify candidate NLP ML models (e.g., based on performance information maintained for each NLP ML model, such as in the metadata). In some embodiments, whether the NLP ML model 423 is available for tuning or other customization can also be a consideration for selection as a candidate if, for example, the performance criteria 415 indicate that the NLP ML model may need tuning.
[0045] In various embodiments, the prompt discovery 224 can implement the prompt and NLP ML evaluation 430 to evaluate the performance of candidate prompts and candidate NLP ML models for recommendation. For example, one or more analysis hosts 431 can be provisioned, and candidate NLP NL models 433 can be deployed on the analysis hosts. Test data 435 can be obtained. The test data 435 can be specified or identified as part of the discovery request 410 (e.g., by identifying a file, storage location, or other path for accessing the test data 435, or the test data 435 can be the sample input 413). In some embodiments, the test data 435 can be maintained by the machine learning service 210 for evaluating the identified NLP tasks. When the test data 435 is obtained, the candidate prompt 432 can be used to generate inferences on the test data 435 using the candidate NLP ML model 433. The results 434 of the candidate prompt can be collected. In some embodiments, the prompt and NLP ML evaluation 430 can perform an initial analysis by, for example, comparing the candidate prompt results 434 with the sample output 413. Also, in some embodiments, a similarity score with the sample output 413 can be generated to determine how well the candidate prompt and candidate NLP ML model perform, and the similarity score can be used to rank or filter the candidate prompts and models. In some embodiments, the test data sample output can be obtained for inclusion in the prompt recommendation 450 (e.g., as indicated at the sample output 453). Other performance metrics can be used for evaluation (e.g., when the labeled data or ground truth data of the test data 435 is used to score the accuracy of the combination of the candidate prompt and candidate NLP ML model).
[0046] In some embodiments, the prompt discovery 224 may implement a prompt recommendation generation 440. The prompt recommendation generation 440 may provide a prompt recommendation 450 that includes a plurality of candidate prompts 451 and an NLP ML model 457 or a single recommended prompt (or, for example, some combination between one prompt and multiple NLP ML models or multiple prompts and one NLP ML model). In some embodiments, the prompt recommendation may include different combinations of prompt features (e.g., one instruction feature with different possible style features). In some embodiments, the prompt recommendation may include context learning prompts to be considered for evaluation, as discussed in detail below with respect to Figure 6 discussed. In some embodiments, the prompt recommendation generation 440 may filter or reduce the number of candidates based on the analysis discussed above. The prompt recommendation generation 440 may generate a prompt recommendation 450 and include various information, such as a prompt 451, a sample output 453 of the prompt, a performance metric 455 of the prompt (e.g., inference time, resources / costs, etc.), the NLP ML model 457 used, and the identified task 459 (e.g., at 410). The prompt recommendation 450 may be presented in various ways using various visualization techniques, including but not limited to sorting, charts, graphs, output displays, etc. In some embodiments, the Figure 5B provides an example of a prompt recommendation.
[0047] Although not depicted, in some embodiments, the various stages of the prompt discovery 224 may be interactive, such that after stages such as NLP task classification 410, prompt and NLP ML candidate selection 420, and prompt and NLP ML evaluation 430, for example, partial or intermediate results may be provided via an interface. Then, the intermediate result may be optimized based on the received feedback, and the prompt discovery process may continue using the optimized intermediate result (e.g., to task classification, candidate prompts, or candidate ML models). Similarly, the prompt recommendation 450 may be used to select or modify the task 459, the NLP ML model 457, or an individual prompt 451 for re-evaluation with the modification.
[0048] Figure 5AShows an example prompt discovery interface for submitting discovery requests. In some embodiments, the prompt discovery interface 510 may be implemented as part of the Prompt-NLP ML model development / management environment 213. The prompt discovery interface 510 may implement different types of user interface elements for receiving information to be included in a prompt discovery request. For example, text input elements may be used to accept a description 512, sample input 514, sample output 516, and performance criteria 518. In some embodiments, various user interface elements are used to translate selections into features of the discovery request, such as a drop-down menu pre-filled with task classifications or a slider / knob for toggling between general information (e.g., fast / high cost, slower / low cost). A submission element 519 may submit the discovery request for evaluation.
[0049] Figure 5B Shows an example prompt discovery interface for providing prompt recommendations. The prompt discovery interface 510 may provide prompt recommendations 520 generated according to the various techniques discussed above. For example, a ranking of recommended prompts may be provided, such as prompts 531, 541, and 551, along with corresponding output information, such as task classifications 533, 543, and 553, models 535, 545, and 555, sample outputs 537, 547, and 557, and performance 539, 549, and 559. In some embodiments, the prompts may be presented in an unordered manner. Note that various other arrangements or interactions may be used to submit discovery requests and receive prompt recommendations. Thus, the examples discussed above with respect to Figure 5A and 5B are not intended to be limiting.
[0050] Figure 6 Is a logical block diagram showing prompt development for NLP tasks according to some embodiments. A development request 610 may be received at prompt development 226. The development request 610 may be received via a graphical interface (e.g., a web console or other user interface implemented as part of the Prompt-NLP ML model development and management 213). For example, the development request 610 may be received as part of an integrated development environment (IDE) for developing prompts for NLP tasks. In some embodiments, a command line or programming interface may also be supported for the development request 610.
[0051] The development request 610 may include various features, such as a prompt 611 (e.g., a text string or identifier designated as a selected prompt from a set of tasks) and an NLP ML model 613 (e.g., specified using a name or identifier). In this way, various tuning jobs can be created to adjust the prompt 611 and / or the NLP ML model 613. The prompt 611 can be one of many prompts supported by the machine learning service 210 selected through a discovery request that browses or identifies prompts for NLP tasks in the set of tasks. In some embodiments, additional information or features can be designated as part of the development request 610. For example, a tuning configuration 615 can be included, which can specify various tuning job features, such as hyperparameters for performing the adjustment, stop criteria, or resource utilization limits, etc. The development request 610 can include identifiers for tuning a dataset 617 (e.g., an identifier, path, location, or upload of a tuned dataset) and a test dataset 619 (e.g., an identifier, path, location, or upload of a test dataset). The test dataset 619 can include ground truth labels, e.g., which can be used to evaluate performance after adjustment.
[0052] Prompt development 226 can implement tuning job generation 620. The tuning jobs can include different types of tasks. Some types of tuning jobs can use ML training techniques to change the weights or other features of a pre-trained NLP ML model to adjust the model performance. Another type of tuning job can include techniques for changing the performance of a pre-trained NLP ML model without changing the NLP model itself, such as in-context learning. The tuning job generation 620 can generate a configuration file or other artifacts for executing the development request 610. For example, the tuning job generation 620 can verify the correctness of the various features of the development request 610 and generate a configuration file for executing the adaptation job according to an appropriate machine learning framework by integrating or invoking various libraries, files, and other information required for the tuning job to tune the pre-trained NLP ML model and / or perform prompt adjustment.
[0053] In various embodiments, the tuning job management 630 can coordinate the execution of tuning jobs to implement different types of adjustments, including prompt adjustment 632, in-context learning 634, and fine-tuning 636. The prompt adjustment 632 can include various adjustment techniques that freeze the core language model parameters of the NLP ML model and instead learn the soft token sequences for each task / domain. In some scenarios, the prompt adjustment can allow the machine learning model to learn task-specific prompt tokens that can match the performance of model fine-tuning (with the underlying model frozen). The prompt adjustment 632 can also be applied at the level of various domains within a given task. For example, learning domain-specific prompts can match the performance of fine-tuning in low-data settings. By freezing the core language model parameters, the prompt adjustment 632 appears to be more resilient to domain shift.
[0054] Another type of adjustment that the adaptation job management 630 can support is fine-tuning 636. Fine-tuning 636 can be an adaptation process that utilizes task-specific data and updates the parameters (usually all or a part of them) of a pre-trained NLP ML model based on the task-annotated data. In some scenarios, a decoding algorithm can be implemented for the input and the output representation of the adopted model to perform the task.
[0055] Another type of adjustment that the adaptation job management 630 can support is in-context learning 634. In-context learning 634 can include the technique of using the text input of a pre-trained language model as a form of task specification: the pre-trained NLP ML model is conditioned on natural language instructions and / or some demonstrations of the task, and then is expected to complete other instances of the task only by predicting what comes next. During unsupervised pre-training, the pre-trained NLP ML model develops general pattern recognition capabilities (e.g., meta-learning). Then, the model uses these capabilities at inference time to quickly adapt to or recognize the desired task. In the absence of prior information about the task itself, for example, if the following is provided to the model during inference:
[0056] Document: John Doe works for ABC
[0057] Augmented document: [John Doe|person] works for [ABC|organization]
[0058] Document: The EU rejected Germany's call to boycott British lamb.
[0059] Augmented document: [The EU|organization] rejected Germany's call to boycott British lamb.
[0060] After in-context learning, the NLP ML model can be adjusted to handle another example:
[0061] Document: Commission chief spokesman Nikolaus van der Pas said at a press conference
[0062] Augmented document: [Commission|organization] chief spokesman [Nikolaus van der Pas|person]
[0063] For in-context learning 634, learning about the task occurs within the forward-pass of each sequence, and no parameter updates are performed. This is in contrast to fine-tuning 636, in which paired source sequences and target sequences are provided, and the NLP ML model learns to generate the target sequence through parameter updates. Therefore, for the adaptation job 631 being performed, the in-context learning evaluation 660 and the analysis host 670 can be the same (because in-context learning can be performed and tested using in-context learning prompts before evaluating the prompts).
[0064] The adaptation job management 630 can support these adaptation techniques by managing or coordinating the execution of the generated adaptation jobs to execute these adaptation techniques on the adaptation host 660. A pre-trained NLP ML model 661 can be obtained and deployed to the adaptation host 660 along with the adaptation dataset 662. The adaptation host 660 can execute the adaptation techniques identified in the adaptation job 631 and then provide the adapted model 633 (or an indication of the generated adapted model) to the adaptation job management.
[0065] The prompting and model evaluation 640 can then use the test data 672 to perform an evaluation using the adapted NLP ML model 671 and the prompt 611 on the analysis host 670. When the test data 672 is obtained, the prompt 611 can be used to generate inferences on the test data 672 using the adapted NLP ML model 671. The results 643 of the adapted NLP model using the prompt can be collected. In some embodiments, the prompting and NLP ML evaluation 640 can perform an analysis by, for example, comparing the results obtained from multiple different adaptation jobs (e.g., one performing prompt adaptation, one performing context learning, and one performing fine-tuning). In some embodiments, a test data sample output can be obtained for inclusion in the adaptation job results 650 (e.g., as indicated at sample output 655). Other performance metrics can be used for evaluation (e.g., when the labeled data or ground truth data of the test data 672 is used to score the accuracy of the prompt and the adapted NLP ML model).
[0066] In various embodiments, the adaptation job results 650 can be provided. The adaptation job results 650 can include inference performance information, such as accuracy information. The adaptation job results 650 can also include computational performance 653. For example, it can include adaptation time, adaptation resource utilization, inference execution time, and inference / adapted NLP ML model resource utilization. In some embodiments, the adaptation job results 650 can include a sample output 655.
[0067] Figure 7A is an example prompt development interface for performing prompt and NLP ML model adaptation according to some embodiments. In some embodiments, the prompt development interface 710 can be implemented as part of the prompt-NLP ML model development / management environment 213. The prompt development interface 710 can implement different types of user interface elements for receiving information to be included in a prompt NLP ML model adaptation or other development requests, as indicated at 720. For example, a prompt 721 (e.g., selected from a set of prompt tasks) and a model 723 (e.g., selected from the supported models) can be identified. In some embodiments, adaptation data 722, adaptation configuration 724, test data 726, and test configuration 728 can also be specified and submitted as indicated at 719.
[0068] Figure 7B is an example prompt development interface for adapting job results according to some embodiments. The tuned job results 730 may be displayed as part of the prompt development interface 710. For example, computational performance metrics 732 and inference performance 734 (e.g., accuracy) may be provided. Various different techniques may be implemented for displaying or visualizing the tuned job results 730. In some embodiments, the performance of different tuning techniques, prompts, and NLP ML models may be compared in a summary or meta-analysis view. Note that various other arrangements or interactions may be used to submit development requests and receive tuned job results. Thus, the examples described above with respect to Figure 7A and 7B are not intended to be limiting.
[0069] Figure 8 is a logical block diagram showing model deployment for tuning an NLP ML model according to some embodiments. The model deployment 228 may manage and execute requests to deploy an NLP ML model discovered and / or tuned using the various techniques discussed above. The model deployment 228 may process a request, such as a deployment request 810, which may identify an NLP ML model 812 for an NLP model and a host configuration 814. For example, an identifier of the NLP ML model 812 may be included such that the model deployment 228 may obtain a model artifact 832 from a storage service 230 that stores the tuned NLP ML model 820 (e.g., after tuning). The host configuration 814 may provide resource allocation or other information (e.g., the amount of processing capacity, network capacity, memory capacity, etc.) to be available on the model host.
[0070] The model deployment 228 may provision a model host 840 according to the host configuration 814 and configure networking to access the model host 840. For example, a model endpoint 852 may be provisioned, which is a publicly accessible network endpoint to which an application 860, for example, may direct requests such as an inference request 861 with a prompt. The model deployment 228 may confirm the completion of the deployment 850 and include the model endpoint 852. Then, the model host 840 may return an inference 863 to the application 860 in response to a request.
[0071] Although described and shown in the context of a provider network implementing machine learning services, Figures 2 - 8 but Figures 2 - 8 the various components shown and described in Figures 2 - 8 can be readily applied to other machine learning systems performing NLP tasks. For example, an NLP service may utilize common NLP ML models but still provide prompt discovery and prompt development features to further optimize performance. Thus,
[0072] Figure 9is a high-level flowchart showing various methods and techniques for generating prompt recommendations for NLP processing tasks. As indicated at 910, according to some embodiments, a request to determine a prompt for an NLP task performed by a pre-trained NLP ML model may be received at a prompt development system. For example, the request may be received as part of a discovery or search interface of the prompt development system, which may provide many different prompts for different NLP tasks. The request may contain various information, such as a description (e.g., keywords, terms, or other search criteria that can be used to discover a prompt), sample input and sample output, performance criteria (e.g., a time range or limit for returning an inferred response, a resource range (or limit) for hosting the NLP ML model to generate an inference, or an expected size of the input) and / or the NLP ML model.
[0073] As indicated at 920, in some embodiments, the task classification of the NLP task may be determined by the prompt development system at least in part based on the request. For example, rule-based techniques for task classification (e.g., using heuristics, rule sets, decision trees, etc.) may be used to determine the NLP task by parsing the description, sample input / output, and / or performance criteria to identify one (or more) applicable classification rules to apply based on keywords found in the parsed information. Keyword search or comparison may be used to identify a task classification rule set (e.g., action words such as "list", "summarize", "explain", "identify", etc., entity words such as "person", "company", "country", "organization"). In some embodiments, ML techniques may be used. For example, an ML model may be trained to generate an encoding of the sample output (or other features received in the request) and return an inference identifying the NLP task classification. In some embodiments, the request may explicitly identify the task category (or a family of related task categories). In some embodiments, the task classification technique may be an optimization of the prompt recommendation technique, and thus in other embodiments, the task classification determination may not be performed (or optionally performed as an option specified in the request), as indicated by the dashed line.
[0074] As indicated at 930, in some embodiments, one or more candidate prompts may be selected from a set of prompt tasks maintained by a prompt development system for an NLP task. For example, keywords may be used to maintain and search an index of the prompts in the task set. In some embodiments, the prompts in the task set may have sample outputs. A comparison of the sample output received in the request with the sample outputs of the items in the prompt may be used. Sample inputs may be used to perform similar techniques (e.g., comparing the sample input text with the sample input text for the prompt). In some embodiments, a combination of the various comparisons (or other analyses) discussed above may be used. In some embodiments, a similarity score may be generated for the prompts in the prompt task set, and the similarity score may be used to determine which (if any) prompt candidates to select. In some embodiments, candidate prompts may be selected and then further filtered by other considerations such as, for example, candidate NLP ML models (or designated models) or performance criteria (e.g., if fast prompt performance is desired, prompts with slow performance may be discarded). In some embodiments, the ML model that generates the similarity score may be used to compare based on the information received in the discovery request and the prompts in the prompt task set. In some embodiments, similar techniques may be implemented to select an NLP ML model. In other embodiments, a default NLP ML model (e.g., a basic pre-trained NLP ML model provided by the prompt development system) may be used, or an NLP ML model specified in the request or corresponding to the task classification may be identified.
[0075] As indicated at 940, in some embodiments, the corresponding prompt results generated using one or more candidate prompts by a pre-trained NLP model may be evaluated. For example, test data for the identified NLP task may be obtained. The candidate prompts may be used to generate inferences on the test data using the pre-trained NLP ML model. Inferences with NLP task results for the candidate prompts may be collected. In some embodiments, the candidate prompt results may be compared with the provided sample output. Similarly, a similarity score with the sample output may be generated to determine how well the candidate prompts perform. In some embodiments, the similarity score may be used to rank or filter the candidate prompts.
[0076] As indicated at 950, in some embodiments, an NLP task prompt recommendation may be returned at least in part based on an evaluation of the corresponding prompt results. For example, a prompt recommendation may be returned that identifies the prompt (or the best-performing prompt), the sample output of the prompt, the identified NLP task classification, and / or various other information about the recommended prompt. As discussed above, in some embodiments, the prompt recommendation may include one or more prompts, as well as different types or features of the prompts, such as instruction features, style features, or in-context learning prompts.
[0077] Figure 101010 is a high-level flow chart illustrating various methods and techniques for performing prompt development and tuning using selected prompts from a set of tasks in accordance with some embodiments. As indicated at 1010, in some embodiments, a request may be received at a prompt development system to perform further tuning of a pre-trained NLP ML model that performs an NLP task according to an input prompt selected from a set of prompt tasks maintained by the prompt development system. For example, the request may be a prompt selected to provide an output to be used in a software application (e.g., integrating a text summarization feature into a cataloging system or indexing application for documents). However, the performance of the pre-trained NLP ML model may be tuned to provide a higher quality of performance relative to a specific use case (e.g., an indexing or cataloging system) for using the prompts and model. Thus, in some embodiments, the request may be for different prompts to be tried and in conjunction with an NLP ML model tuned for a specific application.
[0078] As indicated at 1020, in some embodiments, an adaptation job may be generated to perform the requested further adaptation using the adaptation dataset specified by the request. Different types of adaptations (e.g., adaptation of model changes, such as fine-tuning or prompted adjustment) or changes in model performance without changing the model itself (e.g., contextual learning) may be performed by the adaptation job. For example, a configuration file, manifest, template, or other executable artifact may be created that identifies various data and software artifacts for performing the adaptation job (e.g., ML framework, ML model artifact, adaptation dataset (e.g., adjustment dataset), type of adjustment (e.g., context, fine-tuning, prompted adjustment), hyperparameters, and various other features that control the performance of the adaptation job. As indicated at 1030, in some embodiments, an adaptation job may be performed to adjust a pre-trained NLP ML model.
[0079] As indicated at 1040, in some embodiments, the performance of an adaptation job that adjusts the performance of a pre-trained NLP ML model may be evaluated based on the input prompt to generate results for the adaptation job. For example, test data may be obtained, and the prompt is used to generate inferences on the test data using the adapted NLP ML model. These inferences may then be analyzed by, for example, comparing results obtained from a plurality of different adaptation jobs (e.g., one performing prompt adjustment, one performing context learning, and one performing fine tuning). In some embodiments, a test data sample output may be obtained for inclusion in the adaptation job results. Other performance metrics may be used for evaluation (e.g., when labeled data or ground truth data for the test data is used to score the accuracy of the prompt and the adapted NLP ML model).
[0080] As indicated at 1050, in some embodiments, the results of the adaptation operation may be provided. For example, various adapted NLP ML model performance information may be provided, such as inference performance and computational performance, as discussed in detail above.
[0081] In various embodiments, the methods described herein may be implemented by any combination of hardware and software. For example, in one embodiment, the method may be implemented on one or more computer systems (e.g., Figure 11 the computer systems in) or across one or more computer systems that include one or more processors that execute program instructions stored on one or more computer-readable storage media coupled to the processors. The program instructions may implement the functions described herein (e.g., implement the functions of the various servers and other components of the network-based virtual computing resource provider described herein). The various methods shown in the figures and described herein represent example embodiments of the methods. The order of any method may be changed, and various elements may be added, reordered, combined, omitted, modified, etc.
[0082] Embodiments of the cue discovery and development techniques described herein may be executed on one or more computer systems that may interact with various other devices. Figure 11 One such computer system is shown. In different embodiments, the computer system 2000 may be any of a variety of types of devices, including but not limited to personal computer systems, desktop computers, laptop computers, notebook or netbook computers, mainframe computer systems, handheld computers, workstations, network computers, cameras, set-top boxes, mobile devices, consumer devices, video game consoles, handheld video game devices, application servers, storage devices, peripheral devices such as switches, modems, routers, or generally any type of computing device, computing node, computational node, or electronic device.
[0083] In the illustrated embodiment, computer system 2000 includes one or more processors 1010 coupled to system memory 1020 via an input / output (I / O) interface 1030. Computer system 2000 further includes a network interface 1040 coupled to the I / O interface 1030, and one or more input / output devices 1050, such as a cursor control device 1060, a keyboard 1070, and a display 1080. The display 1080 may include a standard computer monitor and / or other display systems, technologies, or devices. In at least some embodiments, the input / output device 1050 may further include a device having touch or multi-touch capabilities, such as a tablet or tablet computer, through which a user inputs input via a stylus-type device and / or one or more fingers. In some embodiments, it is contemplated that an embodiment may be implemented using a single instance of computer system 2000, while in other embodiments, multiple such systems or multiple nodes comprising computer system 2000 may host different parts or instances of the embodiment. For example, in one embodiment, some elements may be implemented via one or more nodes of computer system 2000 that are different from those hosting other elements.
[0084] In various embodiments, computer system 2000 may be a single-processor system including one processor 1010, or a multi-processor system including several processors 1010 (e.g., two, four, eight, or another suitable number of processors). The processor 1010 may be any suitable processor capable of executing instructions. For example, in various embodiments, the processor 1010 may be a general-purpose or embedded processor implementing any of various instruction set architectures (ISAs), such as the x86, PowerPC, SPARC, or MIPS ISA, or any other suitable ISA. In a multi-processor system, each of the processors 1010 may or may not collectively implement the same ISA.
[0085] In some embodiments, at least one processor 1010 can be a graphics processing unit. A graphics processing unit or GPU can be regarded as a dedicated graphics rendering device for a personal computer, workstation, game console, or other computing or electronic device. Modern GPUs can be very efficient in manipulating and displaying computer graphics, and their highly parallel structure can make them more effective than a typical CPU for a series of complex graphics algorithms. For example, the graphics processor can enable the implementation of multiple graphics primitive operations in a much faster way than directly drawing to the screen with the host central processing unit (CPU). In various embodiments, graphics rendering can be implemented at least in part by program instructions executed on one such GPU or executed in parallel on two or more such GPUs. The GPU can implement one or more application programming interfaces (APIs) that allow programmers to call the functions of the GPU. Suitable GPUs can be commercially obtained from vendors such as NVIDIA Corporation, ATI Technologies (AMD), etc.
[0086] The system memory 1020 can store program instructions and / or data accessible to the processor 1010. In various embodiments, the system memory 1020 can be implemented using any suitable memory technology, such as static random access memory (SRAM), synchronous dynamic RAM (SDRAM), non-volatile / flash-type memory, or any other type of memory. In the illustrated embodiment, the program instructions and data for implementing the desired functions, such as those program instructions and data for implementing the interpretive work of computer vision tasks described above, are shown as being stored in the system memory 1020 as program instructions 1025 and data store 1035, respectively. In other embodiments, the program instructions and / or data can be received, sent, or stored on different types of computer-accessible media or on similar media separate from the system memory 1020 or the computer system 2000. Generally, a non-transitory computer-readable storage medium can include a storage medium or memory medium, such as a magnetic medium or an optical medium, for example, a disk or a CD / DVD-ROM coupled to the computer system 2000 via the I / O interface 1030. The program instructions and data stored via the computer-readable medium can be sent via a transmission medium or signal, such as an electrical signal, an electromagnetic signal, or a digital signal, which can be conveyed via a communication medium such as a network and / or a wireless link, for example, which can be implemented via the network interface 1040.
[0087] In one embodiment, the I / O interface 1030 may coordinate I / O traffic between the processor 1010, the system memory 1020, and any peripheral devices in the device, where the peripheral devices include the network interface 1040 or other peripheral interfaces such as the input / output device 1050. In some embodiments, the I / O interface 1030 may perform any necessary protocol, timing, or other data transformations to convert data signals from one component (e.g., the system memory 1020) into a format suitable for use by another component (e.g., the processor 1010). In some embodiments, for example, the I / O interface 1030 may include support for devices attached via various types of peripheral buses, such as variants of the Peripheral Component Interconnect (PCI) bus standard or the Universal Serial Bus (USB) standard. In some embodiments, the functions of the I / O interface 1030 may be divided into two or more separate components, such as a northbridge and a southbridge. Additionally, in some embodiments, some or all of the functions of the I / O interface 1030, such as the interface to the system memory 1020, may be incorporated directly into the processor 1010.
[0088] The network interface 1040 may allow data to be exchanged between the computer system 2000 and other devices attached to the network (such as other computer systems) or between nodes of the computer system 2000. In various embodiments, the network interface 1040 may support: communication via a wired or wireless general data network (such as any suitable type of Ethernet network); communication via a telecommunications / telephone network (such as an analog voice network or a digital fiber-optic communication network); communication via a storage area network (such as a Fibre Channel SAN); or communication via any other suitable type of network and / or protocol.
[0089] In some embodiments, the input / output device 1050 may include one or more display terminals, keyboards, keypads, touchpads, scanning devices, voice or optical recognition devices, or any other device suitable for inputting or retrieving data through one or more computer systems 2000. Multiple input / output devices 1050 may be present in the computer system 2000 or may be distributed across various nodes of the computer system 2000. In some embodiments, similar input / output devices may be separate from the computer system 2000 and may interact with one or more nodes of the computer system 2000 via a wired or wireless connection, such as via the network interface 1040.
[0090] As Figure 11As shown, the memory 1020 may include: program instructions 1025 that implement the various methods and techniques described herein; and data storage 1035 that includes various data accessible by the program instructions 1025. In one embodiment, the program instructions 1025 may include the software elements of the embodiments described and shown herein. The data storage 1035 may include data that can be used in the embodiments. In other embodiments, other or different software elements and data may be included.
[0091] Those skilled in the art will appreciate that the computer system 2000 is illustrative only and is not intended to limit the scope of the techniques described herein. Specifically, computer systems and devices may include any combination of hardware or software that can perform the indicated functions, including computers, personal computer systems, desktop computers, laptop computers, notebooks or netbook computers, mainframe computer systems, handheld computers, workstations, network computers, cameras, set-top boxes, mobile devices, network devices, Internet appliances, PDAs, wireless handsets, pagers, consumer devices, video game consoles, handheld video game devices, application servers, storage devices, peripheral devices such as switches, modems, routers, etc., or generally any type of computing or electronic device. The computer system 2000 may also be connected to other devices not shown or, alternatively, may operate as a stand-alone system. Additionally, in some embodiments, the functions provided by the illustrated components may be combined in fewer components or distributed among additional components. Similarly, in some embodiments, the functions of some of the illustrated components may not be provided and / or other additional functions may be used.
[0092] Those skilled in the art will also appreciate that although various items are shown as being stored in memory or on a storage device during use, for purposes of memory management and data integrity, these items or portions thereof may be transferred between memory and other storage devices. Alternatively, in other embodiments, some or all of the software components may be executed in the memory of another device and communicate with the illustrated computer system via inter-computer communication. Some or all of the system components or data structures may also be stored (e.g., as instructions or structured data) on a computer-accessible medium or portable article for reading by an appropriate drive, various examples of which are described above. In some embodiments, instructions stored on a non-transitory computer-accessible medium separate from the computer system 2000 may be sent to the computer system 2000 via a transmission medium or signal (e.g., an electrical, electromagnetic, or digital signal conveyed via a communication medium such as a network and / or a wireless link). Various embodiments may also include receiving, sending, or storing instructions and / or data implemented in accordance with the foregoing description on a computer-accessible medium. Thus, the present invention may be practiced with other computer system configurations.
[0093] It should be noted that any distributed system embodiment described herein or any component in its components can be implemented as one or more network services. In some embodiments, a network-based service can be implemented by a software and / or hardware system designed to support interoperable machine-to-machine interactions on a network. A network-based service can have an interface described in a machine-processable format such as, for example, Web Services Description Language (WSDL). Other systems can interact with the network service in the manner specified by the description of the interface of the network-based service. For example, a network-based service can describe various operations that other systems can call, and can describe a specific application programming interface (API), and it can be expected that other systems will follow the specific API when requesting various operations.
[0094] In various embodiments, a network-based service can be requested or invoked by using a message that includes parameters and / or data associated with the network-based service request. Such a message can be formatted according to a specific markup language such as, for example, Extensible Markup Language (XML), and / or a protocol such as, for example, Simple Object Access Protocol (SOAP) can be used to encapsulate such a message. To execute a network service request, a network-based service client can use an Internet-based application layer transport protocol (such as, for example, Hypertext Transfer Protocol (HTTP)) to assemble a message that includes the request and deliver the message to an addressable endpoint corresponding to the network service (such as, for example, a Uniform Resource Locator (URL)).
[0095] In some embodiments, Representational State Transfer (“RESTful”) techniques can be used instead of message-based techniques to implement network services. For example, a network service implemented according to RESTful techniques can be invoked by parameters included within an HTTP method (such as PUT, GET, or DELETE), rather than being encapsulated within a SOAP message.
[0096] The various methods shown in the figures and described herein represent example embodiments of the methods. The methods can be implemented using software, hardware, or a combination thereof. The order of the methods can be changed, and various elements can be added, reordered, combined, omitted, modified, etc.
[0097] Various modifications and changes can be made, which will be obvious to those skilled in the art who benefit from the present disclosure. The present invention is intended to cover all such modifications and changes, and therefore, the above description should be viewed in an illustrative rather than a restrictive sense.
[0098] The embodiments of the present disclosure can be described according to the following terms:
[0099] Clause 1. A system, comprising:
[0100] At least one processor; and
[0101] A memory that stores program instructions which, when executed by the at least one processor, cause the at least one processor to implement a prompt development system for a natural language processing (NLP) machine learning (ML) model, the prompt development system being configured to:
[0102] Receive, via an interface of the prompt development system, a request for further tuning that adjusts the performance of a pre-trained NLP ML model that executes an NLP task based on an input prompt selected from a set of prompt tasks maintained by the prompt development system;
[0103] Generate an adaptation job to perform the requested further tuning of the pre-trained NLP ML model's performance using an adaptation dataset specified by the request;
[0104] Cause the tuning job to be executed;
[0105] Evaluate the performance of the tuning job to generate a result of the tuning job; and
[0106] Return the result of the tuning job via the interface of the prompt development system.
[0107] Clause 2. The system according to clause 1, wherein the tuning job performs prompt adjustment techniques to adjust the pre-trained NLP ML model.
[0108] Clause 3. The system according to any one of clauses 1-2, wherein the tuning job performs context learning techniques to adjust the pre-trained NLP ML model.
[0109] Clause 4. The system according to any one of clauses 1-3, wherein the prompt development system is a machine learning service provided by a provider network, and wherein the prompt development system is further configured to:
[0110] Receive, via the interface, a request to deploy the tuned NLP ML model;
[0111] Provision a host for the tuned NLP ML model; and
[0112] Provide, via the interface, a model endpoint for a client application to submit an inference request to the host to generate a corresponding inference using the tuned NLP ML model.
[0113] Clause 5. A method, comprising:
[0114] The prompt development system receives a request to perform further tuning, where the further tuning adjusts the performance of a pre-trained natural language processing (NLP) machine learning (ML) model that performs natural language processing (NLP) tasks based on input prompts selected from a set of prompt tasks maintained by the prompt development system;
[0115] The prompt development system generates an adaptation job to perform the requested further tuning using an adaptation data set specified by the request;
[0116] The prompt development system evaluates the performance of the tuning job that adjusts the performance of the pre-trained NLP ML model based on the input prompt to generate a result of the tuning job; and
[0117] The prompt development system provides the result of the tuning job.
[0118] Clause 6. The method according to clause 5, wherein the tuning job performs prompt adjustment techniques to adjust the pre-trained NLP ML model.
[0119] Clause 7. The method according to any one of clauses 5-6, wherein the tuning job performs context learning techniques to adjust the pre-trained NLP ML model.
[0120] Clause 8. The method according to any one of clauses 5-7, wherein the tuning job performs fine-tuning techniques to adjust the pre-trained NLP ML model.
[0121] Clause 9. The method according to any one of clauses 5-8, wherein the pre-trained NLP ML model is included in a prompt recommendation provided in response to a discovery request executed by the prompt development system.
[0122] Clause 10. The method according to any one of clauses 5-9, wherein evaluating the performance of the tuning job that adjusts the performance of the pre-trained NLP ML model includes:
[0123] Generating one or more inferences for input test data using the adjusted NLP ML model according to a selected prompt; and
[0124] Determining the inference performance of the one or more inferences based on the ground truth labels of the input test data.
[0125] Clause 11. The method according to any one of clauses 5-10, wherein the result of the tuning job includes computational performance and inference performance.
[0126] Clause 12. The method according to any one of clauses 5-11, which further includes:
[0127] Receive, by the prompt development system, a request to deploy the adjusted NLP ML model;
[0128] Provision, by the prompt development system, a host for the adjusted NLP ML model; and
[0129] Provide a model endpoint for a client application to submit an inference request to the host to generate a corresponding inference using the adjusted NLP ML model.
[0130] Clause 13. The method according to any one of Clauses 5 - 12, further comprising:
[0131] Receive, by the prompt development system, the selected prompt as a prompt submission to be maintained by the prompt development system; and
[0132] Add, by the prompt development system, the selected prompt to the prompt task set.
[0133] Clause 14. One or more non - transitory computer - readable storage media storing program instructions that, when executed on or across one or more computing devices, cause the one or more computing devices to perform the following operations:
[0134] Receive, via an interface of the prompt development system, a request to perform further adaptation that adjusts the performance of a pre - trained natural language processing (NLP) machine learning (ML) model that performs an NLP task based on an input prompt selected from a set of prompt tasks maintained by the prompt development system;
[0135] Generate, by the prompt development system, an adaptation job to perform the requested further adaptation using an adaptation data set specified by the request;
[0136] Evaluate, by the prompt development system, the performance of the adaptation job that adjusts the performance of the pre - trained NLP ML model based on the input prompt to generate a result of the adaptation job; and
[0137] Return, via the interface of the prompt development system, the result of the adaptation job.
[0138] Clause 15. The one or more non - transitory computer - readable storage media according to Clause 14, wherein the adaptation job performs prompt adjustment techniques to adjust the pre - trained NLP ML model.
[0139] Clause 16. The one or more non - transitory computer - readable storage media according to any one of Clauses 14 - 15, wherein the adaptation job performs context learning techniques to adjust the pre - trained NLP ML model.
[0140] Clause 17. One or more non-transitory computer-readable storage media according to any one of Clauses 14 - 16, wherein the adaptation operation performs fine-tuning techniques to adjust the pre-trained NLP ML model.
[0141] Clause 18. One or more non-transitory computer-readable storage media according to any one of Clauses 14 - 17, wherein the selected prompt is included in a prompt recommendation provided in response to a discovery request executed by the prompt development system.
[0142] Clause 19. One or more non-transitory computer-readable storage media according to any one of Clauses 14 - 18, which store additional program instructions that, when executed on or across the one or more computing devices, cause the one or more computing devices to further perform the following operations:
[0143] Receiving, by the prompt development system, the selected prompt as a prompt submission to be maintained by the prompt development system; and
[0144] Adding, by the prompt development system, the selected prompt to the prompt task set.
[0145] Clause 20. One or more non-transitory computer-readable storage media according to any one of Clauses 14 - 19, wherein the prompt development system is a machine learning service provided by a provider network, and wherein the one or more non-transitory computer-readable storage media store additional program instructions that, when executed on or across the one or more computing devices, cause the one or more computing devices to further perform the following operations:
[0146] Receiving, via the interface, a request to deploy the adjusted NLP ML model;
[0147] Provisioning, by the machine learning service, a host for the adjusted NLP ML model; and
[0148] Providing, via the interface, a model endpoint for a client application to submit an inference request to the host to generate a corresponding inference using the adjusted NLP ML model.
[0149] Clause 21. A system, comprising:
[0150] At least one processor; and
[0151] A memory that stores program instructions which, when executed by the at least one processor, cause the at least one processor to implement a prompt development system for a natural language processing (NLP) machine learning (ML) model, the prompt development system being configured to:
[0152] Receive, via an interface of the prompt development system, a request to determine a prompt for an NLP task to be performed by a pre-trained NLP ML model;
[0153] Determine, at least in part based on the request, a task classification of the NLP task;
[0154] Access a set of prompt tasks corresponding to the task classification maintained by the prompt development system and select one or more candidate prompts for the NLP task;
[0155] Cause the one or more candidate prompts to be provided as input to the pre-trained NLP ML model to generate corresponding prompt results produced by the pre-trained NLP model;
[0156] Evaluate the corresponding prompt results produced by one or more candidate prompts using the pre-trained NLP ML model to generate a prompt recommendation for the NLP task; and
[0157] Return, via the interface of the prompt development system, the prompt recommendation for the NLP task.
[0158] Clause 22. The system according to clause 21, wherein, for determining the task classification of the NLP task, the prompt development system is configured to evaluate a sample input and a sample output specified in the request.
[0159] Clause 23. The system according to any one of clauses 21 - 22, wherein the prompt development system is further configured to:
[0160] Receive the one or more candidate prompts as a prompt submission to be maintained by the prompt development system; and
[0161] Add the one or more candidate prompts to the set of prompt tasks.
[0162] Clause 24. The system according to any one of clauses 21 - 23, wherein the prompt development system is implemented as part of a machine learning service provided by a provider network, wherein the set of prompt tasks is one of a plurality of different sets of prompt tasks maintained by the machine learning service, and the set of prompt tasks is a set of prompts submitted to the machine learning service via an interface of the machine learning service.
[0163] Clause 25. A method, comprising:
[0164] Receive, at a prompt development system for a natural language processing (NLP) machine learning (ML) model, a request to determine a prompt for an NLP task to be performed by a pre-trained NLP ML model;
[0165] Select, by the prompt development system, one or more candidate prompts from a set of prompt tasks maintained by the prompt development system for the NLP task;
[0166] Evaluate, by the prompt development system, corresponding prompt results generated by the one or more candidate prompts using the pre-trained NLP ML model; and
[0167] Return, by the prompt development system, a prompt recommendation for the NLP task based at least in part on the evaluation of the corresponding prompt results.
[0168] Clause 26. The method according to clause 25, further comprising determining, by the prompt development system, a task classification of the NLP task based at least in part on the request, wherein the prompt task classification corresponds to the task classification.
[0169] Clause 27. The method according to clause 26, wherein determining the task classification of the NLP task includes identifying the task classification specified in the request.
[0170] Clause 28. The method according to any one of clauses 26-27, wherein determining the task classification of the NLP task includes evaluating sample input and sample output specified in the request.
[0171] Clause 29. The method according to any one of clauses 25-28, wherein the prompt recommendation includes two or more candidate prompts from the candidate prompts having corresponding sample outputs for comparison.
[0172] Clause 30. The method according to any one of clauses 25-29, further comprising selecting, by the prompt development system, one or more candidate pre-trained NLP ML models corresponding to the task classification.
[0173] Clause 31. The method according to any one of clauses 25-30, wherein the prompt recommendation includes instructions and stylistic features.
[0174] Clause 32. The method according to any one of clauses 25-31, wherein the prompt recommendation includes the computational performance of one of the candidate prompts included in the prompt recommendation.
[0175] Clause 33. The method according to any one of Clauses 25-32, wherein the hint recommendation is further generated based on an evaluation of the one or more candidate hints with respect to performance criteria specified in the request.
[0176] Clause 34. The method according to any one of Clauses 25-33, further comprising:
[0177] receiving, by the hint development system, the one or more candidate hints as hint submissions to be maintained by the hint development system; and
[0178] adding, by the hint development system, the one or more candidate hints to the hint task set.
[0179] Clause 35. One or more non-transitory computer-readable storage media storing program instructions that, when executed on or across one or more computing devices, cause the one or more computing devices to perform the following operations:
[0180] receiving, at a hint development system for a natural language processing (NLP) machine learning (ML) model, a request to determine a hint for an NLP task to be performed by a pre-trained NLP ML model;
[0181] selecting, by the hint development system, one or more candidate hints for the NLP task from a hint task set maintained by the hint development system corresponding to the task classification;
[0182] causing, by the hint development system, an evaluation of the corresponding hint results generated by the one or more candidate hints using the pre-trained NLP ML model; and
[0183] returning, by the hint development system, a hint recommendation for the NLP task based at least in part on the evaluation of the corresponding hint results.
[0184] Clause 36. The one or more non-transitory computer-readable storage media according to Clause 35, wherein a task classification is specified in the request, and wherein the hint task set corresponds to the task classification.
[0185] Clause 37. The one or more non-transitory computer-readable storage media according to any one of Clauses 35-36, storing additional program instructions that, when executed on or across the one or more computing devices, further perform an evaluation of sample input and sample output specified in the request to determine a task classification, wherein the hint task set corresponds to the task classification.
[0186] Clause 38. One or more non-transitory computer-readable storage media according to any one of Clauses 35-37, wherein the request does not specify the pre-trained NLP ML model, and wherein the prompt recommendation includes an identification of the pre-trained NLP ML model.
[0187] Clause 39. One or more non-transitory computer-readable storage media according to any one of Clauses 35-38, which store additional program instructions that, when executed on or across the one or more computing devices, cause the one or more computing devices to further perform the following operations:
[0188] Receiving, by the prompt development system, the one or more candidate prompts as a prompt submission to be maintained by the prompt development system; and
[0189] Adding, by the prompt development system, the one or more candidate prompts to the prompt task set.
[0190] Clause 40. One or more non-transitory computer-readable storage media according to any one of Clauses 35-39, wherein the prompt development system is implemented as part of a machine learning service provided by a provider network, and wherein the prompt task set is one of a plurality of different prompt task sets maintained by the machine learning service, and the prompt task set is a set of prompts submitted to the machine learning service via an interface of the machine learning service.
Claims
1. A system, comprising: at least one processor; and a memory storing program instructions that, when executed by the at least one processor, cause the at least one processor to implement a prompt development system for a natural language processing (NLP) machine learning (ML) model, the prompt development system being configured to: receive, via an interface of the prompt development system, a request for performing further adaptation that adjusts the performance of a pre-trained NLP ML model that performs an NLP task based on an input prompt selected from a set of prompt tasks maintained by the prompt development system; generate an adaptation job to perform the requested further adaptation of adjusting the performance of the pre-trained NLP ML model using an adaptation dataset specified by the request; cause the adaptation job to be executed; evaluate the performance of the adaptation job to generate a result of the adaptation job; and return, via the interface of the prompt development system, the result of the adaptation job.
2. The system according to claim 1, wherein the adaptation job performs a prompt adjustment technique to adjust the pre-trained NLP ML model.
3. The system according to any one of claims 1-2, wherein the adaptation job performs a context learning technique to adjust the pre-trained NLP ML model.
4. The system according to any one of claims 1-3, wherein the prompt development system is a machine learning service provided by a provider network, and wherein the prompt development system is further configured to: receive, via the interface, a request to deploy the adjusted NLP ML model; provision a host for the adjusted NLP ML model; and provide, via the interface, a model endpoint for a client application to submit an inference request to the host to generate a corresponding inference using the adjusted NLP ML model.
5. A method, comprising: receiving, by a prompt development system, a request for performing further adaptation that adjusts the performance of a pre-trained NLP machine learning (ML) model that performs a natural language processing (NLP) task based on an input prompt selected from a set of prompt tasks maintained by the prompt development system; generating, by the prompt development system, an adaptation job to perform the requested further adaptation using an adaptation dataset specified by the request; evaluating, by the prompt development system, the performance of the adaptation job that adjusts the performance of the pre-trained NLP ML model based on the input prompt to generate a result of the adaptation job; and providing, by the prompt development system, the result of the adaptation job.
6. The method according to claim 5, wherein the adaptation job performs a prompt adjustment technique to adjust the pre-trained NLP ML model.
7. The method according to any one of claims 5-6, wherein the adaptation job performs a context learning technique to adjust the pre-trained NLP ML model.
8. The method according to any one of claims 5-7, wherein the adaptation job performs a fine-tuning technique to adjust the pre-trained NLP ML model.
9. The method according to any one of claims 5-8, wherein the pre-trained NLP ML model is included in a prompt recommendation provided in response to a discovery request executed by the prompt development system.
10. The method according to any one of claims 5-9, wherein evaluating the performance of the tuning operation that adjusts the performance of the pre-trained NLP ML model includes: Generating one or more inferences for input test data using the tuned NLP ML model according to a selected prompt; And Determining the inference performance of the one or more inferences based on the ground truth labels of the input test data.
11. The method according to any one of claims 5-10, wherein the result of the tuning operation includes computational performance and inference performance.
12. The method according to any one of claims 5-11, further comprising: Receiving, by the prompt development system, a request to deploy the tuned NLP ML model; Provisioning, by the prompt development system, a host for the tuned NLP ML model; And Providing a model endpoint for a client application to submit an inference request to the host to generate a corresponding inference using the tuned NLP ML model.
13. The method according to any one of claims 5-12, further comprising: Receiving, by the prompt development system, the selected prompt as a prompt submission to be maintained by the prompt development system; And Adding, by the prompt development system, the selected prompt to the prompt task set.
14. One or more non-transitory computer-readable storage media storing program instructions that, when executed on or across one or more computing devices, cause the one or more computing devices to perform the following operations: Receiving, via an interface of a prompt development system, a request to perform further tuning that adjusts the performance of a pre-trained natural language processing (NLP) machine learning (ML) model that performs an NLP task according to an input prompt selected from a set of prompt tasks maintained by the prompt development system; Generating, by the prompt development system, a tuning job to perform the requested further tuning using a tuning data set specified by the request; Evaluating, by the prompt development system, the performance of the tuning operation that adjusts the performance of the pre-trained NLP ML model based on the input prompt to generate a result of the tuning operation; And Returning, via the interface of the prompt development system, the result of the tuning operation.
15. The one or more non-transitory computer-readable storage media according to claim 14, storing additional program instructions that, when executed on or across the one or more computing devices, cause the one or more computing devices to further perform the following operations: Receiving, by the prompt development system, a selected prompt as a prompt submission to be maintained by the prompt development system; and Adding, by the prompt development system, the selected prompt to the prompt task set.