Integrating software tools in language model neural network responses through tool embeddings
The integration of software tools with language model neural networks using tool embeddings addresses limitations in task-specific functionality and security risks, achieving efficient and flexible tool addition/removal without re-training, thus enhancing performance and security.
Patent Information
- Application Number
- PCT/EP2024/088641
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-29
- Filing Date
- 2024-12-30
- Publication Date
- 2025-07-03
AI Technical Summary
Existing language model neural networks lack the ability to efficiently integrate software tools, leading to limitations in task-specific functionality and increased computational and storage costs, as well as security risks due to explicit API key exposure.
A system that integrates software tools with language model neural networks using software tool selection and embeddings, allowing for flexible addition and removal of tools without re-learning the neural network parameters, reducing computational and storage costs while enhancing security.
The system achieves high-performance task-specific functionality with reduced computational and storage costs, improved security by encoding API interactions into embeddings, and flexibility in adding or removing tools without re-training the neural network.
Smart Images

Figure EP2024088641_03072025_PF_FP_ABST
Abstract
Description
INTEGRATING SOFTWARE TOOLS IN LANGUAGE MODEL NEURALNETWORK RESPONSES THROUGH TOOL EMBEDDINGSCROSS-REFERENCE TO RELATED APPLICATIONS
[0001] This application claims priority to U.S. Provisional Application No. 63 / 616,230, filed on December 29, 2023, entitled “INTEGRATING SOFTWARE TOOLS IN LANGUAGE MODEL NEURAL NETWORK RESPONSES THROUGH TOOL EMBEDDINGS,” which is incorporated herein by reference in its entirety.BACKGROUND
[0002] This specification relates to integrating software tools with neural networks.
[0003] Neural networks are machine learning models that employ one or more layers of nonlinear units to predict an output for a received input. Some neural networks include one or more hidden layers in addition to an output layer. The output of each hidden layer is used as an input to the next layer in the network, i.e., the next hidden layer or the output layer. Each layer of the network generates an output from a received input in accordance with current value inputs of a respective set of parameters.SUMMARY
[0004] This specification describes a system implemented as computer programs on one or more computers in one or more locations that integrates capabilities of a set of software tools with a language model neural network.
[0005] Some implementations of the system are particularly useful in a parallel or distributed computing environment, as they are specifically adapted to enable the process of learning to select a software tool to be decoupled from a process of using the software tool using, typically, done by a language model, or vision language model neural network. In particular some implementations of the system enable such a language model or vision language model neural network, implemented on a first computer system, to maintain a fixed set of neural network parameters (in particular the trainable parameters, e.g., the weights), whilst allowing the tool use to be learned by a second computer system. This has many advantages, including the possibility of having a shared language model or vision language model neural network on the first computer system and having multiple second computer systems, each of whichcan learn to use a different respective set of software tools, as described later based on one or more different software tool selection embeddings and or one or more different software tool embeddings.
[0006] The language model neural network is a pre-trained neural network trained on training data from various sources and domains. In some cases, the training data does not include specialized information, e.g., task-specific information, up-to-date information on recent events, and so on. To overcome this deficiency in the training data, the system described in this specification relates to providing data and services related to a set of software tools to systems that implement the language model neural network.
[0007] The set of software tools can include software programs that are each operable to execute operations to perform respective functions. For example, the set of software tools can include scheduling software for scheduling a meeting based on a user’s calendar, booking software for booking airline accommodations, transcription software for providing a transcript of a meeting, among others. Incorporating functionality of the software tools to the general functionality of the language model neural network expands the number of practical problems and use cases that a system that implements the language model neural network (e.g., a large language model) can address.
[0008] Some further examples of use of a system as described herein are as follows:
[0009] The software tools can include one or more image processing tools, e.g. tools that are specialized to processing a particular, or different particular types of images. Some examples can be images of the natural environment or of particular types of natural environment, images of urban scenes or types of urban scene, astronomical images, medical images or particular types of medical image, and so forth. The query input can define an image processing task performed on pixels of a still or moving image; and the response to the query input can define a result of the image processing task. The image processing task can be, e.g. an image classification task in which pixels of a still or moving image are processed to generate a response that categorizes a content of the image; or an object detection task in which the image is processed to generate a response that locates coordinates, e.g. bounding box coordinates, of one or more objects or types of object in the image; or an instance detection task in which the image is processed to generate a response that identifies the shapes of one or more objects or types of object in the image.
[0010] The software tools can include one or more materials design tool for designing a physical material. Some examples include a protein design tool (e.g. an AlphaF old-related tool), a retrosynthesis tool (such as described in WO2021 / 263238), or a materials design tool(such as described in Yang et al., “Generative Hierarchical Materials Search”, arXiv: 2409.06762, 2024). The query input can define a materials design task performed, e.g. specifying one or more desired characteristics of the material; and the response to the query input can define a result of the materials design task. Afterwards the material can be physically made (synthesized), e.g. for testing.
[0011] The software tools can include one or more robot control tools for controlling a mechanical robot to perform a physical task. The query input can define a physical task to be performed by the robot; the response to the query input can define a result of the physical task, e.g. it may include one or more observations of a result of the task, i.e. observations captured from the real world that have been returned as an output from the tool and that characterize a result of the physical task (which may be a synthesized material in the case of the preceding example).
[0012] This specification generally describes a prompt tuning approach, in which a system learns and tunes parameters of a soft prompt that corresponds to a particular software tool. The soft prompt is a learned embedded representation of task-specific data in the form of a continuous embedding matrix. After performing prompt tuning, when performing inference, the system appends the learned soft prompt to a hard prompt, and the language model neural network processes the combined prompt. The hard prompt is an embedded representation that includes embeddings generated by processing a sequence of tokens (that includes a query input to the language model neural network, e.g., a prompt, question, etc.) using an embedding layer of the language model neural network. By processing the combined prompt (the soft prompt appended or prepended to the hard prompt), the language model neural network processes task-specific data through a learned embedded representation while keeping the parameters of the neural network fixed.
[0013] Particular embodiments of the subject matter described in this specification can be implemented so as to realize one or more of the following advantages.
[0014] Approaches for providing task-specific data to language model neural networks to integrate the use of software tools into the use of a language model neural network can range from supervised fine-tuning approaches, which demonstrate high performance at a high cost, to few-shot learning approaches, which demonstrate lower performance at a low cost.
[0015] The supervised fine-tuning approach requires re-learning the parameter space of the model. In many cases, the language model neural network has a large number of parameters to re-learn, which is a computationally expensive task. In addition, fine-tuning a model withtraining data related to a particular task with a supervised fine-tuning approach introduces a risk of degrading performance for other tasks.
[0016] The few-shot learning approaches involve processing a few examples of the particular software tool being used as part of each input to the language model neural network. While this approach does not require fine-tuning the language model neural network, the few-shot learning approach typically produces sub-optimal performance in comparison with supervised fine-tuning.
[0017] In addition, the few-shot learning approach is more computationally expensive at inference in comparison with the approach described in the present specification. For the fewshot learning approach, several examples are included in the prompt with each inference, which results in more tokens being processed by the neural network, which increases the amount of required compute resources. In particular, transformer based models compute requirement scales quadratically with the number of tokens processed by the transformer model. Thus, because adding examples to each prompt adds a significant number of tokens to each input to the model, adopting the few-shot learning approach substantially increases how computationally expensive each inference is. In comparison with few-shot learning, the approach described in this specification includes storing embedded representations of prompts and / or examples (e.g., API documentation) instead of the actual prompts and / or examples, which results in a more compact storage approach (e.g., storing compressed data).
[0018] In addition, the approach described in the present specification transmits fewer tokens between computational resources at inference in comparison with the few-shot learning approach. Furthermore, the present approach requires fewer data storage requirements in comparison with the few-shot learning approach. For example, the few-shot learning approach requires storage of data that includes long prompts and / or examples (e.g., API documentation examples). In contrast, the present approach requires storage of embedded representations (e.g., software tool selection embedding and software tool embeddings), which are more compact representations than the associated raw data.
[0019] The approach described in the present specification provides a high-performance solution (e.g., accuracy comparable to supervised fine-tuning) for integrating software tool functionality into the functionality of a language model neural network at a lower computational and economical cost in comparison with the supervised fine-tuning approaches.
[0020] The method, by utilizing a software tool selection embedding and a respective software tool embedding, reduces a risk of API key exposure compared to methods thatrequire explicit inclusion of API keys in a language model output. By encoding API interaction details into embeddings rather than storing them in their original format, the system can reduce the risk of accidentally exposing API keys or other sensitive information in the language model’s output. This improves the security of the system that interacts with the software tools.
[0021] In particular, the techniques described involve a training process that includes learning parameters of a soft prompt (e.g., embedding vectors associated with software tool data or embedding vectors associated with a selection of a software tool of a set of software tools), which in some examples, can include a learned parameter space that is several order of magnitudes smaller than a learned parameter space of the language model neural network. By reducing the number of learned parameters, a training system requires fewer computations, energy consumption, data transmissions, and data storage resources. Furthermore, the described approach demonstrates accuracy of software tool-related outputs (e.g., software tool identification) comparable to the outputs of a supervised fine-tuned model.
[0022] Furthermore, the techniques described offer a flexible approach for adding and removing task-specific functionality to a system that implements the language model neural network. For example, in an example system that implements supervised fine-tuning, the system must re-learn the parameter space of the language model neural network to add training data related to a new software tool. In contrast, the techniques described in the present specification require small and task-specific learning tasks to learn software toolspecific embeddings to provide a soft prompt related to the new software tool. By creating a flexible system that can add and remove software tools without re-learning the parameters of the language model neural network, the system requires fewer computations, energy consumption, data transmissions, and storage resources to add and / or remove functionality of a particular software tool. Similarly, the system can update a particular learned embedding (soft prompt) related to a particular software tool by re-learning the associated embedding in response to a change in the functionality of the particular software tool. A re-learning of parameters of language model neural networks is not required in the event that relevant training data changes.
[0023] In addition to the embodiments described in the present specification, the following embodiments are also innovative:
[0024] In a first aspect, a method is performed by one or more computers, the method including maintaining software tool use data that includes a software tool selection embedding and a respective software tool embedding for each software tool in a set ofsoftware tools. The method includes receiving a query input, generating a software tool selection input sequence that includes the software tool selection embedding and a first set of embeddings characterizing the query input. The method includes processing the software tool selection input sequence using a language model neural network to generate a software tool selection output that identifies a particular software tool from the set of software tools, generating a software tool call input sequence that includes the respective software tool embedding for the particular software tool and a second set of embeddings characterizing the query input.
[0025] In some implementations, the method further includes processing the software tool call input sequence using the language model neural network to generate a software tool call output that specifies an input to the particular software tool, providing the input to the particular software tool, obtaining an output from the particular software tool, and generating a response to the query input from the output obtained from the particular software tool.
[0026] In some implementations, the method includes receiving data specifying a new software tool to be added to the set of software tools, and in response to receiving the data specifying the new software tool to be added to the set of software tools: generating a respective software tool embedding for the new software tool and updating the software tool use data to include the respective software tool embedding for the new software tool.
[0027] In some implementations, the method includes generating a respective software tool embedding for the new software tool includes initializing the respective software tool embedding for the new software tool, receiving first training data comprising a plurality of training software tool call input sequences for the new software tool and, for each training software tool call input sequence, a respective ground truth software tool call output for the new software tool, and performing prompt tuning on the respective software tool embedding for the new software tool using the first training data to update the respective software tool embedding for the new software tool.
[0028] In some implementations, in response to receiving the data specifying the new software tool to be added to the set of software tools, the method includes updating the software tool selection embedding to account for the new software tool being added to the set of software tools, and updating the software tool use data to include the updated software tool selection embedding.
[0029] In some implementations, the method includes updating the software tool selection embedding to account for the new software tool being added to the set of software tools includes receiving second training data comprising multiple training software tool selectioninput sequences and, for each training software tool selection input sequence, data identifying a respective ground truth software tool that should be selected in response to the training software tool selection input sequence, in which, for at least a subset of the training software tool selection input sequences, the respective ground truth software tool is the new software tool. The method further includes performing prompt tuning on the software tool selection embedding using the training data to update the respective software tool embedding for the new software tool.
[0030] In some implementations, the software tool selection embedding has been learned through prompt tuning while holding the language model neural network fixed.
[0031] In some implementations, the respective software tool embeddings for the software tools in the set of software tools have been learned through prompt tuning while holding the language model neural network fixed.
[0032] In some implementations, software tool selection output is an output sequence that specifies an identifier for the particular software tool.
[0033] In some implementations, the software tool selection output is a respective likelihood score for each of one or more of the software tools, and wherein the method further includes selecting the particular software tool based on the respective likelihood scores.
[0034] In some implementations, each software tool receives an input that includes respective values for each of one or more input parameters for the software tool, and wherein the software tool call output is an output sequence that specifies respective particular values for each of the one or more input parameters for the particular software tool.
[0035] In some implementations, in generating a response to the query input from the output obtained from the particular software tool includes generating a software tool output input sequence from at least the output obtained from the particular software tool, and processing the software tool output input sequence using the language model neural network to generate an output sequence that defines the response to the query input.
[0036] In some implementations, the method is implemented in a distributed computing system that includes a first computer system that implements the language model neural network and a second computer system that implements one or both of a tool selection neural network and a tool embedding neural network, the method further including maintaining the software tool use data on the second computer system. The first computer system performs the steps of receiving the query input, processing the software tool selection input sequence, processing the software tool call input sequence, providing the input to the particular software tool, obtaining the output from the particular software tool, and generating the response to thequery input. The second computer system performs the steps of maintaining the software tool use data, generating the software tool selection input sequence, generating the software tool call input sequence.
[0037] In some implementations, the distributed computing system includes a more than one of the second computer systems, in which each of the second computer systems is configured to support a different group of software tools for a different respective group of users, and wherein each of the second computer systems maintains different respective software tool use data for generating a different respective said software tool selection input sequence and software tool call input sequence.
[0038] In some implementations, the method includes learning the software tool use data without modifying learned parameters of the language model neural network.
[0039] In some implementations, the software tools include at least one image processing tool, wherein the query input defines an image processing task, and wherein the response to the query input defines a result of the image processing task.
[0040] In some implementations, the software tools include at least one materials design tool for designing a physical material, in which the query input defines a materials design task, and in which the response to the query input defines a result of the materials design task.
[0041] In some implementations, the software tools include at least one robot control tool for controlling a mechanical robot to perform a physical task, in which the query input defines a physical task, and wherein the response to the query input defines a result of the physical task.
[0042] In another aspect, a method is performed by one or more computers for generating a respective software tool embedding for a new software tool, the method including initializing the respective software tool embedding for the new software tool, receiving first training data including multiple training software tool call input sequences for the new software tool and, for each training software tool call input sequence, a respective ground truth software tool call output for the new software tool. The method further includes performing prompt tuning on the respective software tool embedding for the new software tool using the first training data to update the respective software tool embedding for the new software tool.
[0043] In some implementations, in response to receiving the data specifying the new software tool to be added to the set of software tools, the method includes updating the software tool selection embedding to account for the new software tool being added to the set of software tools and updating the software tool use data to include the updated software tool selection embedding.
[0044] In another aspect, one or more non-transitory computer readable media storing embeddings generated by performing operations including maintaining software tool use data that includes a software tool selection embedding and a respective software tool embedding for each software tool in a set including multiple software tools. The operations include receiving a query input and generating a software tool selection input sequence that includes the software tool selection embedding and a first plurality of embeddings characterizing the query input. The operations include processing the software tool selection input sequence using a language model neural network to generate a software tool selection output that identifies a particular software tool from the set of software tools and generating a software tool call input sequence that includes the respective software tool embedding for the particular software tool and a second plurality of embeddings characterizing the query input.
[0045] While this specification contains many specific implementation details, these should not be construed as limitations on the scope of any invention or on the scope of what may be claimed, but rather as descriptions of features that may be specific to particular embodiments of particular inventions. Certain features that are described in this specification in the context of separate embodiments can also be implemented in combination in a single embodiment. Conversely, various features that are described in the context of a single embodiment can also be implemented in multiple embodiments separately or in any suitable subcombination. Moreover, although features may be described above as acting in certain combinations and even initially be claimed as such, one or more features from a claimed combination can in some cases be excised from the combination, and the claimed combination may be directed to a subcombination or variation of a subcombination.
[0046] The details of one or more embodiments of the subject matter of this specification are set forth in the accompanying drawings and the description below. Other features, aspects, and advantages of the subject matter will become apparent from the description, the drawings, and the claims.BRIEF DESCRIPTION OF THE DRAWINGS
[0047] FIG. l is a diagram of an example system that includes a language model neural network system that interacts with a software tool system.
[0048] FIG. 2 is a diagram of an example system that includes a language model neural network and a software tool system.
[0049] FIG. 3 is a flow diagram of an example process for generating an output in response to a query input to a language model neural network system with a software tools integration.
[0050] FIG. 4 is a flow diagram of an example process for updating software tools use data.
[0051] FIG. 5 is an example plot for comparing training methods of a language model neural network configured to select a software tool.
[0052] FIG. 6 illustrates example plots for comparing training methods of a language model neural network configured to select a software tool.
[0053] FIG. 7 is an example plot for comparing training methods of a language model neural network configured to generate arguments to be processed by a software tool.
[0054] Like reference numbers and designations in the various drawings indicate like elements.DETAILED DESCRIPTION
[0055] This specification describes a system implemented as computer programs on one or more computers in one or more locations that integrates capabilities of a set of software tools with a system that implements a language model neural network. The set of software tools can act as a set of plugins, e.g., extra layers of capability, for the language model neural network. The software tools provide additional functionality to systems that is not available from the traditional outputs of the language model neural network. In some examples, software tools can include software applications that can communicate through an application programming interface (API) that allows other software applications (e.g., the system implementing the language model neural network) to provide inputs to and receive outputs from the software applications associated with the software tools. For example, calendar applications, web browsers, travel planning applications, and internet of things devices, can all communicate through APIs either over the web, local networks, or via local computational resources. In general, software tools that can communicate through a “request” and “response” scheme with predictable and structured input and output formats can provide additional functionality to language model neural networks.
[0056] In some cases, details of particular software tools are not represented in training data of language model neural networks. Many pre-trained language model neural networks are limited to performing general purpose tasks with limited functionality when asked to generate information in relation to recent events, task-specific objectives like mathematical operations, operations that require interacting with external resources, and taking an action. In general,language model neural networks are trained on general purpose information and can perform tasks and reasoning exercises that are closely related to the training data. Re-training and fine-tuning (e.g., changing the parameters (weights) of the neural network) language model neural networks to accommodate specific functionality and information requirements can be computationally expensive and time consuming.
[0057] The description below describes a prompt tuning approach of integrating the functionality of software tools into systems that implement operations of a language model neural network (e.g., a general purpose pre-trained large language model). The prompt tuning approach includes learning embedded representations of software tool-specific data and processing, by the language model neural network, the learned embeddings along with an embedded representation of an input query.
[0058] FIG. 1 is a diagram of an example system 100 that includes a language model neural network system 102 (referred to below as the system 102) that interacts with a software tool system 104. In a general sense, the system 102 can interact with multiple software tool systems.
[0059] The system 102 and the software tool system 104 are examples of systems implemented as computer programs on one or more computers in one or more locations, in which the systems, components, and techniques described below can be implemented.
[0060] The system 102 includes a language model neural network 112 and executes communication operations (e.g., database requests, API calls, local software commands, etc.).
[0061] For example, the system 102 receives a query input 110 from a client device 106. The query input 110 can be a question, prompt, instruction, or any other message to be processed by the system 102. The system 102 tokenizes the query input 110 (e.g., a sequence of words) into a sequence of tokens with a tokenizer, in which each token of the sequence of tokens is mapped to a numerical token identifier by the tokenizer.
[0062] The system 102 processes the sequence of token identifiers with neural network layers of the language model neural network 112 including an embedding layer (e.g., a neural network layer that converts the sequence of token identifiers into one or more embedding vectors). The layers of the language model neural network 112 can also include one or more transformer layers (e.g., self-attention and feedforward networks) that sequentially process the embedding vectors.
[0063] In addition to processing the query input 110 with the language model neural network 112, the system 102 accesses software tool use data 114 via database queries, API calls, localsoftware commands, etc. The system 102 accesses and processes the software tool use data 114 to interact with software tools in response to the query input 110.
[0064] The software tool use data 114 includes embeddings (e.g., one or more embedding vectors). The embeddings include (i) a software tool selection embedding and (ii) a respective software tool embedding for each software tool in a set of multiple software tools.
[0065] The system 102 processes the software tool selection embedding for selecting relevant software tools from the set of software tools in response to the query input 110. The software tool selection embedding is an embedded representation of training data, in which the training data include example query inputs and corresponding software tools that are relevant for addressing the prompt and / or question of the associated query input. An example query input of a training data item is “book a flight for tomorrow from new york to los angeles,” and a corresponding software tool of the example training data item is “airline booking software.”
[0066] For each software tool of the set of multiple software tools, the respective software tool embedding is an embedded representation of instructions for interacting with the respective software tool (e.g., required parameters, expected input data formatting, and API endpoints). Each embedding vector in the software tool selection embedding and each embedding vector in the software tool embeddings has the same dimensions as the embedding vectors generated by an embedding layer of the language model neural network 112 of the system 102.
[0067] To make use of the embeddings in the software tool use data 114, the system 102 appends or prepends embeddings from the software tool use data 114 (e.g., the software tool selection embedding or a software tool embedding) to the embeddings generated by the embeddings layer of the language model neural network 112. For example, the embedding layer of the language model neural network 112 generates an embedded representation of the query input 110. The system 102 appends or prepends the generated embedded representation with an embedding from the software tool use data 114. The remaining layers of the language model neural network 112 (e.g., transformer layers) process the combined embedding structure.
[0068] The system 102 transmits an input message 120 to the software tool system 104 to initiate an execution of operations associated with a respective software tool. For example, the software tool system 104 can execute operations associated with an airline booking software system. The input message 120 can include one or more input parameters for the software tool. In response, the software tool system 104 returns a response message 116 tothe system 102. The response message 116 can be an output sequence that specifies particular values for each of the one or more input parameters for the particular software tool.
[0069] In some implementations, the transmission of the input message 120 and the reception of the response message 116 is one step of multiple data processing steps performed by the system 102 for generating a response 108 to transmitted to the client device 106 in response to the query input 110. For example, the system 102 can receive the response message 116, which is tokenized and then processed by the language model neural network 112 of the system 102 as one step of a multi-step process.
[0070] In some implementations, the system 102 processes the response message 116 from the software tool system 104 to generate the response 108 to be received by the client device 106. In some implementations, the language model neural network 112 processes the response message 116 along with the query input 110 (e.g., both the response message 116 and the query input 110 are included in a single prompt) to generate the response 108 to be received by the client device 106.
[0071] FIG. 2 is a diagram of an example system 200 that includes a language model neural network 112 that interacts with a software tool system 216. The example system 200 is an example of a system implemented as computer programs on one or more computers in one or more locations, in which the systems, components, and techniques described below can be implemented.
[0072] The system 200 receives a query input 202 that initiates a sequence of data processing steps. In some implementations, the system 200 receives the query input 202 via an API connection with a computer that executes operations associated with a user interface (e.g., a user interface of the client device 106 of FIG. 1). In some implementations, the system 200 receives the query input 202 from a system implemented on a common device (e.g., the common device implements the operations of system 200 and operations associated with a user interface that receives and transmits the query input 202). In some other implementations, the system 200 receives the query input 202 as part of an automated data processing pipeline, in which the output of the system 200 can be processed by data processing steps of a subsequent step of the automated pipeline.
[0073] The system 200 includes a software tool selection sequence generator 204. The software tool selection sequence generator 204 processes the query input 202 and a software tool selection embedding 206.
[0074] The system 200 can maintain software tool use data (e.g., the software tool use data 114 of FIG. 1) that includes the software tool selection embedding 206 and a respectivesoftware tool embedding 210 for each software tool in a set of available software tools. In this example, the software tool selection embedding 206 is a learned vector representation of training data that includes example input queries and corresponding indicators of one or more software tools that are relevant to each respective example input query.
[0075] In response to the query input 202, the software tool selection sequence generator 204 can generate a tool selection input sequence 222 that includes the software tool selection embedding 206 and a sequence of embeddings that characterizes the query input 202. In some implementations, the software tool selection sequence generator 204 appends the software tool selection embedding 206 to the query input 202 embedding. In some other implementations, the software tool selection sequence generator 204 prepends the software tool selection embedding 206 to the query input 202 embedding.
[0076] The system 200 can generate the sequence of embeddings characterizing the query input 202 in any appropriate manner, e.g., by tokenizing the query input 202 and generating a respective embedding for each token or, when the query input 202 is a multi-modal query input, generating one or more embeddings of each modality in the query input 202.
[0077] In some implementations, the query input 202 includes one or more modalities including text, image, audio, video, or a combination of one or more modalities. For example, the language model neural network 112 can be a vision language model which can process both visual (image and video) data along with text data as the query input 202. As another example, the language model neural network 112 can include an audio neural network, where the query input 202 includes audio data. In general, the query input 202 can include multiple modalities depending on the specific architecture of the language model neural network 112.
[0078] The system 200 can process the tool selection input sequence 222 using the language model neural network 112 to generate a software tool selection output 224 that identifies a particular software tool from the set of software tools. For example, the output of the language model neural network 112 in response to processing the tool selection input sequence 222 is a sequence of token identifiers. After decoding the sequence of token identifiers, the output can be represented as text that includes an identifier (e.g., an index) that corresponds to a particular software tool, the name of a particular software tool, or a set of likelihood scores that correspond to a set of software tools. Each likelihood score represents how likely a particular software tool is to be relevant to the query input 202. Based on the output of the language model neural network 112, the system 200 determines that a particular software tool is relevant to the query input 202. In some implementations, the system 200 selects a software tool with the highest likelihood score and proceeds to interact (e.g.,transmit a web request) to the selected software tool. In some other implementations, the system 200 can provide a list of software tools with associated likelihood scores to a user via a user interface (e.g., a user interface of the client device 106) for the user to select a software tool of the set of multiple software tools based on the likelihood scores. In some other implementations, the user selects more than one software tool based on the likelihood scores of each software tool (e.g., software tools that are associated with a likelihood score above a threshold value).
[0079] A call input sequence generator 212 can generate a tool call input sequence 226 (obtaining this from the software tool selection output 224) that includes the respective software tool embedding, as retrieved from the software tool embeddings 210 maintained as part of the software tool use data, for the particular software tool and a sequence of embeddings that characterizes the query input 202. Similar to the software tool selection sequence generator 204, the call input sequence generator 212 can either append or prepend the software tool embedding to the embedded representation of the query input 202. The tool call input sequence 226 includes information of the query input 202 (via the embedded representation of the query input 202) as well as information about the requirements of the particular software tool (e.g., parameters and protocol) for the software tool. For example, an example software tool for booking an airline flight might require the tool call input sequence 226 to include an embedded representation of a range of dates, an origin location, and a destination location. In addition, the particular software tool might expect the data to be formatted according to a specific protocol (e.g., JSON or XML). Furthermore, the particular software tool might expect the tool call to be formatted based on a particular communication protocol (e.g., REST or SOAP).
[0080] The language model neural network 112 processes the software tool call input sequence 226 to generate a tool call output 228 that specifies an input to the particular software tool. The tool call output 228 is a sequence of tokens, that when decoded and converted into text, can describe an API call to the particular software tool. In some implementations, additional text processing steps applied to the decoded tool call output 228 are required before transmitting a request to the particular software tool (e.g., formatting into JSON, removing unnecessary text, etc.).
[0081] The system 200 can provide the tool call output 228 as an input to the software tool system 216 associated with the particular software tool through an API, local software command, or any other method of communicating with a software service and obtain an output 230 from the software tool system 216. The system 200 includes a response integrator218 for generating a response 220 to the query input 202 from the output 230 obtained from the particular software tool. In some implementations, the response integrator 218 uses the language model neural network 112 to generate a response to the query input 202 by processing an input sequence that is generated from the response from the particular software tool and in some implementations, the query input 202.
[0082] The integration of software tools with language model neural networks can extend the functionality of a language model neural network. For example, software tools may provide a user an ability to perform actions that can include booking a flight, purchasing an item, requesting transportation, initiating communication, or scheduling a meeting.
[0083] Turning to a specific example approach for augmenting the functionality of the language model neural network 112 through software tool integrations, the system 200 can implement a prompt tuning approach for generating and processing the software tool selection embedding 206 and the software tool embeddings 210 to be processed by the language model neural network 112. That is, prompt tuning can refer to learning an embedding, such as the software tool selection embedding 206 or the software tool embeddings 210.
[0084] Prompt tuning is an approach that includes processing a prompt with the language model neural network 112, in which the prompt includes an embedding input that relates to the query input 202 along with general instructions (i.e., a “hard prompt”), and a learned embedding input (e.g., the software tool selection embedding 206 or an embedding from the software tool embeddings 210). The learned embedding is referred to as a “soft prompt” that provides context to the language model neural network 112 about a specific task (e.g., selecting a software tool or generating a call output related to a particular software tool). In some implementations, the soft prompt relates to data that reflect properties and parameters of a specific software tool, e.g., an API endpoint, API arguments, and / or expected return values. In this disclosure, the soft prompt that corresponds to the software tool selection task is the software tool selection embedding 206 and the soft prompts that correspond to the one or more specific software tools are the software tool embeddings 210.
[0085] In some implementations, the soft prompt (e.g., the software tool selection embedding 206 or a software tool embedding of the software tool embeddings 210) is a tensor with a dimensionality of [length of soft prompt] x [language model neural network embedding length]. The system 200 can define an embedding dimensionality for the language model neural network 112. For example, smaller language model neural networks can have dimensionalities of 256 or 512. Larger language model neural networks can have much largerdimensionalities of 1024 or 2048. The “length of soft prompt” relates to a number of tokens that are embedded in the soft prompt. For example, the number of tokens that make up the soft prompt can be set to 50, 100, 150 or any number. The system 200 can learn a specific soft prompt for each task (e.g., selecting a software tool, generating tool call outputs, etc.) using a prompt tuning technique described below with respect to the description of FIG. 3 in relation to the specific tasks described in this disclosure.
[0086] In some implementations, the system 200 maintains and accesses a collection of embedding sequences in one or more databases associated with software tool use data (e.g., the software tool use data 114 of FIG. 1). In some implementations, the system 200 generates the hard prompt associated with the embedded representation of the query input 202 in response to receiving the query input 202.
[0087] FIG. 3 is a flow diagram of an example process 300 for generating an output in response to a query input to a language model neural network system with a software tools integration. For convenience, the process 300 will be described as being performed by a system of one or more computers located in one or more locations. For example, the system 200 depicted in FIG. 2, appropriately programmed in accordance with this specification, can perform the process 300.
[0088] The system maintains (302) software tool use data that includes a software tool selection embedding and a respective software tool embedding for each software tool in a set including multiple software tools. An embedding can include one or more embedding vectors.
[0089] The system receives (304) a query input. The query input may relate to one of the software tools. For example, the query input may indicate a need for a particular task to be performed that can be performed by one of the software tools. As another example, the query input may include one or more parameters to be processed by one of the software tools. As another example, the query input may include instructions for a software execution environment to execute code, e.g., execute code for interacting with and / or controlling another software system through an API call.
[0090] The query input can include a question, an instruction, a desire, or any communication that may indicate a need for the system to complete a task. Some examples of query inputs include but are not limited to the following:
[0091] In some implementations, the system determines if responding to the query input is likely to require a software tool by using a rules-based system, pattern matching, a machine learning model, or some other approach to determine whether communicating with a software tool is necessary to generate a response to the query input.
[0092] The system generates (306) a software tool selection input sequence that includes the software tool selection embedding and a first sequence of embeddings characterizing the query input. The software tool selection embedding is a sequence of one or more embedding vectors that have the same dimensionality as the embedding vectors of the first sequence of embeddings characterizing the query input. The first sequence of embeddings can be generated from an embedding layer of the language model neural network. The first sequence of embeddings can either be appended or prepended to the software tool selection embedding to generate the software tool selection input sequence. In some implementations, the software tool selection input sequence includes information in addition to the query input. Forexample, the software tool selection input sequence can include information related to a natural language request to select which software tool is most relevant to the query input.
[0093] In some implementations, the language model neural network (e.g., the language model neural network 112 of FIG. 1) processes the software tool selection input sequence and generates an output sequence of tokens, that when decoded into text, indicates that no software tool of the set of one or more software tools is relevant to the received query input, in which case the system can refrain from communicating with a software tool and can instead use only the language model neural network to generate a response to the query input.
[0094] In some implementations, the system generates the software tool selection embedding by processing training data with the language model neural network and adjusting the values of the software tool selection embedding based on an evaluated error between the training data ground truth values and the output of the language model neural network. The training data includes query inputs and ground truth labels that define which specific software tool from the set of one or more available software tools is most relevant to solving the problem or answering the question of the query input. Training data can include entries like the following:, where each training data entry includes a query input and a ground truth software tool output determined to be relevant to the specific query input. In some other implementations, the system can configure the training data in any way that includes a query input and an appropriate label.
[0095] If a new software tool is introduced, the system can generate an updated software tool selection embedding by performing the prompt tuning technique again using a new set of training data that includes examples of the new software tool and relevant queries, as described in relation to the description of FIG. 4.
[0096] The system processes (308) the software tool selection input sequence using the language model neural network to generate a software tool selection output that identifies a particular software tool from the set of software tools. In some implementations, the software tool selection output is a sequence of token identifiers that can be decoded and converted into a sequence of characters.
[0097] In some implementations, the language model neural network outputs a unique identifier that represents a specific software tool that can be used to address the respective query input. For example, the language model neural network can output the name of a software tool, a unique identifier, or a set of likelihood scores for one or more software tools from the set of one or more available software tools. Accordingly, the system can select the appropriate tool based on the distribution of likelihood scores. For example, for the query input, “Use Ticket Vendor 1, no Ticket Vendor 2 to buy concert tickets ”, the query input mentions two product names, e.g., Ticket Vendor 1 and Ticket Vendor 2. However, the target software tool of the query input appears to switch from Ticket Vendor 1 to Ticket Vendor 2. In this case, the system can configure the output of the language model neural network to include “Ticket Vendor 1”, “TixVendrl”, “Ticket Vendor 1 0.9, Ticket Vendor 2 0.5, Online Ticket Marketplace 0.2”, or any other identifying parameter or set of likelihood scores that refer to the one or more available software tools.
[0098] The system generates (310) a software tool call input sequence that includes the respective software tool embedding for the particular software tool and a second sequence of embeddings characterizing the query input.
[0099] The software tool embedding contains information related to the particular software tool that the language model neural network needs to generate an input for the system to send to the software tool. The system appends or prepends the software tool embedding to the query input sequence of embeddings. For example, to access information from the Ticket Vendor 1 tool, the system needs to understand how to send a web request to a computing resource that executes operations associated with the Ticket Vendor 1 tool, how to format the corresponding web request if applicable, and which parameters and information are required to retrieve the information from the resource that can be used to generate a response to the query input.
[0100] In some implementations, the system generates the software tool embedding for each of the one or more software tools in a similar way to the method described above for generating the software tool selection embedding. The system can receive a set of training data that includes examples of input queries and corresponding invocation instructions thatare specific to the software tool and the task implied by the query input. For example, the system can learn a software tool embedding for a Flights tool using the following training data:, where the left column represents the query input and the right column represents a ground truth response for the required parameters that should be extracted from the query input andsent to the software tool, e.g., the Flights API. In some implementations, the system performs a separate execution of the language model neural network to determine values associated with each of the requirements of a particular software tool call (e.g., required parameters, required formatting, expected return parameters, API endpoints, etc.).
[0101] As an example, if the system identifies the Flights tool to be the most appropriate software tool through the software tool selection process using the software tool selection embedding, the system can include the software tool call embedding that is trained on the Flights training data to generate the appropriate parameters and arguments to communicate with the Flights resource or API. Each software tool in the set of available software tools is configured to receive an input that includes respective values for each of one or more input parameters, like the parameters illustrated in the table above.
[0102] In some implementations, the system processes (312) the software tool call input sequence using the language model neural network to generate a software tool call output that specifies an input to the particular software tool. In some implementations, the system processes (312) the software tool call input sequence using a separate computing environment (e.g., a separate server). For example, the software tool call input sequence can be generated on a user device by, e.g., a small language model, then sent to a larger neural network model hosted on a remote computing resource that can access one or more software tools.
[0103] In some implementations, the software tool call output from the language model neural network is an output sequence that specifies respective values for each of the one or more input parameters for the particular software tool. In other words, the system can prompt the language model neural network to generate an output to mimic the right column of the table above along with the respective values extracted from the query input displayed in the left column in the table above. For example, if the query input is “Can you find me a flight from Denver to Chicago for 1 adult and 1 child in Economy Class Seating for one way on 10th September 2023 with Economy Air?”, the output of the language model neural network can include a JSON object that includes both the parameters and the respective values from the query input: {“origin”: “Denver”, “destination”: “Chicago”, “earliest_departure_date”: “2023-09-10”, “airline_codes”; “Economy”, “one_way”: true, “num_adult_passengers”: 1, “num_child_passengers”: 1, “num_infant_in_seat_passengers”: 0}. By configuring the training data to include both the expected parameters and the respective values extracted from the query input, the system can prompt the language modelneural network to format the software tool call output in a format that is understood by the Flights resource.
[0104] In some implementations, the system provides (314) the input to the particular software tool. In some implementations, the system provides the respective input set of arguments which are generated by the language model neural network for the particular software tool to the particular software tool. In some implementations, this can include an API endpoint, web request, or any other means of communicating with the particular software tool.
[0105] In some implementations, the system obtains (316) an output from the particular software tool. For example, in the case of the Flights query input, “Can you find me a flight from Denver to Chicago for 1 adult and 1 child in Economy Class Seating for one way on 10th September 2023 with Economy Air?”, where the language model neural network processed the software tool call input sequence (the software tool embedding for Flights combined with the query input embedding) to generate a software tool call {“origin”: “Denver”, “destination”: “Chicago”, “earliest_departure_date”: “2023-09-10”, “airline_codes”; “Economy”, “one_way”: true, “num_adult_passengers”: 1, “num_child_passengers”: 1, “num_infant_in_seat_passengers”: 0}, the system can obtain a list of flights that meet the requested criteria, as provided by the JSON object, from the Flights API.
[0106] In some implementations, the system generates (318) a response to the query input from the output obtained from the particular software tool. The system can process a software tool output input sequence using the language model neural network to generate an output sequence that defines the response to the query input. In other words, the system can process the output from the particular software tool with the language model neural network to generate a natural language response to the query input. For example, the system can format the list of flights that match the requested criteria from the example above in a list that can be displayed in a chat interface, spreadsheet, or any other appropriate mode of displaying or otherwise providing the data from the respective software tool.
[0107] FIG. 4 is a flow diagram of an example process 400 for updating the software tools use data. For convenience, the process 400 will be described as being performed by a system of one or more computers located in one or more locations. For example, the system 200 depicted in FIG. 2, appropriately programmed in accordance with this specification, can perform the process 400.
[0108] As described in relation to FIGS. 2-3, the software tools use data can include a software tool selection embedding and a software tool embedding for each software tool of the set of multiple software tools represented in the software tool use data.
[0109] The system receives (402) data specifying a new software tool to be added to the set of software tools. The new software tool data can include example queries that are relevant to the new software tool (e.g., example airline trip queries, example calendar event queries, etc.). In addition, the new software tool data can include data related to retrieving data from the software tool (e.g., API format requirements, parameter requirements, API endpoints, among others).
[0110] In response to receiving the new software tool data, the system generates (404) a respective software tool embedding for the new software tool.
[0111] In some implementations, to generate the software tool embedding associated with the new software tool data, the system initializes the respective software tool embedding for the new software tool using a randomized embedding, an existing embedding from a similar software tool, or any other initialization approach.
[0112] The system can receive a set of training data (e.g., training data included in the new software tool data) that is specific to the new software tool which includes multiple examples of software tool call input sequences and respective ground truth software tool call outputs for the new software tool. The system can perform prompt tuning on the respective software tool embedding for the new software tool to update the respective software tool embedding. The system can calculate a loss function, e.g., a next token prediction loss, e.g., a negative log likelihood loss, and so on, for every training step and backpropagate the errors through a language model neural network (e.g., the language model neural network 112 of FIG. 1) and only update the parameters associated with the new software tool embedding while holding the parameters of the underlying language model neural network fixed (in particular the trainable parameters, e.g., the weights).
[0113] The system updates (406) the software tools use data to include the respective software tool embedding for the new software tool. In some implementations, the system updates a database or data repository that stores the software tool use data to include the software tool embedding associated with the new software tool.
[0114] In addition to updating the software tool use data with an embedding associated with the new software tool, the system updates (408) the software tool selection embedding to account for the new software tool being added to the set of software tools. The system includes training data associated with the new software tool and determines anupdated software tool selection embedding using the prompt tuning approach. Furthermore, the system updates (410) the software tool use data to include the updated software tool selection embedding.
[0115] In some implementations, the system updates the set of training data used to learn the software tool selection embedding to include respective training data for the new software tool. For example, if the new software tool is the Online Shopping API, the system can append the training data to include entries such as “Put video game console on the Online Shopping whishlisf ’. The system can perform prompt tuning on the software tool selection embedding with the new set of training data, using the previous software tool selection embedding as the point of initialization. The system can calculate a loss function for every training step and backpropagate the errors through the language model neural network and only update the parameters associated with the software tool selection embedding while holding the parameters of the underlying language model neural network fixed. After the system learns the software tool selection embedding which generated based on training data that includes data related to the new software tool, the system can generate the most likely software tool to solve a problem, answer a question, or perform an action, in which the most appropriate software tool can be the new software tool. For example, before adding a new software tool that provides book reviews to the software tool selection embedding, an input query such as “What is the top-rated spy novel of this year?” would likely not return a suggested software tool. After the new software tool that provides book reviews is added to the software tool selection embedding, the prompt will likely fall within the set of examples provided for the new software tool, and the language model neural network will return the new software tool as a likely source of information to respond to the input query.
[0116] The above described techniques can be implemented in a distributed computing system comprising a first computer system that implements the language model neural network 112 and a second computer system that implements one or both of a tool selection neural network and a tool embedding neural network. The tool selection neural network and the tool embedding neural network can be implemented in the same or different computing systems; the latter approach can facilitate updating just parts of the system, e.g. to add or remove tools. The software tool use data can be maintained on the second computer system(s).
[0117] In such implementations the first computer system can perform one or more of the steps of: receiving the query input; processing the software tool selection input sequence; processing the software tool call input sequence; providing the input to the particularsoftware tool; obtaining the output from the particular software tool; and generating the response to the query input. The second computer system can perform one or more of the steps of maintaining the software tool use data; generating the software tool selection input sequence; and generating the software tool call input sequence. The or each particular software tool can be implemented on one or more third computer systems.
[0118] There can be a plurality of the second computer systems. For example, each of the second computer systems can be configured to support a different group of software tools, e.g. for a different respective group of users. Then each of the second computer systems can maintain different respective software tool use data (i.e. different software tool selection embeddings and or different software tool embeddings) for generating a different respective software tool selection input sequence and software tool call input sequence using each of the second computer systems.
[0119] Implementations of this type are facilitated because the software tool use data, i.e. the software tool selection embeddings and software tool embeddings can be learned, e.g. using prompt tuning, without modifying the learned parameters, e.g. weights, of the language model neural network.
[0120] FIG. 5 is an example plot 500 for comparing training methods of a language model neural network configured to select a software tool based on a query input. The plot 500 includes a horizontal axis 502 that represents a number of training steps for training a neural network and a vertical axis 504 that represents a prediction accuracy of the trained neural network. The plot 500 compares two methods of training a neural network. A first set of data values 506 represents a supervised fine-tuning training method and a second set of data values 508 represents a prompt tuning training method.
[0121] The plot 500 shows accuracy data related to a software tool selection task. For example, the evaluated task of plot 500 relates to operations associated with selecting a software tool from a set of multiple software tools. The vertical axis 504 that represents the prediction accuracy of the respective neural network for predicting (i.e., selecting) the correct software tool based on processing a particular query input and a software tool selection embedding.
[0122] The first set of data values 506 represents the supervised fine-tuning training method of the language model neural network. In other words, the training data related to software tool selection is included in the training data of the language model neural network, and all of the parameters of the network are re-learned based on the updated training data.
[0123] The second set of data values 508 represents the prompt tuning training method, in which the training process represented by the values 508 includes learning the software tool selection embedding as the soft prompt and processing the soft prompt appended to the same hard prompt used for generating the values 506 to predict a relevant software tool from the set of multiple software tools. In contrast to the data values 506, the method related to the data values 508 requires learning only the parameters of the software tool selection embedding (not the parameters of the full language model neural network).
[0124] As illustrated in the comparison of the first set of data values 506 and the second set of data values 508, the prompt tuning training method demonstrates a large accuracy improvement with fewer training steps in comparison with the supervised finetuning training method. However, the prompt tuning training method requires more training steps to reach an asymptotic accuracy level equal to an asymptotic accuracy level of the supervised fine-tuning method. Moreover, the example plot 500 compares the two training approaches, in which the supervised fine-tuning approach requires the system to learn approximately 350 million parameters, whereas the prompt tuning approach requires the system to learn approximately 75,000 parameters, demonstrating a several orders of magnitude reduction of the number of parameters to be learned.
[0125] A system trained both models that are compared in plot 500 with the same training data and provided each model with the same prompt (e.g., a prompt indicative of a query that requires a selection of a software tool). For example, a shared prompt between the two models provided to generate the comparison of plot 500 can be, “The task is to call an API tool from the following tools: ‘Tool A’, ‘Tool B’, ‘Tool C’, ..., ‘Tool N’.” The corresponding training data can include entries like, “Add cat food to my list called Tool A - Tool A,” “Open Tool A, not Tool B,” “Use Tool C and show me the status of my order from the last week,” and “Use Tool C, not Tool N to buy concert tickets.” Each training data entry includes a prompt (the portion of the example prompts before the dash) and a label (the portion of the example prompts after the dash). Other formats for displaying a training data input and a corresponding label are possible.
[0126] The system evaluates the accuracy of the neural networks trained using supervised fine-tuning and prompt tuning by determining exact matches between a training data element label (e.g., “Tool A” or “Tool B”) and an output of a respective trained network.
[0127] FIG. 6 illustrates example plots 600, 650 for comparing training methods of a language model neural network configured to select a software tool. A first example plot 600illustrates a comparison of a set of predicted software tools 602 with a set of corresponding ground truth software tools 604 as represented in training data.
[0128] The example plot 600 shows the accuracy of a language model neural network trained using the prompt tuning approach. A first gradient scale 608 indicates a degree of correlation between the set of predicted software tools 602 and the set of corresponding ground truth software tools 604.
[0129] The example plot 650 includes similar components and represents the accuracy of a language model neural network trained using the supervised fine-tuning approach. The example plots 600, 650 illustrate lightly colored squares as higher correlation data points compared to darker colored squares. A first diagonal correlation line 606 demonstrates a high correlation between the predicted tools 602 with the ground truth tools 604, as represented in the training data. A trained model that demonstrates perfect predictive capabilities would generate a plot similar to the example plot 600 with all diagonal pixels at a maximum value and all off-diagonal pixels with a zero value.
[0130] The example plots 600, 650 depict a correlation after 14,000 training steps, corresponding to a stage of training, as illustrated by the example plot 500, in which the accuracy of outputs associated with neural networks trained by supervised fine-tuning and prompt tuning converge.
[0131] As illustrated by the example plot 600, the network trained with prompt tuning depicts a high correlation between two distinct tools (e.g., Tool C of the set of predicted software tools and Tool J of the set of ground truth software tools). This correlation is depicted as a bright off-diagonal pixel. This correlation can be due to a variety of effects, including the distinct software tools sharing multiple common features. For example, Tool C and Tool J may solve the same problem, use similar technology, share common user bases, etc., such that a neural network processing a prompt-tuned input embedding confuses mentions of Tool C with Tool J in the training data. Contrastingly, the example plot 650, representing the accuracy of the supervised fine-tuned neural network, depicts fewer off- diagonal correlation elements, but depicts a higher number of “no category” predictions. The “no category” predictions occur for outputs of the trained neural network that are indicative of a tool not represented in the training data or an indication of “other”, represented as in the set of predicted software tools 602.
[0132] FIG. 7 is an example plot 700 for comparing training methods of a language model neural network configured to generate arguments to be processed by a software tool.The compared training methods include supervised fine-tuning and prompt tuning, represented by example data values 706 and 708 respectively.
[0133] The plot 700 includes a horizontal axis 702 that represents a number of training steps for training a neural network and a vertical axis 704 that represents a prediction accuracy of the trained neural network in relation to a prediction of software tool arguments.
[0134] The plot 700 includes accuracy data related to a software tool argument generation task. For example, the evaluated task of plot 700 relates to the output of the language model neural network 112 of FIG. 1. The vertical axis 704 that represents the prediction accuracy relates to the accuracy of the respective neural network for predicting the correct arguments to provide to the particular selected software tool based on processing a particular query input.
[0135] The first set of data values 706 represents the supervised fine-tuning training method of the language model neural network. In other words, the training data related to argument generation are included in the training data of the language model neural network, and all of the parameters of the network are re-learned based on the updated training data.
[0136] The second set of data values 708 represents the prompt tuning training method, in which the training process represented by the values 708 includes learning the updated software tool embedding as the soft prompt and processing the soft prompt with the same hard prompt used for generating the values 706 to predict a set of arguments to provide to the selected software tool, as indicated by the hard prompt. In contrast to the data values 706, the method related to the data values 708 does not achieve the same high asymptotic accuracy. However, the method related to the data values 708 only requires learning the embedded representation of the particular software tool, rather than re-learning all of the parameters of the language model neural network.
[0137] In some implementations, the system formats the training data for both training methods in a format that includes a target predictive value of “{tool name: intent},” in which each training data element includes one pair of tool and a corresponding intent. Other application may require multiple arguments from a list of possible arguments to be generated as output from the language model neural network. The predictive accuracy depicted in the example plot 700 can be evaluated in terms of precision, recall, fl -score, or any other metric sensitive to a correctness of the predictive outputs as related to the provided training data.
[0138] This specification describes a system implemented as computer programs, in which at least one computer program implements a language model neural network (e.g., thelanguage model neural network 112). The language model neural network is a generative neural network having parameters and that can be configured through training to process an input sequence that is made up of tokens from a vocabulary in accordance with the parameters to generate, based on the input sequence, an output sequence for a generative task that is made up of tokens from the vocabulary. For example, the input sequence can include a prompt that provides context for the output sequence.
[0139] After training, the training system or another inference system can deploy the generative neural network on one or more computing devices to perform inference for the one or more generative tasks, i.e., to generate new output sequences for the generative tasks based on new input sequences.
[0140] The vocabulary of tokens can include any of a variety of tokens that represent text symbols or other symbols. For example, the vocabulary of tokens can include one or more of characters, sub-words, words, punctuation marks, numbers, or other symbols that appear in a corpus of natural language text and / or computer code.
[0141] Additionally, or alternatively, the vocabulary of tokens can include tokens that can represent data other than text. For example, the vocabulary of tokens can include image tokens that represent a discrete set of image patch embeddings of an image that can be generated by an image encoder neural network based on processing the image patches of the image. As another example, the vocabulary of tokens can include audio tokens that represent code vectors in a codebook of a quantizer, e.g., a residual vector quantizer.
[0142] In some implementations, the generative neural network can be configured as an auto-regressive language model neural network. The language model neural network is referred to as an auto-regressive neural network when the language model neural network auto-regressively generates an output sequence of tokens by generating each particular token in the output sequence conditioned on a current input sequence that includes any (e.g., all) tokens that precede the particular token in the output sequence, i.e., tokens that have already been generated for any previous positions in the output sequence that precede the particular position of the particular token, and the input sequence.
[0143] For example, the current input sequence when generating a token at any given position in the output sequence can include the input sequence and the tokens at any preceding positions that precede the given position in the output sequence. As a particular example, the current input sequence can include the input sequence followed by the tokens at any (e.g., all) preceding positions that precede the given position in the output sequence.Optionally, the input sequence and the current output sequence can be separated by one or more predetermined tokens within the current input sequence.
[0144] More specifically, to generate a particular token at a particular position within an output sequence, the generative neural network can process the current input sequence to generate a score distribution, e.g., a probability distribution, that assigns a respective score, e.g., a respective probability, to each token in the vocabulary of tokens. The generative neural network can then select, as the particular token, a token from the vocabulary using the score distribution. For example, the generative neural network can greedily select the highest-scoring token or can sample, e.g., using nucleus sampling or another sampling technique, a token from the distribution.
[0145] As a particular example, the generative neural network can be or comprise an auto-regressive Transformer-based neural network that includes (i) a sequence comprising a plurality of attention blocks that each apply a self-attention operation and (ii) an output subnetwork that processes an output of the last attention block to generate the score distribution.
[0146] The generative neural network can have any of a variety of Transformer-based language model neural network architectures. Examples of such neural network architectures include those described in Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. Exploring the limits of transfer learning with a unified text-to-text transformer. arXiv preprint arXiv: 1910.10683, 2019; Daniel Adiwardana, Minh-Thang Luong, David R. So, Jamie Hall, Noah Fiedel, Romal Thoppilan, Zi Yang, Apoorv Kulshreshtha, Gaurav Nemade, Yifeng Lu, and Quoc V. Le. Towards a human-like open-domain chatbot. CoRR, abs / 2001.09977, 2020; Tom B Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. Language models are fewshot learners. arXiv preprint arXiv:2005.14165, 2020; Aakanksha Chowdhery, et al. PaLM: Scaling Language Modeling with Pathways, arXiv preprint arXiv: 2204.02311; Rohan Anil, et al. Palm 2 technical report. arXiv preprint arXiv:2305.10403, 2023; and Gemini Team, et al. Gemini: a family of highly capable multimodal models. arXiv preprint arXiv:2312.11805 (2023). arXiv preprint arXiv:2408.00118, 2024. arXiv preprint arXiv:2412.03555.
[0147] Generally, however, the Transformer-based language model neural network includes a sequence of attention blocks, and, during the processing of a given input sequence, each attention block in the sequence receives a respective input hidden state for each input token in the given input sequence. The attention block then updates at least the hidden statefor the last token in a given input sequence at least in part by applying self-attention to generate a respective output hidden state for the last token. The input hidden states for the first attention block are embeddings of the input tokens in the input sequence and the input hidden states for each subsequent attention block are the output hidden states generated by the preceding attention block.
[0148] In this example, the output subnetwork processes the output hidden state generated by the last attention block in the sequence for the last input token in the input sequence to generate the score distribution.
[0149] Additional approaches for implementing a language model neural network include the RecurrentGemma family of language models that combine linear recurrences with local attention, as described in arXiv preprint arXiv:2404.07839 (2024). Similarly, the Griffin neural network model, as described in arXiv preprint arXiv: 2402.19427, is a hybrid model that mixes gated linear recurrences with local attention. Text diffusion models, as described in arXiv preprint arXiv:2211.15089, represent discrete language variables in a continuous space.
[0150] As an example, the generative neural network can generate text sequences, i.e., each output sequence generated by the generative neural network is a sequence of text tokens from a vocabulary of text tokens that includes, e.g., one or more of characters, subwords, words, punctuation marks, numbers, or other symbols that appear in natural language text. For example, the inference system can use the generative neural network to generate text sequences and provide the text sequences for presentation to users.
[0151] As another example, the generative neural network can generate images or videos that have multiple frames (where each frame is an image) by generating images, e.g., either as sequences of pixels or through an iterative denoising process. For example, the output sequence generated by the generative neural network includes a plurality of color values for pixels in an image arranged according to a specified order. As another example, the output sequence generated by the generative neural network includes a plurality of tokens that represent image patch embeddings of an image which can then be processed by a decoder neural network to generate the image. For example, the inference system can use the generative neural network to generate an image or a video conditioned on an input sequence that includes a text description of the content of the image or the video.
[0152] As another example, the input sequence is a sequence of text and the output sequence is another sequence of text, e.g., a completion of the input sequence of text, a paraphrase of the input sequence of text, a response to a question posed in the input sequence,or a sequence of text that is about a topic specified by the input sequence of text. As another example, the input sequence can be an input other than text, e.g., a plurality of pixels included in an image, and the output sequence can be a text sequence that describes the input.
[0153] As another example, the input sequence represents data to be compressed, e.g., image data, text data, audio data, or any other type of data; and the output sequence is a compressed version of the data. The tokens included in the output sequence can include any representation of compressed data, e.g., symbols or embeddings to be decoded by a respective neural network.
[0154] As a particular example, the inference system can be part of a dialog system and the input sequence can include audio or text from the most recent conversational turn submitted by a user of the dialog system during the dialog while the output sequence is the next turn in the conversation, e.g., either text or audio that is a response to the most recent conversational turn. Optionally, the input sequence can also include one or more historical conversational turns that occurred earlier in the conversation.
[0155] As another particular example, the inference system can be part of a machine translation system and the input sequence can include text in a source language while the output sequence can include text in a target language that is a translation of the source text into the target language.
[0156] As another particular example, the inference system can be part of a natural language processing system. For example, if the input sequence is a sequence of words in an original language, e.g., a sentence or phrase, the output sequence can be a summary of the input sequence in the original language, i.e., a sequence that has fewer words than the input sequence but that retains the essential meaning of the input sequence. As another example, if the input sequence is a sequence of words that form a question, the output sequence can be a sequence of words that form an answer to the question.
[0157] As another particular example, the inference system can be part of a computer- assisted medical diagnosis system. For example, the input sequence can be a sequence of data from an electronic medical record and the output sequences can each be a sequence of predicted treatments.
[0158] As another particular example, the inference system can be part of a computer code generation system and the input sequence can include a text description of a desired piece of code or a snippet of computer code in a programming language and the output sequence can include computer code, e.g., a snippet of code that is described by the input sequence or a snippet of code that follows the input sequence in a computer program.
[0159] As another particular example, the inference system can be part of a multimodal system that processes multi-modal input sequences, e.g., both text and image input sequences, or both text and audio input sequences, and generates the output sequences that are either in a single data modality or in multiple data modalities, e.g., combinations of two or more of text, image and audio output sequences. Examples of such multi-modal systems include an image captioning system, a text-based image search system, an image-based question answering system, and so on.
[0160] As another particular example, the inference system can be part of a digital agent, in which an agent autonomously performs specific tasks or assists users by leveraging outputs from the inference system. In addition, the inference system can process live streaming data between the digital agent and a user.
[0161] As another particular example, the inference system can be part of or associated with a robotic control system, i.e., a system for controlling one or more mechanical agents. The input sequence can comprise a natural language description of one or more tasks for a the one or more mechanical agents and the output sequence can comprise a sequence of instructions (e.g., joint angles, torques, velocities, etc.) for the one or more mechanical agents that cause the one or more mechanical agents to perform the one or more tasks described in the input sequence.
[0162] In a similar example, the inference system can be part of or associated with a control system in a manufacturing environment for manufacturing a product, i.e., a system for controlling a manufacturing unit or a machine that operates to manufacture the product. In another similar example, the inference system can be part of or associated with a control system in a service facility comprising a plurality of items of electronic equipment.
[0163] As another particular example, the inference system can be part of or associated with a search system that facilitates searching of resources on the Internet. A resource can be any data that can be provided over the Internet. A resource can be identified by a resource address that is associated with the resource. Resources include web pages, word processing documents, portable document format (PDF) documents, images, video, and news feed sources, to name a few.
[0164] In this particular example, the search system can receive search queries submitted by client devices and, in response, identify resources that are relevant to the search query in the form of search results and return the search results to the user devices in search results pages. A search result page can include search result data generated by the search system that identifies a resource responsive to a search query and includes a link to theresource. The search result page can additionally include a result in the form of an output sequence that is generated by the inference system based on an input sequence derived from the search query.
[0165] The generative neural network is typically trained using a multi-stage approach: a pre-training stage followed by a fine-tuning stage, where at least the fine-tuning stage takes place at a training system. For example, a training system can receive data specifying a pre-trained generative neural network from another system, and then perform the fine-tuning of the pre-trained generative neural network.
[0166] In the pre-training stage, the generative neural network is pre-trained by the training system or another system based on optimizing one or more unsupervised or selfsupervised objective functions, e.g., a maximum-likelihood objective function, on one or more large datasets and then, in some cases, adjusted to the generative tasks, which can include any combination of one or more of the generative tasks mentioned below and possibly other tasks, through fine-tuning adaptation based on supervised learning, reinforcement learning from human feedback (RLHF), reinforcement learning from Al feedback (RLAIF), prompt tuning, instruction tuning, and the like, that use different training objectives, different datasets, or both.
[0167] The one or more large datasets used during the pre-training stage can include a large dataset of text in one or more natural languages, e.g., text that is publicly available from the Internet or another text corpus, a large dataset of computer code in one or more programming languages, e.g., Python, C++, C#, Java, Ruby, PHP, and so on, e.g., computer code that is publicly available from the Internet or another code repository, a large dataset of audio samples, e.g., audio recordings or waveforms that represent the audio recordings, a large dataset of images where each image includes an array of pixels, a large dataset of videos where each video includes a temporal sequence of frames, or a large multi-modal dataset that includes a combination of two or more of these datasets.
[0168] In this specification, the term "configured" is used in relation to computing systems and environments, as well as computer program components. A computing system or environment is considered "configured" to perform specific operations or actions when it possesses the necessary software, firmware, hardware, or a combination thereof, enabling it to carry out those operations or actions during operation. For instance, configuring a system might involve installing a software library with specific algorithms, updating firmware with new instructions for handling data, or adding a hardware component for enhanced processing capabilities. Similarly, one or more computer programs are "configured" to performparticular operations or actions when they contain instructions that, upon execution by a computing device or hardware, cause the device to perform those intended operations or actions.
[0169] The embodiments and functional operations described in this specification can be implemented in various forms, including digital electronic circuitry, software, firmware, computer hardware (encompassing the disclosed structures and their structural equivalents), or any combination thereof. The subject matter can be realized as one or more computer programs, essentially modules of computer program instructions encoded on a tangible non- transitory storage medium for execution by or to control the operation of a computing device or hardware. The storage medium can be a storage device such as a hard drive or solid-state drive (SSD), a storage medium, a random or serial access memory device, or a combination of these. Additionally or alternatively, the program instructions can be encoded on a transmitted signal, such as a machine-generated electrical, optical, or electromagnetic signal, designed to carry information for transmission to a receiving device or system for execution by a computing device or hardware. Furthermore, implementations may leverage emerging technologies like quantum computing or neuromorphic computing for specific applications, and may be deployed in distributed or cloud-based environments where components reside on different machines or within a cloud infrastructure.
[0170] The term "computing device or hardware" refers to the physical components involved in data processing and encompasses all types of devices and machines used for this purpose. Examples include processors or processing units, computers, multiple processors or computers working together, graphics processing units (GPUs), tensor processing units (TPUs), and specialized processing hardware such as field-programmable gate arrays (FPGAs) or application-specific integrated circuits (ASICs). In addition to hardware, a computing device or hardware may also include code that creates an execution environment for computer programs. This code can take the form of processor firmware, a protocol stack, a database management system, an operating system, or a combination of these elements. Embodiments may particularly benefit from utilizing the parallel processing capabilities of GPUs, in a General-Purpose computing on Graphics Processing Units (GPGPU) context, where code specifically designed for GPU execution, often called kernels or shaders, is employed. Similarly, TPUs excel at running optimized tensor operations crucial for many machine learning algorithms. By leveraging these accelerators and their specialized programming models, the system can achieve significant speedups and efficiency gains fortasks involving artificial intelligence and machine learning, particularly in areas such as computer vision, natural language processing, and robotics.
[0171] A computer program, also referred to as software, an application, a module, a script, code, or simply a program, can be written in any programming language, including compiled or interpreted languages, and declarative or procedural languages. It can be deployed in various forms, such as a standalone program, a module, a component, a subroutine, or any other unit suitable for use within a computing environment. A program may or may not correspond to a single file in a file system and can be stored in various ways. This includes being embedded within a file containing other programs or data (e.g., scripts within a markup language document), residing in a dedicated file, or distributed across multiple coordinated files (e.g., files storing modules, subprograms, or code segments). A computer program can be executed on a single computer or across multiple computers, whether located at a single site or distributed across multiple sites and interconnected through a data communication network. The specific implementation of the computer programs may involve a combination of traditional programming languages and specialized languages or libraries designed for GPGPU programming or TPU utilization, depending on the chosen hardware platform and desired performance characteristics.
[0172] In this specification, the term "engine" broadly refers to a software-based system, subsystem, or process designed to perform one or more specific functions. An engine is typically implemented as one or more software modules or components installed on one or more computers, which can be located at a single site or distributed across multiple locations. In some instances, one or more dedicated computers may be used for a particular engine, while in other cases, multiple engines may operate concurrently on the same one or more computers. Examples of engine functions within the context of Al and machine learning could include data pre-processing and cleaning, feature engineering and extraction, model training and optimization, inference and prediction generation, and post-processing of results. The specific design and implementation of engines will depend on the overall architecture and the distribution of computational tasks across various hardware components, including CPUs, GPUs, TPUs, and other specialized processors.
[0173] The processes and logic flows described in this specification can be executed by one or more programmable computers running one or more computer programs to perform functions by operating on input data and generating output. Additionally, graphics processing units (GPUs) and tensor processing units (TPUs) can be utilized to enable concurrent execution of aspects of these processes and logic flows, significantly acceleratingperformance. This approach offers significant advantages for computationally intensive tasks often found in Al and machine learning applications, such as matrix multiplications, convolutions, and other operations that exhibit a high degree of parallelism. By leveraging the parallel processing capabilities of GPUs and TPUs, significant speedups and efficiency gains compared to relying solely on CPUs can be achieved. Alternatively or in combination with programmable computers and specialized processors, these processes and logic flows can also be implemented using specialized processing hardware, such as field-programmable gate arrays (FPGAs) or application-specific integrated circuits (ASICs), for even greater performance or energy efficiency in specific use cases.
[0174] Computers capable of executing a computer program can be based on general- purpose microprocessors, special-purpose microprocessors, or a combination of both. They can also utilize any other type of central processing unit (CPU). Additionally, graphics processing units (GPUs), tensor processing units (TPUs), and other machine learning accelerators can be employed to enhance performance, particularly for tasks involving artificial intelligence and machine learning. These accelerators often work in conjunction with CPUs, handling specialized computations while the CPU manages overall system operations and other tasks. Typically, a CPU receives instructions and data from read-only memory (ROM), random access memory (RAM), or both. The elements of a computer include a CPU for executing instructions and one or more memory devices for storing instructions and data. The specific configuration of processing units and memory will depend on factors like the complexity of the Al model, the volume of data being processed, and the desired performance and latency requirements. Embodiments can be implemented on a wide range of computing platforms, from small embedded devices with limited resources to large- scale data center systems with high-performance computing capabilities. The system may include storage devices like hard drives, SSDs, or flash memory for persistent data storage.
[0175] Computer-readable media suitable for storing computer program instructions and data encompass all forms of non-volatile memory, media, and memory devices.Examples include semiconductor memory devices such as read-only memory (ROM), solid- state drives (SSDs), and flash memory devices; hard disk drives (HDDs); optical media; and optical discs such as CDs, DVDs, and Blu-ray discs. The specific type of computer-readable media used will depend on factors such as the size of the data, access speed requirements, cost considerations, and the desired level of portability or permanence.
[0176] To facilitate user interaction, embodiments of the subject matter described in this specification can be implemented on a computing device equipped with a display device,such as a liquid crystal display (LCD) or an organic light-emitting diode (OLED) display, for presenting information to the user. Input can be provided by the user through various means, including a keyboard), touchscreens, voice commands, gesture recognition, or other input modalities depending on the specific device and application. Additional input methods can include acoustic, speech, or tactile input, while feedback to the user can take the form of visual, auditory, or tactile feedback. Furthermore, computers can interact with users by exchanging documents with a user's device or application. This can involve sending web content or data in response to requests or sending and receiving text messages or other forms of messages through mobile devices or messaging platforms. The selection of input and output modalities will depend on the specific application and the desired form of user interaction.
[0177] Machine learning models can be implemented and deployed using machine learning frameworks, such as TensorFlow or JAX. These frameworks offer comprehensive tools and libraries that facilitate the development, training, and deployment of machine learning models.
[0178] Embodiments of the subject matter described in this specification can be implemented within a computing system comprising one or more components, depending on the specific application and requirements. These may include a back-end component, such as a back-end server or cloud-based infrastructure; an optional middleware component, such as a middleware server or application programming interface (API), to facilitate communication and data exchange; and a front-end component, such as a client device with a user interface, a web browser, or an app, through which a user can interact with the implemented subject matter. For instance, the described functionality could be implemented solely on a client device (e.g., for on-device machine learning) or deployed as a combination of front-end and back-end components for more complex applications. These components, when present, can be interconnected using any form or medium of digital data communication, such as a communication network like a local area network (LAN) or a wide area network (WAN) including the Internet. The specific system architecture and choice of components will depend on factors such as the scale of the application, the need for real-time processing, data security requirements, and the desired user experience.
[0179] The computing system can include clients and servers that may be geographically separated and interact through a communication network. The specific type of network, such as a local area network (LAN), a wide area network (WAN), or the Internet, will depend on the reach and scale of the application. The client-server relationship isestablished through computer programs running on the respective computers and designed to communicate with each other using appropriate protocols. These protocols may include HTTP, TCP / IP, or other specialized protocols depending on the nature of the data being exchanged and the security requirements of the system. In certain embodiments, a server transmits data or instructions to a user's device, such as a computer, smartphone, or tablet, acting as a client. The client device can then process the received information, display results to the user, and potentially send data or feedback back to the server for further processing or storage. This allows for dynamic interactions between the user and the system, enabling a wide range of applications and functionalities.
[0180] While this specification contains many specific implementation details, these should not be construed as limitations on the scope of any invention or on the scope of what may be claimed, but rather as descriptions of features that may be specific to particular embodiments of particular inventions. Certain features that are described in this specification in the context of separate embodiments can also be implemented in combination in a single embodiment. Conversely, various features that are described in the context of a single embodiment can also be implemented in multiple embodiments separately or in any suitable subcombination. Moreover, although features may be described above as acting in certain combinations and even initially be claimed as such, one or more features from a claimed combination can in some cases be excised from the combination, and the claimed combination may be directed to a subcombination or variation of a subcombination.
[0181] Similarly, while operations are depicted in the drawings and recited in the claims in a particular order, this should not be understood as requiring that such operations be performed in the particular order shown or in sequential order, or that all illustrated operations be performed, to achieve desirable results. In certain circumstances, multitasking and parallel processing may be advantageous. Moreover, the separation of various system modules and components in the embodiments described above should not be understood as requiring such separation in all embodiments, and it should be understood that the described program components and systems can generally be integrated together in a single software product or packaged into multiple software products.
[0182] Particular embodiments of the subject matter have been described. Other embodiments are within the scope of the following claims. For example, the actions recited in the claims can be performed in a different order and still achieve desirable results. As one example, the processes depicted in the accompanying figures do not necessarily require theparticular order shown, or sequential order, to achieve desirable results. In some cases, multitasking and parallel processing may be advantageous.
Claims
CLAIMS1. A method performed by one or more computers, the method comprising: maintaining software tool use data that comprises a software tool selection embedding and a respective software tool embedding for each software tool in a set comprising a plurality of software tools; receiving a query input; generating a software tool selection input sequence that comprises the software tool selection embedding and a first plurality of embeddings characterizing the query input; processing the software tool selection input sequence using a language model neural network to generate a software tool selection output that identifies a particular software tool from the set of software tools; and generating a software tool call input sequence that comprises the respective software tool embedding for the particular software tool and a second plurality of embeddings characterizing the query input.
2. The method of claim 1, further comprising: processing the software tool call input sequence using the language model neural network to generate a software tool call output that specifies an input to the particular software tool; providing the input to the particular software tool; obtaining an output from the particular software tool; and generating a response to the query input from the output obtained from the particular software tool.
3. The method of any preceding claim, further comprising: receiving data specifying a new software tool to be added to the set of software tools; and in response to receiving the data specifying the new software tool to be added to the set of software tools: generating a respective software tool embedding for the new software tool; and updating the software tool use data to include the respective software toolembedding for the new software tool.
4. The method of claim 3, wherein generating a respective software tool embedding for the new software tool comprises: initializing the respective software tool embedding for the new software tool; receiving first training data comprising a plurality of training software tool call input sequences for the new software tool and, for each training software tool call input sequence, a respective ground truth software tool call output for the new software tool; and performing prompt tuning on the respective software tool embedding for the new software tool using the first training data to update the respective software tool embedding for the new software tool.
5. The method of claim 3 or claim 4, further comprising: in response to receiving the data specifying the new software tool to be added to the set of software tools: updating the software tool selection embedding to account for the new software tool being added to the set of software tools; and updating the software tool use data to include the updated software tool selection embedding.
6. The method of claim 5, wherein updating the software tool selection embedding to account for the new software tool being added to the set of software tools comprises: receiving second training data comprising a plurality of training software tool selection input sequences and, for each training software tool selection input sequence, data identifying a respective ground truth software tool that should be selected in response to the training software tool selection input sequence, wherein, for at least a subset of the training software tool selection input sequences, the respective ground truth software tool is the new software tool; and performing prompt tuning on the software tool selection embedding using the training data to update the respective software tool embedding for the new software tool.
7. The method of any preceding claim, wherein the software tool selection embedding has been learned through prompt tuning while holding the language model neural networkfixed.
8. The method of any preceding claim, wherein the respective software tool embeddings for the software tools in the set of software tools have been learned through prompt tuning while holding the language model neural network fixed.
9. The method of any preceding claim, wherein the software tool selection output is an output sequence that specifies an identifier for the particular software tool.
10. The method of any one of claims 1-8, wherein the software tool selection output is a respective likelihood score for each of one or more of the software tools, and wherein the method further comprising: selecting the particular software tool based on the respective likelihood scores.
11. The method of any preceding claim, wherein each software tool receives an input that comprises respective values for each of one or more input parameters for the software tool, and wherein the software tool call output is an output sequence that specifies respective particular values for each of the one or more input parameters for the particular software tool.
12. The method of any preceding claim, wherein generating a response to the query input from the output obtained from the particular software tool comprises: generating a software tool output input sequence from at least the output obtained from the particular software tool; and processing the software tool output input sequence using the language model neural network to generate an output sequence that defines the response to the query input.
13. The method of any of claims 1-12, implemented in a distributed computing system comprising a first computer system that implements the language model neural network and a second computer system that implements one or both of a tool selection neural network and a tool embedding neural network, the method further comprising: maintaining the software tool use data on the second computer system; and wherein the first computer system performs the steps of: receiving the query input; processing the software tool selection input sequence;processing the software tool call input sequence; providing the input to the particular software tool; obtaining the output from the particular software tool; and generating the response to the query input; and wherein the second computer system performs the steps of: maintaining the software tool use data; generating the software tool selection input sequence; and generating the software tool call input sequence.
14. The method of claim 13, comprising a plurality of the second computer systems, wherein each of the second computer systems is configured to support a different group of software tools for a different respective group of users, and wherein each of the second computer systems maintains different respective software tool use data for generating a different respective said software tool selection input sequence and software tool call input sequence.
15. The method of claim 13 or 14, further comprising: learning the software tool use data without modifying learned parameters of the language model neural network.
16. The method of any of claims 1-15, wherein the software tools include at least one image processing tool, wherein the query input defines an image processing task, and wherein the response to the query input defines a result of the image processing task.
17. The method of any of claims 1-16, wherein the software tools include at least one materials design tool for designing a physical material, wherein the query input defines a materials design task, and wherein the response to the query input defines a result of the materials design task.
18. The method of any of claims 1-17, wherein the software tools include at least one robot control tool for controlling a mechanical robot to perform a physical task, wherein the query input defines a physical task, and wherein the response to the query input defines a result of the physical task.
19. A method performed by one or more computers for generating a respective software tool embedding for a new software tool to be added to software tool use data, the software tool use data comprising a software tool selection embedding and a respective software tool embedding for each software tool in a set comprising a plurality of software tools, the method comprising: initializing the respective software tool embedding for the new software tool; receiving first training data comprising a plurality of training software tool call input sequences for the new software tool and, for each training software tool call input sequence, a respective ground truth software tool call output for the new software tool; and performing prompt tuning on the respective software tool embedding for the new software tool using the first training data to update the respective software tool embedding for the new software tool.
20. The method of claim 19, further comprising: in response to receiving data specifying the new software tool to be added to the set of software tools: updating the software tool selection embedding to account for the new software tool being added to the set of software tools; and updating the software tool use data to include the updated software tool selection embedding.
21. One or more non-transitory computer readable media storing embeddings generated by performing operations comprising: maintaining software tool use data that comprises a software tool selection embedding and a respective software tool embedding for each software tool in a set comprising a plurality of software tools; receiving a query input; generating a software tool selection input sequence that comprises the software tool selection embedding and a first plurality of embeddings characterizing the query input; processing the software tool selection input sequence using a language model neural network to generate a software tool selection output that identifies a particular software tool from the set of software tools; and generating a software tool call input sequence that comprises the respective softwaretool embedding for the particular software tool and a second plurality of embeddings characterizing the query input.
22. A system comprising one or more computers and one or more storage devices storing instructions that when executed by the one or more computers cause the one or more computers to perform the operations of the respective method of any one of claims 1 to 20.
23. One or more non-transitory computer readable media storing instructions that when executed by one or more computers cause the one or more computers to perform the operations of the respective method of any one of claims 1 to 20.
Citation Information
Patent Citations
Retrosynthesis using neural networks
WO2021263238A1