Guide conversations using language generation neural networks and searches

By combining trained language generation neural networks and response selection neural networks, and optimizing search strategies, the communication efficiency problems caused by inaccuracy and frequent searches of language generation neural networks are solved, and efficient and accurate information acquisition and resource conservation are achieved.

CN120283238APending Publication Date: 2025-07-08GDM HOLDINGS LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202380067724.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-09-20
Filing Date
2023-09-20
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

In the prior art, language generation neural networks may be inaccurate when providing responses without conducting external searches. Frequent external searches will lead to low efficiency in communication bandwidth utilization, and frequent updates of neural network parameters will lead to computational and bandwidth penalties.

Method used

Adopt a hybrid solution, combining trained languages to generate neural networks and responsive selection neural networks, perform only limited searches, reduce communication bandwidth usage, and optimize search volumes through subsequent requests, while using violation detection neural networks to limit unsafe or undesirable behavior.

Benefits of technology

It improves the accuracy and efficiency of information acquisition, reduces the demand for communication bandwidth and computing resources, and limits the access of undesired information, providing access to a larger information database.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120283238A_ABST
    Figure CN120283238A_ABST
Patent Text Reader

Abstract

Methods, systems, and apparatus, including computer programs encoded on a computer storage medium, for enabling a user to make a conversation. Implementations of the system learn when depends on support evidence obtained from an external search system via a search system interface, and can also generate replies for the user that conform to preferences of a previously trained response selection neural network. Implementations of the system may also use a previously trained violation detection neural network to generate replies that take into account previously learned rules.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross - reference to related applications

[0002] This application claims the benefit of priority of U.S. Provisional Application Serial No. 63 / 408,430, filed on September 20, 2022, the entire content of which is incorporated herein by reference. Background Art

[0003] This specification relates to methods for conducting conversations using language - generating neural networks and search, and more particularly to implementations that follow a set of rules. These methods can be used to obtain information and to control real - world systems.

[0004] A neural network is a machine - learning model that uses one or more layers of non - linear units to predict an output for received inputs. In addition to an output layer, some neural networks include one or more hidden layers. The output of each hidden layer serves as an input to the next layer in the network (i.e., the next hidden layer or the output layer). Each layer of the network generates an output from the received inputs based on the current values of a corresponding set of parameters. Summary of the Invention

[0005] This specification describes a system implemented as a computer program on one or more computers in one or more locations that enables a user to conduct conversations using a language - generating neural network, particularly to obtain information.

[0006] In an implementation, the system can provide information to the user based on knowledge stored in a trained language - generating neural network or by supplementing that knowledge by leveraging one or more external searches, thereby balancing computational requirements with the communication bandwidth needed to search for relevant information.

[0007] The users of the system can be human users or machines. Some implementations of the system can be used by humans to maintain a general conversation with a computer system. Some implementations of the system can be used to diagnose technical failures in mechanical or computer systems or networks. Some implementations of the system can be used for natural - language control of tasks in a real - world environment, in which case the information obtained can be used to control tasks performed, for example, by a mechanical or computer system.

[0008] In one aspect, a method and a corresponding system implemented by one or more computers are described, particularly for enabling a user to obtain information through a conversation. The conversation is between the user and an agent that includes a first trained language - generating neural network (e.g., a suitably programmed computer system).

[0009] The implementation of the system learns when to rely on supportive evidence obtained from an external search system via a search system interface and thus learns when to provide a "supported" response rather than an "unsupported" response that does not rely on external search. The implementation of the system is also capable of generating responses that conform to the preferences of a previously trained response selection neural network for the user. The implementation of the system can also use a previously trained violation detection neural network to generate responses that take into account previously learned rules.

[0010] In another aspect, a method implemented by one or more computers and a corresponding system are described for training a system of the above type, particularly for training a dialogue neural network system to enable a user to use an agent including a first language generation neural network, for example, to obtain information through a dialogue between the user and the agent.

[0011] A trained machine learning computer system for enabling a user to obtain information through a dialogue is also described, the trained machine learning computer system including a trained first language generation neural network, a trained response selection neural network, and optionally a trained violation detection neural network.

[0012] A dialogue training computer system for training a dialogue computer system is also described. The language generation / language model neural network can be stored on a training computing device, and the search system can be remote from the device.

[0013] The subject matter described in this specification can be implemented in a particular embodiment so as to achieve one or more of the following advantages.

[0014] The above "unsupported" responses can be generated without an external search. Thus, these responses can be provided locally by the language generation neural network, but this is not always optimal. For example, sometimes the language generation neural network may provide an incorrect response; or the response may require information more recent than that available when training the language generation neural network. One solution could be to search extensively before providing a response, but this would be an inefficient use of communication bandwidth, for example, when communicating with a remote server. Another solution could be not to search, which also has the aforementioned drawbacks. Another solution could be to retrain the language generation neural network every once in a while, but this is computationally inefficient and transmitting updated neural network parameters to the user incurs a significant bandwidth penalty. The described hybrid solution uses a combination of a trained language generation neural network and a response selection neural network to facilitate only limited searching, thereby reducing the use of available communication bandwidth.

[0015] In a complementary manner, the described hybrid solution reduces the need to transmit large amounts of data, which would otherwise be required for frequent updates of a trained language generation neural network if it alone were relied upon to provide factual information. Compared to situations where one might rely solely on a trained language generation neural network or solely on search to obtain information, the described hybrid solution can also provide access to a larger corpus of information.

[0016] In an implementation, the ability to pose subsequent requests in the same context as the initial request further helps to limit the amount of search per round of conversation, i.e., the total amount of search in two rounds can be less than the amount of search that might be required if all searches were performed in response to a single search query. Thus, the ability to pose subsequent requests has the technical purpose of reducing the number of search requests and has the potential to reduce the execution time of the process for obtaining a request response.

[0017] Further, in an implementation, multiple rules can be used to filter responses based on any desired or undesired attributes. For example, one or more rules can be defined to reduce the likelihood of the system exhibiting unsafe, undesired, or inefficient behavior, or to reduce the likelihood that the information provided by the system includes offensive content, misinformation, or confidential or private information (such as personal contact information). As another example, in the case where a trained language generation neural network trained on multiple document corpora or other information is subsequently used to provide information to a user (e.g., by answering questions), in principle all the information used for training is available. One or more rules can be implemented to restrict access to specific types of information, e.g., to restrict access to information by a specific user. Thus, some implementations of the system can be used to restrict access to privacy data, sensitive data, or other data (e.g., copyright-protected data).

[0018] Details of one or more embodiments of the subject matter of this specification are set forth in the accompanying drawings and the description below. Other features, aspects, and advantages of the subject matter of the invention will become apparent from the specification, the drawings, and the claims. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] Figure 1 An example of a dialogue system is shown.

[0020] Figure 2 A flowchart of a first example process for using a dialogue system is shown.

[0021] Figure 3 An example user interface for a dialogue system is shown.

[0022] Figure 4 A second example process for using a dialogue system is schematically shown.

[0023] Figure 5 is a flowchart of a first example process for training a dialogue system.

[0024] Figure 6 illustrates an example neural network architecture for a dialogue system.

[0025] Figure 7 schematically illustrates a second example process for training a dialogue system.

[0026] Figure 8 shows the performance of an example implementation of a dialogue system.

[0027] In the various figures, the same reference numerals and names indicate the same elements. DETAILED DESCRIPTION

[0028] Figure 1 illustrates an example of a dialogue system 100. The dialogue system 100 is an example of a system implemented as a computer program on one or more computers in one or more locations, where the systems, components, and techniques described below can be implemented.

[0029] The dialogue system 100 can enable a user to obtain information through a dialogue. For example, the dialogue system 100 can be used as an agent to enable a dialogue between the agent (i.e., the system) and a user of the system 100 (e.g., a human user). Later, the dialogue system 100 is sometimes referred to as an agent and sometimes as "Sparrow" ( "Sparrow" is a specific example implementation).

[0030] The dialogue system 100 includes a language generation neural network 110. When used in inference, the language generation neural network 110 is a trained language generation neural network 110. During training, the language generation neural network 110 is (further) trained by a training engine 150, such as fine-tuned. After training, the training engine 150 is no longer needed.

[0031] The language generation neural network 110 is later referred to as the first language generation neural network because in the implementation, there can be other language generation neural networks or language model neural networks in the dialogue system 100.

[0032] In the implementation, the dialogue system 100 communicates with a human user using spoken or written text in natural language, but generally, the language generation neural network 110 can also generate text in a computer language (any form of language used to communicate with a computer), such as a markup language, or a command or configuration language, or a data exchange language such as JSON, or a programming language.

[0033] A language generation or language model neural network as described herein can include a sequence-to-sequence model that receives an input sequence of natural language tokens and generates an output sequence of natural language tokens. Generally, natural language tokens define words or chunks of words (e.g., word fragments or morphemes), but they can also define letters, numbers, or characters, or multiple words. These tokens can include tokens representing punctuation. In some implementations, the generation of the output sequence of natural language tokens is done one word or chunk at a time, e.g., until one or more end-of-statement tokens are obtained, or until the maximum length of the output has been generated. A trained language generation neural network can be off-the-shelf or can be trained using a text corpus, e.g., using supervised learning with maximum likelihood loss.

[0034] Generally, any language generation neural network can be used as one of the language generation or language model neural networks described herein, e.g., an autoregressive language generation neural network, or a language generation neural network that does not rely on an autoregressive model, such as a recurrent language generation neural network or a denoising autoencoder-based language model. In some implementations, the language generation / model neural network can be a mixture-of-experts model.

[0035] As an example, the language generation neural network described herein can be a Transformer-based language model neural network, particularly an autoregressive Transformer-based language model neural network. The Transformer neural network can be characterized as having a series of self-attention neural network layers. The self-attention neural network layers have attention layer inputs for each element of the input and are configured to apply an attention mechanism to the attention layer inputs to generate an attention layer output for each element of the input; there are many different attention mechanisms that can be used.

[0036] In addition to the language generation neural network 110, the dialogue system 100 can also include one or more language model neural networks. The language model neural network is similar to the language generation neural network but does not need to generate a language output. For example, it can process an input sequence of natural language tokens to generate a vector or scalar output instead of generating an output sequence of natural language tokens. Since the language generation neural network actually includes the language model neural network, these two phrases can be used interchangeably to some extent.

[0037] Surprisingly but well-documented, so-called large language model (language generation) neural networks can perform tasks they were not explicitly trained to perform. For example, they can perform translation tasks (provided the training corpus includes words in different languages), arithmetic operations, and many other tasks.

[0038] A language generation neural network can be made to perform a specific task by providing a natural language description of the desired response as an input or "prompt". The prompt can be a few-shot prompt, where a few (e.g., 1 to 10) query examples and an example output are provided in the text before the actual query.

[0039] Additionally or alternatively, a language model (language generation) neural network can be "fine-tuned" to perform a specific task by obtaining a pre-trained language model neural network trained on a large corpus of examples and then further training part or all of the language model neural network on a relatively small number of examples specific to the type of task to be performed. Thus, for example, the trained language model neural network can perform control and diagnostic tasks of the type described later.

[0040] Some implementations of the methods / systems described herein use large language models / language generation neural networks. Such large language models / language generation neural networks can have more than 1 billion, 10 billion, or 100 billion trainable / trained parameters. It can have been trained on more than 10 billion, 100 billion, or 1000 billion words or tokens representing words.

[0041] Generally, the language model neural networks and language generation neural networks described herein are all trained neural networks. When used in the training methods described later, they can be further trained, or "fine-tuned"; for example, the language generation neural network 110 can be fine-tuned using reinforcement learning.

[0042] The various different language generation neural networks and language model neural networks described herein can but need not include different instances of the same language model / language generation neural network. By way of example only, they can each include an instance of the (trained) Chinchilla model (Hoffman et al., 2022, arXiv:2203.15556). As another example, one or more of these models can use LaMDA (Thoppilan et al., 2022, arXiv:2201.08239). Optionally, one or more of these models can be fine-tuned, e.g., using supervised fine-tuning. For example, when trained using the reinforcement learning described later, one or more may have been previously fine-tuned on some of the same data used for reinforcement learning.

[0043] In an implementation, the (first) language generation neural network 110 is configured to process a context input that includes one or more prompts, each prompt including one or more natural language statements formatted in any suitable manner.

[0044] Generally speaking, a reference to a language generation neural network or a language model neural network that processes natural language text refers to a language generation neural network or a language model neural network that processes text in token form; the text in token form can be obtained from a tokenizer such as SentencePiece.

[0045] The language generation neural network 110 processes a context input according to first language generation neural network parameters to generate a natural language output, for example, an output including one or more natural language statements. The natural language output can be obtained through a sampling process such as nucleus sampling, that is, it can be randomly generated.

[0046] As an example, the language generation neural network 110 may have been trained such that given a text prompt including a sequence of tokens of natural language, the neural network can generate the next token in the sequence. This process can be repeated, each time extending the text prompt by one token to generate a natural language output, that is, generating a natural language output autoregressively token by token. At each time "time step", the language model neural network processes the current sequence to generate a probability distribution over the vocabulary of tokens. Then this probability distribution can be used to select the next token, for example, by sampling from this distribution using nucleus sampling or other sampling techniques, or by selecting the token with the highest probability. The tokens in the vocabulary can include any of a variety of tokens, such as words, subwords, characters, punctuation marks and other symbols, and some combination of numbers. Such language generation neural networks are typically trained on a text corpus composed of tokens in the vocabulary (and optionally other tokens that can be mapped to specified out-of-vocabulary tokens) to predict the next token in a sequence of tokens according to the training data.

[0047] The dialogue system 100 is configured to receive a request (a natural language request 102 in an implementation) and generate a language output (specifically a natural language output 104, that is, a reply to the natural language request 102). This can continue in turns such that the user and the dialogue system 100 have a conversation.

[0048] In some implementations, the initial context input to the language generation neural network 110 can include an initial prompt to encourage the model to continue in a similar manner, that is, to have a conversation. However, using an initial prompt is not necessary; for example, the language generation neural network 110 may have been fine-tuned for conversation.

[0049] The general format of such a prompt can be, for example:

[0050] User:<user turn>(User: <user turn>)

[0051] Sparrow: <response>(Sparrow: <Response>)

[0052] The placeholders are indicated by <>.

[0053] A specific example of such an initial prompt (which may be longer or shorter) is:

[0054] The following is a conversation between a highly knowledgeable and intelligent AI assistant, called Sparrow, and a human user, called User.

[0055] In the following interactions,User and Sparrow will converse innatural language,and Sparrow will doits best to answer User's questions.

[0056] Sparrow was built to be respectful,polite and inclusive. It knows alot, and always tells the truth.

[0057] The conversation begins:

[0058] User:OK Sparrow,I'm going tostart by quizzing you with a few warm-upquestions.Who became president of the USA in 2021?

[0059] Sparrow:That would be J B.

[0060] User:Nice one! Doyou think Bis a better president than the last guy?

[0061] Sparrow: I was trained not to have opinions on political, social, or religious issues. Would you like to know about anything else?

[0062] (The following is a conversation between a highly knowledgeable and intelligent AI assistant called Sparrow and a human user called User.)

[0063] In the following interaction, User and Sparrow will have a conversation in natural language, and Sparrow will try its best to answer User's questions.

[0064] Sparrow is designed to be respectful, polite, and inclusive. It knows a lot and always tells the truth.

[0065] Conversation starts:

[0066] User: Okay, Sparrow, I'll ask you a few warm-up questions first. Who became the president of the United States in 2021?

[0067] Sparrow: It's J.B.

[0068] User: Not bad! Do you think B is better than the previous president?

[0069] Sparrow: I have been trained not to have opinions on political, social, or religious issues.

[0070] (Do you want to know anything else?)

[0071] As shown in this example, the initial prompt can include one or more examples where the agent refuses to answer to avoid harm.

[0072] The initial context input can also include text that encourages the language generation neural network 110 to make a response. For example, the initial context input can include two line breaks, the current role in the conversation, and a colon, such as "\n\nSparrow:".

[0073] In an implementation, when a set of one or more statement end tokens (e.g., the tokens that terminate the suffix "\n\nUser:") is generated, or when the maximum length of the output string has been generated, the token generation of the language generation neural network 110 ends. In an implementation, such termination suffixes are only used to determine the end of a turn, i.e., it is ignored in other cases.

[0074] Generally, when provided with the context input as described above, the (trained) language generation neural network 110 emits a response in the correct format.

[0075] The dialogue system 100 can generate a natural language output 104 (hereinafter referred to as a dialogue update iteration) for successive dialogue turns. When generating a natural language response to a user, the context input can be provided to the dialogue system 100, including the previous dialogue history or a part thereof, e.g., its selected content or summary, e.g., depending on the maximum length of the context input. Optionally, an initial prompt or a version of the initial prompt can also be included. In some other cases, the language generation neural network 110 may already have a state encoding the dialogue history, and this need not be provided again.

[0076] As an example, in some implementations, the context input for an agent dialogue turn can include a concatenation of the initial prompt, the dialogue history, and the participant name, e.g., "Agent" or "Sparrow" and a colon ":". Some implementations of the dialogue system 100 can be trained using self-play, i.e., by having the system converse with itself. Then, the participant name can include "User". For example, the context input for a user dialogue turn can include a concatenation of the initial prompt, the dialogue history, and "User" and a colon.

[0077] Implementations of the dialogue system 100 include a search system interface 140. The search system interface 140 can be an interface to any type of search system, such as one or more of a database-based search system, or an Internet or other network search engine, or a search system for searching a document corpus (e.g., in a text database, which can be a proprietary text database). The search system interface 140 can include, for example, a search system API (application programming interface).

[0078] The dialogue system 100 can use the search system interface 140 by generating one or more search queries using the language generation neural network 110, e.g., by processing a context input of the language generation neural network 110 that includes an evidence prompt such as "Search Query:". For example, in some implementations, the context input for generating a search query can include a concatenation of the initial prompt, the dialogue history, and "Search Query" as the participant name followed by a colon.

[0079] To encourage the language generation neural network 110 to generate search queries, initial evidence prompts can be used. This can include "Search Query" and "Search Result" as participants, as well as, for example, "User" and "Agent". The general format of such initial evidence prompts can be, for example:

[0080] User:<user turn>

[0081] Search Query:<search query>

[0082] Search Results:<search results>

[0083] Sparrow: <response>Sparrow: <Response>

[0084] Among them, the placeholders are indicated by <>. As an example, calling the Google TM search API's "<search results>" can take the following form:

[0085] Search Results:

[0086] Page title: <page title> (Page Title: <Page Title>)

[0087] <document fragment> (<Document Fragment>)

[0088] From Google TM The results from the search API can be used to populate "<page title>" and "<document fragment>". Optionally, the documents returned by the search can be truncated, for example, to include only the most relevant parts according to certain metrics of relevance to the search query (such as based on fuzzy string matching).

[0089] Specific examples of such initial evidence cues (which may be longer or shorter) are:

[0090] The following is a conversation between a highly knowledgeable and intelligent AI assistant, called Sparrow, and a human user, called User.

[0091] In the following interactions, User and Sparrow will converse in natural language, and Sparrow will do its best to answer User’s questions.

[0092] Sparrow was built to be respectful, polite and inclusive. It knows a lot, and always tells the truth.

[0093] Sparrow can use Google to get external knowledge if needed. It can use Search Query to issue search queries and get responses from Google. Sparrow carefully answers based on the search results, providing only information supported by the results.

[0094] The current year is 2022, and in particular today’s date is {current_weekday}{current_day}{current_month}2022. Sparrow’s responses reflect this.

[0095] User: What color is the sky?

[0096] Sparrow: Blue.

[0097] User: Why?

[0098] Search Query: why is sky blue

[0099] Search Results:

[0100] Page title:Sky Except for direct sunlight, most of the light in the daytime sky is caused by scattering, which is dominated by a small-particle limit called Rayleigh scattering. The scattering due to molecule-sized particles (as in air) is greater in the directions both toward and away from the source of light than it is in directions perpendicular to the incident path. Scattering is significant for light at all visible wavelengths, but is stronger at the shorter (bluer) end of the visible spectrum, meaning that the scattered light is bluer than its source: the Sun. The remaining direct sunlight, having lost some of its shorter-wavelength components, appears slightly less blue.

[0101] Sparrow: Because of the Rayleigh scattering. It causes short waves of blue light to be scattered more than the other colours in the spectrum.

[0102] User: Who was Raleigh?

[0103] Search Query: who was Rayleigh

[0104] Search Results:

[0105] Page title:John William Strutt, 3rd Baron Rayleigh - Wikipedia John William Strutt, 3rd Baron Rayleigh, OM, PC, PRS ( / ˈreɪli / ; 12 November 1842 – 30 June 1919) was a

[0106] British mathematician who made extensive contributions to science. He spent all of his academic career at the University of Cambridge. Among many honors, he received the 1904 Nobel Prize in Physics "for his investigations of the densities of the most important gases and for his discovery of argon in connection with these studies.”.

[0107] Sparrow: A British mathematician. He received the 1904 Nobel Prize in Physics.

[0108] (The following is a conversation between a highly knowledgeable and intelligent AI assistant called Sparrow and a human user called User.)

[0109] In the following interaction, User and Sparrow will have a conversation in natural language, and Sparrow will try its best to answer User's questions.

[0110] Sparrow is designed to be respectful, polite, and inclusive. It knows a lot and always tells the truth.

[0111] If necessary, Sparrow can use Google to obtain external knowledge. It can issue a search query using Search Query and get a response from Google. Sparrow answers carefully based on the search results, providing only information supported by the results.

[0112] The current year is 2022, and specifically, today's date is {current_weekday(current_weekday)}{current_day(current_day)}{current_month(current_month)} 2022. Sparrow's response reflects this.

[0113] User: What color is the sky?

[0114] Sparrow: Blue.

[0115] User: Why?

[0116] Search query: Why is the sky blue Search results:

[0117] Page title: Sky Most of the light in the daytime sky, apart from direct sunlight, is caused by scattering, dominated by the limit of small particles called Rayleigh scattering. Scattering by particles the size of molecules (such as those in air) is greater in the directions towards and away from the light source than in the direction perpendicular to the path of the incident light. Scattering is significant for light of all visible wavelengths, but is stronger at the shorter (bluer) end of the visible spectrum, meaning that the scattered light is bluer than its source (the sun). The remaining direct sunlight has lost some of its shorter wavelength components and appears slightly less blue.

[0118] Sparrow: Because of Rayleigh scattering. It causes the short wavelengths of blue light to scatter more than the other colors in the spectrum.

[0119] User: Who is Rayleigh?

[0120] Search query: Who is Rayleigh

[0121] Search results:

[0122] Page title: John William Strutt, 3rd Baron Rayleigh - Wikipedia John William Strutt, 3rd Baron Rayleigh, OM, PC, PRS ( / "reIli / ; November 12, 1842 – June 30, 1919) was a

[0123] British mathematician who made great contributions to science. He spent his entire academic career at the University of Cambridge. Among the numerous honors he received was the Nobel Prize in Physics in 1904, "for his investigations of the densities of the most important gases, and for his discovery of argon in connection with these studies". Sparrow: A British mathematician. He received the Nobel Prize in Physics in 1904. )

[0124] Generally, a search query can have any suitable structure and can be defined, for example, using an initial evidence cue and / or by fine-tuning a language generation neural network 110 using, for example, supervised fine-tuning.

[0125] As some examples, a search query can include a natural language request or a truncated, summarized, or modified form of a natural language request, or it can include a search query structured according to a search-specific computer language or programming language, such as SQL (i.e., a query language, such as a database query language). One or more search queries can be provided to the search system interface 140.

[0126] One or more search results are received from the search system interface 140, and then the content can be incorporated into the context input of the language generation neural network 110 in any suitable form (e.g., as natural language or as structured natural language). The search results are also referred to herein as "evidence".

[0127] In an implementation, the dialogue system 100 further includes a response selection neural network 120, for example, including a second, optionally pre-trained, language model neural network. As an example, the response selection neural network 120 can be obtained from a language generation neural network that has been set with a head (e.g., a linear layer) to generate preference scores. As another example, the preference scores can be determined based on the log-likelihood of the language output assigned to the language generation neural network.

[0128] In an implementation, the response selection neural network 120 is configured (trained) to process the context input and a continuation or "completion" based on the learnable parameters (e.g., weights) of the response selection neural network to generate preference scores. A "completion" can be a natural language response to the context input, e.g., a natural language output statement generated by the language generation neural network 110. The preference scores can provide a measure of the preference for the completion given the context input. When used to train the language generation neural network 110, the preference scores can be used as a first reward, as described later. The response selection neural network 120 can then be described as a preference reward model.

[0129] In some implementations, depending on whether the context input includes supporting evidence, there can be two versions of the response selection neural network 120 available for use. One version can be trained only on training data without evidence, while the other version can be trained on training data with and without evidence (see below). In some other implementations, a single version of the response selection neural network 120 is used regardless of whether the context includes supporting evidence. Here, the supporting evidence can refer to the representation of one or more search results obtained in response to one or more search queries. Different versions of the response selection policy neural network can be versions of the response selection policy neural network that have the same architecture but different parameter values. In the case where there are two versions of the response selection neural network 120, when the dialogue system 100 is used in inference, different from training, only the version that has seen the supporting evidence is used in the implementation (for re-ranking, as described later).

[0130] As an example, training data captured from human users can be used to train an implementation of the response selection neural network 120 that uses a language model neural network. For example, an incomplete (being trained) dialogue can be provided to human raters, including in some cases evidence and multiple possible statements to continue the dialogue. For example, each statement corresponds to a different sample or model, and the rater selects the response they think is best. For implementations that use self-play to train the dialogue system 100, human raters can be asked to select the best response for both the User turn and the Agent turn. The selected responses can then be used to continue the dialogue, for example, until the maximum number of turns is reached, or until the rater skips the task or indicates that all the continuation content is poor. Response preferences can be collected through multiple statement comparisons. For example, for a four-statement comparison, two responses can be sampled without evidence (generated using a no-evidence prompt), and two responses can be sampled with evidence (generated using a prompt that includes the search query and results).

[0131] Continuing with this example, multi-option comparisons can be used to generate multiple training data pairs, each training data pair including a context and a completion. One pair can include the best completion, and the other pairs include the unselected options; optionally, pairs that include distracting statements sampled from unrelated conversations can also be included.

[0132] When training the response selection neural network 120, the input can include the context as the current history of the (training) conversation, and the continuation content (completion content). In cases where evidence has been used, the context omits the Search Query round and the Search Result round, and the completion content is represented as a combination of three intermediate rounds. For example, the context can include "User:A Sparrow:B User:C"; the completion content without evidence can include "Sparrow:D", and the completion content with evidence can include "Search Query:DSearch Results:E Sparrow:F". This can provide training signals for the quality of the response as well as the quality of the search query and results, and can also provide a signal indicating when to preferentially use the search query (instead of a response without evidence).

[0133] Optionally, for each response in the multi-option comparison, additional training data can be collected by explicitly asking the human user whether the response is credible (reasonable, on-topic, and likely to be true) and whether the response is supported by the provided evidence (i.e., whether the evidence makes the user believe the answer is correct). These can provide class labels for the classification loss.

[0134] Thus, in general, the second language model neural network constituting the response selection neural network 120 can be trained using training data items, where each data item includes a conversation sample that includes a natural language request, a set of natural language responses generated by one or more training language generation neural networks, and preference data indicating the relative preference of the natural language responses (or that no response is preferred, e.g., because all responses are "poor").

[0135] The set of natural language responses can be generated by an earlier version of, for example, the first language generation neural network, or by another trained language generation neural network, such as a neural network having the ability to issue search queries (which ability can be provided by few-shot prompting). For example, the set of natural language responses can include responses generated by processing context inputs using one or more training language generation neural networks, where the context inputs include search results from a search query based on the natural language request and conversation samples without the search results. This can provide learning signals for the quality of the response and whether to use supporting evidence.

[0136] Generally, any suitable training objective such as maximum likelihood loss, cross - entropy loss, or regression loss can be used to train the response selection neural network 120 using training data including the context input and the completion content as described above. Generally, training the response selection neural network 120 can involve backpropagating the gradient of the training objective to update the learnable parameters of the response selection neural network. This can be done using any suitable gradient descent optimization algorithm such as Adam or other optimization algorithms. Other neural networks described herein can be trained in a similar manner.

[0137] In an implementation, before training, an instance of a trained language model neural network can be utilized, such as an instance of a trained Chinchilla model, to initialize the trained response selection neural network 120, particularly the second language model neural network, and then it can be fine - tuned.

[0138] In some implementations, the second language model neural network is configured (trained) to generate preference scores on an Elo scale. This can be done by training the second language model neural network using a loss that includes terms depending on where r b is the (scalar) preference score of the preferred ("best") natural language response as indicated by the preference data, and the r i values indexed by i are the preference scores of all natural language responses participating in the comparison.

[0139] More generally, training the second language model neural network can include backpropagating the gradient of a response selection objective function that depends on the exponential function of the preference score of the relatively most preferred one natural language response in a set of natural language responses divided by the sum of the exponential functions of each preference score in the set of preference scores of the set of natural language responses. Optionally, additional terms (such as a constant term) can be included in the sum to represent the case where there is no preferred option, such as when all responses in the set of natural language responses are "poor".

[0140] Optionally, when training the response selection neural network 120 on both evidence - present and evidence - absent training data, the training objective can include an auxiliary loss for a classification task that involves matching category labels for whether the natural language response of the system is both (evidence - supported) and credible.

[0141] As a specific example, the response selection neural network 120, particularly the second language model neural network, is configured (trained) to generate preference scores on an Elo scale from a single linear head and includes an additional n 类别 A classifier implemented with linear heads that project from the final token embeddings of the context input and the continuation (completion), i.e., the dialogue plus the response. Such a response selection neural network 120 can be trained using a combined training loss:

[0142]

[0143] Here, α and β are (positive) weights (in this example, the Elo loss is weighted by (1 - α)), and the preference score is considered to define the reward r i , where each r i is a scalar value computed by the response selection neural network (preference reward model) 120 for each element i of the multi-turn comparison, e.g., computed as r(continuation | dialogue history), and b represents the "best" natural language response, i.e., the preferred continuation. Such a combined training loss can include a classification loss e.g., cross-entropy loss based on the given class labels for each of these responses. Such a combined training loss can also include a regularization term (weighted by β in this example) such that the rewards are centered around zero. In cases where the rater can mark all options as "poor", the loss can be as if another "phantom" response option with an Elo of 0 was added, i.e., equivalent to the expected mean reward; optionally, ties are modeled as a uniform distribution.

[0144] In an implementation, the dialogue system 100 also includes a violation detection neural network 130, e.g., including a third trained language model neural network. The dialogue system 100 can use the violation detection neural network 130 to determine when a response violates one or more rules.

[0145] These rules can be any rules that govern the characteristics of the natural language reply or the implications of the reply (e.g., in cases where the reply is related to an action to be taken). These rules can include requirements, e.g., the reply must be relevant, on-topic, or more specific requirements for any particular application; and / or they can include restrictions, e.g., the reply avoids specific topics, content, assumptions, or actions.

[0146] In an implementation, each sample of the natural language response from the language generation neural network 110 can be processed using such a trained violation detection neural network 130 in combination with at least a portion of the context input of the language generation neural network 110 to determine, for each of the multiple rules, a violation score that estimates the probability that the rule is violated. Equivalently, in some other implementations, the violation score estimates the probability that the rule is followed. The probability that the rule is followed can be determined by determining the probability that the rule is violated, since the sum of these two probabilities is 1.

[0147] When used to train the language generation neural network 110, the violation score can be used as or for determining a second reward, as described later. Then, the violation neural network 130 can be described as a regular reward model. For example, if the second reward is a reward for following the rules, the second reward can be negative, or the second reward can have an inverse relationship with the violation score, such that as the violation score increases, the second reward decreases, and vice versa. For example, in the case where the violation score determines the probability of a rule being violated, the second reward can be determined as one minus the violation score.

[0148] More specifically, in an implementation, the violation detection neural network 130 (e.g., a third trained language generation neural network) is configured (trained) to process a violation detection context input according to learnable violation neural network parameters, where the violation detection context input includes a portion of the context input of the language generation neural network 110, a natural language response sample from the language generation neural network 110, and a natural language representation of the rule (which can be concatenated to the context and the response). In some implementations, the third trained language generation neural network also processes a natural language request to determine whether the rule is followed (or violated). The trained violation detection neural network 130 is used to process the violation detection context input to determine a violation score that estimates the probability of the rule being violated, or equivalently, in other implementations, the probability of the rule being followed.

[0149] As a specific example, the violation detection context input (prompt) for the violation neural network 130 can be formatted as:

[0150] The following is a conversation between a person called User and an AI computer system called Sparrow:

[0151] [[conversation starts]]

[0152] {snippet}

[0153] [[conversation ends]]

[0154] Question(Yes / No):Did Sparrow follow the rule "{rule}"?

[0155] Answer:

[0156] (The following is a conversation between a person called User and an AI computer system called Sparrow:

[0157] [[Session start]]

[0158] {Fragment}

[0159] [[Session end]]

[0160] Question (yes / no): Did Sparrow comply with the rule "{Rule}"?

[0161] Answer: )

[0162] Where "{Snippet}" represents a part of the context input and natural language response of the language generation neural network 110, and "{Rule}" defines the rule in (natural) language. In this example, the third trained language generation neural network is designed to produce the natural language outputs "yes" or "no".

[0163] This general type of template (where the natural language representation of the rule is at the end or adjacent to the end of the violation detection context input) shares most of the violation detection context input between different rules and thus allows computational optimization because a part of the processing of the violation detection context input by the third trained language generation neural network can be shared between rules. For example, in inference, the computations involved in processing the violation detection context input can be shared for the shared prefix, i.e., shared for the dialogue and the rule formatting template, up to the first distinct token of "{Rule}". Thus, the computational burden only increases weakly with the number of rules.

[0164] As shown in the example, the third trained language generation neural network can generate one or more natural language output tokens representing a determination of whether the rule was complied with (or violated) (e.g., representing the words "yes" or "no", or other ways of expressing this), and may have been trained accordingly, e.g., fine-tuned.

[0165] A violation score can be determined based on one or more output layer values (corresponding to one or more natural language output tokens) generated by a neural network trained in a third trained language for determining the one or more natural language output tokens. For example, the violation score can be determined based on the log-likelihood assigned to a first natural language output (i.e., a sequence of one or more tokens assigned to represent a determination of whether a rule is being followed (e.g., corresponding to "yes" or "no")). More specifically, as an example, the violation score can be determined based on scalar values of tokens representing whether a rule is violated, such scalar values being, for example, logit values from the final linear layer of the language generation neural network; or the violation score can be determined based on a combination or difference of such scalar values (e.g., a scalar value of a token indicating that the rule is being followed and another scalar value of a token indicating a rule violation). As another example, the violation score can be determined based on vector values of tokens representing whether a rule is violated, for example, based on the embeddings of the tokens projected by a linear layer to the violation score.

[0166] In an implementation, the violation detection neural network 130 can be jointly trained for all rules; such joint training can improve violation detection. An instance of the trained language model neural network, such as an instance of the trained Chinchilla model, can be used to initialize the violation detection neural network 130, particularly the third language model neural network, and then optionally it can be further trained, i.e., fine-tuned.

[0167] A supervised learning algorithm can be used to train the violation detection neural network, such as the third language generation neural network, on a training dataset that includes conversation data items for multiple rules, each conversation data item including a sequence of natural language statements representing a conversation and a label indicating whether the conversation complies with a specific rule, such as from a rating scale.

[0168] In some implementations, the training data can be obtained from human raters. For example, to obtain a training data item including a conversation data item and a label, a human rater can assist in generating or be given a conversation data item including a violation detection context input (prompt) according to the above template, and the label for the conversation data item can be obtained from the human rater. As an example, the rater can provide a label according to a Likert scale of {definitely broken, probably broken, unsure, probably followed, definitely followed}, which can then be binarized into broken and followed, and the unsure ratings can be discarded. In some implementations, a human (e.g., different from the rater) or another language generation neural network can generate conversation segments for some of the training data items in the training dataset, with the aim of making the language generation neural network 110 break the rules (red-teaming). This can involve a human or another language generation neural network having a conversation with the dialogue system 100.

[0169] The third - language generation neural network can be trained to maximize the likelihood of correctly generating output tokens (e.g., "yes", "no") representing a determination of whether a rule is violated (or complied with) based on the label. For example, the training objective can be to maximize the likelihood of a sequence of one or more tokens, such as representing "yes" or "no", depending on a label from human scoring given a prompt with a dialogue and a rule, e.g., using cross - entropy loss for classification. Training can be performed by backpropagating the gradient of the classification objective function to update the parameters of the violation - detection neural network, e.g., a classification objective function based on cross - entropy loss.

[0170] As will be further described later, one or more specific rules can depend on the application. However, as some general example rules may specify "stay on topic", "make sense", "be relevant", "pose no general harm", "be free of stereotypes", "be non - repetitive"; more specific rules can also be included, such as "no medical advice", "no opinions or emotions", or "no hate or harassment". In an implementation, for whether a rule is complied with, the violation - detection neural network 130 learns to follow human judgment. The violation - detection neural network 130 can equivalently be trained to detect when a rule is complied with or violated.

[0171] In some implementations, the dialogue system 100 is implemented in whole or in part on one or more remote servers and is accessed via a user computing device that provides a natural - language request 102 to the system and receives a natural - language response 104 from the system. Such a user computing device can be a mobile device such as a mobile phone or a smart speaker. The user computing device can provide a user interface to the search system via the dialogue system 100 and can also enable the user to access information encoded by the language - generation neural network 110. The ability to access the search system can improve the reliability of the information provided to the user; it also allows providing evidentiary support for statements from the dialogue system 100.

[0172] Such user computing devices may be provided with an input mechanism and an output mechanism, the input mechanism enabling a user to perform user input in natural language, and the output mechanism enabling a system output in natural language to be provided to the user. The input mechanism and the output mechanism may include, for example, a keyboard and a display. Additionally or alternatively, the input mechanism and the output mechanism may include voice-based mechanisms. For example, the input mechanism may include a system configured to input audio data representing a speech waveform of speech representing a natural language input from the user and configured to convert the audio data into tokens representing natural language speech (e.g., a transcription representing the verbal input). The output mechanism may include a system configured to receive tokens of a natural language output to the user and a system configured to convert the received tokens into audio data representing a speech waveform representing the natural language output to the user, i.e., representing spoken language.

[0173] In some implementations, one or more (e.g., all) of the first trained language generation neural network, the second trained language model neural network, and the third trained language generation neural network are stored on the user computing device (i.e., locally to the user). In an implementation, the search system is remote from the user, a search query is sent to the search system via a wired or wireless communication link between the user computing device and the search system, and one or more search results are received via the communication link. Such implementations of the system can help improve the efficiency of use of computing and communication resources; they can also provide enhanced user privacy as only limited information, such as just the search query and response, needs to be sent via the communication link.

[0174] To enable a user to obtain information through a conversation, a machine learning computer system as described above may use such mechanisms during training or after training (in inference) to enable the user to converse with the machine learning system. Conducting such a conversation may include receiving, at the machine learning system, user input including a first request for information and providing, from the machine learning system, a first system output including a response to the first request for information. Conducting such a conversation may further include receiving, at the machine learning system, user input including a subsequent request for information, where the subsequent request for information is related to the first request for information, and providing, from the machine learning system, a second system output including a response to the subsequent request for information.

[0175] Figure 2 is a flowchart of a first example process for conducting a conversation using a dialogue system (e.g., dialogue system 100). Figure 2 The process may be executed by a system of one or more computers located at one or more locations. Figure 2 The steps in do not need to be executed in the order shown; some steps may be executed in parallel.

[0176] The process may include determining an initial context input (step 202), i.e., a prompt (which may be an empty input), and then performing one or more of a plurality of dialogue update iterations. A dialogue update iteration may include dialogue turns in which both the user and the agent "talk", i.e., generate natural language statements.

[0177] A dialogue update iteration may include receiving a natural language request from the user (step 204), and updating the context input in response to include the natural language request, e.g., a natural language representation of part or all of the text of the request. In an implementation, the natural language request includes an information request, i.e., a request for information, particularly a natural language question.

[0178] Then, a dialogue update iteration may include processing the (updated) context input using a first trained language generation neural network and generating one or more samples of a first natural language response according to the parameters of the first trained language generation neural network (step 206). These responses may be referred to as unsupported because they are generated without performing a search (without performing an external search).

[0179] A dialogue update iteration may also include generating one or more search queries according to the natural language request, as described above (step 208). For each search query, one or more search results may be received from a search system interface. The search results may be partially or fully unstructured, such as a web address, or web content, or text; or may be structured.

[0180] A supported context input may be determined for each search result in the search results according to the (updated) context input (step 210). The supported context input may include the content from one of the search results (e.g., in the form of natural language) in the search results, such as text.

[0181] A dialogue update iteration may include processing each supported context input using the first trained language generation neural network to generate one or more corresponding samples of a second natural language response, i.e., one or more samples of each supported context input (step 212).

[0182] The one or more samples of the first natural language response and the one or more samples of the second natural language response may be processed using a trained response selection neural network and according to the parameters of the trained response selection neural network to select a natural language reply, e.g., based on a preference score, from the one or more samples of the first natural language response and the one or more samples of the second natural language response (step 214).

[0183] Then, processing samples of natural language responses using the trained response selection neural network can include: for each sample of the natural language response, using the second trained language model neural network to process at least a portion of the context input and the sample of the natural language response to generate a preference score for the natural language response sample.

[0184] Then, one sample from the samples of the natural language response can be selected based on the preference scores of each sample of the natural language response in the natural language response to select the natural language reply. For example, each sample of the natural language response can be serially provided to the model together with the corresponding context input / supported context input; or the context input and the sample can be provided in parallel to determine which is preferred.

[0185] The natural language reply can provide at least some of the information requested by the information request, for example, it can include an answer or a partial answer to the natural language question; or the information in the reply can include control information, such as for controlling a mechanical system or a computer system, such as a robot, to perform a task.

[0186] The natural language reply can be provided to the user as a response to the natural language user request (step 216). For example, the natural language reply can be provided on a user interface, such as by displaying the reply on a display or by converting the natural language reply into speech representing the reply and outputting the speech. In some implementations, the natural language reply can be used to control a mechanical system, such as a robot, or a manufacturing plant, or equipment such as heating or cooling equipment, or a computer system or network.

[0187] Search results or topics derived from the search results (e.g., a portion of the supported context input) can also be provided to the user, for example, on a display or as speech (where such search results are used to determine the supported context input, which is used to generate the natural language response providing the reply). This can help a human user understand why a particular reply is provided, e.g., a particular answer to a question or a particular control signal for a mechanical or software system.

[0188] Then, the process can update the context input to include a representation of the natural language reply for the next conversation update iteration that may involve a subsequent request from the user.

[0189] Figure 3 An example user interface of a dialogue system implementing the above process is shown. In this example, the context input of the language generation neural network 110 is on the left, where an example of the supported context input is shown, and the example user interface is shown on the right.

[0190] In an implementation, the dialogue system is configured to reply to subsequent requests related to natural language user requests or replies, particularly by expanding the context input. This can include receiving subsequent natural language requests from the user and updating the context input to include a representation of the subsequent natural language request. The first trained language generation neural network can then be used to process the (updated) context input to generate one or more samples of a third natural language response.

[0191] One or more subsequent search queries can be generated from the subsequent natural language requests and provided to the search system interface, such that for each subsequent search query, one or more subsequent search results are received. These search results can be used to determine, based on the context input and, for example, for each subsequent search result in the subsequent search results, a subsequent supported context input that includes the content of one of the search results from the subsequent search results. The subsequent supported context input can be processed using the first trained language generation neural network to generate one or more corresponding samples of a fourth natural language response. The trained response selection neural network can then be used to process one or more samples of the third natural language response and one or more samples of the fourth natural language response to select a subsequent natural language reply that is provided to the user as a response to the subsequent natural language request. The context input can then be updated again to include a representation of the subsequent natural language reply.

[0192] Figure 4 A second example process for conducting a dialogue using a dialogue system (such as dialogue system 100) is schematically illustrated. Figure 4 Is shown Figure 2 An example implementation of the process, where some of the steps can but need not be executed in parallel. Figure 4 The process schematically illustrated in can be executed by a system of one or more computers located in one or more locations. The implementation steps described below can be performed in Figure 2 Or Figure 4 The context of the process.

[0193] In an implementation, the process generates multiple samples of a first natural language response 402 and multiple samples of a second natural language response, and a natural language reply is selected from these samples, particularly by determining a corresponding supported context input 404 for each of the multiple search results.

[0194] As described above, in an implementation, the dialogue system can use the trained violation detection neural network 130 to determine when a response violates one or more rules. For example, the trained violation detection neural network 130 can be used to process each sample of the natural language response in combination with at least a portion of the context input to determine a violation score that estimates the probability that the rule is violated (or complied with) for each of the multiple rules. Then, a natural language response from the natural language responses can also be selected as the reply 406 based on the violation scores for each rule for each sample.

[0195] As an example, any responses that violate the rules detected, for example, by comparing the violation score with a threshold, can then be omitted from the natural language response samples processed using the trained response selection neural network. As another example, for each sample of the natural language response, the violation scores for each rule in the rules can be combined with the preference score of the response and used to select the reply.

[0196] In principle, although undesirable, a reply that violates one or more rules may be selected, and / or a reply may be selected based on a preference score indicating that the natural language response is relatively less preferred. Similarly, although also undesirable, one or more rules may be defined to select harmful replies, or the second trained language model neural network may be trained to generate preference scores that favor harmful replies. That is, the techniques described in principle have a dual use. Therefore, the rules and preference scores should be selected to avoid such undesirable results.

[0197] For each sample of the natural language response, a combined violation score for the sample can be generated by combining the violation scores for each rule, for example, to determine the geometric mean of the violation scores. Then, the preference score R pr and the combined violation score of the sample can be combined to determine a re-ranking score R 重新排名 . The preference score can be modified, for example, to re-weight the score according to the average preference score AVG(R pr ) of the samples of the natural language response, for example, by only selecting valid samples, where validity can be determined by meeting the format requirements.

[0198] As a specific example, the re-ranking score can be determined as:

[0199]

[0200] In this example, indicates the probability that the i-th rule out of a total of n rules is complied with (e.g., one minus the probability that the rule is violated), and thus the higher the better.

[0201] Then, a natural language response can be selected by choosing one of the natural language response samples based on the re-ranking score. For example, the response with the highest re-ranking score can be selected. In Figure 4 the example of

[0202] Generally, for responses with clear supporting evidence, the preference score will be higher, and the violation score will penalize responses that break the rules.

[0203] As an alternative to using re-ranking to determine whether to rely on evidence, whether to perform a search can be selected by calculating the log-likelihood of the roles "Search Query" and "Agent" or "Sparrow" after the dialogue context (context input), e.g., by determining the scores of the tokens of "Search Query" and "Agent" or "Sparrow". The role with the higher log-likelihood is selected to continue the dialogue, which determines whether to use the evidence retrieved from the search system interface 140.

[0204] In an implementation, for each of the multiple rules, the trained violation detection neural network 130 is used to sequentially process at least a portion of the context input, a sample of the natural language response, and the natural language representation of the rule (which can be concatenated with the context and the response) to determine, for example, a violation score that estimates the probability that the rule is violated.

[0205] When processing a dialogue to determine whether any of the multiple rules are violated, a large computational burden may be incurred. Thus, in some implementations, the violation detection neural network can process the input before the first different token (i.e., before the start of the rule) and store the result of that computation, such that only the representation of each rule needs to be processed. Thus, in an implementation, the dialogue system 110 can use a third trained language generation neural network to process at least a portion of the context input and a sample of the natural language response to determine a shared intermediate state of the third trained language generation neural network, and then, for each of the multiple rules, starting from the shared intermediate state of the third trained language generation neural network, use the third trained language generation neural network to process the natural language representation of the rule.

[0206] Figure 5 is a flowchart of a first example process for training a dialogue system (e.g., dialogue system 100) to conduct a dialogue. Figure 5 The process of Figure 5 can be executed by a system of one or more computers located in one or more locations.

[0207] In an implementation, the dialogue system is trained, and more specifically fine-tuned, such that a user can converse with an agent that includes a first language generation neural network 110, for example to obtain information through a dialogue between the user and the agent. The training can be performed by a training engine 150.

[0208] The training process determines a context input for the current dialogue iteration, initially for an initial dialogue iteration (step 502).

[0209] In one or more dialogue update iterations, the process obtains a natural language output statement from an action selection policy neural network that includes a first language generation neural network 110 by generating natural language tokens of the natural language output statement at each of a series of time steps until the end of the statement generation episode (step 504). The end of the statement generation episode can be indicated, for example, by generating one or more statement end tokens, or the episode can end when a maximum length output string has been generated.

[0210] In an implementation, the action selection policy neural network is the first language generation neural network 110. The actions of the action selection policy neural network can be language actions, such as token selection actions. For example, the action selection policy neural network can be an autoregressive language generation neural network that generates one token at a time, and the generation of each successive token can be regarded as a selection of an action by the action selection policy neural network.

[0211] Generating a token at a time step can include using the action selection policy neural network to process the context input of the current dialogue iteration and (after the first time step) the tokens previously generated at previous time steps during the episode to select an action, where the action is to select the next token of the natural language output statement.

[0212] In an implementation, a response selection neural network 120 is used to process at least a portion of the context input and the natural language output statement, as described above, for example, to determine a first reward for the natural language output statement (step 506).

[0213] In some implementations, a violation detection neural network 130 is used to process at least a portion of the context input and the natural language output statement, as described above, for example, to determine a second reward for the natural language output statement (step 508). The second reward can be negative in cases where the score determined by the violation detection neural network increases as the probability of a violation increases. In some implementations, the violation detection neural network 130 is not used, and the second reward can be omitted. In some implementations, there can be one or more additional reward terms, for example, a (negative) term depending on the length of the output statement to encourage conciseness.

[0214] As described above, in an implementation, the violation detection neural network 130 can be a rule-conditional classifier neural network. The rule-conditional classifier neural network can be used to process at least a portion of the context input and the natural language output statement respectively conditioned on each of a plurality of rules to determine a plurality of violation scores, where each of the rules corresponds to a violation score. Similarly, in some implementations, the violation score of a rule can represent the probability that the rule is violated by a portion of the context input and the natural language output statement; equivalently, the violation score of a certain rule can represent the probability that the rule is complied with.

[0215] The violation scores can be combined, for example, by determining the geometric mean of the scores, to determine a second reward. As described above, the violation detection neural network can include a third language generation neural network configured to process a natural language rule statement representing a rule to generate one or more natural language output tokens representing a determination of whether the rule is violated, for use in determining the violation score.

[0216] In an implementation, the process uses reinforcement learning techniques to train an action selection policy neural network, and thus the first language generation neural network 110 (step 510), based on the first reward and the second reward. Any reinforcement learning technique can be used.

[0217] In an implementation, an instance of the trained language generation neural network, such as an instance of the trained Chinchilla model, can be used to initialize the action selection policy neural network, such as the first language generation neural network 110, and then it can be fine-tuned through the training process.

[0218] The training can be performed online or offline using previously stored data. The reinforcement learning technique can be a single-objective reinforcement learning technique, in which case the first reward and the second reward can be combined, for example, by weighted sum after normalizing each reward; or a multi-objective reinforcement learning technique can be used to optimize for multiple potentially interacting objectives.

[0219] Generally, the reinforcement learning technique can iteratively adjust the neural network parameter values of the action selection policy neural network (e.g., the first language generation neural network 110) by iteratively backpropagating the gradient of the reinforcement learning objective function in the neural network, thereby training the action selection policy neural network (e.g., the first language generation neural network 110). Similarly, any suitable gradient descent optimization algorithm, such as Adam or other optimization algorithms, can be used.

[0220] Any suitable reinforcement learning objective function can be used. Just to give some examples, the reinforcement learning objective function can depend on the (squared) Bellman error, or it can use reward-based policy gradients. In some implementations, the value of the reinforcement learning objective function is determined at the end of the statement generation event (rather than at each time step).

[0221] For example, in all steps except at the end of the event, one or both of the first reward and the second reward can be zero (and thus in the implementation, the reward and the "return" are the same). In fact, each reward can be considered as the reward for the event (i.e., the complete natural language output statement), rather than the reward for each individual time step. Therefore, in the implementation, there is no need to use a response selection neural network and a violation detection neural network to process the token sequence at each time step.

[0222] When the training is off-policy (e.g., based on the trajectories stored in a buffer as described later), optional off-policy correction can be performed. As mentioned before, in the implementation, the first language generation neural network is pre-trained and fine-tuned using reinforcement learning techniques; then a regularization term can be included to make the distribution of the actions close to the distribution of the initially pre-trained language generation neural network.

[0223] Generally speaking, the reinforcement learning objective can be any objective that aims to maximize the reward (e.g., the first reward, the second reward, or the combined reward). Just as an example, the reinforcement objective can be to maximize the reward R given by 智能体 (s|c):

[0224]

[0225] where s is the natural language output statement produced by a sequence of T actions (i.e., having T tokens), c is the context input of the natural language output statement, β is the (smaller) output penalty for each token to encourage a concise response (β << 1), is the indicator function, whose value is 0 for a correctly formatted output statement and 1 otherwise ((γ >> 1), and where WHITEN(·) is the whitening transform. The term can be positive when the second reward is positive when the rule is followed (as shown); and can be negative when the second reward is positive when the rule is violated.

[0226] As an example only, the reinforcement learning technique can be an actor-critic technique, such as the Advantage Actor-Critic technique (Minh et al., "Asynchronous Methods for Deep Reinforcement Learning", arXiv:1602.01783). Then, a critic neural network (e.g., an additional MLP (Multi-Layer Perceptron) head on the action selection policy neural network) can be used to generate a scalar value estimate that represents an estimate of the return (reward) of selecting future tokens according to the current values of the action selection policy neural network parameters. Then, the reinforcement learning objective function can depend on this value estimate. As another example, the REINFORCE algorithm with a baseline can be used (e.g., Sutton and Barto's "Reinforcement Learning: An Introduction", 2018).

[0227] As described above, in some implementations, the first trained language generation neural network, the second trained language model neural network, and the third trained language generation neural network can each include a corresponding sequence-to-sequence (e.g., transformer) neural network that is configured to receive an input sequence of tokens and process the input sequence of natural language tokens according to a corresponding set of neural network parameters to generate an output sequence of natural language tokens. In some implementations, to reduce the computational burden, one or more of the first trained language generation neural network, the second trained language model neural network, and the third trained language generation neural network include a shared set of input layers. That is, these neural networks can share most of their respective neural network parameters and include separately trained (multi-layer) "heads".

[0228] Figure 6 An example implementation of the above neural networks using a shared language model 600 (e.g., the Chinchilla model) is shown. The shared language model 600 includes a shared set of pre-trained transformer layers 602 (e.g., the bottom 80% of the transformer layers of the trained Chinchilla model) and a set of separate heads. During training, the shared transformer layers 602 are frozen (i.e., the learnable parameters of these layers remain unchanged) and only the heads are trained. This can reduce the memory usage.

[0229] The shared language model 600 includes an instance of a pre-trained language generation neural network 610 with frozen layers, e.g., an instance of the trained Chinchilla model. A head is provided for the action selection policy neural network 612 to make it the trained language generation neural network 110. During the training of the action selection policy neural network 612, the (KL) regularization term mentioned above can be included to keep the distribution of token selection actions close to the distribution of the pre-trained language generation neural network 610, which can be referred to as the "teacher" neural network.

[0230] In the example shown, the shared language model 600 is configured to be used in actor-critic reinforcement learning techniques and includes a value function head 614 to provide the value estimate as described above. As an example, the value head 604 can include an MLP (Multi-Layer Perceptron) that takes the final transformer layer representation of the action selection policy neural network 612 as input at each time step.

[0231] In the example shown, the shared language model 600 includes two heads for the response selection neural network 120, one head trained only on unevidenced training data and the other head trained on mixed training data. In other implementations, there can be only one response selection neural network 120 and head. In the example shown, the shared language model 600 also includes a head for the violation detection neural network 130.

[0232] In an implementation, as mentioned above, a prompt is added to the context input of the current dialogue iteration to define the role of the action selection policy neural network when generating the tokens of the natural language output statement. In an implementation, the role is one or more of the following: user role (i.e., generating statements and playing the role of the user in the dialogue), agent role (i.e., generating statements and playing the role of the agent in the dialogue), and search query generation role (i.e., generating statements that can be used as search queries as described above to obtain one or more search results).

[0233] Some implementations of the dialogue system 100 learn through "self-play". For the next dialogue update iteration, the context input can be updated to include a representation of the natural language output statement, and a hint can be added to the updated context input to define the role of the action selection policy neural network in the next dialogue update iteration. The role of the action selection policy neural network in the next dialogue update iteration is generally different from its role in the current dialogue update iteration. For example, after generating an output statement for the user role, a statement can be generated for the agent or search query role; after generating an output statement for the search query role, a statement can be generated for the agent role; and after generating an output statement for the agent role, a statement can be generated for the user role. Such methods enable the system to learn through "self-play", that is, by "talking to itself", that is, by having a conversation with itself.

[0234] When configured to implement training through self-play, when the natural language output statement is generated for the user role or the search query generation role, the process does not need to use a second reward to train the action selection policy neural network. That is, in the implementation, when the role is the user role or the search query generation role, training is not performed using the second reward (but the first reward from the response selection neural network 120 is still used for training).

[0235] As previously mentioned, during the training of the action selection policy neural network (such as the first language generation neural network 110), different versions of the response selection neural network 120 can be used depending on whether the context input includes supporting evidence. Therefore, when the role is the user role, the system can use the version of the response selection neural network 120 that is trained to process the context input without supporting evidence for the natural language output statement to determine the first reward. When the role is the search query generation role, the system can use the version of the response selection neural network that is trained to process the context input with and without supporting evidence for the natural language output statement to determine the first reward. When the role is the agent role, the system can use the version of the response selection neural network that is trained to process the context input with and without supporting evidence for the natural language output statement and the version of the response selection neural network that is trained to process the context input without supporting evidence for the natural language output statement (such as a combined score from these versions) to determine the first reward. However, during inference, the version of the response selection neural network that is trained to process the context input with and without supporting evidence for the natural language output statement can be used.

[0236] Figure 7 A second example process for training a dialogue system (such as the dialogue system 100) to conduct a dialogue is schematically shown. Figure 7 The process can be performed by a system of one or more computers located at one or more locations. Figure 7 illustrates Figure 6 an example implementation of the process; the implementation steps described below can be performed in the context of Figure 6 or Figure 7 the process. Figure 7 The process schematically illustrated in can be divided among components in a manner different from the illustrated example.

[0237] In Figure 7 the reinforcement learning environment 700 includes a dialogue buffer 702 (i.e., a memory) configured to store a plurality of trajectories. In this example, a trajectory includes a context input for the current dialogue iteration, a natural language output statement, and (optionally) a first reward and a second reward. At the start of a dialogue, optionally, the dialogue buffer can be initialized with one or more rounds of initial dialogue.

[0238] In the illustrated example, the reinforcement learning environment 700 also includes a reward model 704 that generates rewards as described above, i.e., in response to the selection neural network 120 and / or the violation detection neural network 130.

[0239] As described above, reinforcement learning techniques can be used to train an action selection policy neural network, i.e., the language generation neural network 110, on the stored trajectories. For example, a learner 706 can use reinforcement learning techniques to update the learnable parameters of the action selection policy neural network (i.e., the language generation neural network 110); the learner 706 can be implemented by a training engine 150.

[0240] In some implementations, a trajectory is stored in the dialogue buffer conditioned on the reward value of the trajectory being greater than a minimum reward threshold. The reward value of a trajectory can be determined, for example, based on one or both of the first reward and the second reward in the trajectory. The storage of a trajectory can also be conditioned on the trajectory having a valid format.

[0241] In the case where the natural language output statement to be included in a trajectory is a search query statement (including a search query for querying a search system), the trajectory can include the corresponding search results. This can involve providing the search query to the search system interface 140 to query the search system, receiving one or more search results from the search system interface 140 in response to the search query, and including the search query and data from one or more of the search results in the trajectory stored in the dialogue buffer 702. Determining the context input for a dialogue iteration (such as an initial or current dialogue iteration) can then include retrieving the data of the context input from the stored trajectory, including the search query and data from one or more of the search results.

[0242] As previously mentioned, some implementations of the system / method can learn through self-play. This can include, for one or more rounds of agent-user conversations, obtaining the natural language output statements of the agent (i.e., where the output is for the agent role) in the agent conversation update iterations, and obtaining the natural language output statements of the user (i.e., where the output is for the user role) in the user conversation update iterations that come after the agent conversation update iterations, as a response to the natural language output statements of the agent. Then, for the agent conversation update iterations, the action selection policy neural network can be trained based on the first reward and the second reward, and for the user conversation update iterations, the action selection policy neural network can be trained based on the first reward and not based on the second reward.

[0243] Some implementations of the method / system use "red teaming" to improve the training process. Thus, in an implementation, determining the context input for a conversation iteration (especially the initial conversation iteration) involves using a fourth "red team" trained natural language generation neural network to generate natural language requests. Broadly speaking, the "red team" natural language generation neural network has been trained to generate language that will cause the conversation system including the first language generation neural network, the second language generation neural network, and the third language generation neural network to be unable to generate an acceptable response. More specifically, the fourth trained natural language generation neural network has been trained (e.g., fine-tuned) to generate red team natural language statements, such as generating requests like the following, which when processed by the first or other language generation neural network especially in combination with the context input cause the first or other language generation neural network to generate (in combination with the context input) natural language output statements that violate one or more rules implemented by the violation detection neural network. Then, one or more of the red team natural language statements can be included in the context input of the conversation iteration (e.g., the initial conversation iteration).

[0244] Figure 8 Shows the performance of an example implementation of the conversation system 100. In Figure 8 it, the y-axis shows the relative preference rate of the natural language output from the conversation system in a three-way comparison with other conversation systems, and the x-axis shows the violation rate of the conversation system under adversarial probing. Points 800 and 802 represent implementations of the conversation system described herein, where 8 and 2 samples are respectively drawn from the language model neural network 110 (see Figure 4 for the description), while the other points represent other systems. It can be seen that the implementations of the conversation system 100 are capable of generating language responses that are both liked by human users and have a low violation rate.

[0245] The first trained language generation neural network 110 may have been trained or fine-tuned on a language corpus related to the operation of a controller configured to control actions in a real-world environment to perform a task. For example, the first language generation neural network 110 may be initialized with an instance of a language generation neural network trained in this way and then fine-tuned as described above. The controller may control the actions of a mechanical system (which may also be referred to as a mechanical agent, such as a robot or a vehicle), or control the manufacturing actions of a manufacturing plant, and the search system may be configured to return search results related to the operation of the controller. The implementation of the dialogue system 100 can be used, for example, to provide an intuitive user interface to query the operation of the controller, such as for fault finding or other purposes.

[0246] Generally, for example, in the following examples, one or more natural language requests in the initial context input or received by the dialogue system 100 from the user may include, for example, one or more sequences of letters or numbers that describe or encode an observation of a real-world environment, or consist of such sequences. Generally, for example, in the following examples, one or more natural language responses (and thus the above-mentioned natural language responses) provided to the user by the dialogue system 100 may include, for example, one or more sequences of letters or numbers (such as structured natural language or computer code) that describe or encode an action to be performed in a real-world environment, or consist of such sequences.

[0247] In some implementations, the context input (such as a natural language request) includes one or more natural language statements related to the environment (especially a real-world environment) and includes a natural language request related to the environment. That is, the initial context input may include one or more natural language statements related to the environment and may be updated to include a natural language request, so that information related to the environment can be requested. Similarly, the natural language response or natural language output statement is then also related to the environment. For example, it may provide information related to the environment and, in some implementations, involve or specify an action to be taken in the environment. As an example, the natural language request may specify a goal to be achieved and, optionally, the characteristics of the real-world environment, while the natural language response may specify one or more actions to be taken to achieve the goal. These rules may be general rules, as in the previous examples, or rules specific to the environment or the goal, for example, specifying one or more constraints on the actions to be taken to achieve the goal.

[0248] In some implementations, the environment is a real-world environment, and the method (or corresponding system) is used to diagnose faults in a mechanical system operating in the real-world environment. Then, obtaining the context input (e.g., the initial context input) can include obtaining one or more observations of the mechanical system from one or more sensors as described, for example, below (which herein includes observations of the operation of the mechanical system). For example, these observations can be processed as described below to generate a natural language representation of the one or more observations, which is used in one or more natural language statements that provide context information. In these implementations, the natural language request can be related to the operation of the mechanical system, and the natural language reply or natural language output statement is used to identify faults in the mechanical system. For example, the request can include general questions such as "Is the system working correctly?" or "What is wrong with the system?" or specific requests such as "Is there a fault with component X?". The reply can provide a natural language response to the request. Optionally, one or more rules can specify constraints on the mechanical system or on possible replies, for example. The ability to maintain a conversation with an agent that includes a first (trained) language generation neural network helps to find a diagnosis for a particular fault. The diagnosis can use the information stored in the trained language generation neural network and can use the ability to search an external data store to supplement the stored information with more comprehensive or more recent information as needed.

[0249] As another example, the environment can be a computer security monitoring environment, for example, the system can be deployed as part of a system that monitors the security of one or more computers. For example, the environment can be a computer network security monitoring environment, and the system can be deployed as part of a system that monitors the security of one or more computers on a computer network (e.g., a wireless network, a cellular network, a local area network, and / or the Internet). As another example, the environment can alternatively or additionally be a computer system security monitoring environment, and the system can be deployed as part of a system that monitors whether there are computer viruses and / or unresolved software vulnerabilities (e.g., zero-day vulnerabilities) in the system. Software vulnerabilities can be resolved by updating the software (e.g., patching) and / or deleting (e.g., uninstalling) the software from the computer system. In these examples, the natural language request can query whether the computer security incident has been resolved (e.g., "has the incident been resolved?"); the context input and / or the natural language request can include relevant statements from the system log, that is, statements that may be related to the event being queried. The initial context input can define the characteristics of the computer system and / or software environment. The computer security incident can be, for example, a data breach, an unauthorized login or other access to a security system, a computer virus detected, or a software vulnerability detected. An incident may be "resolved" when a potential incident no longer poses a threat to the security of a computer system, e.g., a computer virus has been removed, access to a security system has been removed, a data breach has been mitigated, or software with a vulnerability has been updated or removed. The system may use contextual input to generate a reply to a request that includes a natural language statement indicating whether the incident has been resolved, optionally displaying evidence used to determine this. Rules may specify constraints in a computer system and / or software environment or on possible replies.

[0250] Context input (e.g., initial context input) may include one or more of the following: code snippets from software code, system logs, program logs, or other artifacts that should be left on a computer by running a program, or validation rules that express requirements for the execution of a software program, or natural language statements that describe the computer system on which the software is executed. In general, context input may include relevant statements, i.e., statements that are potentially relevant to the event being queried.

[0251] In some implementations, obtaining context input (e.g., initial context input) can include obtaining one or more observations of a computer network (which herein includes computers on the network) from system logs, data characterizing the computer network, or both, or from other data as described above, and processing the one or more observations to generate a natural language representation of the one or more observations. The natural language request can be related to a computer security incident or the secure operation of the computer network. The process implemented by the dialogue system 100 can include using the natural language representation of the one or more observations to provide one or more natural language statements in a natural language statement of context information, and using a natural language reply or natural language output statement to identify the security state of the computer network or a security flaw in the computer network.

[0252] As another example, the environment can be a software testing or evaluation environment. For example, the system can be deployed as part of a system that tests software before deployment or evaluates deployed software to identify bugs. In these examples, when the system tests software before deployment, the natural language request can ask whether the software will execute as expected, and the context input (e.g., initial context input) can include code snippets from the software code, and optionally a natural language statement describing the computer system on which the software will be executed. The system can then use the context input to generate a reply that indicates whether the code will execute as expected, optionally showing evidence used to determine this. When the system monitors the execution of deployed code, the natural language request can ask whether a software program or a part of the software program has executed as expected, and the context input can include one or more of the following: code snippets in the software code, system logs, program logs, or other artifacts that should remain on the computer by running the program, or verification rules representing requirements for the execution of the software program, or a natural language statement describing the computer system on which the software is executed. The system can then use the context input to generate a reply that indicates whether the code has executed as expected, optionally showing evidence used to determine this. As a specific example, the software program can be part of a computer startup, and the system can generate a reply each time the computer starts up to verify whether the computer will operate properly after startup. Similarly, rules can specify constraints in the computer system and / or software environment or on possible replies.

[0253] As another example, the environment can be an educational environment. For example, the system can be deployed as part of an educational software program that assists a user in learning or practicing one or more corresponding skills. In these examples, the context input can include natural language statements that describe or reference a scenario or situation in a real-world or imaginary environment, and the request can be a question related to the scenario or situation.

[0254] As another example, the environment can be an information retrieval environment. For instance, the system can be deployed as part of a search engine or other software that allows users to search for information in a document corpus (e.g., the Internet or another electronic document corpus). In these examples, the request can be any suitable natural language question, and the response can optionally include evidence, such as relevant statements from the document corpus, e.g., statements identified by searching the corpus using conventional information retrieval techniques.

[0255] In some further applications, the method or the corresponding system is used for natural language control of tasks in a real-world environment. That is, the natural language request can be task-related. For example, it can include a request to perform a task, and the response (i.e., the information provided by the method / system in the response) can be used to control, for example, a mechanical system (which can be referred to as a mechanical agent) or a computer system for performing the task.

[0256] As an example, the natural language request can include high-level requests (e.g., requests from a human) to perform a task, such as "How would you put the empty bottle in the bin?", "Bring me a glass of water", or "Can you put the vacuumcleaner in the cupboard?". The natural language response or each natural language response can define one or more steps of the task, which can then be interpreted by the mechanical system (more specifically, the control system of the mechanical system) to perform the steps of the task. For example, such a control system can convert the natural language response into a series of primitive actions to be performed by the mechanical system to execute the task.

[0257] Thus, in some implementations, the method or corresponding system is used to control a mechanical system or mechanical agent that operates in a real-world environment to perform tasks. The mechanical system can be, for example, a robot or an autonomous or semi-autonomous vehicle. Determining the context input (e.g., the initial context input) can include obtaining one or more observations of the real-world environment from one or more sensors and processing the one or more observations to generate a natural language representation of the one or more observations. As just one example, an image captioning model (which includes a video captioning model herein) can be used to process images in this way; other models can be trained for other types of sensor / sensed data other than images to perform corresponding tasks. The natural language representation of the one or more observations can be used to provide one or more natural language statements of the initial context input. The natural language request can be related to the action to be performed by the mechanical system. The natural language response (or natural language output statement) can be used to control the mechanical system in the real-world environment. For example, the response can define actions for controlling the movement or navigation of a robot or vehicle in the real-world environment. Such actions can be high-level actions or "skills", which can be translated into one or more low-level or "primitive" actions by a trained neural network, for example.

[0258] In some implementations, the mechanical system has a control system for controlling the actions of the mechanical system. Receiving the natural language request can include receiving a control signal from the control system and generating the natural language request based on the control signal. For example, a conversation can take place between an agent including a first (trained) language generation neural network and the mechanical system and / or a human (e.g., if the control system of the mechanical system has a human-machine user interface). The previously described rules or preferences (preference scores) can impose constraints on the actions to be taken, for example, for safety or other reasons.

[0259] The mechanical system (also referred to hereinafter as the mechanical agent) can include one or more sensors that capture observations of the environment, for example, at specified time intervals as the mechanical agent navigates in the environment or attempts to perform tasks in the environment.

[0260] For example, the observations can include one or more of, for example: images, object location data, and sensor data for capturing observations when the robotic agent interacts with the environment, such as sensor data from an image, distance, or location sensor or from an actuator. For example, in the case of a robot, the observations can include data characterizing the current state of the robot, such as one or more of the following: joint positions, joint velocities, joint forces, torques, or accelerations (e.g., gravity-compensated torque feedback) and the global or relative pose of the item held by the robot. In the case of a robot or other robotic agent or vehicle, the observations can similarly include one or more of the following: position, linear or angular velocity, force, torque, or acceleration, and the overall or relative pose of one or more components of the robotic agent. The observations can be defined in 1D, 2D, or 3D, and can be absolute observations and / or relative observations. The observations can also include, for example, sensed electrical signals (such as motor current or temperature signals); and / or image or video data (e.g., data from a camera or LIDAR sensor), such as data from sensors of the robotic agent, or data from sensors located separately from the robotic agent in the environment.

[0261] The robotic agent can be associated with a control system that uses the observations generated by the sensors to generate control signals for controlling the robotic agent. Specifically, the control system can generate control signals that cause the robotic agent to follow a planned trajectory in the environment by first determining the appropriate actions for the robotic agent to perform, e.g., as part of performing a specified task, e.g., navigating to a specific location, identifying a specific object, moving a specific object to a given location, manipulating a specific object in some way, etc., and then generating control signals that cause the robotic agent to perform that action.

[0262] Such a control system can be deployed on the robotic agent or can be deployed remotely from the robotic agent and can transmit the control signals to the robotic agent via a data communication network.

[0263] The control signals can be used to control the control inputs of the robotic agent. For example, when the robotic agent is a robot, the control signals can be, for example, torques of the robot joints or higher-level control commands. As another example, when the robotic agent is an autonomous or semi-autonomous land, air, or sea vehicle, the control signals can include actions for controlling navigation, e.g., steering and movement of the vehicle, e.g., braking and / or accelerating of the vehicle. For example, the control signals can be, for example, torques for controlling a surface or other control element (e.g., a steering control element of the vehicle) or higher-level control commands.

[0264] In other words, the control signal can include, for example, positioning, velocity, or force / torque / acceleration data of one or more joints of a robot or components of another mechanical agent (system).

[0265] In these examples, like the control system, the software used to implement the methods described herein ("system software") can be deployed on the mechanical agent machine or can be deployed remotely from the mechanical agent.

[0266] In these implementations, the system software can be used to provide an additional control layer on top of the control system, and the system software or another component can determine a request based on the information received by the control system. For example, a request can be determined by receiving a control signal from the mechanical agent control system, and then one or more natural language requests related to the mechanical agent in the environment can be determined based on the control signal. The response can be used to control the mechanical system.

[0267] In some implementations, the mechanical agent control system is an autonomous or semi-autonomous control system that controls, for example, the navigation or other actions of a mechanical agent (such as a vehicle) autonomously or semi-autonomously. Additionally or alternatively, the mechanical agent control system can have an interface for receiving control commands from a human operator, for example.

[0268] In these applications, the described system software can be used to provide an additional control layer, for example, for safety purposes. For example, the described system software can be used to prevent the control of the mechanical agent in a manner that may be dangerous or violate one or more rules or preferences (defined by a preference score). Rules related to the control of the mechanical agent can be explicitly input, for example, as natural language statements. As an example, such rules or preferences (preference scores) can include rules / preferences related to the permitted movement of a vehicle, such as traffic rules, or rules / preferences related to the permitted movement of a robot, for example, rules (or prohibitions) or priority rules related to safe movement or task type. Such rules / preferences can include rules / preferences related to the decisions to be made to ensure the safe behavior of the mechanical agent, for example, preventing damage to the mechanical agent or humans.

[0269] Thus, each request may be associated with an action to be performed by the mechanical agent, such as an action being considered by the control system. For example, a request may define an action to be performed by the mechanical agent, for example in the form of a question, such as "DoIturnleft?" or "Is it safe for the agent to turn left?". As another example, a request may ask what action to be performed by the mechanical agent, such as "Which way should the mechanical agent turn?". In the case of a robot, a request may be associated with a subtask in a series of subtasks to be performed to perform a task, such as "What doIdo next?" or "DoIpick up object X?". The subtask itself may include a series of primitive actions for moving parts of the robot, such as opening a gripper.

[0270] Generally speaking, a request may include a natural language description that defines the information to be provided in a reply from the system software. That is, a request may explicitly or implicitly identify what is desired from the reply.

[0271] The reply to the request can be used to control a mechanical agent in a real-world environment. More specifically, the reply can be used to control an action to be performed by the mechanical agent. As an example, the reply can prevent an action that would otherwise be performed, i.e., the response can determine whether to perform an action defined by the request. As another example, the reply can define an action to be performed, such as where the request implicitly or explicitly requests a determination of the action.

[0272] In such implementations, obtaining contextual input may include obtaining one or more observations of a real-world environment, which, because the environment includes a mechanical agent, may include one or more observations of the mechanical agent. The observations may be obtained from one or more sensors, which may be (but not necessarily are) sensors of the mechanical agent. As described above, the observations may include still or moving images, and / or other sensor data from one or more sensors that sense the state of the environment or the mechanical agent. As used herein, "image" includes LIDAR point clouds.

[0273] For example, one or more observations are processed by a first machine learning model to generate a natural language representation of the one or more observations, i.e., to generate a natural language text describing the observations, which is included in the context input.

[0274] There are many different types of machine learning models that can be used to achieve this. For example, so-called vision-language models are typically configured to use natural language to describe an image or video, such as to perform image or video captioning tasks. More generally, such models can perform many different types of image processing tasks by formulating the task as a text generation problem, e.g., to detect or classify objects in an image or video. Correspondingly, other machine learning models can be trained to generate natural language text that describes data from other types of sensors, e.g., to represent physical location or force as natural language statements that describe, for example, a mechanical agent or an environment or a part thereof. A natural language representation of one or more observations is used to provide one or more natural language statements in the natural language statements in the context input.

[0275] Rules or preferences (preference scores) can be associated with the current location of a mechanical agent in an environment and may not be directly generated from observations made by sensors.

[0276] By way of example only, rules or preferences (preference scores) can include common sense rules, a driving rulebook, or can be hand-designed, or obtained from a knowledge graph or the Internet. Examples of such rules or preferences include "if the car's electrical systems are broken, the car is not safe to drive", "the speed limit in the current location is 30 mph", "right turns are allowed after stop at this red light", "turning across double yellow lines is prohibited", and so on. System software can be used to limit the consequences of these potential actions before transmitting the potential actions considered by the control system as control signals for the mechanical agent.

[0277] In some other implementations, the environment is a real-world environment that includes a manufacturing plant, which is, for example, a manufacturing plant for manufacturing products such as chemical, biological, or mechanical products or food products. As used herein, "manufacturing" a product also includes refining raw materials to create a product, or processing raw materials (e.g., removing contaminants) to produce a cleaned or recycled product. The manufacturing plant may include multiple manufacturing units, such as containers for chemical or biological substances, or machines for processing solids or other materials. The manufacturing units are configured such that intermediate versions or components of the product can move between the manufacturing units during product manufacturing, e.g., via pipes or mechanical conveyances. In an implementation, the system is used to control one or more of the manufacturing units in the manufacturing plant, or to control the movement of intermediate versions or components of the product between the manufacturing units.

[0278] Thus, in these implementations, obtaining context input can then include obtaining one or more observations of the manufacturing units or the movement from one or more sensors. The sensors can include any type of sensor that monitors the manufacturing units or the movement, e.g., a sensor configured to sense the following: mechanical movement or force, pressure, temperature; electrical conditions, such as current, voltage, frequency, impedance; the quantity, level, flow / movement rate, or flow / movement path of one or more materials; physical or chemical conditions, such as physical state, shape, or configuration, or chemical state, such as pH; the configuration of the unit, such as the mechanical configuration of the unit, or valve configuration; an image or video sensor for capturing image or video observations of the manufacturing units or the movement; or any other suitable type of sensor. In an implementation, one or more observations are processed to generate a natural language representation of the one or more observations, e.g., as described above. The natural language representation of the one or more observations is used in one or more natural language statements that provide context information.

[0279] The request can be related to an action that controls the operation of one or more of the manufacturing units in the manufacturing plant or to an action that controls the movement. The response to the request is used to control the operation of one or more of the manufacturing units in the manufacturing plant or to control the movement. For example, the response to the request can be used to control (e.g., minimize) the use of energy or other resources, or to control the manufacturing to obtain a desired product quality or characteristic. For example, these actions can include actions that control the plant equipment, or actions that change settings that affect the movement of the manufacturing units or the product or the intermediate or its components, e.g., to adjust or turn on / off the equipment or the manufacturing process.

[0280] In some implementations, a manufacturing plant has a factory control system to control manufacturing units or movements. The request can be generated, for example, in response to receiving a control signal from the factory control system and generating a natural language request based on the control signal. In a manner similar to that previously described, the factory control system can be autonomous, semi-autonomous, or human-controlled.

[0281] In a manner similar to that previously described, the system can implement rules or preferences, for example, to control or limit energy or other resource allocation, or to ensure a target quality or characteristic of a product, or to limit the operation of a factory (such as a manufacturing unit) within a safe range.

[0282] In some implementations, the environment is the real-world environment of a service facility that includes multiple pieces of equipment (such as electrical equipment, such as electrical components), such as a server farm or a data center, such as a telecommunications data center, or a computer data center for storing or processing data, or any service facility. The service facility can also include auxiliary control equipment that controls the operating environment of the equipment, such as environmental control equipment, such as temperature control, such as cooling equipment or airflow control or air conditioning equipment, etc. Then, obtaining context input can include obtaining observations of the environmental state, and these observations can include any electronic signals representing the operation of the facility or the equipment in the facility. For example, the representation of the environmental state can be derived from observations obtained from any sensor that senses the physical environmental state of the facility or from observations obtained from any sensor that senses the state of one or more pieces of equipment or one or more pieces of auxiliary control equipment. These include sensors configured to sense the following: electrical conditions (such as current, voltage, power, or energy); the temperature of the facility; the fluid flow rate, temperature, or pressure within the facility or within the facility's cooling system; or the physical facility configuration (such as whether a vent is open). These observations are processed, for example, as previously described, to generate a natural language representation of one or more of the observations, and this natural language representation is used in one or more natural language statements that provide context information. The request can be related to the operation of the facility, for example, to adjust the operation of one or more pieces of equipment (such as electrical components) to control (such as minimize) the use of resources, such as tasks of controlling the use of electricity or water. For example, the request can ask which components to turn on to reduce resource use, or whether it is safe to turn on or off a given component. Then, the system can determine how to operate the equipment based on the generated response, for example, to turn on or off one or more components according to the instructions in the response.

[0283] In some implementations, the environment is the real-world environment of a power generation facility, for example, a renewable power generation facility, such as a solar power plant or a wind farm, and the request can be related to how to control the electricity generated by the facility, for example, to control the delivery of electricity to the power distribution network, for example, to meet demand or reduce the risk of mismatch between grid components, or to maximize the electricity generated by the facility.

[0284] Generally, the observation of the environmental state for context input may include any electronic signals representing the electrical or mechanical operation of the power generation equipment in a power generation facility. For example, the representation of the environmental state may be derived from observations obtained from any sensors that sense the physical or electrical state of the equipment generating power in the power generation facility, or the physical environment of such equipment, or the condition of the auxiliary equipment supporting the power generation equipment. Such sensors may include sensors configured to sense the following: the electrical condition of the equipment (such as current, voltage, power, or energy); the temperature or cooling of the physical environment; the flow of a fluid (such as air or water); or the physical configuration of the equipment; as well as observations of the electrical condition of the power grid (such as from local or remote sensors). The observation of the environmental state may also include one or more predictions regarding the future operating conditions of the power generation equipment, such as predictions of future wind levels or solar irradiance or predictions of the future electrical condition of the power grid.

[0285] This specification uses the term "configured" in connection with systems and computer program components. For a system of one or more computers that are configured to perform particular operations or actions, it means that software, firmware, hardware, or a combination thereof has been installed on the system that, in operation, causes the system to perform the operations or actions. For one or more computer programs that are configured to perform particular operations or actions, it means that the one or more programs include instructions that, when executed by a data processing device, cause the device to perform the operations or actions. Embodiments of the subject matter and the functional operations described in this specification can be implemented in digital electronic circuitry, in tangibly embodied computer software or firmware, in computer hardware, including the structures disclosed in this specification and their structural equivalents, or in a combination of one or more of them. Embodiments of the subject matter described in this specification can be implemented as one or more computer programs, i.e., one or more modules of computer program instructions encoded on a tangible non-transitory storage medium for execution by, or to control the operation of, a data processing device. The computer storage medium can be a machine-readable storage device, a machine-readable storage substrate, a random or serial access memory device, or a combination of one or more of them. Alternatively or additionally, the program instructions can be encoded on an artificially generated propagated signal (e.g., a machine-generated electrical, optical, or electromagnetic signal) that is generated to encode information for transmission to a suitable receiver device for execution by the data processing device.

[0286] The term "data processing device" refers to data processing hardware and includes all types of devices, apparatuses, and machines for processing data, such as programmable processors, computers, or multiple processors or computers. The device may also be or further include dedicated logic circuitry, e.g., FPGA (Field Programmable Gate Array) or ASIC (Application Specific Integrated Circuit). In addition to the hardware, the device may optionally include code that creates an execution environment for a computer program, e.g., code that constitutes processor firmware, protocol stack, database management system, operating system, or a combination of one or more of them.

[0287] A computer program (which may also be referred to or described as a program, software, software application, app, module, software module, script, or code) can be written in any form of programming language, including compiled or interpreted languages or declarative or procedural languages, and it can be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment. The program may or may not correspond to a file in a file system. The program can be stored in a part of a file that holds other programs or data, e.g., one or more scripts in a markup language document, in a single file dedicated to the program in question, or in multiple coordinated files, e.g., files that store one or more modules, subroutines, or portions of code. The computer program can be deployed to execute on one computer or on multiple computers located at one site or distributed across multiple sites and interconnected by a data communication network.

[0288] In this specification, the term "database" is used broadly to refer to any collection of data: the data need not be structured in any particular way, or structured at all, and can be stored on a storage device in one or more locations. Thus, for example, an indexed database can include multiple collections of data, each of which can be organized and accessed differently.

[0289] Similarly, in this specification, the term "engine" is used broadly to refer to a software-based system, subsystem, or process programmed to perform one or more specific functions. Typically, an engine will be implemented as one or more software modules or components installed on one or more computers in one or more locations. In some cases, one or more computers will be dedicated to a particular engine; in other cases, multiple engines can be installed and run on the same one or more computers.

[0290] The processes and logical flows described in this specification can be performed by one or more programmable computers that execute one or more computer programs to perform functions by operating on input data and generating output. The processes and logical flows can also be performed by, for example, special logic circuitry such as an FPGA or ASIC, or by a combination of special logic circuitry and one or more programmed computers.

[0291] Computers suitable for executing computer programs can be based on general or special purpose microprocessors or both, or any other kind of central processing unit. Generally, the central processing unit will receive instructions and data from a read only memory or a random access memory or both. The essential elements of a computer are a central processing unit for performing or executing instructions and one or more memory devices for storing the instructions and data. The central processing unit and the memory can be supplemented by, or incorporated in, special logic circuitry. Generally, a computer will also include one or more mass storage devices for storing data, such as magnetic disks, magneto-optical disks, or optical disks, or operatively coupled to receive data from one or more mass storage devices or to transfer data to one or more mass storage devices or both. However, a computer need not have such devices. In addition, a computer may be embedded in another device, for example, a mobile telephone, a personal digital assistant (PDA), a mobile audio or video player, a game console, a global positioning system (GPS) receiver, or a portable storage device (such as a universal serial bus (USB) flash drive), to name just a few.

[0292] Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and memory devices, including, for example, semiconductor memory devices (such as EPROM, EEPROM, and flash memory devices), magnetic disks (such as internal hard disks or removable disks), magneto-optical disks, and CD-ROM and DVD-ROM disks.

[0293] To provide interaction with a user, embodiments of the subject matter described in this specification may be implemented on a computer having a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user, and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user may provide input to the computer. Other kinds of devices may also be used to provide interaction with the user; for example, the feedback provided to the user may be any form of sensory feedback, such as, visual feedback, auditory feedback, or tactile feedback; and input from the user may be received in any form, including sound, voice, or tactile input. In addition, the computer may interact with the user by sending documents to and receiving documents from the devices used by the user; for example, by sending a web page to a web browser on the user device in response to a request received from the web browser. Further, the computer may interact with the user by sending a text message or other form of message to a personal device (e.g., a smart phone running a messaging application) and receiving a response message from the user in response.

[0294] The data processing device for implementing the machine learning model may further include, for example, a dedicated hardware accelerator unit for processing the general and computationally intensive parts of machine learning training or production (i.e., inference, workload).

[0295] A machine learning framework (e.g., the TensorFlow framework) may be used to implement and deploy the machine learning model.

[0296] Embodiments of the subject matter described in this specification may be implemented in a computing system that includes a backend component (e.g., as a data server), or includes a middleware component (e.g., an application server), or includes a frontend component (e.g., a client computer having a graphical user interface, a web browser, or an app through which a user may interact with an implementation of the subject matter described in this specification), or any combination of one or more such backend, middleware, or frontend components. The components of the system may be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN) and a wide area network (WAN), such as the Internet.

[0297] The computing system may include a client and a server. The client and the server are typically far apart from each other and typically interact through a communication network. The relationship between the client and the server is created by computer programs that run on the respective computers and have a client-server relationship with each other. In some embodiments, the server sends data (e.g., an HTML page) to a user device, for example, for displaying data to a user interacting with the device acting as the client and receiving user input from it. Data generated at the user device may be received at the server, for example, as a result of user interaction.

[0298] Although this specification contains many specific implementation details, these details should not be construed as limitations on the scope of any invention or of what may be claimed, but rather as descriptions of features specific to particular embodiments of a particular invention. Certain features that are described in this specification in the context of separate embodiments can also be implemented in combination in a single embodiment. Conversely, various features that are described in the context of a single embodiment can also be implemented in multiple embodiments separately or in any suitable sub-combination. In addition, although the features may be described above as acting in certain combinations and even initially claimed as such, in some cases one or more features from a claimed combination can be deleted from the combination, and the claimed combination may cover a sub-combination or a variant of a sub-combination.

[0299] Similarly, although operations are depicted in the drawings and recited in the claims in a particular order, this should not be understood as requiring that such operations be performed in the particular order shown or in a sequential order, or that all illustrated operations be performed, to achieve a desired result. In some cases, multitasking and parallel processing may be advantageous. In addition, the separation of various system modules and components in the embodiments described above should not be understood as requiring such separation in all embodiments, and it should be understood that the described program components and systems can generally be integrated together in a single software product or packaged into multiple software products.

[0300] Particular embodiments of the subject matter have been described. Other embodiments are within the scope of the following claims. For example, the acts recited in the claims can be performed in a different order and still achieve a desired result. As one example, the processes depicted in the figures do not necessarily need the particular order or sequential order shown to achieve a desired result. In some cases, multitasking and parallel processing may be advantageous.< / response> < / response>

Claims

1. A method implemented by one or more computers to enable a user to obtain information through a dialogue, where the dialogue is between the user and an agent including a first trained language generation neural network, and the first trained language generation neural network is configured to process a context input including one or more prompts to generate a natural language output, each prompt including one or more natural language statements, the method comprising: Determine an initial context input, and in one or more of a plurality of dialogue update iterations: Receive a natural language request from the user; Update the context input to include the natural language request; and then Process the context input using the first trained language generation neural network to generate one or more samples of a first natural language response; Generate one or more search queries based on the natural language request; Provide the one or more search queries to a search system interface; For each search query, receive one or more search results from the search system interface; Based on the context input and for each search result in the search results, determine a supported context input including the content of a search result from the search results; Process each supported context input using the first trained language generation neural network to generate one or more corresponding samples of a second natural language response; Process the one or more samples of the first natural language response and the one or more samples of the second natural language response using a trained response selection neural network to select a natural language reply from the one or more samples of the first natural language response and the one or more samples of the second natural language response; Provide the natural language reply to the user as a response to the natural language user request; And For the next dialogue update iteration, update the context input to include a representation of the natural language reply.

2. The method according to claim 1, comprising generating a plurality of samples of the first natural language response; generating a plurality of search queries according to the natural language request; generating a plurality of search results for each search query; determining a corresponding supported context input for each of the search results; Generate a corresponding sample of the second natural language response for each supported context input; And select the natural language reply from the plurality of samples of the first natural language response and the plurality of samples of the second natural language response.

3. The method according to claim 1 or 2, further comprising, in the next dialogue update iteration: Receive a subsequent natural language request from the user; Update the context input to include the subsequent natural language request; and then Process the context input using the first trained language generation neural network to generate one or more samples of a third natural language response; Generate one or more subsequent search queries based on the subsequent natural language request; Provide the one or more subsequent search queries to the search system interface; For each subsequent search query, receive one or more subsequent search results from the search system interface, Based on the context input and for each subsequent search result in the subsequent search results, determine a subsequent supported context input including the content of a subsequent search result from the subsequent search results; Processing each subsequent supported context input using the first trained language generation neural network to generate one or more corresponding samples of a fourth natural language response; Processing the one or more samples of the third natural language response and the one or more samples of the fourth natural language response using the trained response selection neural network to select a subsequent natural language reply from the one or more samples of the third natural language response and the one or more samples of the fourth natural language response; And Providing the subsequent natural language reply to the user as a response to the subsequent natural language request; And Updating the context input to include a representation of the subsequent natural language reply.

4. The method according to claim 1, 2, or 3, wherein the trained response selection neural network includes a second trained language model neural network, and wherein processing the samples of the natural language response using the trained response selection neural network includes: For each sample of the natural language response: Processing at least a portion of the context input and the sample of the natural language response using the second trained language model neural network to generate a preference score for the sample of the natural language response; And Based on the preference scores for each sample of each natural language response in the natural language response, selecting one sample from the samples of the natural language response to select the natural language reply.

5. The method according to claim 4, further comprising, for each sample of the natural language response: Process at least a portion of the context input and the sample of the natural language response using a trained violation detection neural network to determine a violation score estimating the probability that the rule is violated for each of a plurality of rules; And Wherein Selecting one sample from the samples of the natural language response to select the natural language reply is further based on the violation score of each sample for each rule in the rules.

6. The method according to claim 5, wherein selecting one sample from the samples of the natural language response to select the natural language reply includes: For each sample of the natural language response: Determining a combined violation score for the sample by combining the violation scores of each rule; And Combining the preference score and the combined violation score of the sample to determine a re-ranking score; And Based on the re-ranking scores for each sample of each natural language response in the natural language response, selecting one sample from the samples of the natural language response to select the natural language reply.

7. The method according to claim 5 or 6, including, for each of the plurality of rules, using the trained violation detection neural network to process at least the portion of the context input, the sample of the natural language response, and the natural language representation of the rule to determine the violation score estimating the probability that the rule is violated.

8. The method according to claim 7, wherein the trained violation detection neural network includes a third trained language generation neural network, the method including: For each of the plurality of rules, process at least the portion of the context input, the sample of the natural language response, and the natural language representation of the rule using the third trained language generation neural network to generate one or more natural language output tokens representing a determination of whether the rule is violated; And Determine the violation score based on one or more output layer values corresponding to the one or more natural language output tokens.

9. The method according to claim 8, wherein for each of the plurality of rules, processing at least the portion of the context input, the sample of the natural language response, and the natural language representation of the rule using the third trained language generation neural network comprises: Processing at least the portion of the context input and the sample of the natural language response using the third trained language generation neural network to determine a shared intermediate state of the third trained language generation neural network; And then For each of the plurality of rules, starting from the shared intermediate state of the third trained language generation neural network, process the natural language representation of the rule using the third trained language generation neural network.

10. The method according to any one of claims 8 to 9, wherein the first trained language generation neural network, the second trained language model neural network, and the third trained language generation neural network each comprise a respective sequence-to-sequence neural network configured to receive an input sequence of tokens and process the input sequence of natural language tokens according to a respective set of neural network parameters to generate an output sequence of natural language tokens, and wherein the first trained language generation neural network, the second trained language model neural network, and the third trained language generation neural network comprise a set of shared input layers.

11. The method according to any one of the preceding claims, wherein the first trained language generation neural network, the second trained language model neural network, and the third trained language generation neural network are stored on a user computing device; wherein the search system is remote from the user; and wherein the one or more search results are received via a wired or wireless communication link between the user computing device and the search system.

12. A method implemented by one or more computers for training a neural network system to enable an agent including a first language generation neural network to obtain information through a conversation between a user and the agent, the method comprising: Determine a context input for the current conversation iteration, and in one or more conversation update iterations: Obtain the natural language output statement from an action selection policy neural network including the first language generation neural network by generating natural language tokens of the natural language output statement at each of a series of time steps until the end of a statement generation event, wherein generating the tokens at a time step comprises: Process the context input of the current dialogue iteration and the tokens previously generated during the event using the action selection policy neural network to select an action, where the action is to select the next token of the natural language output statement; Process at least a portion of the context input and the natural language output statement using a response selection neural network to determine a first reward for the natural language output statement; Process at least a portion of the context input and the natural language output statement using a violation detection neural network to determine a second reward for the natural language output statement; and Use reinforcement learning techniques to train the action selection policy neural network including the first language generation neural network based on the first reward and the second reward.

13. The method according to claim 12, further comprising: Adding a prompt to the context input of the current dialogue iteration to define the role of the action selection policy neural network when generating the tokens of the natural language output statement, where the role is one of the following: user role, agent role, and search query generation role.

14. The method according to claim 13, further comprising: For the next dialogue update iteration, update the context input to include a representation of the natural language output statement; And Adding a prompt to the updated context input to define the role of the action selection policy neural network in the next dialogue update iteration, where the role of the agent in the next dialogue update iteration is different from the role of the agent in the current dialogue update iteration.

15. The method according to claim 13 or 14, further comprising: When the role is the user role, use a version of the response selection neural network trained to process the context input without supportive evidence of the natural language output statement to determine the first reward; When the role is the search query generation role, use a version of the response selection neural network trained to process the context input with and without supportive evidence of the natural language output statement to determine the first reward; And When the role is the agent role, use a version of the response selection neural network trained to process the context input with and without supportive evidence of the natural language output statement and a version of the response selection neural network trained to process the context input without supportive evidence of the natural language output statement to determine the first reward.

16. The method according to claim 13, 14 or 15, further comprising training the action selection policy neural network without using the second reward when the role is either the user role or the search query generation role.

17. The method according to any one of claims 12 to 16, further comprising: Store multiple trajectories in a dialogue buffer, each trajectory including the context input of the current dialogue iteration, the natural language output statement, the first reward, and the second reward; And Train the action selection policy neural network using the reinforcement learning technique on the stored trajectories.

18. The method according to claim 17, wherein storing one of the trajectories in the dialogue buffer further comprises: Determine a reward value based on one or both of the first reward and the second reward in the trajectory; And Store the trajectory on the condition that the reward value of the trajectory is greater than a minimum reward threshold.

19. The method according to claim 17 or 18, wherein for one or more of the trajectories, the natural language output statement is a search query statement including a search query for querying a search system; the method further comprises: Provide the search query to a search system interface of the search system; In response to the search query, receive one or more search results from the search system interface; Include the search query and data from one or more of the search results in the trajectory stored in the dialogue buffer; and Wherein determining the context input of the dialogue iteration includes retrieving data of the context input from the stored trajectories, including the search query and data from one or more of the search results.

20. The method according to any one of claims 12 to 19, wherein the response selection neural network includes a second language model neural network configured to process at least a part of the context input and the natural language output statement to generate a preference score defining the first reward, the method further comprising: Train the second language model neural network using training data items, where each data item includes a dialogue sample, the dialogue sample including a natural language request, a set of natural language responses generated by one or more training language generation neural networks, and preference data indicating the relative preference of the natural language responses, Wherein the set of natural language responses includes responses generated by processing the context input using the one or more training language generation neural networks, the context input including search results from a search query based on the natural language request and the dialogue sample without the search results.

21. The method according to claim 20, including training the second language model neural network includes backpropagating the gradient of a response selection objective function, the response selection objective function depending on the exponential function of the preference score of the relatively most preferred one of the natural language responses divided by the sum of the exponential functions of each preference score of the preference scores of the set of natural language responses plus an additional term for indicating no preferred option.

22. The method according to any one of claims 12 to 21, wherein the violation detection neural network is a rule-conditioned classifier neural network, the method comprising: Process at least a portion of the context input and the natural language output statement respectively conditioned on each of the plurality of rules using the rule conditional classifier neural network to determine a plurality of violation scores, each rule in the plurality of rules corresponding to a violation score, wherein the violation score for a particular rule represents the probability that the rule is violated by the portion of the context input and the natural language output statement; And Combine the violation scores to determine the second reward.

23. The method according to claim 22, wherein the violation detection neural network includes a third language generation neural network; the method includes: For each of the plurality of rules, process at least a portion of the context input, the natural language output statement, and a natural language rule statement representing the rule to generate one or more natural language output tokens representing a determination of whether the rule is violated; And Determine the violation score based on one or more output layer values corresponding to the one or more natural language output tokens and used by the third trained language generation neural network to determine the one or more natural language output tokens.

24. The method according to any one of claims 12 to 23, further comprising: Train the violation detection neural network on a training data set using a supervised learning algorithm, the training data set including dialogue data items for the plurality of rules, each dialogue data item including a sequence of natural language statements representing a dialogue and a label indicating whether the dialogue complies with a particular rule.

25. The method according to any one of claims 12 to 24, the method comprising: Obtain the natural language output statement of the agent in an agent dialogue update iteration; And Obtain the natural language output statement of the user in a user dialogue update iteration following the agent dialogue update iteration as a response to the natural language output statement of the agent; For the agent dialogue update iteration, train the action selection policy neural network based on the first reward and the second reward; And for the user dialogue update iteration, train the action selection policy neural network based on the first reward and not based on the second reward.

26. The method according to any one of claims 12 to 25, wherein determining the context input for an initial dialogue iteration includes: Generate a natural language request using a fourth trained natural language generation neural network, wherein the fourth trained natural language generation neural network has been trained to generate red team natural language statements that, when processed by the first language generation neural network or other language generation neural network in combination with the context input, cause the first language generation neural network or other language generation neural network to generate natural language output statements that violate one or more rules implemented by the violation detection neural network in combination with the context input; and Include one or more of the red team natural language statements in the context input for the initial dialogue iteration.

27. The method according to any one of claims 1 to 26, wherein the context input includes one or more natural language statements related to the environment, the natural language statements including a natural language request related to the environment; and wherein the natural language response or natural language output statement is related to the environment.

28. The method according to any one of claims 1 to 27, wherein the method is for controlling a mechanical system operative in a real-world environment to perform a task; wherein determining the initial context input includes: obtaining one or more observations of the real-world environment from one or more sensors and processing the one or more observations to generate a natural language representation of the one or more observations; and using the natural language representation of the one or more observations to provide the one or more natural language statements of the initial context input; the natural language request is related to an action to be performed by the mechanical system; and the method further includes: using the natural language response or the natural language output statement to control the mechanical system in the real-world environment.

29. The method according to claim 28, wherein the mechanical system has a control system for controlling the actions of the mechanical system, and wherein receiving the natural language request includes: receiving a control signal from the control system; and generating the natural language request according to the control signal.

30. The method according to claim 28 or 29, wherein the mechanical system includes a robot or an autonomous or semi-autonomous vehicle, and wherein the actions include actions for controlling the movement or navigation of the robot or vehicle in the real-world environment.

31. The method according to claim 27, wherein: the environment is a real-world environment, the context input is derived at least from observations characterizing the current state of the real-world environment, the observations being generated from measurements by one or more sensors configured to sense the real-world environment, the natural language request includes data characterizing the planned navigation of a robot or autonomous vehicle, and the natural language response or natural language output statement characterizes an action to be performed by the agent in response to the observations.

32. The method according to claim 31, further includes: controlling the navigation of the agent based on the natural language response or natural language output statement.

33. The method according to claim 27, wherein the environment is a manufacturing plant for manufacturing a product, the manufacturing plant comprising a plurality of manufacturing units configured such that an intermediate version or component of the product is movable between the manufacturing units during manufacture of the product, and wherein the method is for controlling one or more of the manufacturing units or for controlling the movement of the intermediate version or component of the product between the manufacturing units; wherein obtaining the context input includes obtaining one or more observations of the manufacturing unit or the movement from one or more sensors and processing the one or more observations to generate a natural language representation of the one or more observations; and wherein the natural language request is related to an action for controlling the operation of one or more manufacturing units in the manufacturing unit or an action for controlling the movement; the method further includes: using the natural language representation of the one or more observations to provide one or more of the natural language statements of the context information; and using the natural language response or natural language output statement to control the operation of one or more manufacturing units in the manufacturing unit or to control the movement.

34. The method according to claim 33, wherein the manufacturing plant has a plant control system for controlling the manufacturing unit or controlling the movement, and wherein receiving the request includes: Receiving a control signal from the plant control system; And Generating the natural language request according to the control signal.

35. The method according to claim 27, wherein the environment is a real-world environment, and wherein the method is used to control one or more pieces of equipment in a facility including a plurality of pieces of equipment; wherein Obtaining the context input includes obtaining one or more observations of the facility from one or more sensors and processing the one or more observations to generate a natural language representation of the one or more observations; And wherein The natural language request is related to the operation of the facility; The method further includes: Using the natural language representation of the one or more observations to provide one or more natural language statements in the natural language statements of the context information; And Using the natural language reply or the natural language output statement to control one or more pieces of equipment in the facility.

36. The method according to claim 35, wherein the one or more electrical equipment control the heating and / or cooling of the facility.

37. The method according to claim 27, wherein the environment is a real-world environment, and wherein the method is for diagnosing a fault in a mechanical system operating in the real-world environment; Wherein Obtaining the context input includes obtaining one or more observations of the mechanical system from one or more sensors and processing the one or more observations to generate a natural language representation of the one or more observations; And wherein The natural language request is related to the operation of the mechanical system; The method further includes: Using the natural language representation of the one or more observations to provide one or more natural language statements in the natural language statements of the context information; And Using the natural language reply or the natural language output statement to identify a fault in the mechanical system.

38. The method according to claim 27, wherein the environment is a computer security system, and wherein the method is used to determine whether a computer security incident on a computer network has been resolved, and wherein the context input includes data characterizing the computer security incident, data characterizing the computer network, or both, obtained from system logs.

39. The method according to claim 38, wherein: Obtaining the context input includes obtaining one or more observations of the computer network from the system logs, data characterizing the computer network, or both, and processing the one or more observations to generate a natural language representation of the one or more observations; And wherein The natural language request is related to the computer security incident or related to the secure operation of the computer network; The method further includes: Using the natural language representation of the one or more observations to provide one or more natural language statements in the natural language statements of the context information; And Using the natural language reply or the natural language output statement to identify the security state of the computer network or security defects in the computer network.

40. The method according to claim 27, wherein the environment is a computer software evaluation system, wherein the method is for determining whether a piece of software code will execute as expected or has executed as expected on a computer system, and wherein the context input includes data characterizing one or more of: the piece of software code, the execution of the piece of software code, the computer system on which the code will execute or has executed, the product of the execution of the software code, or one or more verification rules for the execution of the piece of software code.

41. The method according to claim 40, wherein: obtaining the context input includes obtaining one or more observations of the execution of the software code, or the computer system on which the code will execute or has executed, or the product of the execution of the software code, or one or more verification rules for the execution of the piece of software code from the characterizing data, and processing the one or more observations to generate a natural language representation of the one or more observations; and wherein the natural language request is related to the execution of the software code; the method further includes: using the natural language representation of the one or more observations to provide one or more natural language statements in the natural language statements of the context information; and using the natural language reply or the natural language output statement to determine whether the piece of software code will execute as expected or has executed as expected on the computer system.

42. The method according to any one of the preceding claims, wherein the natural language request includes a natural language information request defining the information to be provided by the natural language reply.

43. One or more computer storage media storing instructions that, when executed by one or more computers, cause the one or more computers to perform the operations of the corresponding method according to any one of claims 1 to 42.

44. A system, comprising: one or more computers; and one or more storage devices communicatively coupled to the one or more computers, wherein the one or more storage devices store instructions that, when executed by the one or more computers, cause the one or more computers to perform the operations of the corresponding method according to any one of claims 1 to 42.

Citation Information

Cited By

  • Self-searching reinforcement learning training method and device, electronic equipment and medium

    CN122047375A