Selective Inference Neural Network System
The system addresses the 'black box' issue of neural networks by generating interpretable natural language explanations for its responses, enhancing transparency and trust, especially in safety-critical applications.
Patent Information
- Application Number
- JP2024555988
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-05-19
- Filing Date
- 2023-05-12
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2043-05-12
AI Technical Summary
Neural network-based systems often act as 'black boxes,' making it difficult to infer why specific decisions were made, which is problematic in control decisions for agents or manufacturing plants where transparency is crucial for safety and reliability.
A system that generates responses to queries by alternately performing selection and inference steps, producing trace data that provides interpretable natural language explanations of how the response was generated, ensuring transparency and accountability.
The system provides high-quality responses while offering an interpretable trace of the inferences made, enhancing trust, particularly in safety-critical environments, by providing a clear explanation of decision-making processes.
Smart Images

Figure 2025517593000001_ABST
Abstract
Description
Technical Field
[0001] This specification relates to generating responses to query inputs using neural networks.
Background Art
[0002] A neural network is a machine learning model that utilizes one or more layers of non-linear units to predict an output for a received input. Some neural networks include one or more hidden layers in addition to an output layer. The output of each hidden layer is used as input to the next layer in the network, i.e., the next hidden layer or the output layer. Each layer of the network generates an output from the received input according to the current values of each set of parameters.
Summary of the Invention
Means for Solving the Problems
[0003] This specification describes a system implemented as a computer program on one or more computers in one or more locations that generates responses to queries regarding an environment.
[0004] In particular, the system receives a context input that includes context information regarding the environment and a query input that includes a query regarding the environment.
[0005] The system then uses the context input to generate a response to the query, such as a natural language response.
[0006] The system generates the response by repeatedly and alternately performing selection steps and inference steps. As a result, the system can generate trace data that provides natural language, interpretable information characterizing how the system arrived at its response.
[0007] Certain embodiments of the subject matter described herein may be implemented to realize one or more of the following advantages.
[0008] One problem with using neural network-based techniques to make control decisions, for example, to control an agent or a manufacturing plant, is that it is often difficult to infer why a particular decision was made, i.e., the neural network is a "black box" that is only tasked with performing the task. Similarly, when diagnosing faults in a mechanical system, it can be useful to know the reasons behind the response.
[0009] Embodiments of the described system can address this problem by providing a reasoning trace that can be presented to the user, given to the user, or stored for later review. This can be particularly useful when control decisions or fault diagnoses relate to the safe operation of a machine agent or a manufacturing plant. Embodiments of the described system can provide the reasoning trace as a sequence of logical steps presented as natural language statements that are interpretable by humans, for example, as natural language statements that connect a query to a response within a causal chain.
[0010] More specifically, the described technique generates a response by repeating and alternating between two steps: 1) a selection involving choosing a subset of relevant information sufficient to perform a single step of inference, and 2) an inference that looks only at the limited information provided to the inference by the selection output and uses that information to infer a new intermediate portion of evidence on the way to generating a final answer. This ensures that the intermediate inference output (and, optionally, the selection output) provides an interpretable inference trace that justifies the final answer. Moreover, the inferences generated by the described technique are causal, in that each step follows from and depends on the previous step, and each inference is made independently based only on the limited information provided by the selection output without direct access to the query input or the previous inference step. As a result, a high-quality response is obtained, and at the same time, an interpretable natural language "trace" of the inferences performed by the system while generating the response is provided.
[0011] Details of one or more embodiments of the subject matter of this specification are set forth in the accompanying drawings and the description below. Other features, aspects, and advantages of the subject matter will become apparent from the description, the drawings, and the claims.
Brief Description of the Drawings
[0012]
Fig. 1A
Fig. 1B
Fig. 1C
Fig. 2
Fig. 3
Modes for Carrying Out the Invention
[0013] The same reference numbers and names in the various drawings indicate like elements.
[0014] This specification describes a system implemented as a computer program on one or more computers in one or more locations that uses context information about an environment to generate responses to queries about the environment.
[0015] FIG. 1A shows an exemplary selection-inference system 100. The selection-inference system 100 is an example of a system that can be implemented as a computer program on one or more computers in one or more locations where the systems, components, and techniques described below are implemented.
[0016] The selection-inference system 100 can receive a query input 110 and a context input 120 and use the context input 120 to generate a response 150 to the query input 110.
[0017] The query input 110 includes a query about the environment. For example, each query can be a text fragment in natural language.
[0018] The context input 120 includes context information that provides context about the environment. More specifically, the context information in the context input 120 includes one or more natural language statements each representing a rule or fact about the environment.
[0019] In some implementations, the environment is a real-world environment, and system 100 facilitates reasoning in the real-world environment, for example, to logically and physically control a real-world system within the environment. As described later, in some implementations, system 100 can also provide trace data that includes a causal explanation for why a particular response 150 was generated, which can promote trust in the system, especially in environments where safety is critical.
[0020] For example, in some implementations, the environment is a real-world environment, and response 150 is used to control a machine agent, such as a robot or an autonomous or semi-autonomous vehicle, operating within the real-world environment to perform a task.
[0021] FIG. 1B shows an example of the operation of selection and speculation system 100 when the environment is a real-world environment navigated by machine agent 102.
[0022] In the example of FIG. 1B, agent 102 is shown as a vehicle. However, more generally, machine agent 102 can be any suitable agent that is controlled by a control system when the agent navigates the real-world environment.
[0023] Agent 102 includes one or more sensors 104 that capture observations of the environment when agent 102 navigates the environment, for example, at specified time intervals.
[0024] For example, the observation record may include one or more of, for example, an image, object position data, and sensor data, such as image, distance, or position sensor data, or sensor data from an actuator, for capturing the observation record when an agent interacts with the environment. For example, in the case of a robot, the observation record may include one or more of data characterizing the current state of the robot, such as joint position, joint velocity, joint force, torque or acceleration, such as gravity compensation torque feedback, and the global or relative pose of an item grasped by the robot. In the case of a robot or other mechanical agent or vehicle, the observation record may similarly include one or more of position, linear or angular velocity, force, torque or acceleration, and the global or relative pose of one or more parts of the agent. The observation record may be defined in 1, 2 or 3 dimensions and may be absolute and / or relative observation records. The observation record may include, for example, detected electronic signals such as motor current or temperature signals, and / or, for example, image or video data from a camera or LIDAR sensor, such as data from the agent's sensors or data from sensors placed separately from the agent in the environment.
[0025] Agent 102 is also associated with a control system 106 that generates a control signal for controlling agent 102 using the observation record generated by sensor 104. In particular, control system 106 first determines the appropriate actions for agent 102 to perform, such as, for example, navigating to a specific location, identifying a specific object, moving a specific object to a given location, manipulating a specific object in some way, etc., as part of performing a specified task, and then generates a control signal for causing agent 102 to perform the action, thereby generating a control signal for causing agent 102 to follow a planned trajectory through the environment.
[0026] The control system 106 can be deployed on the agent 102, can be deployed away from the agent 102, and can transmit a control signal to the agent 102 via a data communication network.
[0027] The control signal may be a control input for controlling the agent. For example, when the agent is a robot, the control signal may be, for example, torque for the joints of the robot or a high-level control command. As another example, when the agent is an autonomous or semi-autonomous land, air, or marine vehicle, the control signal may include actions for controlling the navigation of the vehicle, such as steering, and movement, such as braking and / or accelerating the vehicle. For example, the control signal may be, for example, torque to a control surface or other control element, such as a steering control element of the vehicle, or a high-level control command.
[0028] In other words, the control signal may include, for example, position, velocity, or force / torque / acceleration data for one or more joints of a robot or components of another mechanical agent.
[0029] In these examples, like the control system 106, the system 100 can be deployed on the agent 102 and can also be deployed away from the agent 102.
[0030] In these implementations, the system 100 is used to provide an additional control layer on top of the control system 106, and the system 100 or another component can determine the query input 110 based on the information received by the control system.
[0031] That is, the query input 110 may be determined by receiving a control signal from the agent control system 106, and then, for example, determining one or more natural language queries regarding the agent 102 in the environment from the control signal.
[0032] In some implementations, the agent control system 106 is an autonomous or semi-autonomous control system that autonomously or semi-autonomously controls, for example, the navigation of the agent 102.
[0033] In some other implementations, the agent control system 106 has an interface for receiving control commands, for example, from a human operator.
[0034] In both of these applications, the described system 100 may be used to provide an additional control layer, for example, for safety purposes. For example, the described system 100 may be used to prevent the control of the machine agent 102 in a way that may be dangerous or violate one or more rules.
[0035] One or more rules regarding the control of the agent may be explicitly included, for example, as part of the context input, or may be implicit, for example, when training a neural network used by the system 100, especially when each of these includes a language model as described later and may be included as a natural language statement. As an example, such rules may include rules regarding the permitted movement of a vehicle, such as traffic rules. As another example, such rules may include rules regarding decisions to be made to ensure the safe behavior of the machine agent, for example, to prevent damage to the machine agent or to humans.
[0036] Accordingly, each query input 110 may relate to an action to be performed by the agent, for example, an action being considered by the control system 106.
[0037] For example, the query input 110 can define an action to be performed by the machine agent 102 in the form of a question, such as "Should it turn left?" or "Is it safe for the agent to turn left?".
[0038] As another example, the query input 110 can ask what action should be performed by the machine agent, e.g., "Which way should the agent turn?"
[0039] In the case of a robot, the query input 110 often relates to a sub-task out of a series of sub-tasks to be performed to carry out the task, e.g., "What to do next?" or "Pick up object X?" The sub-tasks themselves can include a series of primitive actions for the robot's moving parts to, e.g., open the gripper.
[0040] Generally, the query input 110 includes one or more natural language queries. More specifically, the query input 110 can include a natural language description that defines the information that the response from the system should provide. That is, the query input 110 explicitly or implicitly determines what is required for the response.
[0041] The response 150 to the query input 110 is used to control the machine agent 102 in the real-world environment. More specifically, the response 150 is used to control the actions to be performed by the machine agent 102.
[0042] As an example, the response 150 may prevent an action that would otherwise be performed, i.e., the response can determine whether the action defined by the query input is to be performed.
[0043] As another example, for instance, when the query input implicitly or explicitly requests that an action be determined, the response 150 can define the action to be performed.
[0044] In these implementations, obtaining the context input 120 may include obtaining one or more observational records of the real-world environment, which may include one or more observational records of the agent since the environment includes the agent. The observational records may be obtained from one or more sensors that may be, but need not be, the agent's sensor 104. As described above, the observational records may include still or moving images, which, as used herein, may include LIDAR point clouds and / or other sensor data from one or more sensors that detect the state of the environment or the agent.
[0045] One or more observational records may be processed, for example, by a first machine learning model to generate a natural language representation of the one or more observational records, i.e., to generate natural language text that describes the observational records included in the context input 120.
[0046] There are many different types of machine learning models that can be used to achieve this. For example, so-called vision-language models are generally configured to describe an image or video using natural language, for example, to perform an image or video captioning task. More generally, such models can perform many different types of image processing tasks by formulating the task as a text generation problem, for example, to detect or classify objects in an image or video. Accordingly, other machine learning models can be trained to generate natural language text that describes data from other types of sensors, for example, to represent physical positions or forces as natural language statements that describe the agent or the environment or parts thereof. The natural language representation of the one or more observational records is used to provide one or more of the natural language statements in the context input 120.
[0047] Furthermore, the context input 120 may include one or more rules regarding the current location of the agent 102 in the environment that are not directly generated from the observation records generated by the sensor 104.
[0048] For example, the context input 120 can also have rules of co-notified knowledge input from a trained machine learning model, a driving rule manual weight, which are manually designed or obtained from a knowledge graph or the Internet. Examples of such rules include "The electrical system of a car shorts when wet", "A car reaches where it goes when it travels in any direction", "A car is not safe to drive when its electrical system is broken", etc. Other examples of such rules include "The speed limit at the current location is 30 mph", "A right turn is permitted after stopping at this red signal", "It is prohibited to turn across a double yellow line", etc.
[0049] Thus, in these examples, the system 100 can be used to evaluate the possible course of actions being considered by the control system 106, and then the actions are transmitted as control signals for the agent 102.
[0050] In some other implementations, the environment is a real - world environment including a manufacturing plant, for example, a chemical, bio, or mechanical product manufacturing plant, or a food manufacturing plant. As used herein, "manufacturing" a product includes purifying starting materials to create the product, or processing the starting materials, for example, to remove contaminants, to produce a cleaned or recycled product. A manufacturing plant may include multiple manufacturing units such as containers for chemical or biological substances, or machinery for processing solids or other materials. The manufacturing units are configured such that intermediate versions or components of the product can move between manufacturing units during the manufacture of the product, for example, via pipes or mechanical conveyance. In an implementation, the system is used to control one or more of the manufacturing units or to control the movement of intermediate versions or components of the product between the manufacturing units.
[0051] Thus, in these implementations, obtaining the context input 120 may then include obtaining one or more observational records from one or more sensors of the manufacturing unit or of the movement. The sensors may include sensors configured to detect the configuration of the manufacturing unit or movement, such as mechanical movement or force, pressure, temperature, current, voltage, frequency, impedance, etc. electrical conditions, the amount, level, flow rate / movement rate or flow / movement path of one or more materials, physical or chemical conditions such as physical state, shape or configuration or chemical state such as pH, the mechanical configuration of the unit, etc., the configuration of the unit, or valve configuration, image or video sensors for capturing image or video observational records of the manufacturing unit or of the movement, or any other suitable type of sensor. In an implementation, one or more observational records are processed, for example, as described above, to generate a natural - language expression of the one or more observational records. The natural - language expression of the one or more observational records is used to provide one or more of the natural - language statements of the context information.
[0052] The query input may relate to actions that control the operation or movement of one or more of the manufacturing units. The response to the query input is used to control the operation or movement of one or more of the manufacturing units. For example, the response to the query input may be used to control energy or other resource usage, e.g., to minimize it, or to control manufacturing to obtain the desired quality or characteristics of the product. For example, the actions may include actions that control items of factory equipment or settings that affect the movement of manufacturing units or products or intermediates or their components, e.g., actions that change to adjust or turn on / off items of equipment or the manufacturing process.
[0053] In some implementations, the manufacturing factory has a factory control system for controlling the manufacturing units or for controlling movement. The query input may be generated by receiving a control signal from the factory control system and generating one or more natural language queries for the query input from the control signal. Similar to what has been described above, the factory control system may be autonomous, semi-autonomous, or human-controlled.
[0054] Similar to what has been described above, the system may implement rules for controlling or restricting, e.g., energy or other resource allocation, or for ensuring the target quality or characteristics of the product, or for suppressing, within a safe range, the operation of the factory, e.g., of the manufacturing units.
[0055] In some implementations, the environment is a real-world environment, and the method is used to diagnose faults in a machine system operating in the real-world environment. Then, obtaining the context input may include obtaining from one or more sensors, for example, as described above, one or more observational records of the machine system (here including observational records of the operation of the machine system). These are processed, for example, as described above, to generate natural language expressions of one or more of the observational records used to provide one or more of the natural language statements of the context information. In these implementations, the query input relates to the operation of the machine system, and the response to the query input is used to identify faults in the machine system. For example, the query input may include general queries such as "Is the system operating normally?" or "What's wrong with the system?" or specific queries such as "Is there a fault in component X?".
[0056] As another example, the environment may be an educational environment. For example, the system can be deployed as part of an educational software program that assists a user in learning or practicing one or more corresponding skills. In these examples, the context input 120 may include natural language statements that describe or refer to scenarios or scenes in the real world or an imaginary environment, and the query input 110 may be a question about a scenario or scene that asks for logical inferences. As will be described below, the trace data generated by the system 100 as part of generating the response 150 can be used to provide the user with insights into the logical inferences required to generate the response 150. That is, the trace data generated by the system 100 can be used to generate explanations of questions and corresponding responses to assist the user in learning one or more skills.
[0057] As a specific example, an educational software program can generate (or receive this data from an external source) questions and accompanying context information. An example of such a question could be, for example, "Why do astronauts need an oxygen backpack?" The system 100 (given the context of basic facts) may generate an inference trace as follows to assist the user in understanding the concept. That is, "(1) Humans need oxygen to survive. There is no oxygen in space. Therefore, humans need a supply of oxygen to survive in space. (2) Oxygen can be supplied to astronauts by an oxygen backpack, and humans need a supply of oxygen to survive in space. Therefore, humans can use an oxygen backpack to help astronauts survive in space."
[0058] As another example, the environment may be an information retrieval environment. For example, the system can be deployed as part of a search engine or other software that enables a user to search for information in a corpus of documents, such as the Internet or another electronic document corpus. In these examples, the query input 110 can be any appropriate natural language query, and the context input 120 can include relevant statements identified from the corpus of documents, i.e., by searching the corpus using conventional information retrieval techniques. The system 100 can then use the context input 120 to generate a response 150, i.e., to generate a response 150 to the query input 110 such that the user can view the trace of the logical inferences required to enhance the user's confidence in the accuracy of the response 150 generated by the system 100.
[0059] Returning to the description of FIG. 1A, to generate the response 150, the system 100 uses a selection neural network 130 and a speculation neural network 140.
[0060] More specifically, system 100 generates a response by performing a plurality of update iterations.
[0061] In each update iteration, system 100 performs a selection step by using a selection neural network 130 to select an appropriate subset of the natural language statements in the context input 120.
[0062] System 100 then performs a speculation step using the selected appropriate subset and a speculation neural network 140 to generate a new fact, i.e., a new natural language statement representing the "speculated" fact. The fact is called "speculated" because it does not exist in the context input 120 and is determined by system 100 from the context input 120.
[0063] In each update iteration other than the last update iteration, system 100 updates the context input 120 by using the new fact, i.e., by adding a natural language statement representing the new fact to the context input 120.
[0064] In the last update iteration, system 100 uses the new fact to generate the response 150. For example, system 100 can provide a natural language output derived from the new natural language statement generated in the last update iteration as the response 150.
[0065] Thus, system 100 repeatedly adds new facts to the context input 120 and then uses the new facts generated in the last update iteration to generate the response 150 to the query input 110.
[0066] More specifically, the selection neural network 130 is a neural network configured to process a selection input that includes a context input 120 (at the current update iteration time) and a query input 110, and generate a selection output that includes one or more of the natural language statements from the context input 120. Since the context input 120 is updated in each iteration, the selection output can identify different natural language statements in different update iterations.
[0067] The inference neural network 140 is a neural network configured to process an inference input that includes the selection output and generate an inference output that includes a natural language statement representing a new fact for the update iteration. Since the selection output can be different for different update iterations, the inference output can also be different over the update iterations.
[0068] In some implementations, the inference input does not include either the context input 120 or the query input 110. That is, the inference neural network 140 only has access to an appropriate subset of the context input 120 included in the selection output, and does not have access to the remaining natural language statements in the context input 120 or the query input 110.
[0069] Accordingly, in each update iteration, the system 100 processes a selection input that includes a context input 120 (at the update iteration time) and a query input 110 using the selection neural network 130 to generate a selection output for the update iteration that includes one or more of the natural language statements from the context input 120. These natural language statements can include statements that were included in the original context input obtained by the system 100, statements that were added to the context input in previous update iterations, or both.
[0070] In some implementations, the selection neural network 130 is configured to generate, i.e., regress, a text sequence including one or more of the natural language statements from the context input 120. In some of these implementations, the system 100 ensures that the tokens generated by the neural network 130 are either (i) a valid continuation of some natural language statement in the context input 120, (ii) the start of another natural language statement in the context input 120, or optionally (iii) one or more predetermined tokens that mark, for example, a separator between natural language statements or the end of the regressed text sequence. To this end, constrained sampling can be utilized.
[0071] In some other implementations, the neural network 130 can be configured to generate a text sequence that includes placeholder references to natural language statements in the context input 120. For example, it can generate a sentence such as "Sent 1. Know that 4 was sent", where "Sent 1" is a reference to the first sentence in the context input 120 and "Know that 4 was sent" is a reference to the fourth sentence in the context input 120. The system can then substitute in the actual sentences to generate the selection output.
[0072] In yet other implementations, each selection step includes a plurality of internal steps. In each internal step, the selection neural network 130 is used to add a new statement to the selection output for the selection step. This will be described in more detail below with reference to FIG. 3.
[0073] System 100 then processes the speculative input, including the selected output for the update iteration, using the speculative neural network 140 to generate a speculative output that includes natural language statements representing new facts about the update iteration. As described above, if the update iteration is not the last update iteration, System 100 updates the context input 120 so as to include natural language statements in the speculative output for the update iteration.
[0074] Thanks to the way System 100 generates the response 150, System 100 can also provide, as output, trace data 160 that provides an interpretable natural language summary of the causal inferences performed by System 100 in generating the response 150. In particular, System 100 alternates between two steps: 1) a selection involving choosing a subset of relevant information sufficient to make a single-step speculation, and 2) a speculation that looks only at the limited information provided to the speculation by the selected output and uses that information to speculate about new intermediate portions of evidence on the way to generating the final answer, so that the system ensures that the intermediate speculation output (and, optionally, the selected output) provides an interpretable inference trace to justify the final answer. Moreover, the inferences generated by System 100 are causal, because each step follows from and depends on the previous step, and each speculation is made independently based only on the limited information provided by the selected output without direct access to the query input or the previous inference step.
[0075] In some implementations, System 100 provides the trace data 160 to the user, for example, along with the response 150. In some other implementations, System 100 stores the trace data 160 associated with the data identifying the response 150.
[0076] For example, the trace data 160 can be accessed later by a user who requests to view the "reasoning" performed by System 100 that resulted in a particular control signal being sent to the machine agent.
[0077] As a simplified example, when context input 120 indicates that there is a pedestrian near the lane the agent is on, and query input 110 is "What action should be taken?", for example, either "Continue driving" or "Stop driving", the trace data may include alternative outputs such as "There is a person crossing the road" and "Know that it is not safe to drive when someone is crossing the road", and the inference can be "Therefore, it is not safe to drive".
[0078] As another simplified example, when context input 120 indicates that there is a lake near the vehicle, and query input 110 asks whether the vehicle should drive towards the lake, the trace data 160 may include intermediate inferences such as "Therefore, the vehicle will reach the lake" → "The lake means getting wet" → "Getting wet means short - circuit" → "Short - circuit means not safe", and the response 150 can be "Should not proceed".
[0079] The selection neural network 130 can generally have any suitable architecture such that a neural network is used to map the selection input to an appropriate subset of the context input 120. Similarly, the inference neural network 140 can generally have any suitable architecture such that a neural network is used to map the inference input to a new natural - language statement.
[0080] As a specific example, both the selection neural network 130 and the inference neural network 140 can be respective language model neural networks. Generally speaking, a language model neural network is a neural network that is trained such that, when given a text prompt that includes a sequence of tokens in a natural language, the neural network can generate the next token in the sequence. This process can be repeated to extend the text prompt by one token at a time in order to generate a natural language output, i.e., to autoregressively generate the natural language output one token at a time. At each "time step", the language model neural network processes the current sequence and generates a probability distribution over the vocabulary of tokens. The next token can then be selected using the probability distribution, for example, by sampling from the distribution using nucleus sampling or another sampling technique, or by selecting the most probable token. The tokens in the vocabulary may include any combination of various tokens, such as words, subwords, characters, punctuation marks and other symbols, as well as numbers. Generally speaking, a language model neural network is trained on a corpus of text consisting of tokens from the vocabulary (and optionally, other tokens that can be mapped to tokens not in the specified vocabulary) to predict the next token in a sequence of tokens from the training data.
[0081] It is surprising that large language model neural networks can perform tasks that they were not explicitly trained to perform, but they do so quite well. For example, those neural networks can perform translation tasks (assuming the training corpus included words in different languages), arithmetic, and many other tasks. Language model neural networks can be made to perform a specific task by providing a natural language description of the desired response as an input or "prompt". The prompt can be a few-shot prompt in which a few queries, e.g., 1 to 10 examples and an exemplary output, are given in the text prior to the actual query. Alternatively, or additionally, the language model neural network can be "fine-tuned" to perform a specific task by obtaining a pre-trained language model neural network trained on a large corpus of the aforementioned examples and then further training a portion of all of the language model neural networks with a relatively small number of examples specific to the type of task to be performed.
[0082] Accordingly, trained language model neural networks can perform control and diagnostic tasks of the type described. If the system is to conform to rules in generating a response, the rules can be included in the context information, e.g., in the prompt, and / or as statements in the corpus of training data or in the data used to fine-tune the language model neural network.
[0083] In other words, in some implementations, the selection neural network 130 and the inference neural network 140 are the same, pre-trained language model neural network. In these implementations, each selection input includes a few-shot prompt of a type that causes the language model to generate a selection output, and each inference input includes a different type of few-shot prompt that causes the language model to generate an inference output.
[0084] For example, a few-shot prompt for the inference neural network 140 may include one or more pairs of exemplary inference inputs and inference outputs arranged according to a predetermined syntax (followed by a subset of the context inputs identified by the corresponding selected output), and a few-shot prompt for the selection neural network 130 may include one or more pairs of exemplary selection inputs and selection outputs arranged according to a predetermined syntax (followed by context inputs and query inputs arranged according to the predetermined syntax).
[0085] For example, the few-shot prompt within each selection input may be in the following form. #n-shot prompt #First example. <Context 1> <Query 1> #Exemplary selection <Fact>. Knowing <fact>[and <fact>]*. Therefore, ... #Problem to be solved. <Context> <Query>
[0086] In this example, the statements following # are not included in the prompt (or are optional) and are only included in the example for explanation. <Context> represents the natural language text within the context input, <Query> represents the query input, "..." indicates that one or more additional examples in the same format follow after the first example, each <Fact> is a natural language statement copied from the corresponding context, and [and <Fact>]* means that the system can select multiple facts for each inference step, where the total number of facts in a given inference step is a hyperparameter.
[0087] As another example, the few-shot prompt in each speculative input may be in the following form. #n-shot speculative prompt #First example. Knowing that <fact>. <fact> [and <fact>]*. Therefore, <new fact>. ... #Problem to be solved. <Output of the selection step>. Therefore,
[0088] In this example, the <new fact> in each example in the prompt is an inferred fact that is generated (``inferred'') from the <fact> in the example, and the <output of the selection step> is one or more facts in the selection output for the selection step that are formatted in the same way as the examples in the few-shot prompt.
[0089] The language model neural network may be a large language model neural network, e.g., having more than 1 billion, 10 billion or 100 billion trained parameters. The language model neural network may be trained on more than 100 billion, 1000 billion or 1 trillion words or tokens representing words or other text tokens, e.g., sub-words (also known as "word pieces").
[0090] In some other implementations, both the selection neural network 130 and the inference neural network 140 are language model neural networks having the same architecture and pre-trained on the same large corpus, but the selection neural network 130 is fine-tuned with a first dataset of exemplary selection input and selection output pairs, and the inference neural network 140 is fine-tuned with a second dataset of exemplary inference input and selection input pairs, or both. For example, by fine-tuning the selection neural network 130, the selection neural network 130 may be able to more effectively generate an output that includes placeholder references from the context input to the fact, as described above. In some of these implementations, each selection input and each inference output each include their respective few-shot prompts, and in some other of these implementations, the few-shot prompts are not included.
[0091] In some implementations, the language model neural network is an autoregressive transformer neural network, and the transformer neural network is characterized by having a series of self-attention neural network layers. The self-attention neural network layer has attention layer inputs for each element of the input, and is configured and can be used to apply an attention mechanism to the attention layer inputs to generate attention layer outputs for each element of the input. There are many different attention mechanisms that can be used. In some implementations, the language model neural network can be a mixture of experts model.
[0092] FIG. 1C shows a simplified example of the selection and inference steps implemented when the system 100 is used to control a vehicle, e.g., an autonomous or semi-autonomous vehicle.
[0093] In the example of FIG. 1C, the query input 176 is "Should the vehicle turn left?", and the context input 174 (at the time of the exemplary selection step shown in FIG. 1C) is "There is a lake on the left, the lake has a lot of water, and the water may damage the vehicle." To perform the selection step, system 100 generates a selection input that includes the k-shot prompt 172 (a portion of the exemplary context for one of the examples in the k-shot prompt 172, the exemplary query input for the example, and the exemplary selection output for the example are shown in FIG. 1C), the context input 174, and the query input 176. System 100 uses the selection neural network 130 to process the selection input and generate a selection output 180 stating that "The lake has a lot of water, and thus the water may damage the vehicle."
[0094] System 100 then generates a speculation input that includes the k-shot prompt 182 (one of the examples in the k-shot prompt 182 is shown in FIG. 1C) and the selection output 180. System 100 uses the speculation neural network 140 to process the speculation input and generate a speculation output 182 stating the newly speculated fact, that is, "Entering the lake may damage the vehicle." System 100 then adds the newly speculated fact to the context input 174 at 188 for use in the next selection step.
[0095] FIG. 2 is a flowchart of an exemplary process 200 for generating a response to a query input. For convenience, process 200 is described as being implemented by a system consisting of one or more computers at one or more locations. For example, a selection speculation system appropriately programmed according to this specification, such as the selection speculation system 100 of FIG. 1A, can implement process 200.
[0096] The system obtains a context input that includes context information (step 202). The context information includes one or more natural language statements each representing a fact or rule regarding the environment.
[0097] The system receives a query input that includes one or more queries regarding the environment (step 204).
[0098] The system then generates a response to the query input, such as a natural language output, by performing update iterations until an end criterion is satisfied. For example, the end criterion can be satisfied after a threshold number of iterations are performed. As another example, the end criterion can be satisfied if the same selected output is generated in a threshold number of consecutive update iterations. As another example, the end criterion can be satisfied if the same inferred output is generated in a threshold number of consecutive update iterations. As yet another example, the system can be trained to receive a stop input that includes the query input, the most recent inferred output, and optionally one or more of (i) a pre-generated selected output, inferred output, or both, or (ii) the context input, and to process the stop input to generate an indication as to whether the most recent inferred output is a valid response to the query input. As another example, the system can determine whether the most recent inferred output is a valid response to the query input by applying one or more rules based on, for example, a string match between the inferred output and the query input.
[0099] In each update iteration, the system processes a selection input that includes the context input and the query input using a selection neural network to generate a selection output for the update iteration that includes one or more of the natural language statements from the context input. That is, the selection output generally includes an appropriate subset of the natural language statements in the context input.
[0100] The system then processes the speculative input, including the selected output for the update iteration, using a speculative neural network to generate a speculative output that includes a natural language statement representing a new fact about the update iteration (step 208).
[0101] In each update iteration other than the last update iteration, i.e., when the system determines that the termination criterion is not satisfied after performing an update iteration, the system updates the context input so as to include the natural language statement in the speculative output for the update iteration (step 210).
[0102] In the last update iteration, i.e., when the system determines that the termination criterion is satisfied after performing an update iteration, the system generates a response using the natural language statement in the speculative output for the last update iteration (step 212). The system can also generate an inference trace that includes the speculative output in each update iteration other than the last update iteration, and optionally, the natural language statement in the selected output for the update iteration (either provided as part of the response or stored for future use).
[0103] FIG. 3 is a flowchart showing an exemplary process 300 for generating a selected output in a given update iteration. For convenience, process 300 is described as being performed by a system consisting of one or more computers at one or more locations. For example, a selection speculation system appropriately programmed in accordance with this specification, such as selection speculation system 100 of FIG. 1A, can perform process 300.
[0104] In the example of FIG. 3, the system generates a selected output by selecting a respective natural language statement from the context input in each of a sequence of one or more selection iterations. For example, the system can perform a predetermined number of selection iterations to add a predetermined number of natural language statements to the selected output. As another example, the system can perform selection iterations until some other criterion is satisfied.
[0105] In each selection iteration, the system generates an input for the selection iteration that includes the context input, the query input, and any natural language statements selected in any previous selection iteration preceding the selection iteration in the sequence (step 302). For example, when the selection input to the selection neural network includes a few-shot prompt as described above, the system can modify the few-shot prompt (which already includes the context input and the query input) to include any natural language statements selected in any previous selection iteration preceding the selection iteration in the sequence, at the corresponding location in a given syntax.
[0106] For each set of natural language statements in the context information, the system then uses the selection neural network to process the input for the selection iteration to determine the likelihood assigned to the natural language statement by the selection neural network (step 304).
[0107] For example, the system can select, as the set of natural language statements, the natural language statements in the context input that have not yet been selected in a previous selection iteration.
[0108] As described above, the selection neural network may be a language model neural network that predicts a probability distribution for a text token when the current text token in the sequence is given. In this case, the system can determine the likelihood assigned to a natural language statement by combining, for example, multiplying, the probabilities assigned to each text token in the natural language statement by the selection neural network. The probability assigned to a given token is determined by processing any token preceding the given token in the natural language statement using a selection neural network conditioned on the input to the selection neural network. For example, the probability for a given token in the probability distribution generated by processing a combined sequence that includes the input for the selection iteration and any token preceding the given token in the natural language statement using the selection neural network.
[0109] The system then selects the natural language statement with the highest likelihood from the set of natural language statements (step 306).
[0110] This specification uses the term "configured" in connection with system and computer program components. That a system consisting of one or more computers is configured to perform a particular operation or action means that the system has installed software, firmware, hardware, or a combination thereof that causes the system to perform the operation or action during operation. That one or more computer programs are configured to perform a particular operation or action means that the one or more programs include instructions that, when executed by a data processing apparatus, cause the apparatus to perform the operation or action.
[0111] Embodiments of the subject matter and the functional operations described in this specification can be implemented in digital electronic circuitry, or in tangibly embodied computer software or firmware, or in computer hardware including the structures disclosed in this specification and their structural equivalents, or in one or more of them in combination. Embodiments of the subject matter described in this specification can be implemented as one or more computer programs, i.e., as one or more modules of computer program instructions encoded on a tangible non-transitory storage medium for execution by, or to control the operation of, a data processing apparatus. The computer storage medium may be a machine-readable storage device, a machine-readable storage substrate, a random or serial access memory device, or a combination of one or more of them. Alternatively or in addition, the program instructions may be encoded on an artificially generated propagated signal, e.g., a machine-generated electrical, optical, or electromagnetic signal, that is generated to encode information for transmission to an appropriate receiver apparatus for execution by a data processing apparatus.
[0112] The term "data processing apparatus" refers to data processing hardware and includes, by way of example, any kind of apparatus, device, and machine for processing data, including a programmable processor, a computer, or multiple processors or computers. The apparatus may be, or further include, special purpose logic circuitry, e.g., an FPGA (field programmable gate array) or an ASIC (application specific integrated circuit). The apparatus may optionally include, in addition to hardware, code that creates an execution environment for computer programs, e.g., code that constitutes processor firmware, a protocol stack, a database management system, an operating system, or a combination of one or more of them.
[0113] A computer program may be referred to as, or described as, a program, software, a software application, an app, a module, a software module, a script, or code, and may be written in any form of programming language, including a compiled or interpreted language, or a declarative or procedural language, and may be deployed in any form, including as a stand-alone program, or as a module, component, subroutine, or other unit suitable for use in a computing environment. The program may or may not correspond to a file in a file system. The program may be stored in another program or data, such as in a file portion that holds one or more scripts stored in a markup language document, in a single file dedicated to the program in question, or in multiple cooperating files, such as files that hold one or more modules, subprograms, or portions of code. The computer program may be deployed to be executed on one computer located in one place or on multiple computers, or may be distributed across multiple places and interconnected by a data communication network.
[0114] As used herein, the term "database" is used broadly to refer to any collection of data, and the data can be stored on a storage device in one or more locations, without the need to be structured in any particular way or at all. Thus, for example, an indexed database can contain multiple collections of data, each of which can be organized and accessed in a different way.
[0115] Similarly, as used herein, the term "engine" is used broadly to refer to a software-based system, subsystem, or process that is programmed to perform one or more specific functions. Generally, an engine will be implemented as one or more software modules or components and installed on one or more computers located in one or more locations. In some cases, one or more computers will be dedicated to a particular engine, and in other cases, multiple engines may be installed and operating on the same computer or multiple computers.
[0116] The processes and logical flows described herein may be implemented by one or more programmable computers executing one or more computer programs to perform functions by operating on input data and generating output. The processes and logical flows may also be implemented by special purpose logic circuitry, such as an FPGA or ASIC, or by a combination of special purpose logic circuitry and one or more programmed computers.
[0117] A computer suitable for the execution of a computer program may be based on a general purpose or special purpose microprocessor or both, or any other kind of central processing unit. Generally, the central processing unit will receive instructions and data from read only memory or random access memory or both. The essential elements of a computer are a central processing unit for executing or carrying out instructions, and one or more memory devices for storing the instructions and data. The central processing unit and the memory may be supplemented or incorporated in special purpose logic circuit elements. Generally, a computer also includes one or more mass storage devices for storing data, such as magnetic, magneto-optical disks, or optical disks, or is operatively coupled to a mass storage device for receiving data therefrom, transferring data thereto, or both, although a computer need not have such devices. Moreover, a computer may be incorporated in another device, such as, by way of example only, a mobile telephone, a personal digital assistant (PDA), a mobile audio or video player, a game console, a global positioning system (GPS) receiver, or a portable storage device, such as a universal serial bus (USB) flash drive.
[0118] Computer-readable media suitable for storing computer program instructions and data include, by way of example, all forms of non-volatile memory, media and memory devices, such as semiconductor memory devices, e.g., EPROM, EEPROM, and flash memory devices, magnetic disks, e.g., internal hard disks or removable disks, magneto-optical disks, and CD ROM and DVD-ROM disks.
[0119] To provide interaction with a user, embodiments of the subject matter described herein can be implemented on a computer having a display device, such as a CRT (cathode ray tube) or LCD (liquid crystal display) monitor, for displaying information to the user, as well as a keyboard and a pointing device, such as a mouse or trackball, by which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user. For example, the feedback provided to the user can be any form of sensory feedback, such as visual feedback, auditory feedback, or tactile feedback, and the input from the user can be received in any form, including acoustic, voice, or tactile input. Further, the computer can interact with the user by sending documents to the devices used by the user and receiving documents from the devices, for example, by sending a web page to a web browser on the user's device in response to a request received from the web browser. Also, the computer can interact with the user by sending a text message or other form of message to a personal device, such as a smartphone running a messaging application, and receiving a response message in return from the user.
[0120] A data processing apparatus for implementing a machine learning model can also include, for example, a special-purpose hardware accelerator unit for machine learning training or production, i.e., for processing common and numerical calculation parts of inference and workload.
[0121] The machine learning model can be implemented and deployed using a machine learning framework, such as the TensorFlow framework, the Microsoft Cognitive Toolkit framework, the Apache Singa framework, or the Apache MXNet framework.
[0122] Embodiments of the subject matter described in this specification can be implemented in a computing system that includes back-end components, such as a data server, or middleware components, such as an application server, or front-end components, such as a graphical user interface or a web browser through which a user can interact with an implementation of the subject matter described in this specification, or a client computer having an app, or any combination of one or more such back-end, middleware, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication, such as a communication network. Examples of communication networks include local area networks (LANs) and wide area networks (WANs), such as the Internet.
[0123] A computing system can include clients and servers. Clients and servers are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on respective computers and having a client-server relationship to each other. In some embodiments, a server transmits data, such as HTML pages, for the purpose of displaying the data to a user who interacts with a user device, such as a device acting as a client, and receiving user input from the user. Data generated at the user device, such as the result of a user interaction, can be received at the server from the device.
[0124] This specification includes many specific implementation details, which should not be construed as limitations within the scope of any invention or the scope of what may be claimed, but rather as descriptions of features that may be specific to particular embodiments of a particular invention. Also, some of the features described herein in the context of separate embodiments can be implemented in combination in a single embodiment. Conversely, the various features described in the context of a single embodiment can also be implemented separately, or in any suitable sub-combination, in multiple embodiments. Furthermore, features are described above as acting in some combinations, and may even be initially claimed as such, but one or more features from a claimed combination can, in some cases, be deleted from that combination, and the claimed combination can be directed to a sub-combination, or a variant of a sub-combination.
[0125] Similarly, operations are shown in the drawings and recited in the claims in a particular order, but this should not be understood as requiring that such operations be performed in the particular order shown, or sequentially, or that all of the shown operations be performed to achieve the desired result. In some situations, multitasking and parallel processing may be advantageous. Moreover, the separation of the various system modules and components in the embodiments described above should not be understood as required in all embodiments, and it should be understood that the described program components and systems can generally be integrated together into a single software product, or packaged into multiple software products.
[0126] Certain embodiments of the subject matter have been described. Other embodiments are within the scope of the following claims. For example, the actions recited in the claims may be performed in a different order and still achieve desirable results. As one example, the processes illustrated in the accompanying figures do not necessarily require the particular order or sequence shown to achieve desirable results. In some cases, multitasking and parallel processing may be advantageous.
Description of the Signs
[0127] 100 Selection and Inference System, System 102 Machine Agent, Agent 104 Sensor 106 Control System, Agent Control System 130 Selection Neural Network, Neural Network 140 Inference Neural Network
Claims
1. A method implemented by one or more computers, comprising: obtaining a context input including context information, the context information including one or more natural language statements each representing a fact or rule regarding an environment; receiving a query input including a query regarding the environment; generating a response to the query input by performing update iterations until an end criterion is satisfied, the generating step comprising, in each update iteration: processing a selection input including the context input and the query input using a selection neural network to generate a selection output for the update iteration, the selection output including one or more of the natural language statements from the context input; processing a speculation input including the selection output for the update iteration using a speculation neural network to generate a speculation output including a natural language statement representing a new fact for the update iteration; in each update iteration other than the last update iteration: updating the context input to include the natural language statement in the speculation output for the update iteration; The method as described above.
2. The method according to claim 1, further comprising, as the response: (i) providing a natural language output derived from the natural language statement in the speculation output for the last update iteration; and (ii) providing an inference trace including the natural language statement in the speculation output in each update iteration other than the last update iteration.
3. The method according to claim 2, wherein the inference trace further includes the selection output for the update iteration.
4. The method according to any one of claims 1 to 3, wherein the speculation input does not include the context input or the query input.
5. The environment is a real-world environment, and the method is used to control a machine agent operating in the real-world environment to perform tasks. Obtaining the context input includes obtaining one or more observational records of the real-world environment from one or more sensors and processing the one or more observational records to generate a natural language representation of the one or more observational records. The query input relates to actions to be performed by the agent, The method is, using the natural language expression of the one or more observational records to provide one or more of the natural language statements of the context information, further comprising using the response to the query input to control the machine agent in the real-world environment. The method according to any one of claims 1 to 4.
6. The machine agent has an agent control system for controlling the actions of the machine agent, the query input includes one or more natural language queries, and the step of receiving the query input is receiving a control signal from the agent control system, generating the one or more natural language queries from the control signal. The method according to claim 5.
7. The machine agent includes an autonomous or semi-autonomous vehicle that navigates in the real-world environment, and the action includes an action for controlling the movement of the vehicle in the real-world environment. The method according to claim 5 or 6.
8. The environment is a real-world environment, The context input is derived from observational records characterizing the current state of the real-world environment, which are generated at least from measurement results from one or more sensors configured to detect the real-world environment, The query input includes data characterizing the planned navigation of the agent, The response to the query input characterizes the actions to be performed by the agent in response to the observational records, In particular, the agent is a robot or an autonomous vehicle. The method according to any one of claims 1 to 4.
9. The method according to claim 8, further comprising controlling the navigation of the agent based on the response to the query input.
10. The environment is a manufacturing plant for manufacturing a product, the manufacturing plant comprises a plurality of manufacturing units, the plurality of manufacturing units are configured such that intermediate versions or components of the product are movable between the manufacturing units during manufacture of the product, and the method is used for controlling one or more of the manufacturing units or for controlling the movement of the intermediate versions or components of the product between the manufacturing units. Obtaining the context input includes obtaining one or more observational records of the manufacturing unit or of the movement from one or more sensors, and processing the one or more observational records to generate a natural language representation of the one or more observational records. The query input relates to an action for controlling the operation of one or more of the manufacturing units or for controlling the movement. The method using the natural language representation of the one or more observational records to provide one or more of the natural language statements of the context information; further comprising using the response to the query input to control the operation of one or more of the manufacturing units or to control the movement. The method according to any one of claims 1 to 4.
11. The manufacturing plant has a plant control system for controlling the manufacturing units or for controlling the movement, the query input includes one or more natural language queries, and the step of receiving the query input includes receiving a control signal from the plant control system; generating the one or more natural language queries from the control signal. The method according to claim 10.
12. The environment is a real-world environment, and the method is used for diagnosing a fault in a machine system operating in the real-world environment. The step of obtaining the context input includes obtaining one or more observational records of the machine system from one or more sensors, and processing the one or more observational records to generate a natural language representation of the one or more observational records. The query input relates to the operation of the machine system. The method using the natural language expression of the one or more observation records to provide one or more of the natural language statements of the context information; further comprising using the response to the query input to identify a failure in the machine system; The method according to any one of claims 1 to 4.
13. The method according to any one of claims 1 to 12, wherein the query input includes a natural language description defining the information to be provided by the response.
14. further comprising processing the observation records by using a first machine learning model configured to process the observation records to generate natural language text describing the observation records, thereby generating at least one of the natural language statements in the context input; In particular, further comprising processing the observation records by using a second machine learning model configured to process (i) the observation records, (ii) the natural language text describing the observation records, or (iii) both, to generate natural language text characterizing one or more rules for determining new facts related to the observation records, thereby generating at least one of the natural language statements in the context input; The method according to any one of claims 1 to 13.
15. The method according to any one of claims 1 to 14, wherein the end criterion is satisfied when a threshold number of update iterations have been performed.
16. processing a selection input including the context input and the query input by using a selection neural network to generate a selection output for the update iteration, the selection output including one or more of the natural language statements in the context information; in each of a sequence of one or more selection iterations, selecting a respective natural language statement from the context information, the selecting comprising, in each selection iteration, generating an input for the selection iteration, the input including the context input, the query input, and any natural language statement selected in any previous selection iteration preceding the selection iteration in the sequence; Selecting each of the natural language statements by processing the input for the selection iteration using the selection neural network The method according to any one of claims 1 to 15 **Claim 17** Selecting each of the natural language statements of the context information for the selection iteration comprises For each set of natural language statements in the context information Processing the input for the selection iteration using the selection neural network to determine the likelihood assigned to the natural language statement by the selection neural network Selecting the natural language statement having the highest likelihood from the set of natural language statements The method according to claim 16 **Claim 18** The selection neural network and the speculation neural network are the same neural network The selection input includes a first few-shot prompt The speculation input includes a second different few-shot prompt The method according to any one of claims 1 to 17 **Claim 19** Generating the input for the selection iteration includes modifying the first few-shot prompt to include the natural language statement selected in any previous selection iteration preceding the selection iteration in the sequence, the method according to claim 18 when dependent on claim 16 or claim 17 **Claim 20** The selection neural network and the speculation neural network are the same pre-trained language model neural network, in particular, the language model neural network has more than one billion trained parameters, the method according to any one of claims 1 to 19 **Claim 21** One or more computer-readable storage media storing instructions that, when executed by one or more computers, cause the one or more computers to perform the operations of the method according to any one of claims 1 to 20 **Claim 22** A system comprising one or more computers and one or more storage devices storing instructions, which, when executed by the one or more computers, cause the one or more computers to perform the operations of the method according to any one of claims 1 to 20.
Citation Information
Patent Citations
Method of human-computer interactive interaction based on retrieval data, device, and electronic apparatus
JP2021111334A
Providing a response in a session
US20200202194A1
Cited By
Information processing system and method for vectorizing the meaning of words and searching for reasons, etc., based on vector similarity.
JP7859641B1