System and process for mimic learning to eliminate confounding

By using inference models in imitation learning to eliminate confounding factors, generate environmental beliefs and control the movement of robot equipment, the causal confusion problem is solved, and the accuracy and adaptability of imitation learning is improved.

CN119948488APending Publication Date: 2025-05-06QUALCOMM INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380067551.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-08-31
Filing Date
2023-09-01
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

In imitation learning, it is difficult for imitators to imitate the expert directly because the expert has more information about the task and/or environment than the imitators, resulting in causal confounding problems.

Method used

By using inference models to explain the behavior of experts, imitators’ inference models are trained to eliminate confounding factors, generate beliefs about the environment, and control the movements of the robotic device based on that belief.

Benefits of technology

Improve the accuracy and adaptability of imitation learning, allowing imitators to collect data autonomously and perform actions consistently with experts, thus ensuring more accurate task execution.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119948488A_ABST
    Figure CN119948488A_ABST
Patent Text Reader

Abstract

A processor-implemented method includes observing an environment via one or more sensors associated with a robotic device. The processor-implemented method also includes generating, via an inference model, a belief for the environment based on data associated with a priori action of the robotic device in the environment. The processor-implemented method also includes controlling the robotic device to perform an action in the environment based on generating the belief.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0002] This application claims priority to U.S. patent application No. 18 / 459,258, filed on August 31, 2023, and entitled “SYSTEM AND PROCESS FOR DECONFOUNDED IMITATION LEARNING,” and claims the benefit of U.S. Provisional Patent Application No. 63 / 411,016, filed on September 28, 2022, and entitled “SYSTEM AND PROCESS FOR DECONFOUNDED IMITATION LEARNING,” the disclosures of which are expressly incorporated by reference in their entirety. Technical Field

[0003] Aspects of the present disclosure generally relate to deconfounding imitation learning. Background Art

[0004] An artificial neural network may include interconnected groups of artificial neurons (e.g., a neuron model). An artificial neural network may be a computing device or represented as a method to be performed by a computing device. Artificial neural networks can be used for a variety of tasks, such as image recognition, speech recognition, acoustic scene classification, keyword detection, autonomous driving, and imitation learning.

[0005] In some examples, a machine learning model may learn to perform a task by observing an expert, such as a human, perform the task. The trained machine learning model may be associated with a device, such as (but not limited to) a robotic device, such that the trained machine learning model may control the device to perform the task. In such examples, the expert performing the task may have more information about the task and / or environment than the machine learning model. Therefore, if the machine learning model naively imitates the expert without the same information as the expert, the device may be unable to perform the task. Summary of the invention

[0006] The present disclosure is set out in the independent claims respectively. Some aspects of the present disclosure are described in the dependent claims.

[0007] In some aspects of the present disclosure, a processor-implemented method includes observing an environment via one or more sensors associated with a robotic device. The method also includes generating a belief about the environment based on data associated with a priori actions of the robotic device in the environment via an inference model. The method also includes controlling the robotic device to perform actions in the environment based on generating the belief.

[0008] Some other aspects of the disclosure relate to an apparatus that includes components for observing an environment via one or more sensors associated with a robotic device. The apparatus also includes components for generating beliefs about the environment based on data associated with prior actions of the robotic device in the environment via an inference model. The apparatus also includes components for controlling the robotic device to perform actions in the environment based on generating the beliefs.

[0009] In some other aspects of the present disclosure, a non-transitory computer-readable medium having non-transitory program code recorded thereon is disclosed. The program code is executed by one or more processors and includes program code for observing an environment via one or more sensors associated with a robotic device. The program code also includes program code for generating a belief about the environment based on data associated with a priori actions of the robotic device in the environment via an inference model. The program code also includes program code for controlling the robotic device to perform actions in the environment based on generating the belief.

[0010] Some other aspects of the present disclosure relate to an apparatus having one or more processors and one or more memories coupled to the one or more processors and storing instructions, which when executed by the one or more processors are operable to cause the apparatus to observe an environment via one or more sensors associated with a robotic device. Execution of the instructions also causes the apparatus to generate a belief about the environment based on data associated with a priori actions of the robotic device in the environment via an inference model. Execution of the instructions also causes the apparatus to control the robotic device to perform actions in the environment based on generating the belief.

[0011] Additional features and advantages of the present disclosure will be described below. It will be appreciated by those skilled in the art that the present disclosure can be easily utilized as a basis for modifying or designing other structures for implementing the same purpose as the present disclosure. It will also be appreciated by those skilled in the art that such equivalent constructions do not depart from the teachings of the present disclosure as set forth in the appended claims. The novel features considered to be characteristic of the present disclosure, both in terms of its organization and method of operation, together with further objects and advantages, will be better understood when the following description is considered in conjunction with the accompanying drawings. However, it is to be expressly understood that each of the accompanying drawings is provided for illustration and description purposes only and is not intended to be a definition of limitations on the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] The features, nature and advantages of the present disclosure will become more apparent when the detailed description set forth below is read in conjunction with the accompanying drawings, in which like reference numerals are correspondingly identified throughout.

[0013] Figure 1An example implementation of a neural network using a system on a chip (SOC) including a general purpose processor in accordance with certain aspects of the present disclosure is illustrated.

[0014] Figure 2A , Figure 2B and Figure 2C is a diagram illustrating a neural network according to aspects of the present disclosure.

[0015] Figure 3 is a state diagram illustrating an example of generating a probability model of a Bayesian network for mixed expert data.

[0016] Figure 4 is a flow chart illustrating an example of a process for training an inference model according to various aspects of the present disclosure.

[0017] Figure 5 is a block diagram illustrating an example of an architecture of a latent variable model including an encoder and a decoder according to various aspects of the present disclosure.

[0018] Figure 6 is a diagram illustrating an example of a function for learning an inference model and an imitator strategy for eliminating confusion according to various aspects of the present disclosure.

[0019] Figure 7 is a flow chart illustrating an example of a process for performing actions based on hidden information of an environment using a dynamics model according to various aspects of the present disclosure. DETAILED DESCRIPTION

[0020] The specific embodiments described below in conjunction with the accompanying drawings are intended as descriptions of various configurations and are not intended to represent the only configurations in which the described concepts can be practiced. In order to provide a comprehensive understanding of the various concepts, the specific embodiments include specific details. However, it is obvious to those skilled in the art that these concepts can be practiced without these specific details. In some instances, in order to avoid obscuring such concepts, well-known structures and components are shown in block diagram form.

[0021] Based on the teachings, those skilled in the art will recognize that the scope of the present disclosure is intended to cover any aspect of the present disclosure, whether the aspect is implemented independently of any other aspect of the present disclosure or implemented in combination with any other aspect. For example, a device or a method may be implemented using any number of aspects described. In addition, the scope of the present disclosure is intended to cover such devices or methods that are practiced using other structures, functionality, or structures and functionality that are supplementary to or different from the various aspects of the present disclosure described. It should be understood that any aspect of the present disclosure disclosed may be embodied by one or more elements of a claim.

[0022] The word “exemplary” is used to mean “serving as an example, instance, or illustration.” Any aspect described as “exemplary” is not necessarily to be construed as preferred or advantageous over other aspects.

[0023] Although specific aspects are described, numerous variations and permutations of these aspects fall within the scope of the present disclosure. Although some benefits and advantages of preferred aspects are mentioned, the scope of the present disclosure is not intended to be limited to specific benefits, uses or purposes. On the contrary, various aspects of the present disclosure are intended to be widely applicable to different technologies, system configurations, networks and protocols, some of which are illustrated by way of example in the drawings and the following description of preferred aspects. The specific embodiments and drawings are merely illustrative of the present disclosure and are not limiting, and the scope of the present disclosure is defined by the appended claims and their equivalents.

[0024] As discussed, in some examples, a machine learning model associated with an imitator can learn to perform a task by observing an expert, such as a human, performing the task. For example, a machine learning model can be trained to perform a task, such as lifting a box, by imitating a demonstration of a task performed by an expert. An imitator may also be referred to as an agent (hereinafter used interchangeably). In some examples, an imitator may be a robotic device or another type of autonomous or semi-autonomous device. Learning to perform a task by observing a demonstration may be referred to as imitation learning.

[0025] Imitation learning can be used for various applications such as (but not limited to) autonomous driving and robotics. In most cases, imitation learning uses various functions such as behavior cloning, inverse reinforcement learning, and adversarial methods. One challenge of imitation learning is the distribution mismatch between the expert and the imitator due to the accumulation of errors when implementing the imitator's strategy. This problem caused by the limited range of the expert's actions is different from the problem of hidden confounding factors.

[0026] Several solutions have explored the intersection between causality and imitation learning. Some conventional solutions have focused on scenarios where the state is fully observable but the expert's decision stems from both causal and non-causal factors. Other conventional solutions have focused on the complexity caused by partial models that only utilize segments of the state. Some other conventional solutions explore imitation learning in the context of latent variables influencing the expert's strategy.

[0027] The concept of meta-reinforcement learning (meta-RL) involves training an adaptive imitator to quickly learn new tasks. Some meta-RL functions integrate task encoders and task-conditional policies. However, in contrast to various aspects of the present disclosure, tasks associated with meta-RL functions may have varying reward functions. In the context of multiple environments that differ by latent factors, some solutions suggest that imitators adapt by probing their surroundings to discern these latent factors. However, in contrast to imitation learning methods based on expert demonstrations, these imitators typically operate in a reinforcement learning context with a defined reward.

[0028] As discussed, a particular challenge in imitation learning is when an expert performing a task has more information about the task and / or environment than the imitator. For example, when training an imitator to stack boxes on a shelf, the imitator may not know whether the box is light or heavy, because all boxes may look the same. That is, the appearance of a box may not indicate the weight of the box. However, an expert may know the weight of a box, or may estimate the weight of a box by lifting the box. In some examples, an expert may lift a heavy box with two arms at a time or lift two light boxes at a time (e.g., each arm lifts a light box). In such examples, the imitator cannot determine the weight of each box by simply looking at the box. By naively imitating an expert, the imitator may attempt to lift two heavy boxes at the same time, thus being unable to place a heavy box on a shelf. It may be desirable to improve the imitation learning system to consider situations in which an expert demonstrating a task has more information about the task and / or environment than the imitator.

[0029] Various aspects of the present disclosure relate to a systematic strategy of trial and error for identifying information about a task and / or environment. For example, a simulator can use trial and error to identify the weight of a box, and once the actual weight is identified, behave like an expert. Certain aspects of the subject matter described in the present disclosure can be implemented to achieve one or more of the following potential advantages. In some examples, by using a systematic trial and error solution, a simulator can discern the attributes (e.g., characteristics) of an environment, such as the weight of a box. Once these attributes are determined, the simulator can then emulate the behavior of an expert. The key benefit of this solution is its forward-looking and adaptive nature, allowing the simulator to autonomously collect basic data and then make its actions consistent with the expert, thereby ensuring more accurate task execution.

[0030] Figure 1An example implementation of a system on chip (SOC) 100 is illustrated, which may include a central processing unit (CPU) 102 or a multi-core CPU configured for deconvoluted imitation learning. Variables (e.g., neural signals and synaptic weights), system parameters associated with a computing device (e.g., a neural network with weights), delays, frequency bin information, and task information may be stored in a memory block associated with a neural processing unit (NPU) 108, a memory block associated with CPU 102, a memory block associated with a graphics processing unit (GPU) 104, a memory block associated with a digital signal processor (DSP) 106, a memory block 118, or may be distributed across multiple blocks. Instructions executed at CPU 102 may be loaded from a program memory associated with CPU 102, or may be loaded from memory block 118.

[0031] The SOC 100 may also include additional processing blocks tailored for specific functions, such as a GPU 104, a DSP 106, a connectivity block 110 (which may include fifth generation (5G) connectivity, fourth generation long term evolution (4G LTE) connectivity, Wi-Fi connectivity, USB connectivity, Bluetooth connectivity, etc.), and a multimedia processor 112 that may, for example, detect and recognize gestures. In one specific implementation, the NPU 108 is implemented in the CPU 102, the DSP 106, and / or the GPU 104. The SOC 100 may also include a sensor processor 114, an image signal processor (ISP) 116, and / or a navigation module 120, which may include a global positioning system.

[0032] SOC 100 may be based on the ARM instruction set. In one aspect of the present disclosure, the instructions loaded into the general purpose processor 102 may include code for: observing an environment; generating a belief about the environment based on data associated with an expert's prior actions in the environment via an inference model; and controlling a device to perform an action in the environment based on generating the belief.

[0033] Deep learning architectures can perform object recognition tasks by learning to represent inputs at successively higher levels of abstraction in each layer, thereby building useful feature representations of the input data. In this way, deep learning solves the main bottleneck of traditional machine learning. Before the advent of deep learning, machine learning methods for object recognition problems may rely heavily on human-designed features, possibly in conjunction with shallow classifiers. A shallow classifier can be a two-class linear classifier, for example, in which the weighted sum of feature vector components can be compared to a threshold to predict which class the input belongs to. Human-designed features can be templates or kernels customized for a specific problem domain by engineers with domain expertise. In contrast, although deep learning architectures can learn to represent features similar to those that human engineers may design, they require training. In addition, deep networks can learn to represent and recognize new types of features that humans may not have considered.

[0034] Deep learning architectures can learn hierarchies of features. For example, if presented with visual data, the first layer can learn to recognize relatively simple features in the input stream, such as edges. In another example, if presented with auditory data, the first layer can learn to recognize spectral power in specific frequencies. The second layer, taking the output of the first layer as input, can learn to recognize combinations of features, such as simple shapes for visual data or combinations of sounds for auditory data. For example, higher layers can learn to represent complex shapes in visual data or words in auditory data. Even higher layers can learn to recognize common visual objects or spoken phrases.

[0035] Deep learning architectures perform particularly well when applied to problems that have a natural hierarchical structure. For example, the classification of motorized vehicles can benefit from first learning to recognize wheels, windshields, and other features. These features can be combined in different ways at higher levels to identify cars, trucks, and airplanes.

[0036] Neural networks can be designed to have a variety of connection patterns. In a feedforward network, information is passed from a lower layer to a higher layer, where each neuron in a given layer communicates with a neuron in a higher layer. As described above, hierarchical representations can be constructed in successive layers of a feedforward network. Neural networks can also have loops or feedback (also known as top-down) connections. In a loop connection, the output from a neuron in a given layer can be communicated to another neuron in the same layer. The loop architecture can help identify patterns that span more than one input data block in the input data blocks that are sequentially delivered to the neural network. The connection from a neuron in a given layer to a neuron in a lower layer is called a feedback (or top-down) connection. When the recognition of a high-level concept can assist in discerning specific low-level features of the input, a network with many feedback connections may be helpful.

[0037] The connections between neural network layers can be fully connected or partially connected. Figure 2A An example of a fully connected neural network 202 is illustrated. In a fully connected neural network 202, a neuron in a first layer may communicate its output to each neuron in a second layer, such that each neuron in the second layer will receive input from each neuron in the first layer. Figure 2B An example of a locally connected neural network 204 is illustrated. In the locally connected neural network 204, neurons in a first layer may be connected to a limited number of neurons in a second layer. More generally, the locally connected layers of the locally connected neural network 204 may be configured such that each neuron in the layer will have the same or similar connectivity pattern, but the connection strengths may have different values ​​(e.g., 210, 212, 214, and 216). The locally connected connectivity patterns may produce spatially different receptive fields in higher layers because higher layer neurons in a given region may receive inputs that are tuned through training to the characteristics of a limited portion of the total input to the network.

[0038] An example of a locally connected neural network is a convolutional neural network. Figure 2C An example of a convolutional neural network 206 is illustrated. The convolutional neural network 206 may be configured such that the connection strengths associated with the inputs of each neuron in the second layer are shared (e.g., 208). Convolutional neural networks may be well suited for problems where the spatial location of the inputs is meaningful.

[0039] Deep belief network (DBN) is a probabilistic model including multiple layers of hidden nodes. DBN can be used to extract hierarchical representations of training data sets. DBN can be obtained by stacking the layers of restricted Boltzmann machine (RBM). RBM is a type of artificial neural network that can learn probability distributions through a set of inputs. Because RBM can learn probability distributions without information about the categories to which each input should be classified, RBM is generally used for unsupervised learning. Using an unsupervised and supervised hybrid paradigm, the bottom RBM of DBN can be trained in an unsupervised manner and can be used as a feature extractor, while the top RBM can be trained in a supervised manner (on the joint distribution of the input and target class from the previous layer) and can be used as a classifier.

[0040] A deep convolutional network (DCN) is a network that is a convolutional network configured with additional pooling and normalization layers. DCNs have achieved state-of-the-art performance on many tasks. DCNs can be trained using supervised learning, where both the input and output targets are known for many examples and are used to modify the weights of the network by using a gradient descent method.

[0041] The DCN may be a feed-forward network. In addition, as described above, the connections from a neuron in the first layer of the DCN to a set of neurons in the next higher layer are shared across the neurons in the first layer. The feed-forward and shared connections of the DCN may be used for fast processing. For example, the computational burden of the DCN may be much smaller than that of a similarly sized neural network that includes loops or feedback connections.

[0042] The processing of each layer of the convolutional network can be considered as a spatially invariant template or basis projection. If the input is first decomposed into multiple channels, such as the red, green and blue channels of a color image, the convolutional network trained on the input can be considered to be three-dimensional, with two spatial dimensions along the axis of the image and the third dimension capturing color information. The output of the convolutional connection can be considered to form a feature map in a subsequent layer, where each element in the feature map (e.g., 220) receives input from a certain range of neurons in the previous layer (e.g., feature map 218) and from each channel in the multiple channels. The values ​​in the feature map can be further processed with nonlinearity (such as correction, max(0,x)). The values ​​from adjacent neurons can be further pooled, which corresponds to downsampling and can provide additional local invariance and dimensionality reduction. Normalization corresponding to whitening can also be applied by lateral inhibition between neurons in the feature map.

[0043] In some cases, if the expert has more information about the world than the imitator, and if the expert uses the additional information to select one or more actions, it may be difficult for the imitator to directly imitate the expert. In imitation learning, this problem may be referred to as causal confounding. This type of confounding may be due to one or more factors, such as different sensors between the expert and the imitator, or different robot configurations. In some examples, the expert's additional knowledge of the world can be modeled as hidden variables in a neural network, such as a Bayesian network, that describe the expert's interaction with the world.

[0044] As discussed, imitation learning can be used to train a controller, such as a controller of a robotic device, to perform a task. Specifically, imitation learning is used by an imitator to learn a policy from a data set of expert demonstrations via supervised learning. That is, a machine learning model can be trained to predict the actions of an expert as a supervised learning problem. The expert can be a human or another robotic device. An expert refers to a human or device that performs a task; the expert may not have specific expertise in the task. In some examples, the expert can be associated with a human in a group consisting of a tuple. The strategy to act in the (unrewarded) Markov decision process (MDP) defined by is a state set, is a set of actions, P(s′|s,a) is the transition probability, and P(s0) is the distribution over the initial state. The interaction between the expert and the environment produces a trajectory τ = (s0, a0, …, a T-1,s T ). The expert can maximize the expectation of a reward function. Also, some tasks cannot be expressed using Markov rewards. In some examples of imitation learning, the behavior cloning policy π parameterized by η η (a|s) by minimizing the loss To learn, the parameters Represents a dataset of state-action pairs collected by an expert's policy.

[0045] In some examples, imitation learning can be extended to allow some variables θ∈Θ to be observed by an expert rather than the imitator. A series of Markov decision processes can be defined by the latent space Θ, the distribution P(θ), and for each θ∈Θ, the reward-free MDP We can assume that the expert strategy π exp (a|s,θ) exists for every MDP. Expert strategy π exp (a|s,θ) represents the probability of the expert taking action a when in state s, given the latent variable θ. Specifically, considering the latent variable θ, the expert policy describes how the expert behaves in various states. When the expert policy π exp When (a|s,θ) interacts with the environment, the interaction can generate the following distribution through the trajectory τ:

[0046]

[0047] In formula 1, P(s t+1 ∣s t ,a t ; θ) represents the transition to the next state s for each time step t t+1 The probability distribution of given current state s t , the action taken a t and latent variables θ. That is, the dynamics of the environment can indicate that under the influence of latent variables θ, our current state s t Take action a t Then in a specific state t+1 During imitation learning, the imitator does not observe the hidden θ. Therefore, the imitator can implicitly infer the hidden θ from past transitions. Therefore, in some conventional systems, the imitator can be used with a non-Markov strategy π parameterized by the parameter η. η (a t |s1,a1,…,s t ). That is, in a Markov system, the future state can depend on the current state without depending on the sequence of states that preceded it. However, in a non-Markov system, the future depends on past events. Therefore, the non-Markov strategy π η (a t |s1,a1,…,st ) Consider the entire history of states and actions up to time t to decide the action a at time t t In some examples, the imitator generates the following distribution through the trajectory τ:

[0048]

[0049] Figure 3 is a state diagram illustrating an example of a probability model of a Bayesian network for generating data of mixed experts. Figure 3 In the example of , s represents the initial state of the environment, a represents the action of the expert, and s′ represents the state of the environment after the expert has taken action a. In addition, the latent variable θ represents the hidden information about the environment available to the expert. The latent variable θ affects both the action a and the state s′ of the environment after the expert takes the action a. Conventional imitation learning systems do not eliminate the influence of the latent variable θ on both the action a and the state s′ of the environment after the expert takes the action a.

[0050] In the field of imitation learning, when latent variables θ are used for the policy, one may wish to fine-tune the imitator's parameters η so that during its interaction with the environment, the imitator's decisions closely reflect those of the expert. This consistent mathematical goal is expressed in the following way: maximize the expected distribution of the latent variable Then in this case, we maximize the expectation of the trajectory produced by the imitator's strategy (For example, ). Thus, for each trajectory, the goal is to sum the logarithm of the expert's policy values ​​for all state-action pairs. Essentially, this sum evaluates how likely the expert's policy replicates the decisions found in the imitator's trajectory. If the expert is performing a task optimally (e.g., maximizing some reward function), then aligning the imitator's actions with those of the expert inherently implies that the imitator is also equipped to optimally solve the same task. This approach highlights the central idea of ​​imitation learning in such scenarios: to train the imitator so that, when effective within the environment, it replicates the expert's behavior and thereby achieves similar results.

[0051] Conventional imitation learning systems do not eliminate the influence of hidden variables on the expert's action selection and the influence of hidden variables on the dynamics of the environment. Therefore, when a machine learning model trained using naive imitation learning is deployed in the real world, the machine learning model may incorrectly interpret the previous actions of the associated device as evidence of the hidden variables and thus choose the wrong action to move forward based on this incorrect reasoning.

[0052] Various aspects of the present disclosure relate to improving imitation learning systems by overcoming the above-mentioned limitations of conventional imitation learning systems. Specifically, in some examples, the inference model of the imitator can be trained using hidden information used by experts to make decisions. During training, the inference model can be used to explain the behavior of the expert, and thereby achieve imitation learning. When the imitator is deployed, the inference model incrementally obtains more information about the environment and improves the accuracy of the imitation. In some examples, the inference model can be a variational encoder-decoder that predicts the impact of actions in the environment. Training data for the model is collected using a random strategy. The model generates hidden variables representing beliefs about the environment, which are intended to represent information used by experts to make decisions. The concept of using a learned predictive model to eliminate confusion from expert data is novel.

[0053] Figure 4 4 is a flow chart illustrating an example of a process 400 for training an inference model according to various aspects of the present disclosure. Figure 4 As shown, at box 402, process 400 uses a random strategy for training an inference model to collect data. The data can be observations of the environment and / or observations of the actions of the imitator. In addition, at box 404, process 400 trains the inference model on the data collected by the random strategy. The data may include information about the expert's actions and the environment. The data may be collected from one or more sensors of the imitator. In some examples, the imitator may be a device that implements one or more of the random strategy, the imitation strategy, or the inference model, such as a robotic device. The inference model may also be referred to as a hidden variable prediction model. The inference model receives a trajectory as input and predicts hidden variables. The imitation strategy receives the state of the environment and the predicted hidden variables as input. The dynamics model can be used to train the inference model. The dynamics model receives one or more inputs, including the predicted hidden variables, the state of the environment, and the action selected by the strategy, and predicts the next state.

[0054] At box 406, process 400 trains a random policy for the imitator to perform actions based on expert data that is deconvoluted with the inference model. In some examples, process 400 can estimate the impact of hidden information on the expert and then account for the impact of the hidden information once the inference model has been trained. In some implementations, that is, after the inference model is trained, the expert demonstrations can be labeled with the inferred latent variables. This enriched data set, now free of confounding factors, serves as the basis for conventional behavior cloning. At box 408, the finalized policy generated by the processes associated with boxes 402, 404, and 406 can be deployed with the inference model during real-world testing.

[0055] Figure 55 is a block diagram illustrating an example of an architecture for a latent variable model 500 including an encoder 502 and a decoder 504. The latent variable model 500 may be processed by one or more processors such as reference Figure 1 The described SOC 100 performs. Figure 5 In the example of φ , and the decoder 504 may be a kinetic model p ψ .

[0056] like Figure 5 As shown in the example of o To current time s t Receive observations s of the environment, where t represents a time step counter. The encoder 502 can also start from an initial time a o To the current time a t The action a of the imitator is received. The observation s may be based on the observation of the imitator. Specifically, the encoder 502 may receive the trajectory τ=(s0, a0, ..., s t ). Observation t It may include, for example, the position and posture of all objects in the environment, and / or the speed of different objects. Action a may include, for example, the position of a device (eg, a robotic arm) associated with the imitator, or the position (eg, posture) of the imitator.

[0057] The encoder 502 may implement hidden information for generating an estimate of the environment at time t The probability distribution function of Given an action a 0:t and observations 0:t The encoder 502 selects a possible belief about the environment, which is represented as the estimated hidden variable For example, In some examples, action a t From the strategy π, for example Sampling is performed based on the current state s t and current beliefs (estimated hidden variables ) is a condition for one or more actions a t The imitator can perform the sampled action a t The dynamics of the environment is based on the current state s t , the selected action a t And the actual hidden information θ t To determine the next state s t+1 .

[0058] In some examples, decoder 504 may be used to train encoder 502 and may generate the next state s t+1The estimated probability distribution of given the current state s t , the selected action a t and the estimated hidden information p(s t+1 |s t ,a t ,θ t ) represents the dynamics of the environment (e.g., real-world dynamics). The next state s t+1 It can also be represented as s′. The dynamics can be observed in the real world and stored as training data. The dynamics model can then be trained ( Figure 5 In the ) to predict the dynamics of the environment.

[0059] The encoder 502 and decoder 504 can be trained to minimize the loss

[0060]

[0061] In formula 3, is based on keeping the hidden variables The distribution of beliefs is similar to the prior derived from the variational autoencoder (VAE). To minimize the loss in Formula 3, the encoder 502 learns the hidden All information in the ground truth latent θ used to predict state transitions is encoded in . In some examples, training data for encoder 502 can be collected from the interaction of the exploration strategy with the environment. The exploration strategy can be from the inferred distribution The implicit conditional strategy π η ,sampling Other exploration strategies can be used as long as they explore sufficiently diverse trajectories and do not depend on the actual hidden θ. The trained encoder 502 can be used to infer the hidden in the expert data distribution. These inferences are valid regardless of the distribution shift from the exploration data to the expert data because the transition model is the same. Subsequently, conventional imitation learning can be applied to develop hidden strategies based on the identification of the expert's demonstrations. That is, after training the inference model, the expert demonstrations can be labeled with the inferred hidden variables. This enriched dataset, which is now free of confounding factors, serves as the basis for conventional behavioral cloning. The finalized strategy generated by this process is then deployed together with the inference model during real-world testing.

[0062] At test time, the imitator (e.g., an agent) faces an environment with unknown latents and needs to adapt correct expert behavior. In some examples, the imitator adapts correct expert behavior by alternating between updating posterior beliefs on the latents and acting under the current beliefs. In such examples, the imitator initially starts with a priori beliefs. and actions A hidden is sampled to imitate the expert corresponding to that hidden. The imitator then observes the state transitions and uses the inference network to compute posterior beliefs. Another hidden is sampled from the updated beliefs and the process repeats. Once the inference has converged to match the real hidden of the environment, the real expert of the environment can be consistently imitated.

[0063] Figure 6 is a diagram illustrating an example of a function for learning an inference model and an imitator strategy for eliminating confusion according to various aspects of the present disclosure. Figure 6 In the example, the expert trajectory dataset is composed of , where j represents the index of the dataset and e is the index of a specific expert. The dataset may be obtained from observations of one or more experts performing the task. The imitator may function in a Markov decision process (MDP) which may be defined as a tuple consisting of a state s, an action a, a dynamics probability p, a probability of an initial state p0, and a horizon H defining the number of steps that may be taken in the MDP.

[0064] Figure 6 The function for learning the inference model and eliminating the confusion of the imitator strategy can be a loop initialized by sampling the latent variable θ from the distribution on the latent p(θ), so that the latent variable θ represents the hidden information about the current environment of the imitator. The latent variable θ may also be referred to as the hidden variable θ. In the example of moving boxes, the environment may be a warehouse or a simulated warehouse, and the latent variable θ may represent the weight of the box in the environment. The imitation learning function that eliminates the confusion may then sample the initial state s0 from the probability of the initial state p0. The trajectory τ may be initialized based on the initial state s0, and the counter t may be initialized to zero. The variable t may also refer to the step size in the loop. The loop may continue until the counter t is greater than the horizontal line H. The first step size within the loop uses the trained inference model.

[0065] The learned imitator can be a combination of a reasoning model and an imitator strategy. The learned imitator can be deployed in the environment. The reasoning model can be represented as a function This function generates an estimate of the hidden information of the environment given a trajectory τ after t steps The probability distribution of . Inference model q φ Choose a possible belief about the world, represented by the estimated hidden variable For example, the reasoning model may choose the belief that there are two heavy boxes in the environment. Next, action a t From strategy π η Sampling, the strategy is about taking the current state s t and current beliefs (estimated hidden variables ) is a condition for some actions a t The imitator can perform the sampled action a t The next state s t+1 will be determined by the environmental dynamics based on the current state s t , the selected action a t and the actual hidden information θ t The next state s t+1 and the selected action a t Appended to the trajectory τ, the counter t may then be updated (eg, t+1). If the value of the counter t is less than or equal to the value of the horizontal line H, the loop may repeat.

[0066] In one example, the belief (the estimated hidden variable ) Based on the inference model q φ There may be a light box in the warehouse. Then, the imitator may pick up a box and determine that the box is light. Then, the belief in the state of the environment may be updated. The current example may be an example of using a learned strategy at test time. In some examples, a random strategy is used to collect training data for the inference model. Specifically, the inference model may be updated and the dynamics model ψ to minimize the loss The loss is based in part on the next state s t+1 The logarithmic probability log p ψ (), given the current state s t , current action t and current beliefs about the hidden variables loss Defined in Equation 3.

[0067] like Figure 6 As shown, the corresponding losses of the inference model Φ and the dynamic model ψ can be optimized based on the trajectory τ. In some examples, the inference model Φ parameters and / or the dynamic model ψ parameters can be updated to minimize the corresponding losses. Then, the trained dynamic model ψ can be used to analyze the expert trajectory Specifically, the trained dynamics model ψ can be used to infer the expert trajectory Hidden variables The inference learning function can then be based on the hidden variables The inferred value is used to label the expert trajectory Next, the inference learning function can use the final belief to minimize the loss of the exploration strategy η.

[0068] In some examples, a dynamics model is trained via imitation learning from mixed data without contacting an expert. Conventional systems consult experts during imitation learning. Since machine learning models such as inference models and / or dynamics models require a certain amount of samples, consulting experts can use time and / or resources.

[0069] As discussed, training data for the inference model may be collected via online interaction with the environment. In some examples, the inference model may be trained offline. In such examples, when the amount of data in the expert data set is greater than a threshold, the implicit inference model may be trained only from the expert data without online interaction with the environment.

[0070] Figure 7 is a flow chart illustrating an example of a process 700 for performing an action based on hidden information of an environment using a dynamics model according to various aspects of the present disclosure. The process 700 may be performed as described with respect to Figure 1 The SOC 100 described herein performs the following operations. Figure 7 As shown, process 700 begins at block 702 by observing an environment. At block 704, process 700 generates beliefs about the environment based on data associated with prior actions of experts in the environment via an inference model. At block 706, process 700 controls a device to perform an action in the environment based on the generated belief.

[0071] Specific implementation examples are described in the following numbered clauses:

[0072] Clause 1. A processor-implemented method, the processor-implemented method comprising: observing an environment via one or more sensors associated with a robotic device; generating beliefs about the environment based on data associated with prior actions of the robotic device in the environment via an inference model; and controlling the robotic device to perform actions in the environment based on generating the beliefs.

[0073] Clause 2. The processor-implemented method of clause 1, further comprising training the reasoning model based on the data associated with the expert's prior actions in the environment.

[0074] Clause 3. A processor-implemented method according to any one of clauses 1 to 2, wherein a kinetic model trains the inference model.

[0075] Clause 4. A processor-implemented method as described in any of Clauses 2 to 3, wherein the data associated with the prior action is de-confounded from the reasoning model.

[0076] Clause 5. A processor-implemented method according to any one of clauses 1 to 4, wherein the inference model is a component of a variational encoder-decoder.

[0077] Clause 6. The processor-implemented method of clauses 1 to 5, wherein training the inference model comprises minimizing a first loss of the dynamics model and minimizing a second loss of the inference model.

[0078] Clause 7. The processor-implemented method of any one of clauses 2 to 6, further comprising: observing the prior actions of the agent in the environment via one or more sensors of the device.

[0079] Clause 8. A processor-implemented method as recited in any one of Clauses 2 to 7, wherein the expert is a human or another robotic device.

[0080] Clause 9. A device comprising: one or more processors; one or more memories coupled to the one or more processors; and instructions stored in the one or more memories and operable when executed by the one or more processors to cause the device to perform any one of clauses 1 to 8.

[0081] Clause 10. An apparatus comprising at least one means for performing any of clauses 1 to 8.

[0082] Clause 11. A computer program comprising code for causing an apparatus to perform any of clauses 1 to 8.

[0083] The various operations of the above methods may be performed by any suitable components capable of performing the corresponding functions. These components may include various hardware and / or software components and / or modules, including but not limited to circuits, application specific integrated circuits (ASICs), or processors. In general, where operations are illustrated in the accompanying drawings, these operations may have corresponding paired components plus function components with similar numbers.

[0084] As used, the term "determining" encompasses a wide variety of actions. For example, "determining" may include calculating, computing, processing, deriving, investigating, searching (e.g., searching in a table, database, or another data structure), ascertaining, etc. Additionally, "determining" may include receiving (e.g., receiving information), accessing (e.g., accessing data in a memory), etc. Furthermore, "determining" may include resolving, selecting, choosing, establishing, etc.

[0085] As used, a phrase referring to "at least one of" a list of items refers to any combination of those items, including single members. As an example, "at least one of a, b, or c" is intended to cover: a, b, c, ab, ac, bc, and abc.

[0086] The various illustrative logical blocks, modules, and circuits described in conjunction with the present disclosure may be implemented or performed with a general purpose processor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array signal (FPGA) or other programmable logic device (PLD), discrete gate or transistor logic, discrete hardware components, or any combination thereof designed to perform the functions described. A general purpose processor may be a microprocessor, but in an alternative, the processor may be any commercially available processor, controller, microcontroller, or state machine. The processor may also be implemented as a combination of computing devices, such as a combination of a DSP and a microprocessor, a plurality of microprocessors, one or more microprocessors combined with a DSP core, or any other such configuration.

[0087] The steps or algorithms of the methods described in conjunction with the present disclosure may be directly embodied in hardware, software modules executed by a processor, or a combination of the two. The software module may reside in any form of storage medium known in the art. Some examples of usable storage media include random access memory (RAM), read-only memory (ROM), flash memory, erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), registers, hard disks, removable disks, CD-ROMs, and the like. The software module may include a single instruction or multiple instructions, and may be distributed over several different code segments, distributed between different programs, and distributed across multiple storage media. The storage medium may be coupled to a processor so that the processor can read information from the storage medium and write information to the storage medium. In an alternative, the storage medium may be integral with the processor.

[0088] The disclosed methods include one or more steps or actions for implementing the described methods. The steps and / or actions of the methods may be interchanged with each other without departing from the scope of the claims. In other words, unless a specific order of steps or actions is specified, the order and / or use of specific steps and / or actions may be modified without departing from the scope of the claims.

[0089] The functions described may be implemented in hardware, software, firmware, or any combination thereof. If implemented in hardware, an example hardware configuration may include a processing system in a device. The processing system may be implemented using a bus architecture. Depending on the specific application and overall design constraints of the processing system, the bus may include any number of interconnecting buses and bridges. The bus may link various circuits together, including a processor, a machine-readable medium, and a bus interface. The bus interface may be used to connect a network adapter, etc., to the processing system via the bus. The network adapter may be used to implement signal processing functions. For some aspects, a user interface (e.g., a keypad, a display, a mouse, a joystick, etc.) may also be connected to the bus. The bus may also link various other circuits, such as timing sources, peripherals, voltage regulators, power management circuits, etc., which are well known in the art and will not be described further.

[0090] The processor may be responsible for managing the bus and general processing, including executing software stored on the machine-readable medium. The processor may be implemented using one or more general-purpose processors and / or special-purpose processors. Examples include microprocessors, microcontrollers, DSP processors, and other circuits that can execute software. Software should be broadly interpreted as meaning instructions, data, or any combination thereof, whether referred to as software, firmware, middleware, microcode, hardware description language, or other. By way of example, a machine-readable medium may include a random access memory (RAM), a flash memory, a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), a register, a disk, an optical disk, a hard drive, or any other suitable storage medium, or any combination thereof. The machine-readable medium may be embodied in a computer program product. The computer program product may include packaging materials.

[0091] In a hardware specific implementation, the machine-readable medium can be a part of a processing system separate from the processor. However, as will be readily appreciated by those skilled in the art, the machine-readable medium or any part thereof may be outside the processing system. By way of example, the machine-readable medium may include a transmission line, a carrier wave modulated by data, and / or a computer product separate from the device, all of which may be accessed by the processor through a bus interface. Alternatively or in addition, the machine-readable medium or any part thereof may be integrated into the processor, such as in the case of having a cache and / or a general register stack. Although the various components discussed may be described as having specific locations, such as local components, they may also be configured in various ways, such as certain components being configured as a part of a distributed computing system.

[0092] The processing system can be configured as a general processing system having one or more microprocessors providing processor functionality and an external memory providing at least a portion of a machine-readable medium, all of which are linked together with other supporting circuit systems through an external bus architecture. Alternatively, the processing system may include one or more neuromorphic processors for implementing the described neuron model and neural system model. As another alternative, the processing system may be implemented with an application specific integrated circuit (ASIC) having a processor, a bus interface, a user interface, support circuits, and at least a portion of a machine-readable medium integrated in a single chip, or with one or more field programmable gate arrays (FPGAs), programmable logic devices (PLDs), controllers, state machines, gated logic, discrete hardware components, or any other suitable circuits, or circuits capable of performing the various functionalities described throughout this disclosure. Those skilled in the art will recognize how to best implement the functionality of the processing system depending on the specific application and the overall design constraints imposed on the entire system.

[0093] The machine-readable medium may include multiple software modules. These software modules include instructions that cause the processing system to perform various functions when executed by the processor. The software module may include a sending module and a receiving module. Each software module may reside in a single storage device or be distributed across multiple storage devices. By way of example, when a triggering event occurs, the software module may be loaded from a hard drive into a RAM. During the execution of the software module, the processor may load some of the instructions into a cache to increase access speed. Then one or more cache lines may be loaded into a general register stack for execution by the processor. When the functionality of a software module is mentioned below, it will be understood that such functionality is implemented by the processor when executing instructions from the software module. In addition, it should be understood that various aspects of the present disclosure produce improvements in the functionality of processors, computers, machines, or other systems that implement such aspects.

[0094] If implemented in software, each function may be stored as one or more instructions or codes on or sent through a computer-readable medium. Computer-readable media include both computer storage media and communication media, including any media that facilitates the transfer of computer programs from one place to another. Storage media can be any available media that can be accessed by a computer. By way of example and not limitation, such computer-readable media may include RAM, ROM, EEPROM, CD-ROM or other optical disk storage devices, magnetic disk storage devices or other magnetic storage devices, or any other media that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer. In addition, any connection is also appropriately referred to as a computer-readable medium. For example, if the software is transmitted from a website, server, or other remote source using a coaxial cable, optical cable, twisted pair, digital subscriber line (DSL), or wireless technologies such as infrared (IR), radio, and microwaves, the coaxial cable, optical cable, twisted pair, DSL, or wireless technologies such as infrared, radio, and microwaves are included in the definition of the medium. As used, disks and optical disks include compact disks (CDs), laser disks, optical disks, digital versatile disks (DVDs), floppy disks, and Optical disks, where magnetic disks typically reproduce data magnetically, and optical disks reproduce data optically with lasers. Thus, in some aspects, computer-readable media may include non-transitory computer-readable media (e.g., tangible media). Additionally, for other aspects, computer-readable media may include transient computer-readable media (e.g., signals). Combinations of the above should also be included within the scope of computer-readable media.

[0095] Thus, some aspects may include a computer program product for performing the operations presented. For example, such a computer program product may include a computer-readable medium having stored (and / or encoded) thereon instructions that can be executed by one or more processors to perform the described operations. For some aspects, the computer program product may include packaging materials.

[0096] In addition, it should be understood that the modules and / or other appropriate components for performing the described methods and techniques can be downloaded and / or otherwise obtained by the user terminal and / or base station where applicable. For example, such a device can be coupled to a server to facilitate the transmission of the components for performing the described methods. Alternatively, the various methods described can be provided via a storage component (e.g., RAM, ROM, a physical storage medium such as a compact disc (CD) or a floppy disk) so that once the storage component is coupled to or provided to the device, the user terminal and / or base station can obtain the various methods. In addition, any other suitable technology suitable for providing the described methods and techniques to the device can be used.

[0097] It is to be understood that the claims are not limited to the precise configuration and components illustrated above. Various modifications, changes and variations may be made in the arrangement, operation and details of the methods and apparatus described above without departing from the scope of the claims.

Claims

1. A processor-implemented method, the processor-implemented method comprising: observing an environment via one or more sensors associated with the robotic device; generating, via an inference model, beliefs about the environment based on data associated with prior actions of the robotic device in the environment; and The robotic device is controlled to perform an action in the environment based on generating the belief.

2. The processor-implemented method of claim 1 , further comprising: The reasoning model is trained based on the data associated with an expert's prior actions in the environment.

3. The processor-implemented method of claim 2, wherein a kinetic model trains the inference model. 4 . The processor-implemented method of claim 3 , wherein the data associated with the a priori action is decontaminated from the reasoning model.

5. The processor-implemented method of claim 3, wherein the inference model is a component of a variational encoder-decoder.

6. The processor-implemented method of claim 2, wherein training the inference model comprises minimizing a first loss for the dynamics model and minimizing a second loss for the inference model.

7. The processor-implemented method of claim 2, further comprising: The prior actions of the agent in the environment are observed via one or more sensors of the device.

8. The processor-implemented method of claim 2, wherein the expert is a human or another robotic device.

9. A device, comprising: means for observing an environment via one or more sensors associated with the robotic device; means for generating, via an inference model, beliefs about the environment based on data associated with prior actions of the robotic device in the environment; as well as Means for controlling the robotic device to perform an action in the environment based on generating the belief.

10. The device according to claim 9, further comprising: Means for training the inference model based on the data associated with prior actions of the agent in the environment. The apparatus of claim 10 , wherein a kinetic model trains the inference model.

12. The apparatus of claim 11, wherein the data associated with the a priori action is decontaminated from the inference model.

13. The apparatus of claim 11, wherein the inference model is a component of a variational encoder-decoder.

14. The apparatus of claim 10, wherein the means for training the inference model comprises means for minimizing a first loss of the dynamics model and minimizing a second loss of the inference model.

15. The apparatus according to claim 10, further comprising: Means for observing the a priori actions of the agent in the environment via one or more sensors of the device.

16. The apparatus of claim 10, wherein the expert is a human or another robotic device.

17. A device, comprising: one or more processors; and one or more memories coupled to the one or more processors and storing instructions that, when executed by the one or more processors, are operable to cause the apparatus to: observing an environment via one or more sensors associated with the robotic device; generating, via an inference model, beliefs about the environment based on data associated with prior actions of the robotic device in the environment; and The robotic device is controlled to perform an action in the environment based on generating the belief.

18. The apparatus of claim 17, wherein execution of the instructions further causes the apparatus to train the inference model based on the data associated with prior actions of an agent in the environment.

19. The apparatus of claim 18, wherein a kinetic model trains the inference model.

20. The apparatus of claim 19, wherein the data associated with the a priori action is decontaminated from the inference model.

21. The apparatus of claim 19, wherein the inference model is a component of a variational encoder-decoder.

22. The apparatus of claim 18, wherein execution of the instructions to cause the apparatus to train the inference model further causes the apparatus to minimize a first loss for the dynamics model and to minimize a second loss for the inference model.

23. The apparatus of claim 18, wherein execution of the instructions further causes the apparatus to observe the prior actions of the agent in the environment via one or more sensors of the device.

24. The apparatus of claim 18, wherein the expert is a human or another robotic device.

25. A non-transitory computer readable medium having program code recorded thereon, the program code being executed by one or more processors and comprising: program code for observing an environment via one or more sensors associated with the robotic device; program code for generating beliefs about the environment based on data associated with prior actions of the robotic device in the environment via an inference model; as well as Program code for controlling the robotic device to perform actions in the environment based on generating the belief.

26. The non-transitory computer-readable medium of claim 25, wherein the program code further comprises program code for training the reasoning model based on the data associated with prior actions of an agent in the environment.

27. The non-transitory computer readable medium of claim 26, wherein a kinetic model trains the inference model.

28. The non-transitory computer readable medium of claim 27, wherein the data associated with the a priori action is disambiguated from the reasoning model.

29. The non-transitory computer-readable medium of claim 27, wherein the inference model is a component of a variational encoder-decoder.

30. The non-transitory computer readable medium of claim 26, wherein the program code for training the inference model further comprises program code for minimizing a first loss for the kinetic model and minimizing a second loss for the inference model.