Robotic mobile device and related method
Through the context strategy search method of machine learning, LSTM and Gaussian process generators are used to generate context and policy variable sequences, optimize the loss function, and solve the adaptability problem of robots moving in changing environments, achieving efficient adaptive movement and learning.
Patent Information
- Application Number
- CN202510378186.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2018-09-28
- Publication Date
- 2025-07-04
AI Technical Summary
The prior art is difficult to efficiently adapt to changing environments and learn new skills in robot movement, especially when encountering unforeseen conditions or diverse terrains, and lacks effective robot movement adaptability and learning ability.
Using a machine learning-based context strategy search method, using a long short-term memory network (LSTM) and a Gaussian process sample generator, we calculate the upper-level strategy and optimize the loss function to realize the adaptive movement of the robot in different environments by generating a sequence of context variables and policy variables.
It realizes efficient learning and adaptation in changing environments, improves the robot's mobility ability in simulation and reality, and reduces the computing cost and sample size requirements.
Smart Images

Figure CN120245025A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure generally relates to robots, and more particularly to robotic mobile devices and related methods. Background Art
[0002] Robots can be programmed to perform certain movements. Additionally, artificial neural networks are used to implement robot movement without the robot being programmed for robot movement. Brief Description of the Drawings
[0003] Figure 1 is a block diagram of an example system for performing robot movement in accordance with the teachings of the present disclosure, the example system including an example machine learning context policy searcher.
[0004] Figure 2 is Figure 1 a block diagram of an example machine learning context policy searcher of
[0005] Figure 3 is a flowchart representing machine or computer-readable instructions that may be executed to implement Figure 1 the example system of Figure 1 and Figure 2 the example machine learning context policy searcher of
[0006] Figure 4 is Figure 2 a schematic diagram of an example operation of an example model trainer of
[0007] Figure 5 is Figure 2 a schematic diagram of an example operation of an example model inferencer of
[0008] Figure 6 is a block diagram of an example processing platform configured to execute Figure 3 the instructions of Figure 1 to implement the example system of Figure 1 and Figure 2 the example machine learning context policy searcher of
[0009] The drawings are not drawn to scale. Additionally, generally, the same reference numerals will be used throughout the (one or more) drawings and the accompanying written description to refer to the same or similar components. Detailed Description
[0010] Robot movement, including robot navigation, is important for the utilization of robots to perform specific tasks. Adaptability of robot movement is useful in environments where the robot encounters changing weather, diverse terrains, and / or unanticipated or changing conditions including, for example, collision detection and avoidance and / or loss of functionality (e.g., the robot loses full movement of a leg and needs to continue walking). Additionally, in some examples, it is beneficial to adapt robot movement in cases where the robot is to learn new skills, actions, or tasks.
[0011] Adaptability of robot movement can be achieved through machine learning. Machine learning gives computer systems the ability to gradually improve performance without explicit programming. Example machine learning methods are meta-learning, in which an automatic learning algorithm is applied to metadata. Another example machine learning method is deep learning, which uses deep neural networks or recurrent neural networks to enhance the performance of computer systems based on data representation rather than task-specific algorithms.
[0012] Another example is reinforcement learning, which involves how a robot or other computer system agent should take actions in an environment in order to maximize some notion of long-term reward. Reinforcement learning algorithms attempt to find a policy that maps the state of the world to the actions the robot should take in those states. Thus, based on environmental conditions and / or the conditions or characteristics of the robot, the policy provides the parameters of the actions to be carried out or executed by the robot. With reinforcement learning, the robot interacts with its environment and receives feedback in the form of rewards. The utility of the robot is defined by a reward function, and the robot learns to act in order to maximize the expected reward. Machine or reinforcement learning is based on the observed samples of the results. Reinforcement learning differs from supervised learning in that correct input / output pairs are not presented and sub-optimal actions are not explicitly corrected.
[0013] Reinforcement learning can include step-based policy search or episode-based policy search. Step-based policy search uses exploratory actions at each time step of the learning process. Episode-based policy search changes the parameter vector of the policy at the start of an episode in the learning process.
[0014] As a solution for robotic reinforcement learning, episode-based policy search improves the skill parameters of a robot through trial and error. One of the core challenges of this approach is to generate context policies with high sample efficiency. Bayesian optimization is a sample-efficient method for context policy search, but Bayesian optimization has the drawback of computational burden (which is cubic in the number of samples). Another example method is context covariance matrix adaptation evolution strategy, which uses covariance matrix adaptation evolution strategy to find the best parameters. Context covariance matrix adaptation evolution strategy has much lower sample efficiency than Bayesian optimization context policy search.
[0015] The examples disclosed herein ensure the sample efficiency of context policy search, which has a linear time cost with respect to the number of samples. The examples disclosed herein also have both sample efficiency and computational efficiency. These efficiencies are important in cases where robotic activity simulation episodes can take from about 0.1 seconds to about 1 second per episode, and the learning process can include numerous trials (e.g., hundreds to millions).
[0016] The examples disclosed herein include a training process and a derivation process. The example training process involves a cost function that enables the sampling process to achieve both high values of samples and weighted regression of samples. The samples are weighted regressed to generate an upper policy, which is represented by a parametric function. As disclosed herein, the distance between the upper policy and the ideal upper policy is used as part of the cost consideration in the machine learning process.
[0017] In the example derivation process, a trained long short-term memory (LSTM) model is used to sample in the context policy search process. The derivation process also uses the samples to generate an upper policy through weighted regression. LSTM is a type of recurrent neural network (RNN). An RNN is a network with loops in the network, which allows information to persist, such that for example previous information can be used by the robot for the current task. The loops take the form of a chain of repeating modules of a neural network. In some RNNs, the repeating module will have a simple structure, such as for example a single tanh layer. LSTM also has a chain-like structure, but the repeating module has a more complex structure. Instead of having a single neural network layer, LSTM includes multiple (e.g., four) interacting neural network layers.
[0018] Figure 1 is a block diagram of an example system 100 that implements robotic movement. Example system 100 includes an example robot 102, which includes an example machine learning context policy searcher 104, example sensor(s) 106, and example actuator(s) 108.
[0019] One or more sensors 106 of the robot 102 receive an input 110. The input 110 can be information related to the environment, including, for example, weather information, terrain information, information related to other robots, and / or other information that can be used to evaluate the state of the environment around the robot. Additionally, the input 110 can be information obtained related to the internal functions of the robot 102 or other information related to the robot 102, including, for example, information related to the physical and / or processing functions and / or capabilities of any one of the systems in the robot's system.
[0020] The input 110 is used by the machine learning context policy searcher 104 to determine a policy based on the context (e.g., the expected output). The policy identifies the actions to be taken by the robot 102 based on the input 110 and the context. One or more actuators 108 of the robot 102 are used to deliver an output 112 according to the actions identified by the policy.
[0021] Consider, for example, that the robot 102 holds a ball and controls the robot 102 to throw the ball to a target position. Here, the target position is the context. Different trajectories will be generated to control the robot arm or other actuators 108 according to different target positions or different contexts. The parameters for generating the trajectories are called policy parameters or policies. In this example, the policy is automatically generated by the machine learning context policy searcher 104.
[0022] To facilitate adding new skills to the robot 102, a context-related reward function is defined by the machine learning liaison policy searcher 104 to judge whether the robot 102 performs the activity well. The robot 102 performs multiple attempts to improve the context policy search in simulation, in reality, and / or jointly in simulation and reality. During this process, an upper-level policy is learned, which is a projection from the context to the robot joint trajectory parameters. This learning process is carried out through an optimization process to improve the reward function.
[0023] Figure 2 is Figure 1 A block diagram of an example machine learning context policy searcher 104. The example machine learning context policy searcher 104 includes an example model trainer 202 and an example model inferencer 204. The example model trainer 202 includes an example Gaussian process sample generator 206, an example context training sample generator 208, an example sequence generator 210, an example vector input 212, an example calculator (e.g., an example loss function calculator 214), and an example comparator 216, an example database 218, and an example sequence incrementer 220. The example model inferencer 204 includes an example sequence input 222, an example coefficient input 224, and an example policy calculator 226, and an example database 228.
[0024] Example machine learning context policy searcher 104 and its components form part of a device that moves robot 102 based on a context policy. Machine learning context policy searcher 104 operates in two parts: a training part performed by model trainer 202 and a derivation part performed by model derivator 204. In this example, in the training part, an LSTM model is trained. In other examples, there may be other RNNs besides LSTM, including, for example, a differentiable neural computer (DNC). Again in this example, in the derivation part, the LSTM is used to sample new context policy search tasks, and an upper-level policy is generated according to the sample sequence. Using the upper-level policy, robot 102 can have the ability to obtain a fitful policy according to any context of the task.
[0025] During an example training process, Gaussian process sample generator 206 generates Gaussian process samples. For example, Gaussian process sample generator 206 generates Gaussian process samples: GP i=1...I (x) GP i (x)(d x ) has the same dimension as x (policy parameter vector), and I is the number of training samples.
[0026] A Gaussian process is a description of an unknown random process that only assumes that the random distribution at each time point is a Gaussian distribution, and the covariance between the distributions at every two time points is only related to the time difference between the two time points. The Gaussian distribution describes the unknown point distribution according to the Central Limit theorem.
[0027] Context training sample generator 208 generates context training samples. For example, context training sample generator 208 generates context training samples: CS i=1..I (s, x) CS i=1..I (s, x)(d cs ) has a dimension of: d cs = d x + d s Here, d s is the dimension of context vector s, and CS i=1..I (s, x) is a transformed version of CP i=1..I (x) and the transformation is determined by a randomly generated polynomial function.
[0028] The sequence generator 210 generates sequences of context variable vectors and policy variable vectors. For example, the sequence generator 210 generates sequences of s t and x t When given the LSTM parameters θ, for example, the sampling process can generate sequences of s t and x t The context variable vectors are related to a moving target or target, and the policy variable vectors are related to a moving trajectory. The sequences are based on Gaussian process samples and context training samples. For example, N1 Gaussian process samples can be generated, and for each of the Gaussian process samples, N2 context samples are generated. Each context sample is a polynomial function. In this example, the generated sequences are used for a subset of these N1×N2 samples. The goal of the system and method optimization is to ensure that the sampling process for all or most of the N1×N2 samples converges. Additionally, in some examples, Gaussian process samples and context samples are generated online, which can also increase the score of the model in the RNN for online samples.
[0029] The example vector input 212 inputs multiple inputs into the loss function calculator and into each unit of the RNN. In some examples, the inputs include the latent variable vector h t . Additionally, in some examples, the inputs include the input-context variable vector: and the policy variable vector x t . Additionally, in some examples, the inputs include the return value y t .
[0030] The example model trainer 202 also includes an example loss function calculator 214 that calculates an upper-level policy and a loss function based on the sequences. The upper-level policy indicates the movement of the robot, and the loss function indicates the degree of satisfaction of the movement target or target. In some examples, the loss function calculator 214 calculates the loss function: The loss function consists of two parts: the first part, which encourages sampling better value points: F(s 1:T , x 1:T ) (referred to as BVP); and the second part, which encourages sampling for generating a better upper-level policy: (referred to as BUP).
[0031] The loss function calculator calculates BVP as: where f(s t, x t ) is the value of the t-th point in the sequence. Other methods for defining BVP can be used in other examples.
[0032] The loss function calculator calculates the BUP as: where A (a matrix of size ) is the upper-level policy calculated from the training data according to the sampling point sequence, and is the polynomial that generates the training samples, for example: Here: and then A can be calculated as: A = (Φ T DΦ + λI) -1 Φ T DX Here, D is a diagonal weighting matrix, and the diagonal weighting matrix contains weights d that can be calculated as follows k : d k = ln(T + 0.5) - ln(k) where k is the order of f(s t , x t ) sorted in descending order; Φ is defined as: (N is the length of the sampling sequence), and can be selected as: ( The length of is defined as ); λI is a regularization term. As described above, in some examples, the calculator will use the diagonal weighting of the sequence to calculate the upper-level policy.
[0033] In some examples, can be any - dimensional feature function of the context s. Additionally, in some examples, is selected as a linear generalization of the context, while other examples may have other generalizations. Additionally, in some examples, BUP can also be other forms of matrix distance. Additionally, in some examples, the loss function L cps (θ) can also be other forms of the function, which has F(s 1:T , x 1:T ) and
[0034] In some examples, the LSTM network can be trained to reduce the loss function from the data. For example, the model trainer 202 includes a comparator 216 to determine whether the loss function meets a threshold. If the loss function does not meet the threshold, the sequence incrementer 220 will increment the value of t, and the model trainer 202 will traverse the simulation again, which has updated sequence generation, BVP, BUP, and loss function calculation, etc. For example, for a new task, the RNN generates s_t and x_t, the environment returns y_t, and the RNN then generates s_(t + 1) and x_(t + 1). This process will continue.
[0035] If the comparator 216 determines that the loss function does meet the threshold, the machine learning context policy searcher 104 considers the trained model, and the coefficients of the RNN are set to those calculated during the training phase using the operations of the model trainer 202. In some examples, the coefficients are called the model. The model trainer 202 can store the coefficients, model, samples, and / or other data related to the training process in the database 218 for access in subsequent operations. Using the trained model, the machine learning context policy searcher 104 triggers the operation of the model inferencer 204.
[0036] In the inference phase, the model inferencer 204 inputs or accesses the sequence generated for the new task via the sequence input 222. For example, the sample sequence: (s 1:T ,x 1:T ) The model inferencer 204 also inputs or accesses the coefficients (the trained model) via the coefficient input 224.
[0037] The model inferencer 204 further includes a policy calculator 226, which calculates the upper-level policy. For example, the policy calculator 226 determines the upper-level policy by the following formula: A = (Φ T DΦ + λI) -1 Φ T DX The policy calculator 226 can further determine the corresponding policy for any context s as:
[0038] Using the policy determined after training the model when the loss function meets the threshold, as detailed above, the machine learning context policy searcher 104 can signal the actuator 108 to cause the robot 102 to perform the robot movement of the upper-level policy. The calculated upper-level policy, policy, and / or other data related to the derivation process can be stored by the model derivator 204 for access in subsequent operations. In these examples, the robot movement implemented by the (one or more) actuators 108 according to the policy is first performed by the robot 102 after the sequence generator 210 generates the sequence, and the machine learning context policy searcher 104 operates according to the foregoing teachings.
[0039] The selected context used by the context training sample generator 208 has the same scope as the actual actions to be taken by the robot 102 in the real world. Since any upper-level policy can be approximately calculated by a polynomial function, the training of the randomly generated upper-level policy (context sample) can ensure that the trained result approximates the best upper-level policy, which is the movement taken by the robot 102.
[0040] Although Figure 2 shows an example way to implement Figure 1 the machine learning context policy searcher 104, Figure 2 one or more of the elements, processes, and / or devices shown can be combined, divided, rearranged, omitted, eliminated, and / or implemented in any other way. In addition, the example model trainer 202, example model derivator 204, example Gaussian process sample generator 206, example context training sample generator 208, example sequence generator 210, example vector input 212, example comparator 216, example loss function calculator 214, example database 218, example sequence incrementer 220, example sequence input 222, example coefficient input 224, example policy calculator 226, example database 228, and / or more generally Figure 2The example machine learning context policy searcher 104 can be implemented by hardware, software, firmware, and / or any combination of hardware, software, and / or firmware. Thus, for example, any one of the example model trainer 202, the example model inferencer 204, the example Gaussian process sample generator 206, the example context training sample generator 208, the example sequence generator 210, the example vector input 212, the example comparator 216, the example loss function calculator 214, the example database 218, the example sequence incrementer 220, the example sequence input 222, the example coefficient input 224, the example policy calculator 226, the example database 228, and / or more generally the example machine learning context policy searcher 104 can be implemented by one or more analog or digital circuits, logic circuits, (one or more) programmable processors, (one or more) programmable controllers, (one or more) graphics processing units (GPUs), (one or more) digital signal processors (DSPs), (one or more) application specific integrated circuits (ASICs), (one or more) programmable logic devices (PLDs), and / or (one or more) field programmable logic devices (FPLDs). When reading any of the apparatus or system claims of this patent that cover a pure software and / or firmware implementation, the example model trainer 202, the example model inferencer 204, the example Gaussian process sample generator 206, the example context training sample generator 208, the example sequence generator 210, the example vector input 212, the example comparator 216, the example loss function calculator 214, the example database 218, the example sequence incrementer 220, the example sequence input 222, the example coefficient input 224, the example policy calculator 226, the example database 228, and / or at least one of the example machine learning context policy searcher 104 is hereby expressly defined to include a non-transitory computer-readable storage device or storage disk, such as a memory, a digital versatile disc (DVD), a compact disc (CD), a Blu-ray disc, etc., which includes software and / or firmware. Further, in addition to Figure 2 those elements, processes, and / or devices shown Figure 2 or as an alternative to Figure 1 those elements, processes, and / or devices shown Figure 2 the example machine learning context policy searcher 104 can include one or more elements, processes, and / or devices, and / or can include more than one of any or all of the elements, processes, and devices shown. As used herein, the phrase "in communication" (including variations thereof) encompasses direct communication and / or indirect communication through one or more intermediate components, and does not require direct physical (e.g., wired) communication and / or constant communication, but also includes selective communication at periodic intervals, scheduled intervals, aperiodic intervals, and / or one-time events.
[0041] Figure 3shown to represent Figure 2 a flowchart of example hardware logic, machine or computer readable instructions, a hardware implemented state machine, and / or any combination thereof for implementing an example machine learning context policy searcher 104. The machine readable instructions may be an executable program or part of an executable program for execution by a computer processor (such as the processor 612 shown in the example processor platform 600 discussed below in conjunction with Figure 6 . The program may be embodied as software stored on a non-transitory computer readable storage medium (such as a CD-ROM, floppy disk, hard drive, DVD, Blu-ray disc, or memory associated with the processor 612), but the entire program and / or parts thereof may alternatively be executed by a device other than the processor 612 and / or embodied as firmware or dedicated hardware. Additionally, although the example program is described with reference to the Figure 3 flowchart shown, many other methods for implementing the example machine learning context policy searcher 104 may alternatively be used. For example, the order of execution of the blocks may be changed, and / or some of the blocks may be changed, eliminated, or combined. Additionally or alternatively, any one or all of the blocks may be implemented by one or more hardware circuits (such as discrete and / or integrated analog and / or digital circuits, FPGAs, ASICs, comparators, operational amplifiers (op-amps), logic circuits, etc.) configured to perform the corresponding operations without executing software or firmware.
[0042] As described above, example processes for implementing Figure 3 may be implemented using executable instructions (such as computer and / or machine readable instructions) stored on a non-transitory computer and / or machine readable medium, such as a hard drive, flash memory, read only memory, compact disc, digital versatile disc, cache, random access memory, and / or any other storage device or storage disk in which information is stored for any length of time (such as an extended period of time, permanently, momentarily, temporarily buffered, and / or cached). As used herein, the term "non-transitory computer readable medium" is explicitly defined to include any type of computer readable storage device and / or storage disk and does not include propagating signals and does not include transmission media.
[0043] "Comprising" and "including" (and all of their forms and tenses) are used herein as open-ended terms. Thus, whenever a claim uses any form of "comprising" or "including" (such as including, containing, having, etc.) as a preamble or within any kind of claim recitation, it is to be understood that additional elements, terms, etc. may be present without falling outside the scope of the corresponding claim or recitation. As used herein, when the phrase "at least" is used as a transitional term in, for example, the preamble of a claim, it is open-ended in the same manner as the terms "comprising" and "including" are open-ended. The term "and / or" when used, for example, in the form such as A, B, and / or C, means any combination or subset of A, B, C such as (1) only A, (2) only B, (3) only C, (4) A and B, (5) A and C, (6) B and C, and (7) A and B and C. As used herein in the context of describing a structure, component, item, object, and / or thing, the phrase "at least one of A and B" is intended to mean an implementation that includes (1) at least one A, (2) at least one B, and (3) any one of at least one A and at least one B. Similarly, as used herein in the context of describing a structure, component, item, object, and / or thing, the phrase "at least one of A or B" is intended to mean an implementation that includes (1) at least one A, (2) at least one B, and (3) any one of at least one A and at least one B. As used herein in the context of describing the execution or running of a process, instruction, action, activity, and / or step, the phrase "at least one of A and B" is intended to mean an implementation that includes (1) at least one A, (2) at least one B, and (3) any one of at least one A and at least one B. Similarly, as used herein in the context of describing the execution or running of a process, instruction, action, activity, and / or step, the phrase "at least one of A or B" is intended to mean an implementation that includes (1) at least one A, (2) at least one B, and (3) any one of at least one A and at least one B.
[0044] Figure 3 The program 300 is used to train a model, such as for example an LSTM model. The program 300 includes a Gaussian process sample generator 206 of a model trainer 202 of a machine learning context policy searcher 104 of the robot 102 to generate a policy variable vector, such as for example a Gaussian process sample (block 302). The exemplary program 300 also includes a context training sample generator 208 to generate context training samples (block 304).
[0045] A sequence generator 210 generates a sequence based on the Gaussian process sample and the context training samples (block 306). A vector input 212 inputs a vector into each unit of an RNN of an LSTM model of the sequence (block 308). In some examples, the vector includes a latent variable vector, an input-context vector, a policy vector, and a return value.
[0046] A calculator (such as the loss function calculator 214) calculates the better value point (BVP), the better upper-level policy (BUP), and the loss function (block 310) based on the sequence and the input vector. The comparator 216 determines whether the loss function reaches or meets a threshold (block 312). If the loss function does not meet the threshold, the sequence incrementer 220 increments t by a certain count in the sequence (block 314). As t is incremented, the example program 300 continues, where the sequence generator 210 generates a sequence at the incremented t (block 306). The example program 300 continues to determine the new loss function and so on.
[0047] If the comparator 216 determines that the loss function does meet the threshold (block 312), then the example program 300 has a trained model. For example, a loss function that meets the threshold may indicate a network or LSTM model that meets the expected reduced loss function. In this example, with an acceptably reduced loss function, the robot 102 has learned to meet the context or otherwise take the expected actions.
[0048] The example program 300 continues, where the sequence generator 210 generates a sequence for a new task (block 316), which is accessed or received by the model inferencer 204 via the sequence input 222. The coefficient input 224 of the model inferencer 204 imports coefficients from the trained model (block 318). The policy calculator 226 calculates the BUP (block 320). Additionally, the policy calculator 226 determines a policy based on the BUP (block 322). With the determined policy, the actuator(s) 108 of the robot 102 perform or execute the movement indicated by the policy (block 324).
[0049] Figure 4 Yes Figure 2 A schematic diagram of an example operation of the example model trainer 202, and Figure 5 Yes Figure 2 A schematic diagram of an example operation of the example model inferencer 204. Figure 4 And Figure 5 Shows a sequence (s t and x t ) of Gaussian process samples and context training samples of multiple RNN units. Additionally, the input of the unit consists of three parts: (1) Latent variable vector: h t ; (2) Input-context variable vector and policy variable vector: and x t ; Return value: y t .
[0050] Figure 4 The loss function (labeled "Loss" in the figure) is determined during the training phase. Figure 4 The loss in [X] represents the reward function. The reward function during the training process consists of two parts: (1) the better value (y_t) (BVP); and (2) the regression of the upper-level policy (BUP). Both BVP and BUP are to be calculated and optimized. The input used to determine Figure 4 the loss in [X] is the output of f(x). Figure 4 (During the training phase), the coefficients of the RNN are calculated and recalculated for optimization.
[0051] Figure 5 During the derivation phase, the loss determined during the training phase is used as the input in the RNN unit to determine the upper-level policy (labeled "A" in the figure). The input used to determine Figure 5 A in [X] is s_t and h_t. During Figure 5 the derivation phase of [X], the coefficients of the RNN are fixed.
[0052] Figure 6 is a block diagram of an example processor platform 1000, which is configured to execute Figure 4 the instructions of [X] to implement Figure 1 and Figure 2 the machine learning context policy searcher 104 of [X]. The processor platform 1000 can be, for example, a server, a personal computer, a workstation, a self-learning machine (such as a neural network), a mobile device (such as a cellular phone, a smartphone, a tablet (such as an iPad TM ), a personal digital assistant (PDA), an Internet device, a DVD player, a CD player, a digital camera, a Blu-ray player, a game console, a personal video camera, a set-top box, headphones or other wearable devices, or any other type of computing device.
[0053] The processor platform 600 of the illustrated example includes a processor 612. The processor 612 of the illustrated example is hardware. For example, the processor 612 can be implemented by one or more integrated circuits, logic circuits, microprocessors, GPUs, DSPs, or controllers from any desired family or manufacturer. The hardware processor can be a semiconductor (e.g., silicon-based) device. In this example, the processor 612 implements the example model trainer 202, the example model inferencer 204, the example Gaussian process sample generator 206, the example context training sample generator 208, the example sequence generator 210, the example vector input 212, the example comparator 216, the example loss function calculator 214, the example sequence incrementer 220, the example sequence input 222, the example coefficient input 224, the example policy calculator 226, and / or the example machine learning context policy searcher 104.
[0054] The processor 612 of the illustrated example includes local memory 613 (e.g., a cache). The processor 612 of the illustrated example communicates with a main memory (including volatile memory 614 and non-volatile memory 616) via a bus 618. The volatile memory 614 can be implemented by synchronous dynamic random access memory (SDRAM), dynamic random access memory (DRAM), dynamic random access memory and / or any other type of random access memory device. The non-volatile memory 616 can be implemented by flash memory and / or any other desired type of memory device. Access to the main memory 614, 616 is controlled by a memory controller.
[0055] The processor platform 600 of the illustrated example also includes interface circuitry 620. The interface circuitry 620 can be implemented by any type of interface standard, such as an Ethernet interface, a universal serial bus (USB) interface, a near field communication (NFC) interface, and / or a PCI express interface.
[0056] In the illustrated example, one or more input devices 622, 106, 110 are connected to the interface circuitry 620. The (one or more) input devices 622, 106, 110 permit a user to input data and / or commands into the processor 612. The (one or more) input devices can be implemented by, for example, an audio sensor, a microphone, a photographic device (a still camera or a video camera), a keyboard, a button, a mouse, a touch screen, a trackpad, a trackball, an isopoint, and / or a voice recognition system.
[0057] One or more output devices 624, 108, 112 are also connected to the interface circuit 620 of the illustrated example. The output devices 624, 108, 112 can be implemented, for example, by a display device (such as a light emitting diode (LED), an organic light emitting diode (OLED), a liquid crystal display (LCD), a cathode ray tube display (CRT), an in-plane switching (IPS) display, a touch screen, etc.), a haptic output device, a printer, and / or a speaker. Accordingly, the interface circuit 620 of the illustrated example generally includes a graphics driver card, a graphics driver chip, and / or a graphics driver processor.
[0058] The interface circuit 620 of the illustrated example also includes a communication device (such as a transmitter, a receiver, a transceiver, a modem, a residential gateway, a wireless access point, and / or a network interface) to facilitate the exchange of data with an external machine (such as any kind of computing device) via the network 626. The communication can be carried out via, for example, an Ethernet connection, a digital subscriber line (DSL) connection, a telephone line connection, a coaxial cable system, a satellite system, a line-of-sight wireless system, a cellular phone system, etc.
[0059] The illustrated example of the processor platform 600 also includes one or more mass storage devices 628 for storing software and / or data. Examples of such mass storage devices 628 include a floppy disk drive, a hard disk drive disk, a compact disk drive, a Blu-ray disk drive, a redundant array of independent disks (RAID) system, and a digital versatile disk (DVD) drive.
[0060] Figure 3 The machine-executable instructions 300 and other machine-executable instructions 632 can be stored in the mass storage device 628, the volatile memory 614, the non-volatile memory 616, and / or on a removable non-transitory computer-readable storage medium such as a CD or a DVD.
[0061] It will be appreciated from the foregoing that example devices, systems, manufactured products, and methods have been disclosed that implement robotic movement and, in particular, implement movement learned by a robot outside of the robot's standard or original programming. These examples use inputs such as data collected or otherwise passed to sensors, which data is used in a machine learning context to output a policy for the robot to use to change the robot's activities (including the robot's movement). By enabling a robot to learn new tasks and actions, which allows the robot to adapt to changing environments or changing functional capabilities, the disclosed devices, systems, manufactured products, and methods improve the efficiency of using a computing device. The disclosed devices, systems, manufactured products, and methods accordingly aim at one or more improvements in the capabilities of a computer.
[0062] Because the context policy is a continuous function and both the context and the policy parameters are multi-dimensional, the machine learning context policy searcher disclosed herein can perform the learning process multiple times (e.g., hundreds to millions of times). Even in a simulation setting, the computational cost of such a large number of executions (e.g., about 0.1 second to about 1.0 second per attempt or per simulation) is significant. The examples disclosed herein have a linear computational complexity for achieving the learning of the context policy with high sample efficiency and effective computational cost. Additionally, the examples of the present disclosure provide reasonable time and sample efficiency for robot simulation, which has better performance compared to lower computational capabilities (for cloud computing). Therefore, these examples that enable a robot to effectively adapt to new tasks are useful for edge computing.
[0063] The present disclosure provides example devices, systems, manufactured products, and methods for robot movement. Example 1 includes a robot movement device for a mobile robot, where the device includes a sequence generator that generates a sequence of context variable vectors and policy variable vectors, the context variable vectors being related to a movement target, and the policy variable vectors being related to a movement trajectory. The device further includes a calculator that calculates an upper-level policy and a loss function based on the sequence, the upper-level policy indicating the robot movement, and the loss function indicating the degree of meeting the movement target. Additionally, the device includes: a comparator for determining whether the loss function meets a threshold; and an actuator for causing the robot to perform the upper-level policy's robot movement when the loss function meets the threshold.
[0064] Example 2 includes the robot movement device of Example 1, where the calculator calculates the upper-level policy using diagonal weighting of the sequence.
[0065] Example 3 includes the robot movement device of Example 1 or 2, where the calculator further calculates the loss function based on the upper-level policy.
[0066] Example 4 includes the robot movement device of Examples 1 - 3, where the sequence is a first sequence, the upper-level policy is a first upper-level policy, the robot movement is a first robot movement, and the loss function is a first loss function. The device further includes a sequence incrementer that changes the first sequence into a second sequence when the first loss function does not meet the threshold.
[0067] Example 5 includes the robot movement device of Example 4, where the calculator calculates a second upper-level policy and a second loss function based on the second sequence, the second upper-level policy indicating a second robot movement, and the second loss function indicating the degree of meeting the movement target. The comparator determines whether the second loss function meets the threshold, and the actuator is used to cause the robot to perform the second upper-level policy's second robot movement when the second loss function meets the threshold.
[0068] Example 6 includes the robotic mobile device of Examples 1-5, where the sequence is based on long short-term memory parameters.
[0069] Example 7 includes the robotic mobile device of Examples 1-6, where the calculator will further determine the upper-level policy based on the matrix distance.
[0070] Example 8 includes the robotic mobile device of Examples 1-7, where the robotic movement is first performed by the robot after the sequence generator generates the sequence.
[0071] Example 9 is a robotic mobile device for a mobile robot, where the device includes components for generating a sequence of context variable vectors and policy variable vectors, the context variable vectors are related to the movement target, and the policy variable vectors are related to the movement trajectory. Example 9 also includes components for calculating the upper-level policy and the loss function based on the sequence, the upper-level policy instructs the robotic movement, and the loss function indicates the degree of meeting the movement target. Additionally, Example 9 includes components for determining whether the loss function meets the threshold and components for driving the robotic movement of the robot to execute the upper-level policy when the loss function meets the threshold.
[0072] Example 10 includes the robotic mobile device of Example 9, where the components for calculation will use diagonal weighting of the sequence to calculate the upper-level policy.
[0073] Example 11 includes the robotic mobile device of Example 9 or 10, where the components for calculation will further calculate the loss function based on the upper-level policy.
[0074] Example 12 includes the robotic mobile device of Examples 9-11, where the sequence is the first sequence, the upper-level policy is the first upper-level policy, the robotic movement is the first robotic movement, and the loss function is the first loss function. The device further includes components for changing the first sequence into a second sequence when the first loss function does not meet the threshold.
[0075] Example 13 includes the robotic mobile device of Example 12, where the components for calculation will calculate a second upper-level policy and a second loss function based on the second sequence, the second upper-level policy instructs the second robotic movement, and the second loss function indicates the degree of meeting the movement target. The components for determination will determine whether the second loss function meets the threshold, and the components for actuation will drive the robot to execute the second robotic movement of the second upper-level policy when the second loss function meets the threshold.
[0076] Example 14 includes the robotic mobile device of Examples 9-13, where the sequence is based on long short-term memory parameters.
[0077] Example 15 includes the robotic mobile device of Examples 9-14, where the components for calculation will further determine the upper-level policy based on the matrix distance.
[0078] Example 16 includes the robotic mobile device of Examples 9 - 15, where the robotic movement is first performed by the robot after the component for generating the sequence generates the sequence.
[0079] Example 17 is a non - transitory computer - readable storage medium containing machine - readable instructions that, when executed, cause the machine to at least generate a sequence of context variable vectors and policy variable vectors, where the context variable vectors are related to a movement target and the policy variable vectors are related to a movement trajectory. The instructions further cause the machine to calculate an upper - level policy and a loss function based on the sequence, where the upper - level policy indicates robotic movement and the loss function indicates the degree of meeting the movement target. Additionally, the instructions cause the machine to determine whether the loss function meets a threshold, and when the loss function meets the threshold, drive the robot to perform the robotic movement of the upper - level policy.
[0080] Example 18 includes the storage medium of Example 17, where the instructions cause the machine to calculate the upper - level policy using diagonal weighting of the sequence.
[0081] Example 19 includes the storage medium of Example 17 or 18, where the instructions cause the machine to further calculate the loss function based on the upper - level policy.
[0082] Example 20 includes the storage medium of Examples 17 - 19, where the sequence is a first sequence, the upper - level policy is a first upper - level policy, the robotic movement is a first robotic movement, and the loss function is a first loss function. The instructions further cause the machine to change the first sequence into a second sequence when the first loss function does not meet the threshold.
[0083] Example 21 includes the storage medium of Example 20, where the instructions further cause the machine to calculate a second upper - level policy and a second loss function based on the second sequence, where the second upper - level policy indicates a second robotic movement and the second loss function indicates the degree of meeting the movement target. Additionally, the instructions cause the machine to determine whether the second loss function meets the threshold, and when the second loss function meets the threshold, drive the robot to perform the second robotic movement of the second upper - level policy.
[0084] Example 22 includes the storage medium of Examples 17 - 21, where the sequence is based on long - short - term memory parameters.
[0085] Example 23 includes the storage medium of Examples 17 - 22, where the instructions further cause the machine to determine the upper - level policy further based on matrix distance.
[0086] Example 24 includes the storage medium of Examples 17 - 23, where the robotic movement is first performed by the robot after the instructions cause the machine to generate the sequence.
[0087] Example 25 is a method for a mobile robot, the method including generating a sequence of context variable vectors and policy variable vectors, the context variable vectors being related to a movement target, and the policy variable vectors being related to a movement trajectory. The method further includes calculating an upper-level policy and a loss function based on the sequence, the upper-level policy instructing the robot to move, and the loss function indicating the degree of meeting the movement target. Additionally, the method includes determining whether the loss function meets a threshold, and when the loss function meets the threshold, driving the robot to perform the upper-level policy's robot movement.
[0088] Example 26 includes the method of Example 25, further including calculating the upper-level policy using diagonal weighting of the sequence.
[0089] Example 27 includes the method of Example 25 or 26, further including further calculating the loss function based on the upper-level policy.
[0090] Example 28 includes the method of Examples 25-27, where the sequence is a first sequence, the upper-level policy is a first upper-level policy, the robot movement is a first robot movement, and the loss function is a first loss function, the method further including changing the first sequence into a second sequence when the first loss function does not meet the threshold.
[0091] Example 29 includes the method of Example 28, and further includes calculating a second upper-level policy and a second loss function based on the second sequence, the second upper-level policy instructing a second robot movement, and the second loss function indicating the degree of meeting the movement target. The exemplary method further includes determining whether the second loss function meets the threshold, and when the second loss function meets the threshold, driving the robot to perform the second upper-level policy's second robot movement.
[0092] Example 30 includes the method of Examples 25-29, where the sequence is based on long short-term memory parameters.
[0093] Example 31 includes the method of Examples 25-30, further including further determining the upper-level policy based on matrix distance.
[0094] Example 32 includes the method of Examples 25-31, where the robot movement is first performed by the robot after the generation of the sequence.
[0095] Although certain example methods, devices, and manufactured products have been disclosed herein, the scope of coverage of this patent is not limited thereto. On the contrary, this patent covers all methods, devices, and manufactured products that fall entirely within the scope of the claims of this patent.
Claims
1. A memory includes machine-readable instructions to cause at least one processor circuit to: Train a reward function using reinforcement learning, the reward function defining the activities of a robot in an environment; Deploy the reward function in the robot, and after deploying the reward function in the robot, cause the robot to move in the environment; Receive reward feedback based on the movement of the robot; And Process the reward feedback to update the reward function.
2. The memory according to claim 1, wherein, The reinforcement learning is episode-based.
3. The memory according to claim 1, wherein, The instructions cause the at least one processor circuit to store data related to the reward feedback in a database.
4. A method includes Training a reward function using reinforcement learning, the reward function defining the activities of a robot in an environment; Deploying the reward function in the robot, and after deploying the reward function in the robot, causing the robot to move in the environment; Receiving reward feedback based on the movement of the robot; And Processing the reward feedback to update the reward function.
5. The method according to claim 4, wherein, The reinforcement learning is episode-based.
6. The method according to claim 4, further comprising storing data related to the reward feedback in a database.
7. An apparatus includes: Means for training a reward function using reinforcement learning, the reward function defining the activities of a robot in an environment; Means for deploying the reward function in the robot, and means for causing the robot to move in the environment after deploying the reward function in the robot; Means for receiving reward feedback based on the movement of the robot; And Means for processing the reward feedback to update the reward function.
8. The device according to claim 7, wherein The reinforcement learning is episode-based.
9. The apparatus according to claim 7, further comprising means for storing data related to the reward feedback in a database.
10. A computer program product comprising instructions that, when executed by at least one processor, cause the at least one processor to perform the method according to any one of claims 4-6.