Offline Training Method for Control Strategy Based on Model Uncertainty and Behavioral Priors

By training integrated dynamic model and variational autoencoder on offline data of the robot arm, combined with the weighted Bellman update framework, the problem of data distribution mismatch in offline reinforcement learning is solved, and more stable and efficient robot arm control strategy training is achieved, and the success rate of robot arm assembly tasks is improved.

CN115972211BActive Publication Date: 2025-08-05NANJING UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310064893.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-02-06
Publication Date
2025-08-05
Estimated Expiration
2043-02-06

AI Technical Summary

Technical Problem

The existing offline reinforcement learning technology cannot effectively train a good-performing strategy due to mismatch in robotic arm control strategy training, and the existing solutions fail to effectively utilize the differences in robotic arm operating data, which limits the performance of the control strategy.

Method used

By training integrated dynamic model and variational autoencoder on offline data of the robot arm, uncertainty measurement and behavior priors are built, combined with the weighted Bellman update framework, robot arm control strategies are trained, and the data is used to interact with the model and data to expand the data set and optimize weights to improve training stability and efficiency.

Benefits of technology

It realizes more stable and efficient robotic arm control strategy training under offline data, improves the performance of the control strategy, and can successfully complete part assembly in robotic arm assembly tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115972211B_ABST
    Figure CN115972211B_ABST
Patent Text Reader

Abstract

The present invention discloses an offline training method for a control strategy based on model uncertainty and behavioral priors. This method constructs an uncertainty measure for the robot arm data samples by training an integrated dynamics model on offline robot arm operation data. A variational autoencoder is then used to fit the behavioral prior strategy collected from the robot arm's offline data. Within the framework of weighted Bellman updates, the robot arm's control strategy is trained using only the robot arm's offline data. This method enables the robot arm's control strategy to selectively utilize the robot arm's offline dataset during offline training, reducing the impact of unreliable robot arm data samples on strategy training while ensuring that reliable robot arm data samples still have a positive impact on strategy training. This method can further stabilize the offline learning process of the robot arm's control strategy and improve its performance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an offline training method for a control strategy based on model uncertainty and behavior priors, which is used for learning the control strategy of a robotic arm. Background Art

[0002] Reinforcement learning is an important branch of machine learning. Using reinforcement learning methods, intelligent agents can interact with their environment to receive reward or penalty signals and, based on these signals, learn strategies that maximize rewards within that environment. However, reinforcement learning methods typically require continuous interaction with the environment to gain learning experience. For robotic arm tasks, these interactions with the operating environment consume significant time and financial resources.

[0003] Offline reinforcement learning provides a new approach to solving this problem. It learns strategies from a previously collected robotic arm operation dataset without interacting with the environment, thus eliminating the time and economic costs required for sampling in the environment.

[0004] However, due to a distribution mismatch between the behavioral strategies used in the collected robotic arm operation data and the control strategies to be learned, it's impossible to train a high-performing strategy directly from offline robotic arm operation data. Recent technical solutions have largely relied on strategy distribution constraints or conservative value estimates, without carefully considering the differences in different robotic arm operation data. For example, the robotic arm operation data may contain some erroneous operation data, which is detrimental to learning the robotic arm control strategy and limits the performance of the robotic arm control strategy after offline learning using this data. Summary of the Invention

[0005] Purpose of the invention: In response to the problems and shortcomings of existing offline reinforcement learning technologies in learning robotic arm control strategies, the present invention provides an offline training method for control strategies based on model uncertainty and behavioral priors. By training an integrated dynamics model and a variational autoencoder on the offline data of the robotic arm, confidence differentiation of the robotic arm operation data is provided. The control strategy of the robotic arm is trained offline in the framework of weighted Bellman update, which can make the offline learning process of the robotic arm control strategy more stable and improve the performance of the robotic arm control strategy.

[0006] Technical solution: An offline training method for control strategies based on model uncertainty and behavioral priors. The integrated dynamics model is trained on the robot arm's offline data to construct an uncertainty measure for the robot arm data samples, and a variational autoencoder is used to fit the behavioral prior strategy for collecting the robot arm's offline data. The robot arm control strategy continuously interacts with the integrated dynamics model to obtain more robot arm operation data. Only the robot arm's offline data and model data are used to train the robot arm's control strategy under the framework of weighted Bellman update.

[0007] The steps include:

[0008] Step 1: Train an integrated dynamics model on a robotic arm assembly operation dataset. The resulting model can simulate the real robotic arm operation environment.

[0009] Step 2: Train a variational autoencoder on a dataset of robotic assembly operations. The resulting behavioral prior model can simulate the behavioral strategy used to collect this data.

[0010] Step 3: Start training the actor-critic policy network. The actor-critic policy network is the robot arm control strategy. The control strategy interacts with the integrated dynamics model to generate operation samples of the robot arm and stores them in the model dataset.

[0011] Step 4: Sample a small batch of robotic arm operation samples from the mixed dataset, calculate the model uncertainty and decoder reconstruction probability of the samples, and calculate the Bellman update weight of the samples;

[0012] Step 5: Use the sampled small batch of robot operation samples to perform weighted Bellman update training value function, target value function and control strategy;

[0013] Step 6: Repeat steps 3-5 until the control strategy training reaches convergence and the training process is completed.

[0014] The robot arm operating environment that the robot arm control strategy faces is modeled to obtain an integrated dynamic model. The robot arm control strategy can interact with the integrated dynamic model to expand the robot arm's data set and provide uncertainty estimates of the robot arm's state-action pairs based on the integrated dynamic model errors.

[0015] Model the behavior strategy of the robot arm's offline data collection to obtain a behavior prior model. The behavior prior model can provide the probability of the robot arm's state-action pair occurring under the behavior strategy.

[0016] The actor-critic based policy network is the robotic arm control strategy that needs to be learned. During the learning phase, it is trained using a pre-collected offline dataset of the robotic arm. The training process adopts weighted Bellman update, and the weights are jointly constructed by the integrated dynamics model and the behavior prior model.

[0017] The above integrated dynamics model, behavior prior model and actor-critic based policy network can be trained in an end-to-end manner.

[0018] Specifically, the integrated dynamic model is represented by N fully connected neural networks with the same architecture but different initializations, aiming to simulate the robot arm operation environment. The robot arm operation environment E can be modeled as a Markov decision process<S,A,P,R,γ> In this environment, the robot control strategy receives state information s∈S at each decision step. The state information includes individual information of the robot, such as the angles of each joint, the readings of various sensors, the images captured by the camera on the robot, and relevant information about the assembly task within the field of view. The robot control strategy selects an executable action a from the action space A to make a decision. The action space includes the actions performed by the robot, such as movement and gripping of the robot. The dynamic function P of the robot operation environment will transfer to the next state s′~P(s,a) after receiving the action, and the reward function R will give an immediate reward R(s,a), for example, a reward when the robot grips the target object. Each neural network is modeled using a Gaussian distribution, that is, The input is the current state s and action a of the robot arm, and the output is the next state s′ and reward r of the robot arm, where Represents the Gaussian distribution, φ represents the parameters of the neural network, μ and Σ represent the mean and standard deviation of the Gaussian distribution, respectively. Each neural network in the integrated dynamics model can be trained based on the following minimization loss function L(φ), which is expressed as follows:

[0019]

[0020] Where D is an offline dataset that stores experience samples of robotic arm operations, where s, a, s′, and r represent the robotic arm’s motion state, executed action, next state, and reward, respectively.

[0021] Specifically, the interaction process between the robot control strategy and the integrated dynamics model includes the following steps:

[0022] Step 21: Sample a state from the robot arm offline dataset D as the current state of the robot arm;

[0023] Step 22: The control strategy of the robot arm samples an action based on the current state of the robot arm;

[0024] Step 23: Randomly select a fully connected neural network in the integrated dynamics model to generate the next state and reward of the robot arm based on the current state and action of the robot arm;

[0025] Step 24: Take the next state as the current state of the robot arm and repeat steps 22-23 until the given rollout length is reached. All generated robot arm interaction data are stored in the model dataset.

[0026] Specifically, the uncertainty u(s,a) of each manipulator state-action pair (s,a) can be estimated by integrating the dynamics model, and the calculation formula is as follows:

[0027]

[0028] in represents the Gaussian mean of the output of the i-th dynamics model (also known as a fully connected neural network); the rewards in the robotic arm operation data generated by the dynamics model are all subject to an uncertainty penalty, that is, r is replaced by r-ku(s,a), and k is a hyperparameter.

[0029] Specifically, the behavioral prior model is modeled using a variational autoencoder, aiming to model the behavioral strategy for collecting robot arm operation data. It consists of two parts: an encoder, which maps the robot arm's state-action pairs into a latent space; and a decoder, which maps the latent space vectors into the state-action space, aiming to reconstruct the previously input robot arm state-action pairs from the latent space vectors. Both the encoder and decoder are multi-layer fully connected neural networks, trained based on the following minimization loss function L(α), which is mathematically expressed as:

[0030]

[0031] in Denotes the encoder, D α2 represents the decoder, z represents the latent variable output by the encoder, is the standard normal distribution, D KL [·||·] is the relative entropy.

[0032] Specifically, the actor-critic strategy network refers to a strategy for controlling a robot in a robot operation scenario, which can perform actions such as moving and gripping in the robot operation environment, and can complete the parts assembly task through a series of actions. The offline training method of the control strategy based on model uncertainty and behavioral prior can learn the robot control strategy offline through the historical operation data of the robot. The robot control strategy is constructed using the actor-critic model, where the actor is the strategy π θ , is a random policy modeled by a Gaussian distribution, from which actions are sampled each time the policy is executed in the robot operation environment; the critic is the value function, including the value function Q ψ With the target value function in It is used to improve the training efficiency and stability of the network to be trained Q ψ The parameters of the target network are identical and will be periodically updated to the parameters of the network to be trained. Both the policy and value functions are composed of multi-layer fully connected neural networks.

[0033] Specifically, the value function Q of the manipulator control strategy is ψ The training process uses weighted Bellman update and is based on the following minimization loss function L(ψ). The mathematical expression of the minimization loss function L(ψ) is:

[0034]

[0035] Where w(s,a) is the weight, is the expected return, γ is the decay factor, and π θ (·|s ′ ) represents the action taken by the strategy with θ as parameter in the robot state s′, so Represents ψ - The target value function of the parameter In the robot state s′ and strategy π θ The output value of the action is taken, and d f It is a mixed dataset formed by the offline dataset of the robot arm and the model dataset in proportion f. The weight of the robot arm sample w(s,a) is constructed using both model uncertainty and the reconstruction probability of the behavior prior. Its calculation formula is as follows:

[0036]

[0037] Where c(s,a)=exp(-u(s,a)), exp() is the exponential function, is the reconstruction probability of the encoder, and λ∈[0,1] is a hyperparameter to adjust the coefficients of these two weight factors.

[0038] A computer device includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the control strategy offline training method based on model uncertainty and behavior prior is implemented.

[0039] A computer-readable storage medium stores a computer program for executing the above-mentioned control strategy offline training method based on model uncertainty and behavior prior. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] Figure 1 is a flow chart of a method according to an embodiment of the present invention;

[0041] Figure 2 Schematic diagram of training and interaction of the integrated dynamics model in an embodiment of the present invention;

[0042] Figure 3 2 is a schematic diagram of training a behavioral prior model according to an embodiment of the present invention;

[0043] Figure 4 This is the verification result of the offline training method for the control strategy based on model uncertainty and behavior prior in a simulation environment described in an embodiment of the present invention. DETAILED DESCRIPTION

[0044] The present invention is further illustrated below with reference to specific examples. It should be understood that these examples are only used to illustrate the present invention and are not used to limit the scope of the present invention. After reading the present invention, modifications of various equivalent forms of the present invention made by those skilled in the art all fall within the scope defined by the claims attached to this application.

[0045] As mentioned above, online reinforcement learning technology requires continuous interaction with the environment to gain learning experience for robotic arm-related tasks, which consumes a lot of time and economic costs. Offline reinforcement learning can learn control strategies using only offline datasets of robotic arm operations. However, due to the distribution mismatch between the behavioral strategies used to collect robotic arm operation data and the control strategies to be learned, it is impossible to train a high-performing control strategy directly from the offline robotic arm operation data. In response to this, technical solutions in recent years have mostly been based on strategy distribution restrictions or conservative value estimates, without carefully considering the differences between different robotic arm operation data. For example, there may be some erroneous operation data in the robotic arm operation data, which is not conducive to learning the control strategy and limits the performance of the control strategy after offline learning using this data.

[0046] In light of this, we propose an offline training method for control strategies based on model uncertainty and behavioral priors. This method, while learning a dynamics model, generates more robotic arm manipulation samples to expand the dataset. Furthermore, by integrating the uncertainty estimates of the dynamics model and behavioral priors as weights for weighted Bellman updates, we can better utilize the robotic arm manipulation samples, making the control strategy training process more stable and improving both strategy learning efficiency and final performance. For robotic arm assembly scenarios, where the task requires the robotic arm to successfully assemble parts, this method can train the robotic arm to successfully complete part assembly using only offline robotic arm manipulation data. This method is not limited to robotic arm tasks but can also be applied to any other control tasks.

[0047] The offline training method of control strategy based on model uncertainty and behavior prior includes the following steps:

[0048] Step 1: Train an integrated dynamics model on a robotic arm assembly operation dataset. The resulting model can simulate the real robotic arm operation environment.

[0049] Step 2: Train a variational autoencoder on the robotic arm assembly operation dataset. The resulting behavioral prior model can simulate the behavioral strategy for collecting this data.

[0050] Step 3: Start training the actor-critic policy network. The actor-critic policy network is the control strategy for the robotic arm. The control strategy interacts with the integrated dynamics model to generate operation samples of the robotic arm and store them in the model dataset.

[0051] Step 4: Sample a small batch of robotic arm operation samples from the mixed dataset, calculate the model uncertainty and decoder reconstruction probability of the samples, and calculate the Bellman update weights of the samples.

[0052] Step 5: Use the sampled small batch of robotic arm operation samples to perform weighted Bellman update training value function, target value function and control strategy.

[0053] Step 6: Repeat steps 3-5 until the control strategy training reaches convergence and the training process is completed.

[0054] like Figure 1 As shown, it includes three parts: integrated dynamics model, behavioral prior model and actor-critic based strategy network:

[0055] An integrated dynamics model is used to model the robot's operating environment. The strategy can interact with the integrated dynamics model to expand the robot's data set and provide uncertainty estimates of the robot's state-action pairs based on the integrated dynamics model errors.

[0056] The behavioral prior model is used to model the behavioral strategy for collecting offline data of the robot arm and provide the probability of occurrence of the robot arm state-action pair under the behavioral strategy;

[0057] Based on the actor-critic policy network, this is the robotic arm control strategy that needs to be learned. During the learning phase, it is trained using a pre-collected offline dataset of the robotic arm. The training process adopts weighted Bellman update, and the weights are jointly constructed by the integrated dynamics model and the behavior prior model.

[0058] The proposed integrated dynamics model, behavior prior model, and actor-critic based policy network can be trained in an end-to-end manner.

[0059] The robot assembly operation environment E faced by the control strategy is modeled as a Markov decision process<S,A,P,R,γ> In this environment, the control strategy receives state information s∈S at each decision step. This state information includes individual information about the robot arm, such as the angles of each joint, readings from various sensors, images captured by the robot's camera, and information about the assembly task within its field of view. The robot arm control strategy then selects an executable action a from the action space A, which includes robot movement, gripping, and other actions. Upon receiving an action, the environment's dynamics function P transitions to the next state s′~P(s,a), and the reward function R provides an immediate reward R(s,a), for example, a reward when the robot grasps the target object. The offline dataset consists of historical data from the robot arm's operation scenarios, representing trajectory samples generated by the behavioral strategy during assembly. The offline dataset is denoted as D = {(s,a,s′,r)}, where s,a,s′,r represent the robot arm's motion state, executed action, next state, and reward, respectively.

[0060] Integrated dynamics models such as Figure 2 As shown in Figure 1, the integrated dynamics model aims to fit the transfer function P(s′|s,a) and reward function R(s,a) of the environment, which is obtained by training on the offline dataset D of the manipulator operation. The integrated dynamics model aims to simulate the manipulator operation scenario and is represented by a multi-layer fully connected neural network. Each neural network is modeled with a Gaussian distribution, that is, The input is the current state s and action a of the robot arm, and the output is the next state s′ and reward r of the robot arm, where Represents Gaussian distribution, φ represents the parameters of the neural network, μ and Σ represent the mean and standard deviation of the Gaussian distribution respectively, and an integrated dynamics model is composed of N multi-layer fully connected neural networks with the same structure. Different initialization methods are used to initialize these N neural networks. The robot arm operation dataset D is divided into a training set and a test set in a certain ratio. The integrated dynamics model is trained on the training set. Each neural network in the integrated dynamics model can be trained based on the following minimization loss function L(φ). The mathematical expression of minimizing the loss function L(φ) is:

[0061]

[0062] At each iteration, a batch of manipulator operation samples are sampled from the training set, and the stochastic gradient descent method is used to optimize the above loss function on the batch samples. When the error of the integrated dynamics model on the manipulator operation test set no longer decreases, the model training is completed. Where D is the above-mentioned manipulator operation offline dataset, which stores the experience samples of the manipulator operation process. The samples include the manipulator's motion state, execution action, next state and reward obtained. The interaction process between the manipulator control strategy and the integrated dynamics model each time is as follows: Figure 2 As shown in the figure. First, the state of a robot arm is sampled from the offline dataset D as the current state. The strategy samples an action such as moving or gripping based on the current state of the robot arm. Then, a model is randomly selected from the integrated dynamics model. The next state and reward are generated based on the current state of the robot arm and the action taken by the strategy. The next state is then used as the current state of the robot arm. The interaction is repeated until the given interaction length is reached. All generated robot arm operation data are stored in the model dataset. The uncertainty u(s,a) of each robot arm state-action pair (s,a) can be estimated through the integrated dynamics model. The calculation formula is as follows:

[0063]

[0064] in Represents the Gaussian mean of the output of the i-th dynamics model; the rewards in the robotic arm operation data generated by the model are all subject to an uncertainty penalty, that is, r is replaced by r-κu(s,a), and κ is a hyperparameter.

[0065] Behavioral prior models such as Figure 3 As shown in the figure, the variational autoencoder is used to model the behavior strategy of collecting robot operation data. It consists of two parts, one of which is the encoder. The encoder maps the state-action pair of the robot arm to the latent space, and the other part is the decoder The decoder reconstructs the state-action pair of the robot arm based on the latent space vector Both the encoder and decoder are multi-layer fully connected neural networks. The variational autoencoder is trained based on the following minimization loss function L(α). The mathematical expression of the minimization loss function L(α) is:

[0066]

[0067] in represents the encoder, represents the decoder, z represents the latent variable output by the encoder, is the standard normal distribution, D KL[·||·] is the relative entropy. At each iteration, a batch of samples is sampled from the offline dataset of the robot’s operations. Stochastic gradient descent is used on these batches to optimize the loss function. Training ends when a given number of optimization rounds is reached.

[0068] The robot control strategy refers to the strategy for controlling the robot in the robot operation scenario. It can perform actions such as moving and gripping in the robot operation environment and complete the parts assembly task through a series of actions. The control strategy is constructed using the actor-critic model. The actor is the strategy π θ , is a random policy modeled by a Gaussian distribution, from which actions are sampled each time the policy is executed in the robot operation environment; the critic is the value function, including the value function Q ψ and the objective value function It is used to improve the training efficiency and the network to be trained Q ψ The parameters of the target network are identical and will be periodically updated to the parameters of the network to be trained. The policy and value functions are both composed of multi-layer fully connected neural networks. When training the policy and value functions, the reinforcement learning algorithm SAC is used. Initialization strategy π θ , value function Q ψ With the target value function The value function Q of the robot control strategy ψ The training process uses weighted Bellman update and is based on the following minimization loss function L(ψ). The mathematical expression of the minimization loss function L(ψ) is:

[0069]

[0070] Where w(s,a) is the weight, is the expected maximum return, γ is the decay factor,

[0071] π θ (·|s′) represents the action taken by the strategy with θ as parameter in the robot state s′, so Represents ψ - The target value function of the parameter In the robot state s′ and strategy π θ The output value of the action is taken, and d f It is a mixed dataset formed by the offline dataset of the robot arm and the model dataset in proportion f. The weight of the robot arm sample w(s,a) is constructed using both model uncertainty and the reconstruction probability of the behavior prior. Its calculation formula is as follows:

[0072]

[0073] Where c(s,a)=exp(-u(s,a)), exp() is the exponential function, is the reconstruction probability of the decoder, and λ∈[0,1] is a hyperparameter to adjust the coefficients of these two weight factors.

[0074] The offline training method of control strategy based on model uncertainty and behavior prior is verified on the medium playback dataset of Half Cheetah simulation environment. Figure 4 The verification results of the offline training method of the control strategy based on model uncertainty and behavior prior and other recent related offline reinforcement learning technology solutions MOPO and UWAC in this simulation environment and dataset are demonstrated. The experimental results show that this method can achieve better strategy performance than existing offline reinforcement learning technology solutions in this simulation environment and dataset.

[0075] Strategy π θ and integrated dynamics models Interactively expand the robotic arm operation dataset, the process is as follows Figure 2 As shown, the following steps are included:

[0076] Step 41: Sample a state from the robot arm operation offline dataset D as the current state s of the robot arm;

[0077] Step 42: Strategy π θ Sample an action a~π according to the current state s of the robot arm θ (s);

[0078] Step 43: Integrating the Dynamics Model Randomly select a model from Generate the next state s′ and reward r based on the current state s of the robot and the action a taken by the strategy:

[0079] Step 44: Take the next state s′ as the current state s of the robot arm, repeat steps 42-43 until the given interaction length is reached, and store all generated robot arm operation data (s, a, s′, r) into the model dataset D model middle.

[0080] Step 45: Repeat steps 41-44 until a given number of samples are collected.

[0081] Offline dataset D and model dataset D operated by a robotic arm model Train the policy and value function.

[0082] Step 51: Offline dataset D of robot arm operation and model dataset D modelUpdate samples {(s,a,s′,r)} in the mixed sampling mini-batch

[0083] Step 52: Calculate the model uncertainty u(s,a) of the manipulator state-action pair (s,a) in each sample by integrating the dynamics model. The calculation formula is as follows:

[0084]

[0085] in Represents the Gaussian mean of the output of the i-th dynamics model. An uncertainty penalty is imposed on the rewards of the manipulator operation samples generated by the integrated dynamics model, that is, r is replaced by r-ku(s,a), and k is a hyperparameter.

[0086] Step 53: Calculate the model confidence c(s,a) based on the calculated model uncertainty u(s,a). The calculation formula is as follows:

[0087] c(s,a)=exp(-u(s,a))

[0088] Step 54: Calculate the updated weight w(s,a) for each robotic arm operation sample pair, which is constructed using both the model confidence and the reconstruction probability of the variational autoencoder. The calculation formula is as follows:

[0089]

[0090] in is the reconstruction probability of the decoder in the variational autoencoder, and λ∈[0,1] is a hyperparameter to adjust the coefficients of these two weight factors.

[0091] Step 55: Calculate the objective function value Q target , the calculation formula is as follows:

[0092]

[0093] where a′~π θ (s′) is the action sampled by the strategy in the robot state s′, is the target value function network, and α is the entropy coefficient in SAC.

[0094] Step 56: Update the value function Q ψ , using the weighted Bellman update method, the update formula is as follows:

[0095]

[0096] where λ Q is the learning rate of the value function, and w(s,a) is the weight of the robot operation sample calculated in step 54.

[0097] Step 57: Update policy π θ , the update formula is as follows:

[0098]

[0099] where λ π is the learning rate of the strategy, α is the entropy coefficient in SAC, and w(s,a) is the weight of the robot operation sample calculated in step 54.

[0100] Step 58: Update the target value function Use soft update method, using the current value function Q ψ Parameters and objective value function The convex combination of the parameters is used to update the target value, so that the change of the target value is smoother and a certain stability is maintained. The update formula is as follows:

[0101] ψ - ←τψ+(1-τ)ψ -

[0102] where τ is the coefficient of the soft update.

[0103] After the control strategy training of the robotic arm reaches convergence by alternately repeating steps 41 and 58, the training process is completed.

[0104] Obviously, those skilled in the art should understand that the various steps of the offline training method for control strategies based on model uncertainty and behavioral priors of the above-mentioned embodiment of the present invention can be implemented using a general-purpose computing device. They can be concentrated on a single computing device or distributed on a network composed of multiple computing devices. Optionally, they can be implemented using program codes executable by the computing device, so that they can be stored in a storage device and executed by the computing device. In some cases, the steps shown or described can be performed in a different order than herein, or they can be made into individual integrated circuit modules, or multiple modules or steps therein can be made into a single integrated circuit module for implementation. Thus, the embodiments of the present invention are not limited to any specific combination of hardware and software.

Claims

1. A control strategy offline training method based on model uncertainty and behavior priors, characterized by: The steps include: Step 1: Train an integrated dynamics model on a robotic arm assembly operation dataset. The resulting model simulates the real robotic arm operation environment. Step 2: Train a variational autoencoder on a dataset of robotic assembly operations. The resulting behavioral prior model simulates the behavioral strategy used to collect this data. Step 3: Start training the actor-critic policy network. The actor-critic policy network is the robot arm control strategy. The control strategy interacts with the integrated dynamics model to generate operation samples of the robot arm and stores them in the model dataset. Step 4: Sample a small batch of robotic arm operation samples from the mixed dataset, calculate the model uncertainty and decoder reconstruction probability of the samples, and calculate the Bellman update weight of the samples; Step 5: Use the sampled small batch of robot operation samples to perform weighted Bellman update training value function, target value function and control strategy; Step 6: Repeat steps 3-5 until the control strategy training reaches convergence, completing the training process; The integrated dynamics model is represented by N fully connected neural networks with the same architecture but different initializations, each neural network is modeled by a Gaussian distribution; the uncertainty of each state-action pair of the manipulator is estimated by the integrated dynamics model; The behavioral prior model is built using a variational autoencoder to model the behavioral strategy for collecting robot arm operation data. It consists of two parts: an encoder that maps the robot arm's state-action pairs into a latent space; and a decoder that maps the latent space vectors into the state-action space and reconstructs the previously input robot arm state-action pairs from the latent space vectors. The actor-critic strategy network uses the historical operation data of the robot arm to learn the robot arm control strategy offline. The robot arm control strategy is constructed using the actor-critic model. The actor is a stochastic strategy modeled by a Gaussian distribution. Each time the strategy is executed in the robot arm operation environment, the action is sampled from the Gaussian distribution. The critic is the value function. Both the strategy and the value function are composed of a multi-layer fully connected neural network. The value function training process of the robotic arm control strategy adopts weighted Bellman update and is trained based on the minimization loss function.

2. The control strategy offline training method based on model uncertainty and behavior prior according to claim 1 is characterized in that: The robot arm's operating environment is modeled to obtain an integrated dynamics model. The robot arm control strategy can interact with the integrated dynamics model to expand the robot arm's data set and provide uncertainty estimates of the robot arm's state-action pairs based on the integrated dynamics model errors. Model the behavior strategy of the robot arm's offline data collection to obtain a behavior prior model. The behavior prior model can provide the probability of the robot arm's state-action pair occurring under the behavior strategy. The actor-critic based policy network is the robotic arm control strategy that needs to be learned. During the learning phase, it is trained using a pre-collected offline dataset of the robotic arm. The training process adopts weighted Bellman update, and the weights are jointly constructed by the integrated dynamics model and the behavior prior model.

3. The control strategy offline training method based on model uncertainty and behavior prior according to claim 1 is characterized in that: The integrated dynamics model is represented by N fully connected neural networks with the same architecture but different initializations. Each neural network is modeled by a Gaussian distribution, i.e. The input is the current state s and action a of the robot arm, and the output is the next state s′ and reward r of the robot arm, where Represents Gaussian distribution, φ represents the parameters of the neural network, μ and Σ represent the mean and standard deviation of the Gaussian distribution respectively; each neural network in the integrated dynamics model is trained based on the following minimization loss function L(ω), and the mathematical expression of the minimization loss function L(φ) is: Where D is the offline dataset of the robotic arm, which stores the experience samples of the robotic arm operation, where s, a, s′, and r represent the motion state, execution action, next state, and reward of the robotic arm, respectively.

4. The control strategy offline training method based on model uncertainty and behavior prior according to claim 1 is characterized in that: The interaction process between the manipulator control strategy and the integrated dynamics model includes the following steps: Step 21: Sample a state from the robot arm offline dataset D as the current state of the robot arm; Step 22: The control strategy of the robot arm samples an action based on the current state of the robot arm; Step 23: Randomly select a dynamic model from the dynamic models and generate the next state and reward of the robot arm based on the current state and action of the robot arm; Step 24: Use the next state as the current state of the robot arm and repeat steps 22-23 until the given rollout length is reached. Store all generated robot arm interaction data in the model dataset.

5. The control strategy offline training method based on model uncertainty and behavior prior according to claim 3 is characterized in that: The uncertainty u(s,a) of each manipulator state-action pair (s,a) can be estimated by integrating the dynamics model, and the calculation formula is as follows: in represents the Gaussian mean of the output of the i-th dynamic model.

6. The control strategy offline training method based on model uncertainty and behavior prior according to claim 5 is characterized in that: The rewards in the robotic arm operation data generated by the dynamics model are all subject to an uncertainty penalty, that is, r is replaced by r-κu(s,a), and k is a hyperparameter.

7. The control strategy offline training method based on model uncertainty and behavior prior according to claim 3 is characterized in that: The behavioral prior model is modeled using a variational autoencoder, which aims to model the behavioral strategy for collecting robot arm operation data. It consists of two parts: an encoder that maps the state-action pairs of the robot arm into a latent space; and a decoder that maps the latent space vectors into the state-action space and reconstructs the previously input robot arm state-action pairs from the latent space vectors. Both the encoder and the decoder are multi-layer fully connected neural networks, trained based on the following minimization loss function L(α), the mathematical expression of which is: in represents the encoder, represents the decoder, z represents the latent variable output by the encoder, is the standard normal distribution, D KL [·||·] is the relative entropy.

8. The control strategy offline training method based on model uncertainty and behavior prior according to claim 1 is characterized in that: The actor-critic strategy network is used to implement the strategy of manipulator control in the manipulator operation scenario. It can perform actions in the manipulator operation environment and complete the parts assembly task through a series of actions. The manipulator control strategy is learned offline through the historical operation data of the manipulator. The manipulator control strategy is constructed using the actor-critic model. The actor is the strategy π θ , is a random policy modeled by a Gaussian distribution, from which actions are sampled each time the policy is executed in the robot operation environment; the critic is the value function, including the value function Q ψ and the target value function Q ψ- ; where Q ψ- It is used to improve the training efficiency and stability of the network to be trained Q ψ The parameters of the target network are exactly the same and will be periodically updated to the parameters of the network to be trained. Both the strategy and value functions are composed of multi-layer fully connected neural networks.

9. The control strategy offline training method based on model uncertainty and behavior prior according to claim 3 is characterized in that: The value function Q of the manipulator control strategy ψ The training process uses weighted Bellman update and is based on the following minimization loss function L(ψ). The mathematical expression of the minimization loss function L(ψ) is: Where w(s,a) is the weight, y=r+γQ ψ -(s′,π θ (·|s′)) is the expected return, γ is the decay factor, and π θ (·|s′) represents the action taken by the strategy with θ as parameter in the robot state s′, so Q ψ -(s′,π θ (·|s′)) represents the - The target value function Q is the parameter ψ - In the robot state s′ and strategy π θ The output value of the action is taken, and d f It is a mixed dataset formed by the offline dataset of the robot arm and the model dataset in proportion f. The weight of the robot arm sample w(s,a) is constructed using both model uncertainty and the reconstruction probability of the behavior prior. Its calculation formula is as follows: Where c(s,a)=exp(-u(s,a)), exp() is the exponential function, is the reconstruction probability of the encoder, λ∈[0,1] is a hyperparameter to adjust the coefficients of the two weight factors, and z represents the latent variable output by the encoder.

10. A computer device, characterized in that: The computer device includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the control strategy offline training method based on model uncertainty and behavior prior according to any one of claims 1 to 9 is implemented.

Citation Information

Patent Citations

  • Robust target tracking method and system based on hierarchical decision network

    CN112802061A

  • Multi-unmanned-aerial-vehicle and multi-unmanned-ship inspection control system based on reinforcement learning

    CN113671994A