A Probability-Based Strategy Transfer Method

By combining the Q-function estimator of the probability and the policy gradient optimization, the problem of low migration efficiency from virtual environment to real environment is solved, and the generalization and robustness of policy migration is improved, and it is suitable for a variety of continuous control tasks.

CN114781645BActive Publication Date: 2025-07-18BEIJING INST OF CONTROL ENG
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210255129.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-03-15
Publication Date
2025-07-18
Estimated Expiration
2042-03-15

AI Technical Summary

Technical Problem

The prior art is inefficient in strategy migration from virtual environment to real environment, especially in space robot operation tasks, it is difficult to effectively use the virtual environment to approach the real environment, resulting in limited policy learning and operation performance.

Method used

A Q-function estimator for probability is constructed and combined with policy gradient optimization. A Q-function estimator for probability is constructed through Monte Carlo dropout. Combined with policy gradient optimization, it realizes identification and decomposition of environmental uncertainty and improves policy migration performance.

Benefits of technology

It improves the generalization ability and robustness of policy migration, improves the reliability of policy operation, and is suitable for learning and transfer of multiple continuous control tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114781645B_ABST
    Figure CN114781645B_ABST
Patent Text Reader

Abstract

A probability-based policy transfer method belongs to the field of artificial intelligence technology. The environment of continuous control tasks such as robot operation is affected by high dynamics, uncertainty, etc. It is actually very difficult to approximate the real environment using a virtual environment. The method of the present invention includes: constructing a probability-based Q-function estimator through Monte Carlo dropout and combining it with policy gradient optimization, so that the algorithm has the ability to identify environmental uncertainty. Specifically, through virtual environment training data collection, uncertainty decomposition and inference, policy gradient optimization, and real environment operation performance evaluation, the decomposition and measurement of environmental uncertainty are realized, and the policy learning efficiency and policy operation performance are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a probability-based policy transfer method, belonging to the technical field of artificial intelligence. Background Art

[0002] The poor performance of virtual-real policy transfer is an important factor restricting the in-depth application of reinforcement learning. For general continuous control learning problems, the common solution is to learn and train in a virtual environment and transfer the trained policy network to the real environment at the cost of zero-shot or few-shot, involving two environments. For problems such as space robot operation, since a large number of ground reliability tests are required, at least three-environment transfer is needed, namely virtual environment, ground test environment, and real space environment. The environments of such tasks are affected by high dynamics, uncertainty, etc., and it is actually very difficult to use the virtual environment to approximate the real environment, which restricts the further improvement of the policy learning efficiency and policy operation performance. Therefore, studying how to evaluate the differences between environments, measure the uncertainty within and between environments, and propose a probability-based virtual-real policy transfer method will be beneficial to improving the policy transfer performance. Summary of the Invention

[0003] The technical problem solved by the present invention is: overcoming the deficiencies of the prior art, providing a probability-based policy transfer method, by constructing a probability-based Q-function estimator and combining it with policy gradient optimization, enabling the algorithm to have the ability to identify environmental uncertainty, and forming a probability-based virtual-real policy transfer method. The constructed method helps to improve the generalization and transfer performance of the algorithm, thereby improving the robustness and reliability of policy operation, and has practical engineering significance.

[0004] The technical solution of the present invention is: a probability-based policy transfer method, including the following steps:

[0005] Construct a policy network and a Q-function estimator;

[0006] The virtual environment receives the output of the policy network and decides whether to receive action exploration according to a preset policy, generating a virtual environment output; the virtual environment is a simulation model corresponding to the entity system;

[0007] Decide whether to superimpose environmental perturbation on the virtual environment output according to a preset policy, generating training data;

[0008] The policy network and the Q-function estimator are updated using the training data. At the same time, the policy network is updated using a preset policy gradient optimization method according to the output of the Q-function estimator; the update stops when and only when the training end condition is reached;

[0009] Deploy the trained policy network to the entity system corresponding to the virtual environment to implement the corresponding system functions.

[0010] Further, the generation of the training data specifically includes:

[0011] Define the system state of the virtual environment as s, the reward at this moment as r, and the state at the next moment as s'; given s, sampling the virtual environment obtains s' ~ p(s,a);

[0012] Define the policy network π, with the state s as the input and the action a as the output;

[0013] Define the Q-function estimator, with s and a as the inputs, and output the expected cumulative reward of applying the action a in the s state;

[0014] Collect data {<s,a,s′,r> t,i} t=0:T.i=0:N , to form the training data; where s is the state at the current moment, a is the action at the current moment, s′ is the state at the next moment, r is the reward at the current moment, t is the time, i is the sampling trajectory number, T is the total time length of a single sampling trajectory, and N is the total number of sampling trajectories.

[0015] Further, the preset policy is: by controlling variables, the action exploration and environment perturbation are respectively set in the uncertainty estimation of the Q-function estimator.

[0016] Further, the setting of the action exploration and environment perturbation in the uncertainty estimation of the Q-function estimator specifically is:

[0017] Step 2.1, apply the action exploration e, without applying the environment perturbation Δ, and push forward N rounds through the Q-function estimator to estimate the aleatoric uncertainty σ ale (s,a);

[0018] Step 2.2, apply the environment perturbation Δ, without applying the action exploration e, and also push forward N rounds through the Q-function estimator to estimate the epistemic uncertainty σ epi (s,a);

[0019] Step 2.3, if the termination condition is satisfied, end; otherwise repeat Step 2.1 and Step 2.2.

[0020] Further, the specific form of the policy gradient estimation in the preset policy gradient optimization method is:

[0021]

[0022] where s t is the state at the current moment, a t is the action at the current moment, is the estimated value of the state-action value function at the current moment, that is, the output value of the Q-function estimator defined in claim 2, πθ is the policy network, θ are the parameters of the policy network, and π θ (a t (i) |s t (i) ) is the output of the policy network, t is the time, i is the sampling trajectory number, T is the total time length of a single sampling trajectory, and N is the total number of sampling trajectories.

[0023] Furthermore, the virtual environment is a simulation system constructed according to the real robot operation task, corresponding to the physical robot operation scenario, and is used to generate training data, specifically including: 1) a multi-rigid-body dynamics calculation model of the robotic arm, which takes the current state and current action signal of the robotic arm as inputs and numerically calculates and propagates forward to obtain the next state of the robotic arm; 2) an operation object attribute calculation model, which is used to simulate the physical and visual attributes of the operation object.

[0024] Furthermore, the specific way of applying the environmental perturbation is: adjusting the parameters of the operation object attribute calculation model according to claim 6 to realize the simulation of the changes in the physical and visual attributes of the operation object.

[0025] A probability-based policy transfer system, comprising:

[0026] A model construction module, which is used to construct a policy network and a Q-function estimator;

[0027] A virtual environment module, which is used to receive the output of the policy network and decide whether to receive action exploration according to a preset policy, and generate a virtual environment output; the virtual environment is a simulation model corresponding to the physical system;

[0028] A data module, which decides whether to superimpose an environmental perturbation on the virtual environment output according to a preset policy to generate training data;

[0029] A training module, the policy network and the Q-function estimator are updated using the training data, and at the same time, the policy network is updated using a preset policy gradient optimization method according to the output of the Q-function estimator; the update stops only when the training end condition is reached; the trained policy network is used to be deployed to the physical system corresponding to the virtual environment;

[0030] The generation of the training data specifically includes:

[0031] Define the system state of the virtual environment as s, the reward at this moment as r, and the state at the next moment as s'; given s, sampling the virtual environment obtains s'~p(s,a);

[0032] Define the policy network π, which takes the state s as the input and the action a as the output;

[0033] Define a Q - function estimator that takes s and a as inputs and outputs the expected cumulative reward for applying action a in state s;

[0034] Collect data {<s,a,s′,r> t,i} t=0:T.i=0:N to form training data;

[0035] The preset policy is as follows: By means of controlling variables, action exploration and environmental perturbation are respectively set in the uncertainty estimation of the Q - function estimator.

[0036] The setting of action exploration and environmental perturbation respectively in the uncertainty estimation of the Q - function estimator is specifically as follows:

[0037] Step 2.1, Apply action exploration e and do not apply environmental perturbation Δ. Push forward through the Q - function estimator for N rounds to estimate the aleatoric uncertainty σ ale (s,a);

[0038] Step 2.2, Apply environmental perturbation Δ and do not apply action exploration e. Similarly, push forward through the Q - function estimator for N rounds to estimate the epistemic uncertainty σ epi (s,a);

[0039] Step 2.3, If the termination condition is satisfied, end; otherwise, repeat Step 2.1 and Step 2.2;

[0040] The specific form of the policy gradient estimation in the preset policy gradient optimization method is:

[0041]

[0042] where s t is the state at the current moment, a t is the action at the current moment, is the estimated value of the state - action value function at the current moment, π θ is the policy network, θ is the parameter of the policy network, π θ (a t (i) |s t (i) ) is the output of the policy network, t is the time, i is the sampling trajectory number, T is the total time length of a single sampling trajectory, and N is the total number of sampling trajectories;

[0043] The virtual environment is a simulation system constructed according to real robot operation tasks, corresponding to the physical robot operation scenario, and is used to generate training data, specifically including: 1) a multi-rigid body dynamics calculation model of the robotic arm, which takes the current state and the current action signal of the robotic arm as inputs and numerically calculates forward to obtain the next state of the robotic arm; 2) an operation object attribute calculation model, whose function is to simulate the physical and visual attributes of the operation object.

[0044] A computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the method for probability-based policy transfer are implemented.

[0045] A probability-based policy transfer device includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the method for probability-based policy transfer are implemented.

[0046] The advantages of the present invention compared with the prior art are as follows:

[0047] (1) Without assuming the structure and model of the environment and the controlled object, it can be applied to the learning and transfer of various continuous control strategies, and has broad engineering applicability.

[0048] (2) Without assuming the way of policy optimization, either same-policy optimization or different-policy optimization can be used, and it has broad algorithm applicability;

[0049] (3) By constructing a probability-based Q-function estimator through Monte Carlo dropout and combining it with policy gradient optimization, the algorithm has the ability to identify environmental uncertainties. BRIEF DESCRIPTION OF THE DRAWINGS

[0050] Figure 1 is the schematic diagram of the method of the present invention;

[0051] Figure 2 is the flowchart of the method of the present invention;

[0052] Figure 3 is the algorithm flowchart of uncertainty decomposition and inference in step 2 of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0053] In order to better understand the above technical solutions, the technical solutions of the present application will be described in detail below with reference to the drawings and specific embodiments. It should be understood that the specific features in the embodiments of the present application and the embodiments are detailed descriptions of the technical solutions of the present application, rather than limitations on the technical solutions of the present application. Without conflict, the technical features in the embodiments of the present application and the embodiments can be combined with each other.

[0054] The following further elaborates on a probability-based policy transfer method provided by an embodiment of the present application in conjunction with the accompanying drawings of the specification. The specific implementation methods may include (as Figures 1 to 3 shown):

[0055] Step 1, virtual environment training data collection. The virtual environment S refers to establishing a generative model that can generate data: define the system state as s, the reward at this moment as r, and the state at the next moment as s'; given s, sampling the generative model can obtain s' ∼ p(s,a). Define the policy network π, with the state s as the input and the action a as the output. Define the Q-function estimator, with s and a as the inputs, and output the expected cumulative reward of applying the action a to the state s. Collect data {<s,a,s′,r> t,i} t=0:T.i=0:N , and form the dataset D1.

[0056] Step 2, uncertainty decomposition and inference. The purpose of this step is to make the action exploration e and the environmental perturbation Δ be respectively reflected in the uncertainty estimation of the Q-function estimator by controlling variables. First, apply the action exploration e and do not apply the environmental perturbation Δ, and push forward N rounds through the Q-function estimator to estimate the aleatoric uncertainty. Then, apply the environmental perturbation Δ and do not apply the action exploration e, and also push forward N rounds through the Q-function estimator to estimate the epistemic uncertainty. Repeat the above process at a certain period until the termination condition is met (the sampling scale is greater than a certain set value).

[0057] Furthermore, the specific steps of uncertainty decomposition and inference in Step 2 are as follows:

[0058] Step 2.1, apply the action exploration e and do not apply the environmental perturbation Δ, and push forward N rounds through the Q-function estimator to estimate the aleatoric uncertainty σ ale (s,a);

[0059] Step 2.2, apply the environmental perturbation Δ and do not apply the action exploration e, and also push forward N rounds through the Q-function estimator to estimate the epistemic uncertainty σ epi (s,a);

[0060] Step 2.3, if the termination condition is met, end; otherwise repeat Step 2.1 and Step 2.2

[0061] Step 3, policy gradient optimization. For a general policy gradient algorithm, the output of the Q-function estimator is the weighted term of the log policy likelihood function. After introducing the uncertainty estimation, use the Q-function estimation uncertainty σ epi (s,a) obtained in Step 2 to modulate this weighted term, and perform parameter optimization through the policy gradient.

[0062] Step 4, Real - environment running performance evaluation. Run the policy network optimized in Step 3 in the real environment R, apply the same environmental perturbation as in Step 2.2, and evaluate the performance of the policy network π.

[0063] In the solution provided by the embodiments of the present application, as Figure 1 shown, the present invention provides a probability - based virtual - real policy transfer method, which constructs a probability Q - function estimator through Monte Carlo dropout and combines it with policy gradient optimization, enabling the algorithm to have the ability to identify environmental uncertainties. The method is implemented in two stages, namely the training stage and the running stage.

[0064] The implementation process of the training stage is as follows: The virtual environment S, a generative model capable of generating data; collect data {<s,a,s′,r> t,i} t=0:T.i=0:N in the virtual environment S to form a data set D1; the environmental perturbation Δ includes the initial state of the controlled object in the environment, imaging characteristics, measurement noise, surrounding interference, etc.; the policy network π, with the state s as the input and the action a as the output, can be formed as MLP, LSTM, GRU, etc.; the action exploration e refers to adding Gaussian noise to the output space of the policy network π to increase the exploration ability of the system; the Q - function estimator, with s and a as the inputs, outputs the expected cumulative reward of applying the action a in the s state, and its function is to evaluate the performance of the policy network π; the policy gradient optimization O is an algorithm for optimizing the policy network π / Q - function estimator, which can be formed as policy gradient optimization algorithms such as PPO, SAC, etc.

[0065] The implementation process of the running stage is as follows: Collect data {<s,a,s′,r> t,i} t=0:T.i=0:N in the real environment R to form a data set D2; run the policy network optimized in Step 3 in the real environment R, apply the same environmental perturbation Δ as in the training stage, and evaluate the performance of the policy network π; where s is the state at the current moment, a is the action at the current moment, s′ is the state at the next moment, r is the reward at the current moment, t is the time, i is the sampling trajectory number, T is the total time length of a single sampling trajectory, and N is the total number of sampling trajectories.

[0066] As Figure 2 shown, the present invention provides a probability - based virtual - real policy transfer method, which is mainly implemented through four steps: virtual - environment training data collection (21), uncertainty decomposition and inference (22), policy gradient optimization (23), and real - environment running performance evaluation (24).

[0067] As Figure 3As shown, the present invention provides a virtual-real policy transfer method based on probability. The specific algorithm implementation process of step 2, uncertainty decomposition and inference (22), is as follows:

[0068] 31 Initialize the virtual environment S and the Q-function estimator.

[0069] 32 Apply action exploration e.

[0070] 33 Randomly set part of the network weights of the Q-function estimator to zero and repeat for N rounds to estimate the aleatoric uncertainty σ ale (s,a).

[0071] 34 Apply environmental perturbation Δ.

[0072] 35 Randomly set part of the network weights of the Q-function estimator to zero and repeat for N rounds to estimate the aleatoric uncertainty σ ale (s,a).

[0073] 36 Data batch sampling.

[0074] 37 Policy gradient optimization

[0075]

[0076] where s t is the state at the current moment, a t is the action at the current moment, is the estimated value of the state-action value function at the current moment, that is, the output value of the Q-function estimator defined in claim 2, π θ is the policy network, θ is the parameter of the policy network, π θ (a t (i) |s t (i) ) is the output of the policy network, t is the time, i is the sampling trajectory number, T is the total time length of a single sampling trajectory, and N is the total number of sampling trajectories.

[0077] In a possible implementation, the virtual environment is a simulation system constructed according to the real robot operation task, corresponding to the entity robot operation scenario, and is used to generate training data, specifically including: 1) A multi-rigid body dynamics calculation model of the robotic arm, which takes the state of the robotic arm at the current moment (the state refers to the joint angular displacement and joint angular velocity of the robotic arm) and the action signal at the current moment as inputs, and numerically calculates and propagates forward to obtain the state of the robotic arm at the next moment; 2) An operation object property calculation model, whose function is to simulate the physical properties (mass, moment of inertia, surface friction, etc.) and visual properties (surface texture, surface reflection, surface scattering, etc.) of the operation object.

[0078] Further, in a possible implementation, the specific way to apply environmental perturbation is: adjusting the parameters of the operation object attribute calculation model to simulate the changes in the physical and visual attributes of the operation object.

[0079] The present application provides a computer-readable storage medium storing computer instructions, which, when run on a computer, cause the computer to execute Figure 1 the method described above.

[0080] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage and optical storage, etc.) containing computer-usable program code.

[0081] The present application is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or block in the flowcharts and / or block diagrams, as well as the combination of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for implementing the specified functions in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.

[0082] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device that implements the specified functions in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.

[0083] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, so that the instructions executed on the computer or other programmable device provide steps for implementing the specified functions in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.

[0084] Obviously, those skilled in the art can make various changes and modifications to this application without departing from the spirit and scope of this application. Thus, if these modifications and variations of this application fall within the scope of the claims of this application and their equivalent technologies, this application is also intended to include these modifications and variations.

[0085] The content not described in detail in the specification of the present invention belongs to the well-known technology of those skilled in the art.

Claims

1. A probability-based strategy transfer method, characterized in that, Including: Constructing a policy network and a Q-function estimator; The virtual environment receives the output of the policy network and decides whether to receive action exploration according to a preset policy, generating a virtual environment output; the virtual environment is a simulation model corresponding to the physical system; Deciding whether to superimpose environmental perturbations on the virtual environment output according to a preset policy to generate training data; The policy network and the Q-function estimator are updated using the training data. At the same time, the policy network is updated using a preset policy gradient optimization method according to the output of the Q-function estimator; the update stops if and only if the training end condition is reached; Deploying the trained policy network to the physical system corresponding to the virtual environment to implement the corresponding system functions; The generating of the training data specifically includes: Defining the system state of the virtual environment as s, the reward at this moment as r, and the state at the next moment as s'; given s, sampling the virtual environment obtains s'~p(s,a); Defining a policy network π, with state s as the input and action a as the output; Defining a Q-function estimator, with s and a as the inputs, and outputting the expected cumulative reward of applying action a to state s; Collect data {<s,a,s′,r> t,i} t=0:T.i=0:N , and form training data; where s is the state at the current moment, a is the action at the current moment, s′ is the state at the next moment, r is the reward at the current moment, t is the time, i is the sampling trajectory number, T is the total time length of a single sampling trajectory, and N is the total number of sampling trajectories; The preset policy is: by controlling variables, action exploration and environmental perturbations are respectively set in the uncertainty estimation of the Q-function estimator; The setting of action exploration and environmental perturbations respectively in the uncertainty estimation of the Q-function estimator is specifically: Step 2.1, apply the action exploration e, without applying the environmental perturbation Δ, and perform forward propagation for N rounds through the Q-function estimator to estimate the aleatoric uncertainty σ ale (s,a); Step 2.2, apply environmental perturbation Δ and do not apply action exploration e. Also, perform forward propagation through the Q-function estimator for N rounds to estimate the cognitive uncertainty σ epi (s,a); Step 2.3, if the termination condition is satisfied, end; otherwise repeat Step 2.1 and Step 2.2; The virtual environment is a simulation system constructed according to the real robot operation task, corresponding to the physical robot operation scenario, and is used to generate training data, specifically including: 1) A multi-rigid-body dynamics calculation model of the robotic arm, taking the current state and current action signal of the robotic arm as inputs, and numerically calculating and pushing forward to obtain the next state of the robotic arm; 2) An operation object attribute calculation model, used to simulate the physical and visual attributes of the operation object.

2. The method for policy transfer based on probability according to claim 1, characterized in that: The specific form of the policy gradient estimation in the preset policy gradient optimization method is: where s t is the state at the current moment, a t is the action at the current moment, is the estimated value of the state-action value function at the current moment, π θ is the policy network, θ are the parameters of the policy network, π θ (a t (i) |s t (i) ) is the output of the policy network, t is the time, i is the sampling trajectory number, T is the total time length of a single sampling trajectory, and N is the total number of sampling trajectories.

3. A probability-based strategy transfer method according to claim 1, characterized in that: The specific way of applying the environmental perturbation is: adjusting the parameters of the operation object attribute calculation model to simulate the changes in the physical and visual attributes of the operation object.

4. A probability-based policy transfer system, characterized in that, Including: A model construction module, used to construct a policy network and a Q-function estimator; A virtual environment module, used to receive the output of the policy network and decide whether to receive action exploration according to a preset policy, generating a virtual environment output; the virtual environment is a simulation model corresponding to the physical system; A data module, deciding whether to superimpose environmental perturbations on the virtual environment output according to a preset policy to generate training data; A training module, used for the policy network and the Q-function estimator to be updated using the training data. At the same time, the policy network is updated using a preset policy gradient optimization method according to the output of the Q-function estimator; the update stops if and only if the training end condition is reached; the trained policy network is used to be deployed to the physical system corresponding to the virtual environment; The generating of the training data specifically includes: Define the system state of the virtual environment as s, the reward at this moment as r, and the state at the next moment as s'; given s, sampling the virtual environment obtains s' ~ p(s,a); Define the policy network π, which takes the state s as input and the action a as output; Define the Q-function estimator, which takes s and a as input and outputs the expected cumulative reward for applying the action a in the s state; Collect data {<s,a,s′,r> t,i} t=0:T.i=0:N , and form training data; where s is the state at the current moment, a is the action at the current moment, s′ is the state at the next moment, r is the reward at the current moment, t is the time, i is the sampling trajectory number, T is the total time length of a single sampling trajectory, and N is the total number of sampling trajectories; The preset policy is: by controlling variables, set action exploration and environmental perturbation respectively in the uncertainty estimation of the Q-function estimator; The setting of action exploration and environmental perturbation respectively in the uncertainty estimation of the Q-function estimator is specifically: Step 2.1, apply the action exploration e, without applying the environmental perturbation Δ, and perform forward propagation for N rounds through the Q-function estimator to estimate the aleatoric uncertainty σ ale (s,a); ale (s,a); Step 2.2, apply environmental perturbation Δ, without applying action exploration e, and also perform forward inference for N rounds through the Q-function estimator to estimate the cognitive uncertainty σ epi (s,a); epi (s,a); Step 2.3, if the termination condition is satisfied, end; otherwise repeat Step 2.1 and Step 2.2; The virtual environment is a simulation system constructed according to the real robot operation task, corresponding to the entity robot operation scenario, and is used to generate training data, specifically including: 1) The multi-rigid-body dynamics calculation model of the robotic arm, which takes the current state and current action signal of the robotic arm as input and numerically calculates the next state of the robotic arm by forward inference; 2) The operation object attribute calculation model, whose function is to simulate the physical and visual attributes of the operation object.

5. The probabilistic-based policy transfer system according to claim 4, wherein: The specific form of the policy gradient estimation in the preset policy gradient optimization method is: where s t is the state at the current moment, a t is the action at the current moment, is the estimated value of the state-action value function at the current moment, π θ is the policy network, θ are the parameters of the policy network, π θ (a t (i) |s t (i) ) is the output of the policy network, t is the time, i is the sampling trajectory number, T is the total time length of a single sampling trajectory, and N is the total number of sampling trajectories.

6. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the steps of the method according to any one of claims 1 to 3.

Citation Information

Patent Citations

  • Unmanned aerial vehicle autonomous navigation method based on transfer learning

    CN112783199A

  • Reinforcement learning for multi-access traffic management

    US20220014963A1