Robot Skill Learning Method Based on Knowledge Data-Driven Hierarchical Reinforcement Learning
Through the combination of layered reinforcement learning and motion primitives, complex robot operation tasks are decomposed into sub-tasks, and the dimensionality reduction learning strategy parameters are solved, and the existing reinforcement learning algorithms are slow convergence speed and dimensionality disaster problems in high-dimensional states and action spaces are achieved, and the acceleration and generalization performance of robot skill learning is improved.
Patent Information
- Application Number
- CN202310000085.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-01-01
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2043-01-01
AI Technical Summary
When existing reinforcement learning algorithms deal with robot operation tasks in high-dimensional, continuous states and actions, it is easy to cause slow convergence speed and dimensional disasters, making it difficult to achieve ideal learning effects.
The hierarchical reinforcement learning method is adopted to decompose complex tasks into multiple subtasks, design the task structure and strategy structure through human knowledge, and use motion primitives to represent subtask strategies, realize dimensionality reduction of state space and action space, and optimize the strategy parameters of each subtask based on classic reinforcement learning algorithms.
Through layered reinforcement learning, accelerate the speed of robot skills learning, reduce the search space of control strategies, improve the convergence speed of learning and the robustness of algorithms, and enable robots to be applied in a wider scenario.
Smart Images

Figure CN116306896B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of robotics, and particularly relates to a method for robot skill learning. Background Art
[0002] Robots have broad application scenarios in the future. Applying reinforcement learning algorithms to robot problems enables robots to learn skills by interacting with the environment, making them better applicable in unstructured environments. Different from typical reinforcement learning benchmark problems, many practical problems in the field of robotics have high-dimensional and continuous states and actions. As the dimension increases, the computational amount grows exponentially, which leads to the problem of the curse of dimensionality. Therefore, the convergence speed of the learning system will slow down, and even an ideal learning effect cannot be achieved. Utilizing human knowledge to introduce abstract concepts, designing task structures and policy structures, and constructing a hierarchical reinforcement learning framework is a feasible solution.
[0003] When using reinforcement learning algorithms to handle large-scale problems, state-action pairs will lead to problems such as slow convergence speed and the curse of dimensionality when representing behavioral strategies. Many complex systems in nature have a hierarchical feature, that is, large-scale systems can be divided into several smaller-scale subsystems. Hierarchical reinforcement learning decomposes difficult tasks into simpler subtasks and completes the entire task by learning the subtasks. Based on this idea, hierarchical reinforcement learning introduces hierarchical and abstract ideas to reduce the dimensionality of the state space and action space, accelerate the convergence speed of learning, and uses classical reinforcement learning methods to learn in each constrained hierarchical task to search for appropriate execution strategies for each subtask. The execution strategy of a specific subtask is expressed using motion primitives. Based on the policy structure expression of motion primitives, the search space of control strategies can be greatly reduced, making it feasible to obtain the desired strategy using learning methods in a high-dimensional motion system, which is beneficial to the design and online adjustment of the control system. The hierarchical optimal policy is the maximum cumulative reward policy applicable to all hierarchical structures, and the subtasks need to be associated with their upper and lower level subtasks. The recursive optimal policy acts optimally on the strategies of all subtasks.
[0004] Human-machine hybrid intelligence is one of the core research directions of the new generation of artificial intelligence. Combining human knowledge and the computing power of agents is an implementation form of hybrid intelligence. Using human knowledge to design task structures and policy structures can construct an interpretable task model and improve the interpretability of intelligent systems. Usually, there will be some redundant information in the subtasks after task stratification. Introducing abstract ideas lies in simplifying the problem and removing redundant information, thereby accelerating the learning process. By combining hierarchical reinforcement learning and motion primitives, the robustness of reinforcement learning algorithms applied to robot problems can be improved. Summary of the Invention
[0005] To overcome the deficiencies of the prior art, the present invention provides a robot skill learning method based on knowledge data-driven hierarchical reinforcement learning. First, establish the dynamic equation of the robot model and construct a robot operation skill library; then, based on human knowledge, split the robot operation task into multiple subtasks, design the policy structure of the subtasks, and realize the dimensionality reduction of the state space and action space; next, select appropriate motion primitives from the robot operation skill library, and design the state space, action space, and reward function of the robot operation task; then, optimize the policy parameters of each subtask based on the traditional reinforcement learning algorithm; finally, use the optimized policy parameters of each subtask to implement the robot task. The present invention overcomes the limitations of the existing reinforcement learning in robot operation application problems, realizes the skill learning of real robots, and enables robots to be applied in a wider range of scenarios.
[0006] The technical solution adopted by the present invention to solve its technical problems includes the following steps:
[0007] Step 1: Establish the dynamic equation of the robot model and construct a robot operation skill library;
[0008] Use the impedance control method to achieve compliant control of the robot in the Cartesian space. The impedance controller in the Cartesian space is shown in Equation (1).
[0009]
[0010] Where, M d , C d , K d respectively represent the environmental mass, damping, and stiffness parameters, x d represents the desired trajectory of the robot, x represents the actual trajectory of the robot, and F e represents the contact force between the robot and the contacting object;
[0011] The robot operation skill library includes three basic robot operation skills: trajectory tracking, contact operation, and hitting skills, and these three skills are represented by dynamic motion primitives, compliant motion primitives, and hitting motion primitives;
[0012] Step 2: Based on human knowledge, split the robot operation task into multiple subtasks, design the policy structure of the subtasks, and realize the dimensionality reduction of the state space and action space;
[0013] Step 3: Select motion primitives from the robot operation skill library, and design the state space, action space, and reward function of the robot operation task;
[0014] For different types of task characteristics, select different types of motion primitives for policy expression; the state and action spaces of the system are composed of policy parameters; the cost function of the robot operation task is as follows:
[0015]
[0016] Among them, J is the cost of the trajectory τ i within a finite time, including the final cost That is, the instantaneous cost r t and the instantaneous control cost t i 、t N respectively represent the initial time and the final time of the finite time;
[0017] Step 4: Optimize the policy parameters of each subtask based on the reinforcement learning algorithm;
[0018] Using the policy improvement method based on path integral, explore M times and execute in each iteration process; Based on the policy parameters and cost function of M explorations, update the policy, and its parameter update rule is as follows, until the algorithm converges:
[0019]
[0020]
[0021]
[0022]
[0023]
[0024]
[0025] θ←θ + δθ (9)
[0026] Among them, τ i represents the trajectory of each trial, {τ i} k represents the trajectory of the Kth exploration, S({τ i} k ) represents the cost of each trial trajectory, N represents the number of time steps, represents the instantaneous cost, R represents the positive semi - definite weight matrix of the quadratic control cost, θ represents the policy parameter, represents the noise parameter, represents the feature vector of the policy, λ represents the scale, K represents the number of explorations, represents the Gaussian basis function, [δθ] j represents the policy average value of each time step;
[0027] Step 5: Use the optimized policy parameters of each subtask to implement the robot task.
[0028] Preferably, the dynamic motion primitive is a non - linear dynamic system, including a second - order system and a trajectory shape learner, as shown in Equation (10) and Equation (11); a second - order dynamic system with self - stability is used to construct an attractor point model, and the final state of the system is changed through this attractor point, so as to achieve the purpose of modifying the target position of the trajectory. The trajectory shape learner fits various trajectory shapes through the normalized linear superposition of multiple non - linear basis functions; the dynamic motion primitive is used to represent the trajectory - level strategy:
[0029]
[0030]
[0031] where α y and β y are constants, equivalent to the P parameter and D parameter in the PD controller, g represents the target state, y represents the system state, ψ i (t) represents the Gaussian basis function, and ω i represents the weight.
[0032] The beneficial effects of the present invention are as follows:
[0033] 1. The present invention adopts the method of hierarchical reinforcement learning, which can accelerate the speed of robot skill learning while ensuring the completion of operation tasks.
[0034] 2. The present invention uses different motion primitives to represent various robot operation skills, constructs a skill library including various robot operation skills, and improves the generalization performance of the robot skill learning method.
[0035] 3. The present invention splits the robot operation task into multiple subtasks based on human knowledge and experience, and designs the strategy structure of the subtasks.
[0036] 4. The present invention realizes the dimensionality reduction of the state space and action space by means of task stratification and representing subtask strategies with motion primitives, effectively reducing the search space of the reinforcement learning algorithm and accelerating the convergence speed of learning.
[0037] 5. The present invention uses the classical reinforcement learning method to learn in each constrained hierarchical task, searches for appropriate execution strategies for each subtask. By introducing the ideas of stratification and abstraction, and designing an interpretable task model, it can improve the interpretability of the algorithm and the transparency of the algorithm to operators. Brief Description of the Drawings
[0038] Figure 1 is the flowchart of the method of the present invention.
[0039] Figure 2 is the schematic diagram of task decomposition of the method of the present invention based on human knowledge.
[0040] Figure 3 Schematic diagram of the robot operation skill library for the method of the present invention.
[0041] Figure 4 Schematic diagram of parameter update based on the reinforcement learning algorithm for the method of the present invention.
[0042] Figure 5 Simulation model diagram of the robotic arm jacking task in the embodiment of the present invention.
[0043] Figure 6 Cost function curve graph in the embodiment of the present invention. Detailed implementation manners
[0044] The present invention will be further described below in conjunction with the accompanying drawings and embodiments.
[0045] Since robot operation tasks have the characteristics of high-dimensional and continuous state and action spaces, when a robot conducts skill learning, the agent needs to learn a large number of parameters and requires a large storage space. Therefore, it is difficult to quickly obtain the optimal policy parameters and achieve an ideal learning effect. Therefore, a robot skill learning technology is proposed that uses human knowledge and experience to construct a skill library, design a task structure, and accelerate the learning speed based on hierarchical reinforcement learning.
[0046] The purpose of the present invention is to provide a method that can be applied to robot tasks with high-dimensional and continuous state and action spaces. This method can enable a robot to quickly and efficiently achieve skill learning during complex operations. At the same time, by constructing a robot operation skill library, the generalization performance of the robot skill learning method is enhanced.
[0047] Hierarchical reinforcement learning adopts a way similar to human cognition and decision-making. By coherently combining simple operations into a sequence of subtasks, these subtask sequences work together to achieve the overall task and goal. Using the knowledge and experience of human users, the task is divided into different execution stages, appropriate low-level controllers are selected, and simple behaviors are provided by selecting appropriate forms of motion primitives to achieve the task by completing the sequence of subtasks. Design the task hierarchical structure based on human knowledge, construct the policy structure, design a parameterized policy, and optimize the policy parameters through reinforcement learning to obtain the learning of robot skills.
[0048] The present invention proposes a robot skill learning method combining hierarchical reinforcement learning and motion primitives. First, the operator designs abstract concepts based on knowledge and experience, decomposes the task into multiple subtasks, and constructs a policy structure suitable for the corresponding subtasks, such as Figure 2As shown in the figure. The strategy is modeled in the form of motion primitives to achieve the effect of reducing the state and action spaces, constructing a robot operation skill library, and improving the generalization performance of the robot skill learning method. In the present invention, as Figure 3 , it is considered that the robot operation skill library includes three basic robot operation skills: trajectory tracking, contact operation, and hitting skills, and these three skills are represented by dynamic motion primitives, compliant motion primitives, and hitting motion primitives. The optimal parameters of each subtask strategy are learned using a reinforcement learning algorithm, as Figure 4 shown. The method for robot skill learning proposed in the present invention combines human knowledge and the reinforcement algorithm using hierarchical reinforcement learning. While ensuring the convergence of the reinforcement learning algorithm, it speeds up the convergence rate of the reinforcement learning algorithm by reducing the dimensions of the state and action spaces, making it feasible to apply in high-dimensional problems represented by robot operation problems. The objective of the present invention is to fuse human intelligence and machine intelligence using hierarchical reinforcement learning to improve the interpretability and transparency of the algorithm, overcome the limitations of existing reinforcement learning in robot operation application problems, and achieve the skill learning of real robots, enabling robots to be applied in a wider range of scenarios. The method proposed in the present invention can be applied to various complex scenarios of robot skill learning. Specific embodiments:
[0050] Model the robot skill learning task, design the control algorithm of the robot, study key technologies such as the construction of the robot operation skill library, task decomposition of the operation task, policy parameterization, reward function construction, and reinforcement learning algorithm, and construct a robot skill learning framework based on hierarchical reinforcement learning. The specific implementation steps of the present invention are as follows:
[0051] First step: Use the impedance control method to achieve compliant control of the robot in the Cartesian space. The impedance controller in the Cartesian space is shown in Equation (1).
[0052]
[0053] Among them, M d , C d , K d respectively represent the environmental mass, damping, and stiffness parameters, x d represents the desired trajectory of the robot, x represents the actual trajectory of the robot, and F e represents the contact force between the robot and the contacted object.
[0054] Step 2: Based on human knowledge, decompose the robot operation tasks into multiple stages, construct a motion skill library based on multiple motion primitives, and use the motion primitives to represent the strategies for each stage. Taking the dynamic motion primitive as an example, the dynamic motion primitive is a non-linear dynamic system, including a second-order system (Equation (2)) and a trajectory shape learner (Equation (3)). Use the second-order dynamic system with self-stability to construct an "attracting point" model, and change the final state of the system through this "attracting point" to achieve the purpose of modifying the target position of the trajectory. The trajectory shape learner fits various trajectory shapes through the normalized linear superposition of multiple non-linear basis functions. The dynamic motion primitive can be used to represent the trajectory-level strategy. The compliant motion primitive and the striking motion primitive have similar representation forms to the dynamic motion primitive.
[0055]
[0056]
[0057] Step 3: For different types of task characteristics, different types of motion primitives can be selected for strategy expression. The state and action space of the system are composed of policy parameters. Design the cost function of the robot operation task as follows:
[0058]
[0059] where J is the cost of the trajectory τ i within a finite time, including the final cost the immediate cost r t and the immediate control cost
[0060] Step 4: Use the path integrals policy improvement method. Explore M times and execute in each iteration process. Update the policy based on the policy parameters and cost function of M explorations. The parameter update rule is as follows. Until the algorithm converges.
[0061]
[0062]
[0063]
[0064]
[0065]
[0066]
[0067] θ ← θ + δθ
[0068] The method for realizing robot skill learning by using human knowledge and experience based on the hierarchical reinforcement learning algorithm is as follows Figure 1 shown
[0069] The method proposed by the present invention has been verified in the simulation environment. A simulation model of the robotic arm jacking task was constructed in the Mujoco simulation environment, as Figure 5 shown. Taking the robotic arm jacking task as an example, the robotic arm jacking task is simplified into two subtasks: approaching and inserting. The compliant motion primitive is used as the task representation form of the two subtasks. The policy of the robotic arm jacking task includes the trajectory in the Cartesian space and the robotic arm impedance parameters. In the approaching stage, the policy consists of 6 compliant motion primitives, including the trajectories in the X, Y, and Z directions and the robotic arm impedance parameters; in the inserting stage, the movement of the robotic arm is restricted within the Z direction. Therefore, the policy only includes two motion primitives: the trajectory in the Z direction and the robotic arm impedance parameters
[0070] Due to the overly large exploration space, it is unrealistic to directly use the reinforcement learning algorithm to learn the jacking task. By using human knowledge to divide the task stages and simplify the task parameters, the robotic arm can quickly achieve the jacking task. Through learning, the robotic arm can achieve the effect of maintaining a small contact force while jacking, and its cost function is as Figure 6 shown
Claims
1. A robot skill learning method based on knowledge data-driven hierarchical reinforcement learning, characterized in that, it includes the following steps: Step 1: Establish the dynamic equation of the robot model and construct a robot operation skill library; Use the impedance control method to achieve compliant control of the robot in the Cartesian space. The impedance controller in the Cartesian space is shown in Equation (1): Among them, M d , C d , K d respectively represent environmental quality, damping, and stiffness parameters, x d represents the desired trajectory of the robot, x represents the actual trajectory of the robot, and F e represents the contact force between the robot and the contacting object; The robot operation skill library includes three basic robot operation skills: trajectory tracking, contact operation, and hitting skills, which are represented by dynamic motion primitives, compliant motion primitives, and hitting motion primitives; Step 2: Split the robot operation task into multiple subtasks based on human knowledge, design the policy structure of the subtasks, and realize the dimensionality reduction of the state space and action space; Step 3: Select motion primitives from the robot operation skill library, and design the state space, action space, and reward function of the robot operation task; For different types of task characteristics, select different types of motion primitives for policy expression; the state and action space of the system are composed of policy parameters; the cost function of the robot operation task is as follows: where J is the cost of trajectory τ i over a finite time, including the terminal cost the instantaneous cost r t and the instantaneous control cost t i and t N represent the initial and final times of the finite time, respectively; Step 4: Optimize the policy parameters of each subtask based on the reinforcement learning algorithm; Use the policy improvement method based on path integral, explore M times and execute in each iteration process; based on the policy parameters and cost function of M times of exploration, update the policy, and its parameter update rule is as follows, until the algorithm converges: θ←θ+δθ (9) Among them, τ i represents the trajectory of each trial, {τ i} k represents the trajectory of the K-th exploration, S({τ i} k ) represents the cost of each trial trajectory, N represents the number of time steps, represents the immediate cost, R represents the positive semi-definite weight matrix of the quadratic control cost, θ represents the policy parameter, represents the noise parameter, represents the eigenvector of the policy, λ represents the scale, K represents the number of explorations, represents the Gaussian basis function, [δθ] j represents the policy average value of each time step; Step 5: Use the optimized policy parameters of each subtask to implement the robot task.
2. The robot skill learning method based on knowledge data-driven hierarchical reinforcement learning according to claim 1, characterized in that, The dynamic motion primitive is a nonlinear dynamic system, including a second-order system and a trajectory shape learner, as shown in Equation (10) and Equation (11); use a second-order dynamic system with self-stability to construct an attractor model, and change the final state of the system through this attractor, so as to achieve the purpose of modifying the trajectory target position. The trajectory shape learner fits various trajectory shapes through the normalized linear superposition of multiple nonlinear basis functions; use the dynamic motion primitive to realize the representation of the trajectory-level policy: where α y , β y are constants, corresponding to the P parameter and D parameter in the PD controller, g represents the target state, y represents the system state, ψ i (t) represents the Gaussian basis function, and ω i represents the weight.
Citation Information
Patent Citations
Hexapod robot impedance control method based on reinforcement learning
CN112743540A
For hiearchical decomposition deep reinforcement learning for an artificial intelligence model
WO2018236674A1