Robot reinforcement learning method, device, equipment and medium for ensuring contact safety
By generating Cartesian spatial variable impedance control parameters and pose constraint actions in the robot reinforcement learning method, and compensating the process when the collision occurs, the dangerous behavior problems of the robot in unknown scenarios or when encountering interference are solved, and the safety of task space and joint space is achieved.
Patent Information
- Application Number
- CN202210744763.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-27
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2042-06-27
AI Technical Summary
Existing robot reinforcement learning methods are prone to dangerous behaviors in unknown scenarios or when encountering interference, and fail to effectively consider the contact safety of joint space.
By acquiring the robot target task, the parameters of Cartesian spatial variable impedance control are generated using a preset reinforcement learning strategy, and constraint actions are generated based on the pose constraints. If a collision occurs, calculate the compensation amount of variable impedance control, correct the attitude constraints, and generate the compensation execution actions to ensure contact safety.
Effectively avoid accidental collisions, ensure that the robot does not cause dangerous behavior in unknown scenarios or encounters interference, realizes the safety of task space and joint space, and has good scalability.
Smart Images

Figure CN115042178B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of intelligent robot technology, and in particular to a robot reinforcement learning method, device, equipment and medium that ensures contact safety. Background Art
[0002] Reinforcement learning applied to robot manipulation tasks can handle complex tasks, has the advantages of strong generalization performance and high efficiency, and provides a way to solve complex manipulation tasks using a model-free approach. Safety issues are a major obstacle that prevents reinforcement learning from being widely used in real-world tasks. The common strategy of reinforcement learning in the exploration phase is to adopt random behavior or explore unknown states as much as possible, but these methods can easily cause robots to behave dangerously. Even when the training process of the agent has converged, the agent may produce some unexpected and dangerous behaviors in unknown scenarios or when encountering interference.
[0003] Depending on where the collision occurs, collision safety is divided into task space safety and joint space safety. Task space safety requires that the contact force between the robot end effector and the target object it performs the task on is within a safe range. Joint space safety focuses on potentially dangerous collisions with the robot, which are common in human-robot interaction tasks or unstructured environments.
[0004] The common method for task space safety is Cartesian space variable impedance control. Joint space safety mainly focuses on avoiding collisions. However, avoiding all collisions requires accurate object geometry information, which is very difficult to obtain in actual scenarios. In addition, in chaotic or unknown environments, accidental collisions are generally inevitable. In some cases, it is even necessary to contact the environment to better complete the task. In these cases, joint space safety should not be achieved by avoiding collisions, but by limiting the size of the contact force.
[0005] However, none of the existing methods considers the contact safety of such joint spaces when using reinforcement learning scenarios. Summary of the invention
[0006] The present application provides a robot reinforcement learning method, device, equipment and medium that ensure contact safety, so as to solve the problem that the existing methods cannot avoid accidental collisions, resulting in some unexpected dangerous behaviors in unknown scenarios or when encountering interference, and do not consider issues such as contact safety in joint space.
[0007] The first aspect of the present application provides a robot reinforcement learning method for ensuring contact safety, comprising the following steps: obtaining a robot target task; using a preset reinforcement learning strategy to generate parameters of the Cartesian space variable impedance control of the target task, and generating a constraint action based on a posture constraint; if the robot does not collide, a first execution action is generated based on the parameters of the Cartesian space variable impedance control and the constraint action, and the robot is controlled to perform the target task according to the first execution action; otherwise, a variable impedance control compensation amount is calculated based on the collision force, and the parameters of the Cartesian space variable impedance control are compensated by using the variable impedance control compensation amount, while the posture constraint is corrected to generate a compensation constraint action, a second execution action is generated based on the compensated Cartesian space variable impedance control parameters and the compensation constraint action, and the robot is controlled to perform the target task according to the second execution action.
[0008] Optionally, in one embodiment of the present application, before using the preset reinforcement learning strategy to generate the parameters of the Cartesian space variable impedance control of the target task, it also includes: taking the robot target task as input and the parameters of the Cartesian space variable impedance control of the target task as output, performing simulation training to obtain the preset reinforcement learning strategy, wherein the parameters of the Cartesian space variable impedance control include the balance point, stiffness matrix and damping matrix corresponding to the robot target task.
[0009] Optionally, in one embodiment of the present application, a preset reinforcement learning strategy is used to generate parameters of the Cartesian space variable impedance control of the target task, wherein the parameters of the Cartesian space variable impedance control are:
[0010]
[0011] Among them, K 1 is the stiffness matrix, D 1 is the damping matrix, x d is the balance point, is the task space quality matrix, is the dynamically consistent task Jacobian matrix, x is the position of the end effector in Cartesian space, is the velocity of the end effector in Cartesian space, are the Coriolis force and centripetal force vectors, p t (q) is the gravity vector, τ ext is the torque on the robot joint, q is the robot's configuration in generalized coordinates, is the velocity of the robot in generalized coordinates;
[0012] Generate a constraint action according to the posture constraint, wherein the constraint action is:
[0013]
[0014] Among them, K 2 is the joint stiffness matrix, D 2 is the damping matrix, q pose is the desired configuration of the robot.
[0015] Optionally, in one embodiment of the present application, the first execution action is generated according to the parameters and constraint actions of the Cartesian space variable impedance control, wherein the first execution action is:
[0016]
[0017] in, is the robot's task space Jacobian matrix, is the null space projection matrix corresponding to the Jacobian matrix of the robot task space.
[0018] Optionally, in one embodiment of the present application, the variable impedance control compensation amount is used to compensate the parameters of the Cartesian space variable impedance control, wherein the compensated parameters of the Cartesian space variable impedance control are:
[0019]
[0020] Among them, γ is the residual vector.
[0021] Optionally, in one embodiment of the present application, the second execution action is generated according to the parameters of the compensated Cartesian spatial variable impedance control and the compensation constraint action, wherein the second execution action is:
[0022]
[0023] in, The projection matrix to the task-consistent null space for the degenerate contact Jacobian.
[0024] The second aspect of the present application provides a robot reinforcement learning device that ensures contact safety, including: an acquisition module, used to acquire a robot target task; a generation module, used to generate parameters of Cartesian space variable impedance control of the target task using a preset reinforcement learning strategy, and generate a constraint action based on a posture constraint; a processing module, used to generate a first execution action based on the parameters and constraint action of the Cartesian space variable impedance control when the robot has not collided, and control the robot to perform the target task according to the first execution action; when the robot collides, calculate the variable impedance control compensation amount according to the collision force, use the variable impedance control compensation amount to compensate for the parameters of the Cartesian space variable impedance control, and correct the posture constraint to generate a compensation constraint action, generate a second execution action based on the compensated Cartesian space variable impedance control parameters and the compensation constraint action, and control the robot to perform the target task according to the second execution action.
[0025] Optionally, in one embodiment of the present application, the generation module is specifically used to perform simulation training with the robot target task as input and the parameters of the Cartesian space variable impedance control of the target task as output before using the preset reinforcement learning strategy to generate the parameters of the Cartesian space variable impedance control of the target task, so as to obtain the preset reinforcement learning strategy, wherein the parameters of the Cartesian space variable impedance control include the balance point, stiffness matrix and damping matrix corresponding to the robot target task.
[0026] Optionally, in one embodiment of the present application, the generating module is specifically used to generate the parameters of the Cartesian space variable impedance control of the target task using a preset reinforcement learning strategy, wherein the parameters of the Cartesian space variable impedance control are:
[0027]
[0028] Among them, K 1 is the stiffness matrix, D 1 is the damping matrix, x d is the balance point, is the task space quality matrix, is the dynamically consistent task Jacobian matrix, x is the position of the end effector in Cartesian space, is the velocity of the end effector in Cartesian space, are the Coriolis force and centripetal force vectors, p t (q) is the gravity vector, τ ext is the torque on the robot joint, q is the robot's configuration in generalized coordinates, is the velocity of the robot in generalized coordinates;
[0029] Generate a constraint action according to the posture constraint, wherein the constraint action is:
[0030]
[0031] Among them, K 2 is the joint stiffness matrix, D 2 is the damping matrix, q pose is the desired configuration of the robot.
[0032] Optionally, in an embodiment of the present application, the processing module is specifically configured to generate the first execution action according to the parameters and constraint actions of the Cartesian space variable impedance control, wherein the first execution action is:
[0033]
[0034] in, is the robot's task space Jacobian matrix, is the null space projection matrix corresponding to the Jacobian matrix of the robot task space.
[0035] Optionally, in one embodiment of the present application, the processing module is specifically used to compensate the parameters of the Cartesian space variable impedance control by using the variable impedance control compensation amount, wherein the compensated parameters of the Cartesian space variable impedance control are:
[0036]
[0037] Among them, γ is the residual vector.
[0038] Optionally, in one embodiment of the present application, the processing module is specifically configured to generate the second execution action according to the compensated Cartesian spatial variable impedance control parameter and the compensation constraint action, wherein the second execution action is:
[0039]
[0040] in, The projection matrix to the task-consistent null space for the degenerate contact Jacobian.
[0041] A third aspect of the present application provides an electronic device, comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to perform a robot reinforcement learning method for ensuring contact safety as described in the above embodiment.
[0042] The fourth aspect of the present application provides a computer-readable storage medium having a computer program stored thereon, which is executed by a processor to perform the robot reinforcement learning method for ensuring contact safety as described in the above embodiment.
[0043] Therefore, this application has at least the following beneficial effects:
[0044] By obtaining the robot's target task, the reinforcement learning strategy is used to generate the parameters of the Cartesian space variable impedance control of the target task, and the constraint action is generated according to the posture constraint; the behavior is switched according to whether a collision occurs at the joint. If no collision occurs, the controller executes the Cartesian space variable impedance control target given by the reinforcement learning strategy; and through the zero space projection method, a posture task is executed under the premise of ensuring that the main task is not disturbed to maintain a specific posture, thereby reducing the problem of uncontrolled joints that may be caused when only the main task is executed. If a collision occurs, the compliant controller will estimate the magnitude of the external force currently received and take this force into account to compensate for the disturbance; at the same time, the zero space projection matrix at this time is modified, the purpose of which is to make the robot apply zero force in the direction of the collision without violating the collision and interfering with the main task. Therefore, by actively detecting the external collision force received by the robot and adjusting its own behavior according to the collision force, the safety of the task space and the joint space is guaranteed. At the same time, since the reinforcement learning strategy only constrains the action space, it also has good scalability. This solves the problem that existing methods cannot avoid accidental collisions, resulting in some unexpected dangerous behaviors in unknown scenarios or when encountering interference, and do not consider issues such as contact safety in joint space.
[0045] Additional aspects and advantages of the present application will be given in part in the description below, and in part will become apparent from the description below, or will be learned through the practice of the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] The above and / or additional aspects and advantages of the present application will become apparent and easily understood from the following description of the embodiments in conjunction with the accompanying drawings, in which:
[0047] Figure 1 A flowchart of a robot reinforcement learning method for ensuring contact safety provided according to an embodiment of the present application;
[0048] Figure 2 A framework diagram of a robot reinforcement learning method for ensuring contact safety provided according to an embodiment of the present application;
[0049] Figure 3 A schematic block diagram of a robot reinforcement learning device for ensuring contact safety provided according to an embodiment of the present application;
[0050] Figure 4A schematic diagram of the structure of an electronic device provided for an application embodiment.
[0051] Explanation of reference numerals: acquisition module-100, generation module-200, processing module-300, memory-401, processor-402, communication interface-403. DETAILED DESCRIPTION
[0052] Embodiments of the present application are described in detail below, and examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to be used to explain the present application, and should not be construed as limiting the present application.
[0053] The following describes a robot reinforcement learning method, device, equipment and medium for ensuring contact safety in an embodiment of the present application with reference to the accompanying drawings. In view of the problems mentioned in the above background technology center, the present application provides a robot reinforcement learning method for ensuring contact safety, in which the robot target task is obtained; the parameters of the Cartesian space variable impedance control of the target task are generated using a preset reinforcement learning strategy, and a constraint action is generated according to the posture constraint; if the robot does not collide, a first execution action is generated according to the parameters and constraint action of the Cartesian space variable impedance control, and the robot is controlled to perform the target task according to the first execution action, otherwise the variable impedance control compensation amount is calculated according to the collision force, and the parameters of the Cartesian space variable impedance control are compensated by the variable impedance control compensation amount, and the posture constraint is corrected to generate a compensation constraint action, and a second execution action is generated according to the compensated Cartesian space variable impedance control parameters and the compensation constraint action, and the robot is controlled to perform the target task according to the second execution action. Thus, the existing method cannot avoid accidental collisions, resulting in some unexpected dangerous behaviors in unknown scenarios or when encountering interference, and the contact safety of the joint space is not considered.
[0054] Specifically, Figure 1 A schematic flow chart of a robot reinforcement learning method for ensuring contact safety provided in an embodiment of the present application.
[0055] like Figure 1 As shown, the robot reinforcement learning method for ensuring contact safety includes the following steps:
[0056] In step S101, the robot target task is obtained.
[0057] In an embodiment of the present application, the target task of the robot is first obtained so that whether a collision occurs can be subsequently identified based on the target task of the robot, and the position and posture of the robot can be adjusted while ensuring the safety of the task space and the safety of the joint operation space.
[0058] In step S102, the parameters of the Cartesian space variable impedance control of the target task are generated using a preset reinforcement learning strategy, and a constrained action is generated according to the posture constraint.
[0059] In order to ensure contact safety, the embodiments of the present application use reinforcement learning methods to ensure the safety of the task space and the safety of the joint operation space. Figure 2 As shown, the robot reinforcement learning method for ensuring contact safety in the embodiment of the present application is divided into two parts: reinforcement learning strategy and compliance controller. The first is the reinforcement learning strategy part: output the balance point, stiffness matrix and damping matrix of Cartesian space variable impedance control (main task), and input the information required to complete the specific task. Compliant controller part: switch the behavior according to whether a collision occurs at the joint. The compliance controller part is introduced in detail in the following embodiments.
[0060] The embodiments of the present application can design reinforcement learning strategies and training algorithms according to the specific requirements of the task and the actual situation, and ensure that the output space of the reinforcement learning strategy is the parameter of Cartesian space variable impedance control. Among them, the reinforcement learning strategy should at least include: the input of the reinforcement learning agent (for example: selecting pictures and the posture of the robot end effector as input); the composition of the reinforcement learning agent neural network (for example: using variational autoencoders as feature extractors, and using fully connected layers to convert the obtained features into network outputs); the training algorithm used by the reinforcement learning strategy (for example: PPO Proximal Policy Optimization), etc.
[0061] It should be noted that the output of the reinforcement learning strategy must be: the parameters of the Cartesian space variable impedance control, for example: the equilibrium point x d , stiffness matrix K 1 and the damping matrix D 1 The amount of change.
[0062] Furthermore, the parameters of the Cartesian space variable impedance control of the target task are generated using a preset reinforcement learning strategy, wherein the parameters of the Cartesian space variable impedance control are:
[0063]
[0064] Among them, K 1 is the stiffness matrix, D 1 is the damping matrix, x d is the equilibrium point, which is determined by the output of the reinforcement learning strategy. is the task space quality matrix, is the dynamically consistent task Jacobian matrix, x is the position of the end effector in Cartesian space, is the velocity of the end effector in Cartesian space, are the Coriolis force and centripetal force vectors, p t (q) is the gravity vector, τ ext is the torque on the robot joint, q is the robot's configuration in generalized coordinates, is the velocity of the robot in generalized coordinates.
[0065] Since the direct use of Cartesian space variable impedance control produces unpredictable motion in free space, we need an additional attitude task to impose certain constraints and generate constraint actions based on the attitude constraints. The constraint actions are:
[0066]
[0067] Among them, K 2 is the joint stiffness matrix, D 2 is the damping matrix, q pose is the desired configuration of the robot, q pose Used x d Calculated using inverse kinematics.
[0068] Optionally, in one embodiment of the present application, before using a preset reinforcement learning strategy to generate the parameters of the Cartesian space variable impedance control of the target task, it also includes: taking the robot target task as input and the parameters of the Cartesian space variable impedance control of the target task as output, performing simulation training to obtain the preset reinforcement learning strategy, wherein the parameters of the Cartesian space variable impedance control include the balance point, stiffness matrix and damping matrix corresponding to the robot target task.
[0069] It can be understood that after obtaining the target task, the embodiment of the present application takes the target task as input, and the balance point, stiffness matrix and damping matrix corresponding to the robot target task as output, and performs training in simulation to obtain a preset reinforcement learning strategy.
[0070] It should be noted that the embodiments of the present application do not have any limitations on the details of the training method or deployment. For example, real machine fine-tuning training models, offline training and other technologies can be used according to specific needs.
[0071] In step S103, if the robot does not collide, a first execution action is generated according to the parameters and constraint actions of the Cartesian space variable impedance control, and the robot is controlled to perform the target task according to the first execution action. Otherwise, the variable impedance control compensation amount is calculated according to the collision force, and the parameters of the Cartesian space variable impedance control are compensated by the variable impedance control compensation amount. At the same time, the posture constraint is corrected to generate a compensation constraint action, and a second execution action is generated according to the compensated Cartesian space variable impedance control parameters and the compensation constraint action, and the robot is controlled to perform the target task according to the second execution action.
[0072] When there is no collision, the controller executes the Cartesian space variable impedance control target given by the reinforcement learning strategy; and through the zero space projection method, it performs a posture task to maintain a specific posture while ensuring that the main task is not interfered with, thereby reducing the problem of uncontrolled joints that may be caused when only the main task is executed. When a collision occurs, the compliant controller estimates the magnitude of the external force currently received, takes this force into account, and compensates for the disturbance; at the same time, it modifies the zero space projection matrix at this time, with the purpose of making the robot apply zero force in the direction of the collision without violating the collision and interfering with the main task.
[0073] In the embodiments of the present application, the specific implementation codes may be different depending on the specific model of the robot used and the simulation environment used, but the ideas and the same principles are the same. It should be emphasized that the robot must include redundant degrees of freedom and be a force-controlled robot.
[0074] The dynamic equation of the robot can be written as:
[0075]
[0076] in is the robot generalized coordinate, is the robot's mass matrix, are the Coriolis and centripetal forces of the robot, is the gravity term, and the right and They are respectively the torque output by the robot driver and the torque received by the robot joints.
[0077] Express the dynamic equations in task space:
[0078]
[0079] The task space quality matrix is defined as: The generalized inverse of the dynamically consistent task Jacobian matrix is defined as: J t (q) is the task Jacobian matrix, μ t and p t is the projection of the Coriolis force, centripetal force and gravity term in the task space, and the torque τ output by the robot m pass to calculate.
[0080] In the embodiment of the present application, when the robot does not collide, a first execution action is generated according to the parameters and constraint actions of the Cartesian space variable impedance control of the target task, and the robot is controlled to execute the target task according to the first execution action; wherein the first execution action is:
[0081]
[0082] in, is the robot's task space Jacobian matrix, is the null space projection matrix corresponding to the Jacobian matrix of the robot task space.
[0083] When the robot collides, the robot should exhibit compliant behavior to ensure contact safety. The embodiment of the present application uses a momentum observer to detect whether an unexpected collision occurs, and switches to a contact sensing controller when a collision occurs to exhibit compliant behavior.
[0084] The generalized momentum of the robot is defined as:
[0085]
[0086] Then the momentum observer can be obtained as follows:
[0087]
[0088] Among them, γ is called the residual vector, which is a stable, linear, decoupled external joint torque τ ext The first-order estimation method of K I The gain of the residual vector; F eff is the force received by the end effector, which can usually be directly measured by a force-torque sensor; Yes An estimate of .
[0089] In free space and ideally, the residual vector γ should be zero. However, in practice, it is difficult to obtain an accurate estimate of the robot model information, so it is assumed that a collision occurs when the residual is greater than a predefined threshold δ.
[0090] In order to avoid interfering with the main task, it is necessary to offset the influence of these unexpected contact forces on the main task. In the embodiment of the present application, the variable impedance control compensation amount is used to compensate the parameters of the Cartesian space variable impedance control, wherein the parameters of the Cartesian space variable impedance control after compensation are:
[0091]
[0092] To further ensure compliant behavior in the joint space, the goal at this point is to have the robot apply zero force in the joint contact direction without violating the contact constraints and interfering with the previous task. When in contact with the environment, the robot is subject to additional constraints, which causes the robot to lose a degree of freedom in the motion space. To prevent the robot from violating the constraints, this degree of freedom needs to be removed from the posture task.
[0093] In an embodiment of the present application, if the posture task is projected into the intersection of the null space of the task Jacobian matrix and the null space of the degenerate contact Jacobian matrix (i.e., the task consistent null space of the degenerate contact Jacobian matrix), the posture task will not violate the contact constraint and will not affect the main task, thereby ensuring that the robot exhibits compliant behavior while ensuring that the main task is not affected.
[0094] In practice, the projection matrix of the task-consistent null space of the degenerate contact Jacobian matrix can be calculated as:
[0095]
[0096] Where I is the identity matrix, is weighted by the inverse of the mass matrix γ T N t The generalized inverse of .
[0097] When a collision occurs, a second execution action is generated according to the compensated Cartesian space variable impedance control parameters and the compensation constraint action, wherein the second execution action is:
[0098]
[0099] in, The projection matrix to the task-consistent null space for the degenerate contact Jacobian.
[0100] The following is a detailed description of the robot reinforcement learning method for ensuring contact safety of the present application using a specific embodiment. The specific steps are as follows:
[0101] 1) Implement the compliant controller and deploy the entire compliant controller to the simulation environment and the real robot according to the formula and robot type.
[0102] The control rate of the compliant controller can be written as follows:
[0103] (a) No collision (free space controller):
[0104] (b) When a collision occurs (contact sensing controller):
[0105] In the actual execution process, the compliant controller may experience jitter behavior, which is caused by the oscillation of the residual vector and the switching between the two sub-controllers. The embodiment of the present application can alleviate the jitter behavior caused by the oscillation of the residual vector by adding a low-pass filter and setting the gain of the residual vector to be relatively small; by stopping the switching between controllers until the execution time of the contact perception controller is greater than the set threshold t cThe controller is switched only when the time comes to alleviate the jittery behavior caused by switching between the two sub-controllers.
[0106] 2) In the reinforcement learning strategy design phase, specific tasks are determined, and reinforcement learning strategies and training algorithms are designed based on the tasks. This ensures that the output space of the reinforcement learning strategy is a parameter of Cartesian space variable impedance control.
[0107] 3) In the training and application phase, the entire framework of phases 1) and 2) is trained in simulation and deployed on a real robot after training.
[0108] Furthermore, after completing the training part, the model is deployed to the robot and the output of the upper-level reinforcement learning strategy is connected to the input of the compliant controller.
[0109] It should be noted that, depending on the difficulty and requirements of the specific task, it may be necessary to, for example, fine-tune the real machine, perform offline training, etc. These can be solved by slightly modifying the training part in stage 3), but this does not affect the use of the entire reinforcement learning method at all, which further reflects the versatility of this application.
[0110] According to the robot reinforcement learning method for ensuring contact safety proposed in the embodiment of the present application, by obtaining the robot target task, the parameters of the Cartesian space variable impedance control of the target task are generated by the reinforcement learning strategy, and the constraint action is generated according to the posture constraint; the behavior is switched according to whether a collision occurs at the joint. If no collision occurs, the controller executes the Cartesian space variable impedance control target given by the reinforcement learning strategy; and through the zero space projection method, a posture task is executed under the premise of ensuring that the main task is not interfered with, so as to maintain a specific posture and reduce the problem of uncontrolled joints that may be caused when only the main task is executed. If a collision occurs, the compliant controller will estimate the size of the external force currently received, and take this force into account to compensate for the disturbance; at the same time, the zero space projection matrix at this time is modified so that the robot applies zero force in the direction of the collision without violating the collision and interfering with the main task first. By actively detecting the external collision force received by the robot and adjusting its own behavior according to the collision force, the safety of the task space and the joint space is guaranteed. At the same time, since the reinforcement learning strategy part only constrains the action space, it also has good scalability. This solves the problem that existing methods cannot avoid accidental collisions, resulting in some unexpected dangerous behaviors in unknown scenarios or when encountering interference, and do not consider issues such as contact safety in joint space.
[0111] Next, a robot reinforcement learning device for ensuring contact safety proposed in accordance with an embodiment of the present application will be described with reference to the accompanying drawings.
[0112] Figure 3It is a block diagram of a robot reinforcement learning device that ensures contact safety according to an embodiment of the present application.
[0113] like Figure 3 As shown, the robot reinforcement learning device 10 for ensuring contact safety includes: an acquisition module 100 , a generation module 200 and a processing module 300 .
[0114] Among them, the acquisition module 100 is used to obtain the robot's target task. The generation module 200 is used to generate the parameters of the Cartesian space variable impedance control of the target task using a preset reinforcement learning strategy, and generate a constraint action according to the posture constraint. The processing module 300 is used to generate a first execution action according to the parameters and constraint action of the Cartesian space variable impedance control when the robot does not collide, and control the robot to perform the target task according to the first execution action; when the robot collides, the variable impedance control compensation amount is calculated according to the collision force, and the parameters of the Cartesian space variable impedance control are compensated by the variable impedance control compensation amount, while the posture constraint is corrected to generate a compensation constraint action, and a second execution action is generated according to the compensated Cartesian space variable impedance control parameters and the compensation constraint action, and the robot is controlled to perform the target task according to the second execution action.
[0115] Optionally, in one embodiment of the present application, the generation module 200 is specifically used to perform simulation training with the robot target task as input and the parameters of the Cartesian space variable impedance control of the target task as output before using the preset reinforcement learning strategy to generate the parameters of the Cartesian space variable impedance control of the target task, so as to obtain the preset reinforcement learning strategy, wherein the parameters of the Cartesian space variable impedance control include the balance point, stiffness matrix and damping matrix corresponding to the robot target task.
[0116] Optionally, in one embodiment of the present application, the generation module 200 is specifically used to generate the parameters of the Cartesian space variable impedance control of the target task using a preset reinforcement learning strategy, wherein the parameters of the Cartesian space variable impedance control are:
[0117]
[0118] Among them, K 1 is the stiffness matrix, D 1 is the damping matrix, x d is the balance point, is the task space quality matrix, is the dynamically consistent task Jacobian matrix, x is the position of the end effector in Cartesian space, is the velocity of the end effector in Cartesian space, are the Coriolis force and centripetal force vectors, p t (q) is the gravity vector, τ extis the torque on the robot joint, q is the robot's configuration in generalized coordinates, is the velocity of the robot in generalized coordinates;
[0119] Generate constraint actions based on posture constraints, where the constraint actions are:
[0120]
[0121] Among them, K 2 is the joint stiffness matrix, D 2 is the damping matrix, q pose is the desired configuration of the robot.
[0122] Optionally, in one embodiment of the present application, the processing module 300 is specifically configured to generate a first execution action according to the parameters and constraint actions of the Cartesian space variable impedance control, wherein the first execution action is:
[0123]
[0124] in, is the robot's task space Jacobian matrix, is the null space projection matrix corresponding to the Jacobian matrix of the robot task space.
[0125] Optionally, in one embodiment of the present application, the processing module 300 is specifically used to compensate the parameters of the Cartesian space variable impedance control by using the variable impedance control compensation amount, wherein the compensated parameters of the Cartesian space variable impedance control are:
[0126]
[0127] Among them, γ is the residual vector.
[0128] Optionally, in one embodiment of the present application, the processing module 300 is specifically used to generate a second execution action according to the compensated Cartesian spatial variable impedance control parameters and the compensation constraint action, wherein the second execution action is:
[0129]
[0130] in, The projection matrix to the task-consistent null space for the degenerate contact Jacobian.
[0131] It should be noted that the above explanation of an embodiment of a robot reinforcement learning method for ensuring contact safety is also applicable to a robot reinforcement learning device for ensuring contact safety in this embodiment, which will not be repeated here.
[0132] According to the robot reinforcement learning device for ensuring contact safety proposed in the embodiment of the present application, by obtaining the robot target task, the parameters of the Cartesian space variable impedance control of the target task are generated by the reinforcement learning strategy, and the constraint action is generated according to the posture constraint; the behavior is switched according to whether a collision occurs at the joint. If no collision occurs, the controller executes the Cartesian space variable impedance control target given by the reinforcement learning strategy; and through the zero space projection method, a posture task is executed under the premise of ensuring that the main task is not interfered with, so as to maintain a specific posture and reduce the problem of uncontrolled joints that may be caused when only the main task is executed. If a collision occurs, the compliant controller will estimate the size of the external force currently received, and take this force into account to compensate for the disturbance; at the same time, the zero space projection matrix at this time is modified so that the robot applies zero force in the direction of the collision without violating the collision and interfering with the main task first. By actively detecting the external collision force received by the robot and adjusting its own behavior according to the collision force, the safety of the task space and the joint space is guaranteed. At the same time, since the reinforcement learning strategy part only constrains the action space, it also has good scalability. This solves the problem that existing methods cannot avoid accidental collisions, resulting in some unexpected dangerous behaviors in unknown scenarios or when encountering interference, and do not consider issues such as contact safety in joint space.
[0133] Figure 4 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. The electronic device may include:
[0134] Memory 401 , processor 402 , and a computer program stored in the memory 401 and executable on the processor 402 .
[0135] When the processor 402 executes the program, the robot reinforcement learning method for ensuring contact safety provided in the above embodiment is implemented.
[0136] Furthermore, the electronic device further comprises:
[0137] The communication interface 403 is used for communication between the memory 401 and the processor 402 .
[0138] The memory 401 is used to store computer programs that can be executed on the processor 402 .
[0139] The memory 401 may include a high-speed RAM memory, and may also include a non-volatile memory (non-volatile memory), such as at least one disk memory.
[0140] If the memory 401, the processor 402 and the communication interface 403 are implemented independently, the communication interface 403, the memory 401 and the processor 402 can be connected to each other through a bus and communicate with each other. The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component (PCI) bus or an Extended Industry Standard Architecture (EISA) bus. The bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 4 Only one thick line is used in the diagram, but this does not mean that there is only one bus or only one type of bus.
[0141] Optionally, in a specific implementation, if the memory 401, the processor 402 and the communication interface 403 are integrated on a chip, the memory 401, the processor 402 and the communication interface 403 can communicate with each other through an internal interface.
[0142] The processor 402 may be a central processing unit (CPU), or an application specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present application.
[0143] This embodiment also provides a computer-readable storage medium having a computer program stored thereon, wherein the program, when executed by a processor, implements the above-mentioned robot reinforcement learning method for ensuring contact safety.
[0144] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" etc. means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or N embodiments or examples in a suitable manner. In addition, those skilled in the art may combine and combine the different embodiments or examples described in this specification and the features of the different embodiments or examples, without contradiction.
[0145] In addition, the terms "first" and "second" are used for descriptive purposes only and should not be understood as indicating or implying relative importance or implicitly indicating the number of technical features indicated. Therefore, a feature defined as "first" or "second" may explicitly or implicitly include at least one of the features. In the description of this application, "N" means at least two, such as two, three, etc., unless otherwise clearly and specifically defined.
[0146] Any process or method description in a flowchart or otherwise described herein may be understood to represent a module, fragment or portion of code comprising one or more executable instructions for implementing the steps of a custom logical function or process, and the scope of the preferred embodiments of the present application includes alternative implementations in which functions may not be performed in the order shown or discussed, including performing functions in a substantially simultaneous manner or in reverse order depending on the functions involved, which should be understood by technicians in the technical field to which the embodiments of the present application belong.
[0147] It should be understood that the various parts of the present application can be implemented by hardware, software, firmware or a combination thereof. In the above-mentioned embodiment, the N steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, it can be implemented by any one of the following technologies known in the art or their combination: a discrete logic circuit having a logic gate circuit for implementing a logic function for a data signal, a dedicated integrated circuit having a suitable combination of logic gate circuits, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.
[0148] A person skilled in the art may understand that all or part of the steps in the method for implementing the above-mentioned embodiment may be completed by instructing related hardware through a program, and the program may be stored in a computer-readable storage medium, which, when executed, includes one or a combination of the steps of the method embodiment.
Claims
1. A robot reinforcement learning method to ensure contact safety, It is characterized in that The following steps are involved: Get the robot's target task; The parameters of the Cartesian space variable impedance control of the target task are generated by using a preset reinforcement learning strategy, and a constraint action is generated according to the posture constraint, wherein the parameters of the Cartesian space variable impedance control are: Among them, K 1 is the stiffness matrix, D 1 is the damping matrix, x d is the equilibrium point, Λ t (q) is the task space quality matrix, is the dynamically consistent task Jacobian matrix, x is the position of the end effector in Cartesian space, is the velocity of the end effector in Cartesian space, are the Coriolis force and centripetal force vectors, p t (q) is the gravity vector, τ ext is the torque on the robot joint, q is the robot's configuration in generalized coordinates, is the velocity of the robot in generalized coordinates; The constraint actions are: Among them, K 2 is the joint stiffness matrix, D 2 is the damping matrix, q pose is the desired configuration of the robot; If the robot does not collide, a first execution action is generated according to the parameters of the Cartesian space variable impedance control and the constraint action, and the robot is controlled to perform the target task according to the first execution action; otherwise, a variable impedance control compensation amount is calculated according to the collision force, and the parameters of the Cartesian space variable impedance control are compensated by the variable impedance control compensation amount, and the posture constraint is corrected to generate a compensation constraint action, and a second execution action is generated according to the compensated Cartesian space variable impedance control parameters and the compensation constraint action, and the robot is controlled to perform the target task according to the second execution action; wherein the first execution action is: in, is the robot's task space Jacobian matrix, is the null space projection matrix corresponding to the Jacobian matrix of the robot task space; Before using the preset reinforcement learning strategy to generate the parameters of the Cartesian space variable impedance control of the target task, it also includes: taking the robot target task as input and the parameters of the Cartesian space variable impedance control of the target task as output, performing simulation training to obtain the preset reinforcement learning strategy, wherein the parameters of the Cartesian space variable impedance control include the balance point, stiffness matrix and damping matrix corresponding to the robot target task.
2. The method according to claim 1, It is characterized in that The variable impedance control compensation amount is used to compensate the parameters of the Cartesian space variable impedance control, wherein the compensated parameters of the Cartesian space variable impedance control are: Among them, γ is the residual vector.
3. The method according to claim 2, It is characterized in that The second execution action is generated according to the compensated Cartesian spatial variable impedance control parameter and the compensation constraint action, wherein the second execution action is: in, The projection matrix to the task-consistent null space for the degenerate contact Jacobian.
4. A robot reinforcement learning device that ensures contact safety, It is characterized in that include: Acquisition module, used to obtain the robot's target task; A generation module is used to generate the parameters of the Cartesian space variable impedance control of the target task by using a preset reinforcement learning strategy, and generate a constraint action according to the posture constraint, wherein the parameters of the Cartesian space variable impedance control are: Among them, K 1 is the stiffness matrix, D 1 is the damping matrix, x d is the equilibrium point, Λ t (q) is the task space quality matrix, is the dynamically consistent task Jacobian matrix, x is the position of the end effector in Cartesian space, is the velocity of the end effector in Cartesian space, are the Coriolis force and centripetal force vectors, p t (q) is the gravity vector, τ ext is the torque on the robot joint, q is the robot's configuration in generalized coordinates, is the velocity of the robot in generalized coordinates; The constraint actions are: Among them, K 2 is the joint stiffness matrix, D 2 is the damping matrix, q pose is the desired configuration of the robot; The processing module is used to generate a first execution action according to the parameters and constraint actions of the Cartesian space variable impedance control when the robot does not collide, and control the robot to perform the target task according to the first execution action; when the robot collides, calculate the variable impedance control compensation amount according to the collision force, use the variable impedance control compensation amount to compensate the parameters of the Cartesian space variable impedance control, correct the posture constraint to generate a compensation constraint action, generate a second execution action according to the compensated Cartesian space variable impedance control parameters and the compensation constraint action, and control the robot to perform the target task according to the second execution action; wherein the first execution action is: in, is the robot's task space Jacobian matrix, is the null space projection matrix corresponding to the Jacobian matrix of the robot task space; The generation module is specifically used to perform simulation training with the robot target task as input and the parameters of the Cartesian space variable impedance control of the target task as output before using the preset reinforcement learning strategy to generate the parameters of the Cartesian space variable impedance control of the target task, so as to obtain the preset reinforcement learning strategy, wherein the parameters of the Cartesian space variable impedance control include the balance point, stiffness matrix and damping matrix corresponding to the robot target task.
5. The device according to claim 4, It is characterized in that The processing module is specifically used to compensate the parameters of the Cartesian space variable impedance control by using the variable impedance control compensation amount, wherein the compensated parameters of the Cartesian space variable impedance control are: Among them, γ is the residual vector.
6. The device according to claim 5, It is characterized in that The processing module is specifically configured to generate the second execution action according to the compensated Cartesian spatial variable impedance control parameter and the compensation constraint action, wherein the second execution action is: Among them, is the projection matrix of the task-consistent null space of the degenerate contact Jacobian matrix.
7. An electronic device, It is characterized in that include: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the robot reinforcement learning method for ensuring contact safety as described in any one of claims 1 to 3.
8. A computer-readable storage medium having a computer program stored thereon, It is characterized in that The program is executed by a processor to implement a robot reinforcement learning method for ensuring contact safety as described in any one of claims 1 to 3.
Citation Information
Patent Citations
Path planning method and device, and electronic equipment
CN111515953A
Method and device for controlling robot, robot and storage medium
CN113814978A