Robotic arm control method, device, electronic device and storage medium based on reinforcement learning

By constructing a reinforcement learning model for the robotic arm and secondary adjustment of impedance parameters, the problems of insufficient flexibility in robotic arm control and low efficiency in operating state adjustment are solved, automatic flexibility compensation for the robotic arm is achieved, and control reliability and efficiency are improved.

CN115972186BActive Publication Date: 2025-09-30GUANGZHOU AIMUYI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202310037530.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-01-10
Publication Date
2025-09-30
Estimated Expiration
2043-01-10

AI Technical Summary

Technical Problem

Existing technologies are unable to quickly and effectively compensate for the external forces acting on the end of the robotic arm, resulting in insufficiently smooth control of the robotic arm and low efficiency in adjusting the operating state.

Method used

By pre-building a reinforcement learning model for the robotic arm, the optimal impedance parameter set is obtained, and secondary adjustments are made based on the current operating status and force data to determine the target operating state of the robotic arm, and automatic compliance compensation is achieved using the reinforcement learning method.

Benefits of technology

The reliability and efficiency of the robot arm's compliance compensation control are improved, ensuring the stable operation of the robot arm under external forces.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115972186B_ABST
    Figure CN115972186B_ABST
Patent Text Reader

Abstract

The present application discloses a method, device, electronic device and storage medium for controlling a robotic arm based on reinforcement learning, and the present application belongs to the field of control technology. The method includes: obtaining an optimal impedance parameter set for controlling the robotic arm based on a pre-built robotic arm reinforcement learning model; obtaining the current operating state of the robotic arm, and determining the optimal impedance parameters of the robotic arm in the current operating state according to the current operating state; obtaining the force data of the robotic arm in the current operating state, and performing secondary adjustment on the optimal impedance parameters according to the force data to obtain the final impedance parameters; determining the mechanical parameters of the resistance of the robotic arm according to the final impedance parameters, and determining the target operating state of the robotic arm according to the mechanical parameters of the resistance. This technical solution can achieve the effect of automatically performing compliance compensation on the robotic arm, thereby improving the reliability and efficiency of the compliance compensation control of the robotic arm.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application belongs to the field of control technology, and specifically relates to a robotic arm control method, device, electronic device and storage medium based on reinforcement learning. Background Art

[0002] As the application areas of robotic arms continue to expand, using them to perform tasks more reliably is becoming increasingly important. Since the tools mounted on the end of the robotic arm are subject to external forces during the execution of tasks, causing deformation and thus reducing work efficiency, the flexible compensation of the forces acting on the tools and the control and adjustment of the operating state of the robotic arm have become a hot topic of research.

[0003] In existing technologies, the main method for controlling the compliance of a robotic arm is to predict and simulate the robotic arm's trajectory, calculate the external force data applied to the end of the robotic arm, and then compensate the robotic arm for the force based on this data and adjust the robotic arm's operating state. Alternatively, a force sensor is installed at the end of the robotic arm to sense the external force, and then the compliance compensation is performed based on the external force data to adjust the robotic arm's operating state. However, using existing technologies, it is impossible to quickly compensate for the external force applied to the end of the robotic arm, resulting in insufficient compliance control of the robotic arm and low efficiency in adjusting the robotic arm's operating state. Summary of the Invention

[0004] The purpose of the embodiments of the present application is to provide a robotic arm control method, device, electronic device and storage medium based on reinforcement learning, which can solve the problems of insufficient flexibility of robotic arm control and low efficiency of adjusting the operating state of the robotic arm. By pre-building a robotic arm reinforcement learning model, an optimal impedance parameter set for controlling the robotic arm is obtained, and the optimal impedance parameters of the robotic arm are determined according to the current operating state of the robotic arm. The optimal impedance parameters are adjusted secondary to determine the target operating state of the robotic arm. The effect of automatic compliance compensation of the robotic arm can be achieved, thereby improving the reliability and efficiency of the compliance compensation control of the robotic arm.

[0005] In a first aspect, an embodiment of the present application provides a method for controlling a robotic arm based on reinforcement learning, the method comprising:

[0006] Based on a pre-built reinforcement learning model of the robotic arm, an optimal impedance parameter set for controlling the robotic arm is obtained; wherein the impedance parameters in the optimal impedance parameter set are associated with the operating state;

[0007] Acquiring a current operating state of the robotic arm, and determining an optimal impedance parameter of the robotic arm in the current operating state according to the current operating state;

[0008] Obtaining force data of the robotic arm in a current operating state, and performing secondary adjustment on the optimal impedance parameter according to the force data to obtain a final impedance parameter;

[0009] The mechanical parameters of the resistance of the robot arm are determined according to the final impedance parameters, and the target operating state of the robot arm is determined according to the mechanical parameters of the resistance.

[0010] Furthermore, force data of the robot arm in the current operating state is obtained, and the optimal impedance parameter is adjusted secondary according to the force data to obtain the final impedance parameter, including:

[0011] Obtaining force data collected by the robotic arm;

[0012] Calculating a discount factor based on the force data and a preset force data critical value;

[0013] The optimal impedance parameter is adjusted secondary according to the discount factor to obtain a final impedance parameter.

[0014] Furthermore, a discount factor is calculated based on the force data and a preset force data critical value, including using the following formula:

[0015]

[0016] Wherein, α is the discount factor, f is the force data collected by the robot arm, and f t is the preset critical value of the force data, and f t <0;

[0017] The optimal impedance parameter is adjusted twice according to the discount factor to obtain the final impedance parameter, including calculation using the following formula:

[0018]

[0019] Where α is the discount factor; K and B are the two optimal impedance parameters of the manipulator in the current operating state. They are the final impedance parameters after secondary adjustment of K and B respectively.

[0020] Furthermore, determining the resistance mechanical parameters of the robotic arm according to the final impedance parameters, and determining the target operating state of the robotic arm according to the resistance mechanical parameters, includes:

[0021] Read the vector of the joint angle of the robotic arm in the current state, and obtain the inertia matrix, Coriolis effect, gravity torque vector and Jacobian matrix of the robotic arm in the current state;

[0022] Calculating the mechanical parameters of the resistance of the current state of the robotic arm based on the vector of the joint angle, the inertia matrix, the Coriolis effect, the vector of the gravity moment and the Jacobian matrix, and the final impedance parameter;

[0023] The mechanical parameters of the resistance, the inertia matrix and the final impedance parameters are substituted into the robot dynamics equation to determine the target operating state of the robot arm.

[0024] Furthermore, based on the vector of the joint angle, the inertia matrix, the Coriolis effect, the vector of the gravity moment and the Jacobian matrix, and the final impedance parameter, the mechanical parameters of the resistance of the current state of the robotic arm are calculated, including calculation using the following formula:

[0025]

[0026] Wherein, q is the current operating state parameter of the manipulator, M is the inertia matrix, C is the Coriolis effect, g is the vector of the gravitational torque, J is the Jacobian matrix describing the velocity kinematics, and τ is the resistance mechanical parameter of the current state of the manipulator;

[0027] Substituting the mechanical parameters of the resistance, the inertia matrix, and the final impedance parameters into the robot dynamics equation, the target operating state of the robot arm is determined, including calculation using the following formula:

[0028]

[0029]

[0030] Among them, q d is the target operating state parameter of the robotic arm.

[0031] Furthermore, obtaining the optimal impedance parameter set for controlling the robotic arm based on the pre-built robotic arm reinforcement learning model includes:

[0032] Constructing a reward function for the current state of the robotic arm based on the state parameter values ​​of the robotic arm at different times and preset weights;

[0033] Obtaining a value function for the control input of the robotic arm according to the accumulated reward function;

[0034] Using the value function approximation strategy and the error minimization principle, the relationship between the current state of the robotic arm and the optimal control input is obtained;

[0035] An optimal impedance parameter set for controlling the robotic arm is obtained according to the relationship between the current state of the robotic arm and the optimal control input.

[0036] In a second aspect, an embodiment of the present application provides a robotic arm control device based on reinforcement learning, the device comprising:

[0037] An impedance parameter set acquisition module is used to acquire an optimal impedance parameter set for controlling the robotic arm based on a pre-built robotic arm reinforcement learning model; wherein the impedance parameters in the optimal impedance parameter set are associated with the operating state;

[0038] An impedance parameter determination module is used to obtain the current operating state of the robotic arm and determine the optimal impedance parameter of the robotic arm in the current operating state according to the current operating state;

[0039] an impedance parameter adjustment module, configured to obtain force data of the robotic arm in a current operating state, and perform secondary adjustment on the optimal impedance parameter according to the force data to obtain a final impedance parameter;

[0040] The target operating state determination module is used to determine the resistance mechanical parameters of the robot arm according to the final impedance parameters, and to determine the target operating state of the robot arm according to the resistance mechanical parameters.

[0041] Furthermore, the impedance parameter adjustment module is specifically used to:

[0042] Obtaining force data collected by the robotic arm;

[0043] Calculating a discount factor based on the force data and a preset force data critical value;

[0044] The optimal impedance parameter is adjusted secondary according to the discount factor to obtain a final impedance parameter.

[0045] In a third aspect, an embodiment of the present application provides an electronic device comprising a processor, a memory, and a program or instruction stored in the memory and executable on the processor, wherein the program or instruction, when executed by the processor, implements the steps of the method described in the first aspect.

[0046] In a fourth aspect, an embodiment of the present application provides a readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by a processor, the steps of the method described in the first aspect are implemented.

[0047] In a fifth aspect, an embodiment of the present application provides a chip, which includes a processor and a communication interface, the communication interface and the processor are coupled, and the processor is used to run programs or instructions to implement the method described in the first aspect.

[0048] In an embodiment of the present application, based on a pre-constructed reinforcement learning model of the robotic arm, an optimal impedance parameter set for controlling the robotic arm is obtained; wherein, the impedance parameters in the optimal impedance parameter set are correlated with the operating state; the current operating state of the robotic arm is obtained, and the optimal impedance parameters of the robotic arm in the current operating state are determined according to the current operating state; the force data of the robotic arm in the current operating state is obtained, and the optimal impedance parameters are secondary adjusted according to the force data to obtain final impedance parameters; based on the final impedance parameters, the resistance mechanical parameters of the robotic arm are determined, and, based on the resistance mechanical parameters, the target operating state of the robotic arm is determined. The above-mentioned reinforcement learning-based robotic arm control method can solve the problems of insufficient flexibility of robotic arm control and low efficiency of adjusting the operating state of the robotic arm. By pre-building a robotic arm reinforcement learning model, the optimal impedance parameter set for controlling the robotic arm is obtained, and the optimal impedance parameters of the robotic arm are determined according to the current operating state of the robotic arm. The optimal impedance parameters are adjusted secondary to determine the target operating state of the robotic arm. This can achieve the effect of automatic compliance compensation of the robotic arm, thereby improving the reliability and efficiency of compliance compensation control of the robotic arm. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] Figure 1 1 is a flow chart of a method for controlling a robotic arm based on reinforcement learning according to the first embodiment of the present application;

[0050] Figure 2 This is a flow chart of a robotic arm control method based on reinforcement learning provided in Example 2 of the present application;

[0051] Figure 3 Schematic diagram of the structure of the robot arm control device based on reinforcement learning provided in Example 3 of the present application;

[0052] Figure 4 This is a schematic diagram of the structure of the electronic device provided in Example 4 of the present application. DETAILED DESCRIPTION

[0053] In order to make the purpose, technical solutions and advantages of the present application clearer, the specific embodiments of the present application are further described in detail below in conjunction with the accompanying drawings. It is understood that the specific embodiments described herein are only used to explain the present application and are not intended to limit the present application. It should also be noted that, for ease of description, only parts related to the present application, not all of the contents, are shown in the accompanying drawings. Before discussing the exemplary embodiments in more detail, it should be mentioned that some exemplary embodiments are described as processes or methods depicted as flow charts. Although the flow charts describe each operation (or step) as a sequential process, many of the operations therein can be implemented in parallel, concurrently or simultaneously. In addition, the order of the operations can be rearranged. The process can be terminated when its operation is completed, but can also have additional steps not included in the accompanying drawings. The process can correspond to a method, function, procedure, subroutine, subprogram, etc.

[0054] The following will be combined with the accompanying drawings in the embodiments of the present application to clearly describe the technical solutions in the embodiments of the present application. Obviously, the embodiments described are part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field are within the scope of protection of this application.

[0055] The terms "first," "second," and the like in the specification and claims of this application are used to distinguish similar objects, and are not used to describe a specific order or precedence. It should be understood that the terms used in this manner are interchangeable where appropriate, so that the embodiments of this application can be implemented in an order other than that illustrated or described herein, and that the objects distinguished by "first," "second," and the like are generally of the same type, and do not limit the number of objects; for example, the first object can be one or more. In addition, the term "and / or" in the specification and claims refers to at least one of the connected objects, and the character " / " generally indicates that the objects connected are in an "or" relationship.

[0056] The following, in conjunction with the accompanying drawings, describes in detail the reinforcement learning-based robotic arm control method, device, equipment and medium provided in the embodiments of the present application through specific embodiments and their application scenarios.

[0057] Example 1

[0058] Figure 1 This is a flow chart of the robot arm control method based on reinforcement learning provided in Example 1 of this application. Figure 1 As shown, the specific steps include:

[0059] S101, based on a pre-built reinforcement learning model of the robotic arm, obtaining an optimal impedance parameter set for controlling the robotic arm; wherein the impedance parameters in the optimal impedance parameter set are associated with the operating state;

[0060] First, this solution can be used in scenarios where compliance compensation control of a robotic arm is required. Specifically, it can be used in scenarios where the external forces acting on the robotic arm cause deformation of the executing tool, thereby affecting the task performance. Based on the reinforcement learning model, the optimal impedance parameters of the robotic arm are adjusted in real time and then adjusted again according to the external forces acting on the robotic arm. This can achieve the effect of automatic compliance compensation of the robotic arm, improving the reliability and efficiency of the compliance compensation control of the robotic arm.

[0061] Based on the above usage scenarios, it can be understood that the execution subject of this application can be software installed in the robotic arm control system, which has data acquisition, calculation and model building functions, etc., and no further restrictions are made here.

[0062] In this proposal, reinforcement learning, a paradigm and methodology within machine learning, can be used to describe and solve the problem of learning strategies to maximize rewards or achieve specific goals during a robotic arm's interaction with its environment. The biggest difference between reinforcement learning and supervised learning is that it lacks the pre-prepared training data output of supervised learning. Reinforcement learning only has a reward value, but unlike supervised learning, this reward value is not given in advance but rather delayed. For example, the robotic arm receives a reward only after interacting with the outside world. Furthermore, each step of reinforcement learning is closely tied to a temporal sequence. The process of constructing a robotic arm reinforcement learning model involves treating the robotic arm control system as the algorithm's individual execution entity. Based on a preliminary assessment of the external environment, the robotic arm selects an appropriate action to interact with the environment. After the robotic arm executes the action, the external environment changes state, and feedback is provided to the robotic arm control system regarding the action. This feedback can be positive feedback, representing a reward for the action, or negative feedback, representing a penalty for the action. The robotic arm control system then selects the next action for the robotic arm based on this feedback, and this cycle repeats.

[0063] In this solution, the impedance parameter can be used to control the operating state of the robotic arm or to generate a force on the robotic arm, thereby offsetting changes in the operating state of the robotic arm caused by the influence of the external environment. The optimal impedance parameter can be an impedance parameter that can flexibly compensate for the operating state of the robotic arm and reduce the error to a minimum. Since the force exerted by the external environment on the robotic arm may be different at different times, the optimal impedance parameter for controlling the operating state of the robotic arm at different times is also different. Specifically, all impedance parameters of the robotic arm in various operating states can be obtained based on a pre-constructed reinforcement learning model of the robotic arm, thereby forming an optimal impedance parameter set for controlling the robotic arm.

[0064] S102, obtaining a current operating state of the robotic arm, and determining an optimal impedance parameter of the robotic arm in the current operating state according to the current operating state;

[0065] In this solution, the current operating state of the robotic arm can be represented by operating parameters. For example, the operating parameters may include parameters such as the operating speed, displacement, and target position of the robotic arm. Specifically, the current operating state of the robotic arm may be the operating state reached by the robotic arm after an external environmental force acts on the robotic arm. Based on the current operating state of the robotic arm and the obtained optimal impedance parameter set, the impedance parameters of the robotic arm are adjusted in real time to determine the optimal impedance parameters for the robotic arm in the current operating state.

[0066] S103, obtaining force data of the robotic arm in the current operating state, and performing secondary adjustment on the optimal impedance parameter according to the force data to obtain a final impedance parameter;

[0067] In this solution, the force data of the robotic arm mainly refers to the force data of the external environment on the execution tool during the execution of the task. Since the execution tool is connected to the end of the robotic arm, the force data of the robotic arm in the current operating state can be read by installing a sensor on the end of the robotic arm. The optimal impedance parameter is adjusted twice according to the force data to obtain the final impedance parameter. Specifically, the secondary adjustment of the optimal impedance parameter can be due to the different effects of different force data on the execution tool installed at the end of the robotic arm. Therefore, if the force data is small, the impedance parameter of the robotic arm can be increased to avoid changes in the running direction of the robotic arm; if the force data is large, in order to avoid deformation of the execution tool and thus affect the execution effect, the impedance parameter can be reduced to make the robotic arm follow compliantly, and the impedance parameter after the secondary adjustment is used as the final impedance parameter.

[0068] Based on the above embodiment, optionally, obtaining force data of the robotic arm in the current operating state, and performing secondary adjustment on the optimal impedance parameter according to the force data to obtain a final impedance parameter, includes:

[0069] Obtaining force data collected by the robotic arm;

[0070] Calculating a discount factor based on the force data and a preset force data critical value;

[0071] The optimal impedance parameter is adjusted secondary according to the discount factor to obtain a final impedance parameter.

[0072] In this solution, the force data collected by the robotic arm can be obtained by installing a sensor on the end of the robotic arm. The force data can be the interaction force data between the external environment and the execution tool at the end of the robotic arm when the execution tool at the end of the robotic arm performs a task. The discount factor is calculated based on the force data and the preset force data critical value. Specifically, the force data critical value can be a parameter value that is pre-set based on the external environment in which the current robotic arm is operating and the properties of the execution tool, and is used to calculate the discount factor. The optimal impedance parameter is secondary adjusted according to the discount factor to obtain the final impedance parameter, and the robotic arm control system controls the operating state of the robotic arm according to the final impedance parameter.

[0073] In this solution, by obtaining the force data collected by the robotic arm and the preset force data critical value, calculating the discount factor, and performing secondary adjustment on the optimal impedance parameter according to the discount factor, the final impedance parameter is obtained, which can improve the reliability and efficiency of controlling the robotic arm using the impedance parameter.

[0074] Based on the above embodiment, optionally, a discount factor is calculated according to the force data and a preset force data critical value, including using the following formula:

[0075]

[0076] Wherein, α is the discount factor, f is the force data collected by the robot arm, and f t is the preset critical value of the force data, and f t >0;

[0077] The optimal impedance parameter is adjusted twice according to the discount factor to obtain the final impedance parameter, including calculation using the following formula:

[0078]

[0079] Where α is the discount factor; K and B are the two optimal impedance parameters of the manipulator in the current operating state. They are the final impedance parameters after secondary adjustment of K and B respectively.

[0080] In this solution, the discount factor can be calculated using the formula: Calculation results show that, where α is the discount factor, f is the force data collected by the robot arm, and f t is the preset critical value of the force data, and f t >0; Based on the discount factor of the current force state of the manipulator and the optimal impedance parameter of the current operating state, use the formula: The optimal impedance parameter is adjusted twice to obtain the final impedance parameter, where α is the discount factor under the current force state of the manipulator, and K and B are the two optimal impedance parameters of the manipulator under the current operating state. They are the final impedance parameters after secondary adjustment of K and B respectively.

[0081] In this solution, the discount factor is calculated using a formula based on the relationship between the force data collected by the robotic arm and the preset force data critical value, and the optimal impedance parameter is adjusted secondary according to the discount factor to obtain the final impedance parameter, which can improve the reliability and efficiency of controlling the robotic arm using the impedance parameter.

[0082] S104 , determining the resistance mechanical parameters of the robotic arm according to the final impedance parameters, and determining the target operating state of the robotic arm according to the resistance mechanical parameters.

[0083] In this solution, the mechanical parameters of the resistance of the robotic arm may be mechanical parameters resulting from changes in the operating state of the robotic arm caused by the force exerted by the external environment on the executing tool at the end of the robotic arm, such as parameters such as moment, torque, and direction of force. Based on the final impedance parameters, the mechanical parameters of the resistance of the robotic arm are determined, and then based on the mechanical parameters of the resistance, the error value between the target operating state and the current operating state of the robotic arm can be calculated, and the target operating state of the robotic arm can be calculated based on the error value and the current operating state. The target operating state of the robotic arm may be the operating state after the robotic arm system performs compliance compensation on the robotic arm. Specifically, it may be the final state reached by controlling the robotic arm to eliminate operating errors. The target operating state may be the target speed, target displacement, target position, and target angle of the robotic arm, etc., which are not limited in detail here.

[0084] Based on the above embodiment, optionally, determining the resistance mechanical parameters of the robotic arm according to the final impedance parameters, and determining the target operating state of the robotic arm according to the resistance mechanical parameters, includes:

[0085] Read the vector of the joint angle of the robotic arm in the current state, and obtain the inertia matrix, Coriolis effect, gravity torque vector and Jacobian matrix of the robotic arm in the current state;

[0086] Calculating the mechanical parameters of the resistance of the current state of the robotic arm based on the vector of the joint angle, the inertia matrix, the Coriolis effect, the vector of the gravity moment and the Jacobian matrix, and the final impedance parameter;

[0087] The mechanical parameters of the resistance, the inertia matrix and the final impedance parameters are substituted into the robot dynamics equation to determine the target operating state of the robot arm.

[0088] In this solution, the joint angle vector of the robotic arm in its current state can be read by installing an angle sensor on the robotic arm to read the magnitude of the joint angle vector and determining the direction of the joint angle vector based on the direction of movement of the robotic arm. Alternatively, the joint angle vector can be obtained based on the position of the end of the robotic arm in its previous operating state and its current position, etc., without further limitation here. By reading the vector table of the joint angles in the robotic arm's current state and obtaining the inertia matrix, Coriolis effect, vector of gravitational torque, Jacobian matrix, and final impedance parameters of the robotic arm in its current state, the mechanical parameters of the resistance force applied to the robotic arm in its current state are calculated. The mechanical parameters of the resistance force applied, the inertia matrix, and the final impedance parameters are substituted into the robot's dynamics equations to determine the target operating state of the robotic arm.

[0089] In this solution, by reading the vector of the joint angle in the current state of the robotic arm, and obtaining the inertia matrix, Coriolis effect, vector of gravity torque, Jacobian matrix and the final impedance parameters in the current state of the robotic arm, the mechanical parameters of the resistance force in the current state of the robotic arm are calculated, and the mechanical parameters of the resistance force, inertia matrix and final impedance parameters are substituted into the robot dynamics equation to determine the target operating state of the robotic arm, thereby improving the timeliness and reliability of the control of the robotic arm.

[0090] Based on the above embodiment, optionally, the mechanical parameters of the resistance force applied to the current state of the robotic arm are calculated based on the vector of the joint angle, the inertia matrix, the Coriolis effect, the vector of the gravity moment and the Jacobian matrix, and the final impedance parameter, including calculation using the following formula:

[0091]

[0092] Wherein, q is the current operating state parameter of the manipulator, M is the inertia matrix, C is the Coriolis effect, g is the vector of the gravitational torque, J is the Jacobian matrix describing the velocity kinematics, and τ is the resistance mechanical parameter of the current state of the manipulator;

[0093] Substituting the mechanical parameters of the resistance, the inertia matrix, and the final impedance parameters into the robot dynamics equation, the target operating state of the robot arm is determined, including calculation using the following formula:

[0094]

[0095]

[0096] Among them, q d is the target operating state parameter of the robotic arm.

[0097] In this solution, the mechanical parameters of the resistance of the current state of the manipulator can be calculated using the formula: Calculate, where q is the current operating state parameter of the manipulator, M is the inertia matrix, C is the Coriolis effect, g is the vector of the gravitational torque, J is the Jacobian matrix describing the velocity kinematics, and τ is the mechanical parameter of the resistance of the manipulator in its current state. Substitute the mechanical parameter, inertia matrix, and final impedance parameter into the robot dynamics equation, and obtain the formula: Calculate the target operating state of the robotic arm, where q d is the target operating state parameter of the robotic arm.

[0098] In this solution, by reading the vector of the joint angle in the current state of the robotic arm, and obtaining the inertia matrix, Coriolis effect, vector of gravity torque, Jacobian matrix and the final impedance parameters in the current state of the robotic arm, the mechanical parameters of the resistance force in the current state of the robotic arm are calculated, and the mechanical parameters of the resistance force, inertia matrix and final impedance parameters are substituted into the robot dynamics equation to determine the target operating state of the robotic arm, thereby improving the timeliness and reliability of the control of the robotic arm.

[0099] The technical solution provided in the embodiment of the present application obtains an optimal impedance parameter set for controlling the robotic arm based on a pre-built robotic arm reinforcement learning model; wherein the impedance parameters in the optimal impedance parameter set are correlated with the operating state; the current operating state of the robotic arm is obtained, and the optimal impedance parameters of the robotic arm in the current operating state are determined according to the current operating state; the force data of the robotic arm in the current operating state is obtained, and the optimal impedance parameters are secondary adjusted according to the force data to obtain the final impedance parameters; the resistance mechanical parameters of the robotic arm are determined according to the final impedance parameters, and, according to the resistance mechanical parameters, the target operating state of the robotic arm is determined. The above-mentioned reinforcement learning-based robotic arm control method can solve the problems of insufficient flexibility of robotic arm control and low efficiency of adjusting the operating state of the robotic arm. By pre-building a robotic arm reinforcement learning model, the optimal impedance parameter set for controlling the robotic arm is obtained, and the optimal impedance parameters of the robotic arm are determined according to the current operating state of the robotic arm. The optimal impedance parameters are adjusted secondary to determine the target operating state of the robotic arm. This can achieve the effect of automatic compliance compensation of the robotic arm, thereby improving the reliability and efficiency of compliance compensation control of the robotic arm.

[0100] Example 2

[0101] Figure 2 This is a flow chart of the robot arm control method based on reinforcement learning provided in the second embodiment of this application. Figure 2 As shown, the specific steps include:

[0102] S201, based on a pre-built reinforcement learning model of the robotic arm, obtaining an optimal impedance parameter set for controlling the robotic arm; wherein the impedance parameters in the optimal impedance parameter set are associated with the operating state;

[0103] Based on the above embodiment, optionally, obtaining an optimal impedance parameter set for controlling the robotic arm based on a pre-built robotic arm reinforcement learning model includes:

[0104] Constructing a reward function for the current state of the robotic arm based on the state parameter values ​​of the robotic arm at different times and preset weights;

[0105] Obtaining a value function for the control input of the robotic arm according to the accumulated reward function;

[0106] Using the value function approximation strategy and the error minimization principle, the relationship between the current state of the robotic arm and the optimal control input is obtained;

[0107] An optimal impedance parameter set for controlling the robotic arm is obtained according to the relationship between the current state of the robotic arm and the optimal control input.

[0108] In this solution, the state parameter value can be the operating state value of the robot arm at a certain moment during its operation, and can include parameter values ​​such as operating speed, acceleration, displacement, and operating error. Since the operating state parameter values ​​of the robot arm at different moments may be different, and different state parameters have different effects on the reinforcement learning model of the robot arm, different weights can be set for different state parameters of the robot arm, and based on the state parameter values ​​of the robot arm at different moments and the preset weights, a reward function for the current state of the robot arm can be constructed. Specifically, the formula can be used: Construct a reward function for the current state of the robotic arm, where is the actual displacement of the robot arm, x d is the expected displacement of the robot arm, i.e., the target displacement, is the robot arm operation error, Q1, Q2, R r Represents different weights, is the running speed of the robotic arm, u r The robot arm reinforcement learning model and the robot arm operating state parameters input during reinforcement learning training of the robot arm system are represented. A value function for the robot arm control input is obtained by accumulating the reward function of the robot arm at the current moment and after that moment. A value function approximation strategy and the error minimization principle are used to obtain the relationship between the current state of the robot arm and the optimal control input. Based on this relationship between the current state of the robot arm and the optimal control input, an optimal impedance parameter set for controlling the robot arm is obtained.

[0109] In this solution, by using a pre-built robotic arm reinforcement learning model and training the robotic arm control system based on the model, the optimal impedance parameter set for controlling the robotic arm is obtained, which can achieve the purpose of automatic control of the robotic arm and improve the efficiency of controlling the robotic arm.

[0110] S202, obtaining a current operating state of the robotic arm, and determining an optimal impedance parameter of the robotic arm in the current operating state according to the current operating state;

[0111] S203, obtaining force data of the robotic arm in the current operating state, and performing secondary adjustment on the optimal impedance parameter according to the force data to obtain a final impedance parameter;

[0112] S204 : Determine the resistance mechanical parameters of the robot arm according to the final impedance parameters, and determine the target operating state of the robot arm according to the resistance mechanical parameters.

[0113] The technical solution provided in the embodiment of the present application constructs a reward function for the current state of the robotic arm based on the state parameter values ​​of the robotic arm at different times and preset weights; obtains a value function for the control input of the robotic arm by accumulating the reward function; utilizes the value function approximation strategy and the error minimization principle to obtain the relationship between the current state of the robotic arm and the optimal control input; obtains the optimal impedance parameter set for controlling the robotic arm based on the relationship between the current state of the robotic arm and the optimal control input, thereby achieving the purpose of automatic control of the robotic arm and improving the efficiency of controlling the robotic arm.

[0114] Example 3

[0115] Figure 3 This is a schematic diagram of the structure of the robot arm control device based on reinforcement learning provided in Example 3 of this application. Figure 3 As shown, specifically including the following:

[0116] An impedance parameter set acquisition module is used to acquire an optimal impedance parameter set for controlling the robotic arm based on a pre-built robotic arm reinforcement learning model; wherein the impedance parameters in the optimal impedance parameter set are associated with the operating state;

[0117] An impedance parameter determination module is used to obtain the current operating state of the robotic arm and determine the optimal impedance parameter of the robotic arm in the current operating state according to the current operating state;

[0118] an impedance parameter adjustment module, configured to obtain force data of the robotic arm in a current operating state, and perform secondary adjustment on the optimal impedance parameter according to the force data to obtain a final impedance parameter;

[0119] The target operating state determination module is used to determine the resistance mechanical parameters of the robot arm according to the final impedance parameters, and to determine the target operating state of the robot arm according to the resistance mechanical parameters.

[0120] Furthermore, the impedance parameter adjustment module includes:

[0121] A force data acquisition unit, used to acquire the force data collected by the robotic arm;

[0122] a discount factor calculation unit, configured to calculate a discount factor based on the force data and a preset force data critical value;

[0123] An impedance parameter adjustment unit is used to perform secondary adjustment on the optimal impedance parameter according to the discount factor to obtain a final impedance parameter.

[0124] Furthermore, the discount factor calculation unit includes calculating using the following formula:

[0125]

[0126] Wherein, α is the discount factor, f is the force data collected by the robot arm, and f t is the preset critical value of the force data, and f t <0;

[0127] The impedance parameter adjustment unit includes calculating using the following formula:

[0128]

[0129] Where α is the discount factor; K and B are the two optimal impedance parameters of the manipulator in the current operating state. They are the final impedance parameters after secondary adjustment of K and B respectively.

[0130] Furthermore, the target operating state determination module is specifically configured to:

[0131] Read the vector of the joint angle of the robotic arm in the current state, and obtain the inertia matrix, Coriolis effect, gravity torque vector and Jacobian matrix of the robotic arm in the current state;

[0132] Calculating the mechanical parameters of the resistance of the current state of the robotic arm based on the vector of the joint angle, the inertia matrix, the Coriolis effect, the vector of the gravity moment and the Jacobian matrix, and the final impedance parameter;

[0133] The mechanical parameters of the resistance, the inertia matrix and the final impedance parameters are substituted into the robot dynamics equation to determine the target operating state of the robot arm.

[0134] Furthermore, the target operating state determination module is specifically configured to calculate using the following formula:

[0135]

[0136] Wherein, q is the current operating state parameter of the manipulator, M is the inertia matrix, C is the Coriolis effect, g is the vector of the gravitational torque, J is the Jacobian matrix describing the velocity kinematics, and τ is the resistance mechanical parameter of the current state of the manipulator;

[0137] The target operating state determination module is specifically configured to calculate using the following formula:

[0138]

[0139]

[0140] Among them, q d is the target operating state parameter of the robotic arm.

[0141] Furthermore, the impedance parameter set acquisition module is specifically used to:

[0142] Constructing a reward function for the current state of the robotic arm based on the state parameter values ​​of the robotic arm at different times and preset weights;

[0143] Obtaining a value function for the control input of the robotic arm according to the accumulated reward function;

[0144] Using the value function approximation strategy and the error minimization principle, the relationship between the current state of the robotic arm and the optimal control input is obtained;

[0145] An optimal impedance parameter set for controlling the robotic arm is obtained according to the relationship between the current state of the robotic arm and the optimal control input.

[0146] The technical solution provided in this embodiment includes an impedance parameter set acquisition module, which is used to obtain the optimal impedance parameter set for controlling the robotic arm based on a pre-built robotic arm reinforcement learning model; wherein the impedance parameters in the optimal impedance parameter set are correlated with the operating state; an impedance parameter determination module, which is used to obtain the current operating state of the robotic arm, and determine the optimal impedance parameters of the robotic arm in the current operating state according to the current operating state; an impedance parameter adjustment module, which is used to obtain the force data of the robotic arm in the current operating state, and perform secondary adjustment on the optimal impedance parameters according to the force data to obtain the final impedance parameters; and a target operating state determination module, which is used to determine the resistance mechanical parameters of the robotic arm according to the final impedance parameters, and determine the target operating state of the robotic arm according to the resistance mechanical parameters. The above-mentioned reinforcement learning-based robotic arm control device can solve the problems of insufficient flexibility of robotic arm control and low efficiency of adjusting the operating state of the robotic arm. By pre-building a robotic arm reinforcement learning model, the optimal impedance parameter set for controlling the robotic arm is obtained, and the optimal impedance parameters of the robotic arm are determined according to the current operating state of the robotic arm. The optimal impedance parameters are adjusted secondary to determine the target operating state of the robotic arm. This can achieve the effect of automatic flexibility compensation of the robotic arm, thereby improving the reliability and efficiency of the flexibility compensation control of the robotic arm.

[0147] The reinforcement learning-based robotic arm control device in the embodiment of the present application can be a device, or a component, integrated circuit, or chip in a terminal. The device can be a mobile electronic device or a non-mobile electronic device. For example, the mobile electronic device can be a mobile phone, a tablet computer, a laptop computer, a PDA, an in-vehicle electronic device, a wearable device, an ultra-mobile personal computer (UMPC), a netbook, or a personal digital assistant (PDA), etc. The non-mobile electronic device can be a server, a network attached storage (NAS), a personal computer (PC), a television (TV), an ATM, or an kiosks, etc., which are not specifically limited in the embodiment of the present application.

[0148] The reinforcement learning-based robotic arm control device in the embodiment of the present application may be a device having an operating system. The operating system may be an Android operating system, an iOS operating system, or other possible operating systems, which are not specifically limited in the embodiment of the present application.

[0149] The robot arm control device based on reinforcement learning provided in the embodiment of the present application can achieve Figures 1 to 2 To avoid repetition, the various processes implemented in the method embodiment are not described here.

[0150] Example 4

[0151] like Figure 4 As shown, an embodiment of the present application also provides an electronic device 400, including a processor 401, a memory 402, and a program or instruction stored in the memory 402 and executable on the processor 401. When the program or instruction is executed by the processor 401, the various processes of the above-mentioned embodiment of the reinforcement learning-based robotic arm control method are implemented, and the same technical effect can be achieved. To avoid repetition, they will not be described here.

[0152] It should be noted that the electronic devices in the embodiments of the present application include the mobile electronic devices and non-mobile electronic devices mentioned above.

[0153] Example 5

[0154] An embodiment of the present application also provides a readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by a processor, the various processes of the above-mentioned reinforcement learning-based robotic arm control method embodiment are implemented, and the same technical effect can be achieved. To avoid repetition, it will not be repeated here.

[0155] The processor is the processor in the electronic device described in the above embodiment. The readable storage medium includes a computer-readable storage medium, such as a computer read-only memory (ROM), random access memory (RAM), a magnetic disk, or an optical disk.

[0156] Example 6

[0157] An embodiment of the present application further provides a chip, which includes a processor and a communication interface, wherein the communication interface is coupled to the processor, and the processor is used to run programs or instructions to implement the various processes of the above-mentioned reinforcement learning-based robotic arm control method embodiment, and can achieve the same technical effect. To avoid repetition, it will not be repeated here.

[0158] It should be understood that the chip mentioned in the embodiments of the present application can also be called a system-level chip, a system chip, a chip system or a system-on-chip chip, etc.

[0159] It should be noted that, in this article, the terms "comprise", "include" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, an element defined by the statement "comprises a ..." does not exclude the presence of other identical elements in the process, method, article or device comprising the element. In addition, it should be noted that the scope of the methods and devices in the embodiments of the present application is not limited to performing functions in the order shown or discussed, and may also include performing functions in a substantially simultaneous manner or in the opposite order according to the functions involved. For example, the described method may be performed in an order different from that described, and various steps may also be added, omitted, or combined. In addition, the features described with reference to certain examples may be combined in other examples.

[0160] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus the necessary general hardware platform, and of course can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art can be embodied in the form of a computer software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), including a number of instructions for enabling a terminal (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods described in each embodiment of the present application.

[0161] The embodiments of the present application are described above in conjunction with the accompanying drawings, but the present application is not limited to the above-mentioned specific implementation methods. The above-mentioned specific implementation methods are merely illustrative and not restrictive. Under the guidance of this application, ordinary technicians in this field can also make many forms without departing from the purpose of this application and the scope of protection of the claims, all of which are within the protection of this application.

[0162] The above are only preferred embodiments of the present application and the technical principles employed. The present application is not limited to the specific embodiments described herein, and various obvious changes, readjustments, and substitutions that are possible for those skilled in the art will not depart from the scope of protection of the present application. Therefore, although the present application has been described in more detail through the above embodiments, the present application is not limited to the above embodiments and may include more other equivalent embodiments without departing from the concept of the present application. The scope of the present application is determined by the scope of the claims.

Claims

1. A robotic arm control method based on reinforcement learning, characterized in that: The method comprises: Based on a pre-built reinforcement learning model of the robotic arm, an optimal impedance parameter set for controlling the robotic arm is obtained; wherein the impedance parameters in the optimal impedance parameter set are associated with the operating state; Acquiring a current operating state of the robotic arm, and determining an optimal impedance parameter of the robotic arm in the current operating state according to the current operating state; Obtaining force data of the robotic arm in a current operating state, and performing secondary adjustment on the optimal impedance parameter based on the force data to obtain a final impedance parameter, wherein the method includes: obtaining force data collected by the robotic arm, calculating a discount factor based on the force data and a preset force data critical value, and performing secondary adjustment on the optimal impedance parameter based on the discount factor to obtain a final impedance parameter, wherein the secondary adjustment of the optimal impedance parameter includes: increasing the impedance parameter of the robotic arm to prevent the running direction of the robotic arm from changing, and reducing the impedance parameter to enable the robotic arm to follow compliantly; According to the final impedance parameters, the mechanical parameters of the resistance of the robotic arm are determined, and according to the mechanical parameters of the resistance, the target operating state of the robotic arm is determined, and the target operating state includes the target speed, target displacement, target position and target angle of the robotic arm.

2. The method according to claim 1, characterized in that Calculating a discount factor based on the force data and a preset force data critical value includes calculating using the following formula: ; in, is the discount factor, is the force data collected by the robotic arm, is the preset stress data critical value, and ; The optimal impedance parameter is adjusted twice according to the discount factor to obtain the final impedance parameter, including calculation using the following formula: ; in, is the discount factor; 、 are the two optimal impedance parameters of the robotic arm in the current operating state, 、 Respectively for 、 Final impedance parameters after secondary adjustment.

3. The method according to claim 1, characterized in that Determining the resistance mechanical parameters of the robotic arm according to the final impedance parameters, and determining the target operating state of the robotic arm according to the resistance mechanical parameters, includes: Read the vector of the joint angle of the robotic arm in the current state, and obtain the inertia matrix, Coriolis effect, gravitational torque vector and Jacobian matrix of the robotic arm in the current state; Calculating the mechanical parameters of the resistance of the current state of the robotic arm based on the vector of the joint angle, the inertia matrix, the Coriolis effect, the vector of the gravity moment and the Jacobian matrix, and the final impedance parameter; The mechanical parameters of the resistance, the inertia matrix and the final impedance parameters are substituted into the robot dynamics equation to determine the target operating state of the robot arm.

4. The method according to claim 3, characterized in that Calculate the mechanical parameters of the resistance of the current state of the manipulator based on the vector of the joint angle, the inertia matrix, the Coriolis effect, the vector of the gravity moment and the Jacobian matrix, and the final impedance parameter, including using the following formula: ; in, is the current operating state parameter of the robotic arm, is the inertia matrix, is the Coriolis effect, is the vector of gravitational torque, is the Jacobian matrix describing the velocity kinematics, is the mechanical parameter of the resistance of the manipulator in its current state; Substituting the mechanical parameters of the resistance, the inertia matrix, and the final impedance parameters into the robot dynamics equation, the target operating state of the robot arm is determined, including calculation using the following formula: ; ; in, is the target operating state parameter of the robotic arm.

5. The method according to claim 1, wherein The obtaining of an optimal impedance parameter set for controlling the robotic arm based on a pre-built robotic arm reinforcement learning model includes: Constructing a reward function for the current state of the robotic arm based on the state parameter values ​​of the robotic arm at different times and preset weights; Obtaining a value function for the control input of the robotic arm according to the accumulated reward function; Using the value function approximation strategy and the error minimization principle, the relationship between the current state of the robotic arm and the optimal control input is obtained; An optimal impedance parameter set for controlling the robotic arm is obtained based on a relationship between the current state of the robotic arm and the optimal control input.

6. A robotic arm control device based on reinforcement learning, characterized in that: The device comprises: An impedance parameter set acquisition module is used to acquire an optimal impedance parameter set for controlling the robotic arm based on a pre-built robotic arm reinforcement learning model; wherein the impedance parameters in the optimal impedance parameter set are associated with the operating state; An impedance parameter determination module is used to obtain the current operating state of the robotic arm and determine the optimal impedance parameter of the robotic arm in the current operating state according to the current operating state; an impedance parameter adjustment module, configured to obtain force data of the robotic arm in a current operating state, and perform secondary adjustment on the optimal impedance parameter based on the force data to obtain a final impedance parameter, wherein the module comprises: obtaining force data collected by the robotic arm, calculating a discount factor based on the force data and a preset force data critical value, and performing secondary adjustment on the optimal impedance parameter based on the discount factor to obtain a final impedance parameter, wherein the secondary adjustment of the optimal impedance parameter comprises: increasing the impedance parameter of the robotic arm to prevent the running direction of the robotic arm from changing, and reducing the impedance parameter to enable the robotic arm to follow compliantly; A target operating state determination module is used to determine the resistance mechanical parameters of the robotic arm based on the final impedance parameters, and to determine the target operating state of the robotic arm based on the resistance mechanical parameters, wherein the target operating state includes the target speed, target displacement, target position and target angle of the robotic arm.

7. The device according to claim 6, characterized in that The impedance parameter adjustment module is specifically used to: Obtaining force data collected by the robotic arm; Calculating a discount factor based on the force data and a preset force data critical value; The optimal impedance parameter is adjusted secondary according to the discount factor to obtain a final impedance parameter.

8. An electronic device, characterized in that: The invention comprises a processor, a memory, and a program or instruction stored in the memory and executable on the processor, wherein the program or instruction, when executed by the processor, implements the steps of the reinforcement learning-based robotic arm control method as described in any one of claims 1 to 5.

9. A readable storage medium, characterized in that The readable storage medium stores a program or instruction, and when the program or instruction is executed by the processor, the steps of the reinforcement learning-based robotic arm control method according to any one of claims 1 to 5 are implemented.