Robot joint error compensation method and device, computer device and storage medium
By constructing a link state prediction network and a reinforcement learning system, and combining feedforward and feedback control, the problem of inaccurate robot joint error compensation was solved, and robot joint error compensation and vibration suppression were achieved with high precision and high speed.
Patent Information
- Application Number
- CN202211367676.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-03
- Publication Date
- 2026-01-02
- Estimated Expiration
- 2042-11-03
AI Technical Summary
In existing technologies, robot joint error compensation is inaccurate, resulting in poor vibration suppression, especially in high-precision, high-speed motion, where it fails to meet technical requirements.
By constructing a link state prediction network and a reinforcement learning system, and combining feedforward control and feedback control, error compensation is performed using sensor feedback values and joint planning values, and compensation parameters are trained to reduce joint errors.
It achieves error compensation and vibration suppression for robot joints under high-precision and high-speed motion, thereby improving the motion accuracy and stability of robot joints.
Smart Images

Figure CN115816443B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of industrial machinery, in particular to a robot joint error compensation method and device, a computer device and a storage medium. BACKGROUND
[0002] With the development of science and technology, the industrial robot arm is the most widely used robot in the field of industrial machinery. The robot arm is composed of joints, connecting rods and reducers, and is rigidly connected. Due to the higher requirements for the high-precision and high-speed motion characteristics of the robot arm, when the robot arm works and contacts with the outside world, the reducer itself will exhibit partial flexibility characteristics, i.e. insufficient stiffness, resulting in vibration and error, which cannot meet the technical requirements in some cases.
[0003] In the prior art, error compensation and vibration suppression can be divided into two categories: feedback control and feedforward control. The feedback control usually needs an additional sensor to measure the output of the system, and the actual position error of the connecting rod is generally taken as the input, and the motor torque for compensation is taken as the output. The feedforward control needs an accurate dynamic model. However, due to the complex relationship between the joints and the joints, and the inability to determine the specific linear relationship between the joint and other joints connected in series, the error compensation is not accurate, and thus the vibration suppression cannot be achieved. SUMMARY
[0004] The technical problem to be solved by the present application is to predict and compensate for the robot joint error without adding an additional sensor during work. However, due to the complex relationship between the joints and the joints, and the inability to determine the specific linear relationship between the joint and other joints connected in series, the error compensation is not accurate, and thus the vibration suppression cannot be achieved.
[0005] To solve the above problems, the present application provides a robot joint error compensation method, wherein the joint is driven by a motor, the joint is connected to a connecting rod through a reducer, and the method comprises:
[0006] performing feedforward control on the target joint according to the joint planning value and the dynamic model to obtain a first joint parameter;
[0007] performing negative feedback adjustment on the first joint parameter according to the sensor feedback value to obtain a second joint parameter;
[0008] obtaining a predicted joint parameter according to the joint planning value, the sensor feedback value and a connecting rod state prediction network;
[0009] obtaining a compensation parameter according to the sensor feedback value, the predicted joint parameter, the joint planning value and a reinforcement learning system;
[0010] According to the compensation parameter, the second joint parameter is positively fed back to realize error compensation.
[0011] Optionally, before the predicted joint parameter is obtained according to the joint planning value, the sensor feedback value and the link state prediction network, the method further comprises:
[0012] The link state prediction network is constructed, and the construction of the link state prediction network comprises:
[0013] An initial link state prediction network is obtained.
[0014] A training data set is obtained, and the training data set comprises a planned link joint parameter, a motor feedback parameter and an actual link parameter.
[0015] The training data set is input into the initial link state prediction network for training, wherein the input of the initial link state prediction network is the planned link joint parameter and the motor feedback parameter, and the output of the initial link state prediction network is a predicted link joint parameter.
[0016] The initial link state prediction network after training is taken as the link state prediction network.
[0017] Optionally, the training data set further comprises an actual link parameter, and the construction of the link state prediction network further comprises:
[0018] A loss function is obtained through the predicted link joint parameter and the actual link parameter.
[0019] Whether the initial link state prediction network ends training is determined according to the loss function.
[0020] Optionally, before the compensation parameter is obtained according to the sensor feedback value, the predicted joint parameter, the joint planning value and a reinforcement learning system, the method further comprises constructing the reinforcement learning system,
[0021] The construction of the reinforcement learning system comprises:
[0022] A reinforcement learning neural network is obtained.
[0023] The reinforcement learning neural network is trained by using a reinforcement learning method, and the trained reinforcement learning neural network is taken as a reinforcement learning system.
[0024] When the training reinforcement learning neural network, the planning link joint parameter, the actual link parameter and the motor feedback parameter are taken as states, the compensation parameter is taken as an action, the reinforcement learning neural network is taken as a policy, a reference parameter is taken as a reward, the policy is optimized by the action and the states to obtain the reward, and whether the training is ended is judged according to the reward.
[0025] The reference parameter is related to the associated joint of the target joint.
[0026] Optionally, the reference parameter is:
[0027]
[0028] wherein j is the total number of joints, the total number of joints is the sum of the number of target joints and the number of associated joints of the target joints, x[i]=[p[i],v[i]] T , x′[i]=[p′[i],v′[i]] T , p[i] is the actual position of the i-th joint, p′[i] is the planning position of the i-th joint, v[i] is the actual speed of the i-th joint, and v′[i] is the planning speed of the i-th joint, is the compensation amount of the i-th joint.
[0029] Optionally, whether the training is ended is judged according to the reward, including:
[0030] whether the reward is greater than or equal to a preset correction threshold value, if greater than or equal to the preset correction threshold value, the training is ended.
[0031] Optionally, the target joint contains a link degree of freedom, and the forward control of the target joint according to the joint planning value and the dynamics model to obtain the first joint parameter, including:
[0032] the link degree of freedom of the target joint is forward controlled according to the joint planning value and the dynamics model to obtain the first joint parameter.
[0033] The robot joint error compensation method provided by the application eliminates the error caused by the flexibility of the reducer of the mechanical arm in high-speed motion and the error caused by the mutual influence of other joints connected in series with the joint, and performs error compensation on the planning joint parameter, so that the vibration of the robot joint is suppressed in the working state of high precision and high speed.
[0034] The application further provides a robot joint error compensation device, including:
[0035] A feedforward control unit is configured to perform feedforward control on the target joint according to the joint planning value and a dynamics model to obtain a first joint parameter;
[0036] An adjustment unit is configured to perform negative feedback adjustment on the first joint parameter according to a sensor feedback value to obtain a second joint parameter;
[0037] A prediction unit is configured to perform prediction on the joint planning value, the sensor feedback value and a link state prediction network to obtain a predicted joint parameter;
[0038] A reinforcement learning unit is configured to perform compensation on the sensor feedback value, the predicted joint parameter, the joint planning value and a reinforcement learning system to obtain a compensation parameter;
[0039] The adjustment unit is further configured to perform positive feedback adjustment on the second joint parameter according to the compensation parameter to realize error compensation.
[0040] The robot joint error compensation device provided by the application eliminates the error caused by the flexible characteristics of the speed reducer of the robot arm in high-speed motion and the error caused by the mutual influence of other joints connected in series with the joint, performs error compensation on the planning joint parameter, and finally realizes vibration suppression of the robot joint in a high-precision and high-speed working state.
[0041] The application further provides a computer device, which comprises a memory, a processor and a computer program stored in the memory and executable on the processor, and the processor realizes the robot joint error compensation method according to any one of the above when executing the computer program.
[0042] The computer device provided by the application eliminates the error caused by the flexible characteristics of the speed reducer of the robot arm in high-speed motion and the error caused by the mutual influence of other joints connected in series with the joint, performs error compensation on the planning joint parameter, and finally realizes vibration suppression of the robot joint in a high-precision and high-speed working state.
[0043] The application further provides a computer readable storage medium, which stores a computer program, and the computer program is executable on a processor to realize the robot joint error compensation method according to any one of the above.
[0044] The computer readable storage medium of the present application eliminates the error caused by the flexible characteristics of the reducer of the mechanical arm in high-speed motion and the error caused by the mutual influence of other joints connected in series with the joint by training the connecting rod state prediction network and the reinforcement learning compensation strategy system and taking them as an important part of error compensation, and finally realizes the vibration suppression of the robot joint in the working state of high precision and high speed. BRIEF DESCRIPTION OF DRAWINGS
[0045] Figure 1 The robot joint error compensation method flowchart in the embodiment of the present application;
[0046] Figure 2 The robot joint error compensation method schematic diagram in the embodiment of the present application;
[0047] Figure 3 The robot joint error compensation method flowchart in the embodiment of the present application;
[0048] Figure 4 The robot joint error compensation method flowchart in the embodiment of the present application;
[0049] Figure 5 The connecting rod state prediction network schematic diagram of the robot joint error compensation method in the embodiment of the present application;
[0050] Figure 6 The robot joint error compensation method flowchart in the embodiment of the present application;
[0051] Figure 7 The robot joint error compensation device schematic diagram in the embodiment of the present application;
[0052] Figure 8 The robot joint error compensation device schematic diagram in the embodiment of the present application;
[0053] Figure 9 The computer equipment schematic diagram in the embodiment of the present application. DETAILED DESCRIPTION
[0054] In order to make the purpose, technical scheme and advantages of the embodiments of the present application clearer, the technical scheme of the embodiments of the present application will be described clearly and completely below in combination with the drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor fall within the scope of protection of the present application.
[0055] In combination with Figure 1 and Figure 2As shown, the embodiment provides a robot joint error compensation method, the joint is driven by a motor, the joint is connected with a connecting rod through a reducer, and the method comprises:
[0056] S3: performing feedforward control on the target joint according to the joint planning value and the dynamics model to obtain first joint parameters;
[0057] S4: performing negative feedback adjustment on the first joint parameters according to the sensor feedback value to obtain second joint parameters;
[0058] S5: obtaining predicted joint parameters according to the joint planning value, the sensor feedback value and a connecting rod state prediction network;
[0059] S6: obtaining compensation parameters according to the sensor feedback value, the predicted joint parameters, the joint planning value and a reinforcement learning system;
[0060] S7: performing positive feedback adjustment on the second joint parameters according to the compensation parameters to realize error compensation.
[0061] In S3, after receiving the joint planning value, the corresponding feedforward torque is calculated according to the robot inverse dynamics, the initial planning parameters are positively fed back and adjusted according to the feedforward torque, and the first joint parameters are obtained.
[0062] Wherein, the robot inverse dynamics calculation can adopt Newton-Euler iteration or Lagrange equation, or other same or similar algorithms, or other feasible methods, to perform inverse dynamics calculation to obtain the corresponding feedforward torque.
[0063] In S4, the first joint parameters are negatively fed back and adjusted through the sensor feedback value, and the adjusted first joint parameters are taken as the second joint parameters;
[0064] In S5, the sensor feedback value and the joint planning value are obtained, the connecting rod state is predicted through the model to obtain the predicted joint parameters, and the predicted joint parameters are output, the predicted connecting rod joint parameters are obtained by previously acquiring the robot end pose through the sensor and calculating through inverse kinematics.
[0065] In S6, after the reinforcement learning system obtains the sensor feedback value, the predicted joint parameters and the joint planning value, the compensation parameters are obtained through reinforcement learning, and the compensation parameters are output, wherein the second joint parameters are obtained through the feedback controller, and the sensor feedback value is obtained according to the memory system in the robot system;
[0066] In S7, the second joint parameters are positively fed back and adjusted through the compensation parameters and output to the robot for action;
[0067] The joint planning value comprises a joint planning position / speed, the first joint parameter comprises a first joint position / speed, the second joint parameter comprises a second joint position / speed, and the predicted joint parameter comprises a predicted joint position / speed.
[0068] The robot joint error compensation method of the application trains a connecting rod state prediction network and a reinforcement learning compensation strategy system, and uses them as an important part of error compensation. The error is caused by multiple factors, especially the error caused by the flexibility of the reducer and the error caused by the mutual influence of other joints connected in series with the joint. The error compensation is performed on the planned joint parameters, and finally the vibration of the robot joint is suppressed in the high-precision and high-speed working state.
[0069] In combination with FIGS. 1 to 3, Figure 3 and Figure 4 Before the predicted joint parameter is obtained according to the joint planning value, the sensor feedback value and the connecting rod state prediction network, the method further comprises the following steps in the embodiment of the application:
[0070] S1: constructing the connecting rod state prediction network, the constructing the connecting rod state prediction network comprising:
[0071] S11: obtaining an initial connecting rod state prediction network;
[0072] S12: obtaining a training data set, the training data set comprising a planned connecting rod joint parameter, a motor feedback parameter and an actual connecting rod parameter;
[0073] S13: inputting the training data set into the initial connecting rod state prediction network for training, wherein the input of the initial connecting rod state prediction network is the planned connecting rod joint parameter and the motor feedback parameter, and the output of the initial connecting rod state prediction network is a predicted connecting rod joint parameter;
[0074] S14: using the trained initial connecting rod state prediction network as the connecting rod state prediction network.
[0075] In the embodiment, the initial connecting rod state prediction network is obtained, the initial connecting rod state network does not have prediction capability, and a training data set needs to be constructed. When the training data set is constructed, a large number of planned connecting rod positions / speeds, motor feedback positions / speeds and actual connecting rod positions / speeds are needed. The actual connecting rod position / speed is calculated by using a sensor to obtain the position and posture of the connecting rod end. The above three types of parameters are used as input and output data of the initial connecting rod state prediction network to train the network. When the training is completed, the obtained initial connecting rod state prediction network is the connecting rod state prediction network.
[0076] The robot joint error compensation method of the application, by training a neural network, i.e. an initial link state prediction network, after deep learning, the joint movement is predicted to replace the sensor used in the traditional technology to calculate and measure the joint movement, when the initial link state prediction network training is completed, it is introduced into the system as a link state prediction network, when the robot works, no additional sensor is needed, saving the working space.
[0077] In combination Figure 5 As shown in the embodiment of the application, the training data set further includes actual link parameters, and the construction of the link state prediction network further includes:
[0078] S131: Obtain a loss function by the predicted link joint parameters and the actual link parameters;
[0079] S132: Determine whether the initial link state prediction network is trained according to the loss function.
[0080] In the embodiment, the training data set includes actual joint position / velocity, and the actual link position / velocity can be recorded by Laser tracker (high-precision laser tracker) or CompuGauge (wire-type sensor) on the robot end pose trajectory, and the encoder (motor encoder) data is recorded synchronously, the detection equipment is used to detect the robot end position / velocity, and the actual link pos / vel (actual link position / velocity) is calculated through IK (robot inverse dynamics), and the motor encoder data is recorded synchronously actual motor pos / vel (actual position / velocity of motor end), wherein the motor feedback parameter is the actual position / velocity of motor end;
[0081] The error between the predicted link joint position / velocity and the actual link position / velocity is taken as the loss function, when the value of the loss function is less than the preset threshold value, the training completion condition is reached at this time, the training of the initial link state prediction network is stopped, and the DNN (link state prediction network) is obtained; the link state prediction network can output the actual link position / velocity by giving a series of motor end actual position / velocity and (cmd link pos / vel) joint planning position / parameters.
[0082] The loss function expression is:
[0083] Wherein, j is the total number of joints, the total number of joints is the sum of the number of target joints and the number of associated joints of the target joints, x[i]=[p[i],v[i]] T x′[i]=[p′[i],v′[i]] T, p[i] is the actual position of the i th joint, p'[i] is the predicted position of the i th joint, v[i] is the actual speed of the i th joint, and v'[i] is the predicted speed of the i th joint.
[0084] The robot joint error compensation method of the application introduces a loss function to optimize the initial link state prediction network in training, so that the output prediction value loss is smaller and more accurate, further increasing the accuracy of robot work.
[0085] In combination Figure 6 As shown in the figure, before the compensation parameter is obtained according to the sensor feedback value, the second joint parameter, the predicted joint parameter, the joint planning value and the reinforcement learning system in the embodiment of the application, the method further comprises:
[0086] S2: constructing the reinforcement learning system, the construction of the reinforcement learning system comprising:
[0087] S21: obtaining a reinforcement learning neural network;
[0088] S22: using a reinforcement learning method to train the reinforcement learning neural network, taking the trained reinforcement learning neural network as a reinforcement learning system, and taking the planning link joint parameter, the actual link parameter and the motor feedback parameter as a state, taking the compensation parameter as an action, taking the reinforcement learning neural network as a policy, and taking the reference parameter as a reward, optimizing the policy by the action and the state to obtain the reward,
[0089] S23: judging whether the training is completed according to the reward;
[0090] The reference parameter is related to the associated joint of the target joint.
[0091] In the embodiment, the reinforcement learning neural network is obtained and trained, that is, data-driven training is performed by using reinforcement learning, and for the reinforcement learning system, states, actions, policies and rewards are included, and interaction is performed based on a set of interaction objects, that is, an agent and an environment. In the embodiment, the states include planned link joint position / speed, actual joint position / speed and motor feedback position / speed, the actions include compensation parameters, the policy includes the reinforcement learning neural network, and the reward is a reference parameter. The compensation parameters are obtained by inputting the planned link joint position / speed, the actual joint position / speed and the motor feedback position / speed to the reinforcement learning neural network, the reference parameter is calculated by using the compensation parameters, and then whether the training is completed is determined according to the reference parameter. If the training is completed, the compensation parameters are output. Since the error compensation of the target joint affects other associated joints, the reference parameter is set to be related to the associated joints of the target joint. In the embodiment, the reinforcement learning neural network can be an initial neural network or a reference neural network used in an intermediate link, and details are not described herein.
[0092] In the embodiment, the associated joint of the target joint refers to other joints that affect the movement of the target joint. For example, if the target joint is located on a mechanical arm, other joints on the mechanical arm will affect the movement of the target joint. In another embodiment, if the target joint is located at the starting end of multiple mechanical arms, multiple joints on the multiple mechanical arms will affect the target joint.
[0093] The robot joint error compensation method of the application can minimize the error and improve the compensation accuracy by selecting the best value of the compensation parameter for joint error compensation on the basis of considering the influence of other joints.
[0094] In the embodiment, the reference parameter is:
[0095]
[0096] wherein j is the total number of joints, the total number of joints is the sum of the number of target joints and the number of associated joints of the target joints, and x[i]=[p[i],v[i]] T x′[i]=[p′[i],v′[i]] T p[i] is the actual position of the i th joint, p′[i] is the planned position of the i th joint, v[i] is the actual speed of the i th joint, and v′[i] is the planned speed of the i th joint. is the compensation amount of the i th joint.
[0097] In the embodiment, it is judged whether the training is ended, and the evaluation criterion is the reference parameter, wherein, a value is obtained by adding the compensation parameter obtained by the reinforcement learning neural network to the absolute value of the difference between the actual joint position / speed in each joint and the planned joint position / speed, the same calculation is performed on the joints associated with the joint, a value is obtained, the values of all associated joints and the target joint are added to obtain a value set, and the reciprocal of the value set is taken as the reference parameter.
[0098] The robot joint error compensation method fully trains the reinforcement learning neural network by introducing the reference parameter in the reinforcement learning system as the evaluation criterion for ending the reinforcement learning training.
[0099] In the embodiment, the training is judged to be ended according to the reward, including:
[0100] It is judged whether the reward is greater than or equal to a preset correction threshold, and if yes, the training is ended.
[0101] In the embodiment, when the joint error is minimized and the compensation parameter reaches the optimal value, the reference parameter should be the maximum value, and it is determined whether the compensation parameter output by the reinforcement learning neural network in the strategy reaches the optimal value and whether the joint error is minimized by judging whether the value of the reference parameter is greater than or equal to the preset correction threshold.
[0102] The robot joint error compensation method sets the preset correction threshold to judge whether the reinforcement learning neural network is sufficient, so that the robot can output the optimal compensation parameter in the working process and ensure the accuracy of error compensation.
[0103] In the embodiment, the target joint includes a motor degree of freedom and a connecting rod degree of freedom, and the first joint parameter is obtained by performing feedforward control on the target joint according to the joint planning value and the dynamics model, including:
[0104] The connecting rod degree of freedom of the target joint is controlled by feedforward according to the joint planning value and the dynamics model, and the first joint parameter is obtained.
[0105] In the embodiment, the target joint is driven by a motor, the joint drives a connecting rod to work, thus the actual speed / position of the motor is related to the motor degree of freedom, at the beginning of the robot joint error compensation, a planned click position / speed is obtained according to a user instruction or upper planning software, the motor degree of freedom of the target joint is controlled by feedforward according to the planned motor position / speed, the inverse dynamics is calculated by the robot Newton-Euler iteration or Lagrange equation according to the joint planned position / speed, the feedforward torque is obtained, and the connecting rod degree of freedom of the target joint is controlled by feedforward to obtain the first joint position / speed.
[0106] The robot joint error compensation method of the application, at the beginning of the robot joint error compensation, the joint is controlled by feedforward first, by which the joint is controlled by the motor and the connecting rod is driven to work to obtain the first joint position / speed, and the preliminary control is performed before the connecting rod state prediction network and the reinforcement learning system work, which can meet the requirements for the robot joint error compensation with low requirements, and can be used as the basis for the subsequent feedback control for the robot working with high precision and high speed.
[0107] The error compensation method in the prior art usually only considers the error compensation of one degree of freedom in each joint of the robot, while in the embodiment of the application, two degrees of freedom of a target joint, i.e., the motor degree of freedom and the connecting rod degree of freedom, are considered, and the errors possibly caused by the two degrees of freedom are compensated, so that the accuracy of the error compensation can be improved.
[0108] In combination with Figure 7 the application also provides a robot joint error compensation device 100, which comprises:
[0109] a feedforward control unit 110, configured to control a target joint by feedforward according to a joint planning value and a dynamics model to obtain a first joint parameter;
[0110] an adjusting unit 120, configured to adjust the first joint parameter by negative feedback according to a sensor feedback value to obtain a second joint parameter;
[0111] a prediction unit 130, configured to obtain a predicted joint parameter according to the joint planning value, the sensor feedback value and a connecting rod state prediction network;
[0112] a reinforcement learning unit 140, configured to obtain a compensation parameter according to the sensor feedback value, the predicted joint parameter, the joint planning value and a reinforcement learning system;
[0113] The adjusting unit 120 is further configured to adjust the second joint parameter by positive feedback according to the compensation parameter to realize error compensation.
[0114] Combining Figure 8 As shown in the drawings, in one embodiment of the present application, the robot joint error compensation device 100 further comprises a construction unit 150 for constructing the continuous rod state prediction network;
[0115] The construction unit 150 further comprises an acquisition module 1501, a training module 1502, a function module 1503 and a judgment module 1504.
[0116] The acquisition module 1501 is configured to acquire an initial continuous rod state prediction network, and acquire a training data set, wherein the training data set comprises planned continuous rod joint parameters, motor feedback parameters and actual continuous rod parameters.
[0117] The training module 1502 is configured to input the training data set into the initial continuous rod state prediction network for training, wherein the input of the initial continuous rod state prediction network is the planned continuous rod joint parameters and the motor feedback parameters, and the output of the initial continuous rod state prediction network is predicted continuous rod joint parameters; and the initial continuous rod state prediction network after training is taken as the continuous rod state prediction network.
[0118] The function module 1503 is configured to acquire a loss function through the predicted continuous rod joint parameters and the actual continuous rod parameters.
[0119] The judgment module 1504 is configured to judge whether the initial continuous rod state prediction network ends training according to the loss function.
[0120] The construction unit 150 is further configured to construct the reinforcement learning system,
[0121] The acquisition module 1501 is further configured to acquire a reinforcement learning neural network.
[0122] The training module 1502 is further configured to train the reinforcement learning neural network by using a reinforcement learning method, and take the trained reinforcement learning neural network as a reinforcement learning system.
[0123] The judgment module 1504 is further configured to, when training the reinforcement learning neural network, take the planned continuous rod joint parameters, the actual continuous rod parameters and the motor feedback parameters as states, take a compensation parameter as an action, take the reinforcement learning neural network as a policy, take a reference parameter as a reward, optimize the policy by the action and the states to obtain the reward, and judge whether training ends according to the reward; wherein the reference parameter is related to the associated joint of the target joint.
[0124] The judgment module 1504 is further configured to judge whether the reward is greater than or equal to a preset correction threshold, and if greater than or equal to the preset threshold, end training.
[0125] The feedforward control unit 110 further comprises an acquisition module 1101 and a control module 1102;
[0126] The acquisition module 1101 is configured to acquire a planned motor parameter according to a sensor feedback value.
[0127] The control module 1102 is configured to perform feedforward control on a link degree of freedom of the target joint according to the joint planning value and the dynamics model, to obtain a first joint parameter.
[0128] The robot joint error compensation device provided by the application eliminates the error caused by the flexible characteristics of the reducer of the mechanical arm in high-speed motion and the error caused by the mutual influence of other joints connected in series with the joint, and compensates for the error of the planned joint parameter, so that the vibration of the robot joint is suppressed in the working state of high precision and high speed.
[0129] In combination with Figure 9 The application further provides a computer device, which comprises a memory, a processor, and a computer program stored in the memory and executable on the processor, and the processor implements the following steps when executing the computer program:
[0130] S3: performing feedforward control on a target joint according to a joint planning value and a dynamics model, to obtain a first joint parameter;
[0131] S4: performing negative feedback adjustment on the first joint parameter according to a sensor feedback value, to obtain a second joint parameter;
[0132] S5: obtaining a predicted joint parameter according to the joint planning value, the sensor feedback value, and a link state prediction network;
[0133] S6: obtaining a compensation parameter according to the sensor feedback value, the predicted joint parameter, the joint planning value, and a reinforcement learning system;
[0134] S7: performing positive feedback adjustment on the second joint parameter according to the compensation parameter, to realize error compensation.
[0135] The application further provides a computer readable storage medium, which stores a computer program, and the computer program is executed by a processor to implement the following steps:
[0136] S3: performing feedforward control on a target joint according to a joint planning value and a dynamics model, to obtain a first joint parameter;
[0137] S4: negative feedback adjustment is performed on the first joint parameter according to the sensor feedback value, and a second joint parameter is obtained;
[0138] S5: a predicted joint parameter is obtained according to the joint planning value, the sensor feedback value and a connecting rod state prediction network;
[0139] S6: a compensation parameter is obtained according to the sensor feedback value, the predicted joint parameter, the joint planning value and a reinforcement learning system;
[0140] S7: positive feedback adjustment is performed on the second joint parameter according to the compensation parameter, and error compensation is realized.
[0141] The computer readable storage medium of the present application eliminates the error caused by the flexible characteristics of the reducer of the mechanical arm in high-speed motion and the error caused by the mutual influence of other joints connected in series with the joint, performs error compensation on the planning joint parameter, and finally realizes vibration suppression of the robot joint in a high-precision and high-speed working state.
[0142] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiments can be completed by a computer program instructing related hardware, and the program can be stored in a non-volatile computer readable storage medium. When the program is executed, it can include the processes of the above-mentioned embodiments. Any reference to memory, storage, database or other medium used in the embodiments provided by the present application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. As an illustration but not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM) and memory bus dynamic RAM (RDRAM).
[0143] It has to be noted that, in the present document, relational terms are intended only to convey a possible relationship between elements or
[0144] The foregoing is considered as illustrative only of the principles of the application. Numerous modifications and changes will readily occur to those skilled in the art, which modifications and changes are to be understood as intended to be encompassed by the general scope of the application. Accordingly, the application is not to be limited to the above described or illustrated embodiments that are merely given by way of example. It is also be understood that various combinations of the above described embodiments and variations thereof are encompassed by the application.
Claims
1. A robot joint error compensation method, characterized by, The joint is driven by a motor, the joint is connected with a connecting rod through a reducer, and the method comprises: According to the joint planning value and the dynamics model, the target joint is subjected to feedforward control to obtain first joint parameters; According to the sensor feedback value, the first joint parameters are subjected to negative feedback adjustment to obtain second joint parameters; A connecting rod state prediction network is constructed, and the construction of the connecting rod state prediction network comprises: According to the joint planning value, the sensor feedback value and the connecting rod state prediction network, predicted joint parameters are obtained; A reinforcement learning system is constructed, and the construction of the reinforcement learning system comprises: According to the sensor feedback value, the predicted joint parameters, the joint planning value and the reinforcement learning system, compensation parameters are obtained; According to the compensation parameters, the second joint parameters are subjected to positive feedback adjustment to realize error compensation.
2. The method of claim 1, wherein, The training data set further comprises actual connecting rod parameters, and the construction of the connecting rod state prediction network further comprises: A loss function is obtained through the predicted connecting rod joint parameters and the actual connecting rod parameters; Whether the initial connecting rod state prediction network ends training is determined according to the loss function.
3. The method of claim 1, wherein, The reference parameters are: ; wherein j is the total number of joints, which is the sum of the number of target joints and the number of associated joints of the target joints, , , is the actual position of the i-th joint, is the planned position of the i-th joint, is the actual velocity of the i-th joint, is the planned velocity of the i-th joint, is the compensation amount of the i-th joint.
4. The method of claim 1, wherein, The determination of whether the training ends according to the reward comprises: Whether the reward is greater than or equal to a preset correction threshold is determined, and if the reward is greater than or equal to the preset correction threshold, the training ends.
5. The method of claim 1, wherein, The target joint comprises a connecting rod degree of freedom, The feedforward control of the target joint according to the joint planning value and the dynamics model to obtain the first joint parameters comprises: According to the joint planning value and the dynamics model, the connecting rod degree of freedom of the target joint is subjected to feedforward control to obtain the first joint parameters.
6. A robot joint error compensation apparatus, characterized by, It comprises: A feedforward control unit is configured to perform feedforward control on the target joint according to the joint planning value and the dynamics model to obtain first joint parameters; An adjustment unit is configured to perform negative feedback adjustment on the first joint parameters according to the sensor feedback value to obtain second joint parameters; A constructing unit is configured to construct a connecting rod state prediction network and a reinforcement learning system. The constructing unit further comprises an obtaining module, a training module and a judging module. The obtaining module is configured to obtain an initial connecting rod state prediction network and a training data set, wherein the training data set comprises planned connecting rod joint parameters, motor feedback parameters and actual connecting rod parameters. The training module is configured to input the training data set into the initial connecting rod state prediction network for training, wherein the input of the initial connecting rod state prediction network is the planned connecting rod joint parameters and the motor feedback parameters, and the output of the initial connecting rod state prediction network is predicted connecting rod joint parameters; and the initial connecting rod state prediction network after training is taken as the connecting rod state prediction network. The training module is further configured to train the reinforcement learning neural network by using a reinforcement learning method, and take the trained reinforcement learning neural network as the reinforcement learning system. The judging module is further configured to, when training the reinforcement learning neural network, take the planned connecting rod joint parameters, the actual connecting rod parameters and the motor feedback parameters as states, take a compensation parameter as an action, take the reinforcement learning neural network as a policy, take a reference parameter as a reward, optimize the policy by the action and the states to obtain the reward, and judge whether the training is completed according to the reward, wherein the reference parameter is related to an associated joint of the target joint. A prediction unit is configured to obtain predicted joint parameters according to the joint planning value, the sensor feedback value and the connecting rod state prediction network. A reinforcement learning unit is configured to obtain a compensation parameter according to the sensor feedback value, the predicted joint parameters, the joint planning value and the reinforcement learning system. The adjusting unit is further configured to perform positive feedback adjustment on the second joint parameters according to the compensation parameter to realize error compensation.
7. A computer device comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, The processor executes the computer program to realize the steps of the method in any one of claims 1 to 5.
8. A computer-readable storage medium having stored thereon a computer program, characterized in that, The computer program is executed by the processor to realize the steps of the method in any one of claims 1 to 5.
Citation Information
Patent Citations
Robot feedforward force moment compensation method
CN108393892A
Robot inverse dynamics feedforward control method and system based on fusion model
CN115042172A