Robot skill learning method, device, mechanical arm robot and storage medium
Through block timing imitation learning model and reinforced learning environment interaction technology, the problem of increasing compound error during long-term prediction of traditional imitation learning models is solved, the quality and efficiency of robot skills learning is improved, and the control accuracy and fluency are improved.
Patent Information
- Application Number
- CN202411561141.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-04
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2044-11-04
AI Technical Summary
The composite error of traditional imitation learning models increases during long-term prediction, resulting in poor learning effect of robot skills and low control accuracy.
The robot control training data of the original robot performing the task to be learned is collected through preset robot control strategies, and the learning model is trained using block timing imitation, and the model is optimized to improve the quality and efficiency of skill learning.
It improves the robot's perception of the environment, alleviates the impact of cumulative errors of traditional imitation learning algorithms, ensures the robot's skill learning quality and efficiency of new tasks, and improves control accuracy and fluency.
Smart Images

Figure CN119283030B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of robot control and artificial intelligence technology, and in particular to a robot skill learning method, device, mechanical arm robot and storage medium. Background Art
[0002] With the rapid development of robotics technology, robotic arms have shown great application prospects in many fields such as manufacturing, service and logistics due to their excellent flexibility and adaptability. The industry is committed to improving the professional skills of these robotic arms in their respective application scenarios and enhancing their precise operation capabilities in actual operations.
[0003] Manipulator robots mainly rely on preset motion paths to perform some simple, repetitive tasks. They lack the ability to perceive and adapt to the environment, that is, they lack the ability to learn skills and cannot autonomously plan action paths. The prediction results of traditional imitation learning models at each moment need to be calculated at the previous moment. This iterative prediction method will accumulate errors in long-term predictions. The increase in compound errors leads to poor skill learning effects and poor control accuracy. Summary of the invention
[0004] The present application provides a robot skill learning method, device, robotic arm robot and storage medium, which are used to solve the problems of increased compound error, poor skill learning effect and low control accuracy in traditional imitation learning models during long-term prediction.
[0005] The first aspect of the present application provides a robot skill learning method, comprising: collecting robot control training data when an original robot performs a task to be learned through a preset robot control strategy;
[0006] Training a preset initial block timing imitation learning model according to the robot control training data to obtain a target block timing imitation learning model;
[0007] Inputting the first robot control data of the original robot at the current moment into the target block timing imitation learning model, and performing robot control according to each action timing block output until the task to be learned is completed;
[0008] Each action timing block is evaluated according to a preset reward function, and the model of the original robot is optimized according to the evaluation result to obtain a target robot.
[0009] A second aspect of the present application provides a robot skill learning device, comprising: an acquisition module, for collecting robot control training data when an original robot performs a task to be learned through a preset robot control strategy;
[0010] A training module, used for training a preset initial block timing imitation learning model according to the robot control training data to obtain a target block timing imitation learning model;
[0011] An execution module, used for inputting the first robot control data of the original robot at the current moment into the target block timing imitation learning model, and performing robot control according to each action timing block outputted until the task to be learned is completed;
[0012] The optimization module is used to evaluate the execution result of the task to be learned according to a preset reward function, and optimize the model of the original robot according to the evaluation result to obtain a target robot.
[0013] The third aspect of the present application provides a robotic arm robot, comprising: a memory and at least one processor, wherein the memory stores instructions; the at least one processor calls the instructions in the memory so that the robotic arm robot executes the above-mentioned robot skill learning method.
[0014] A fourth aspect of the present application provides a computer-readable storage medium, wherein the computer-readable storage medium stores instructions, which, when executed on a computer, enable the computer to execute the above-mentioned robot skill learning method.
[0015] In the technical solution provided by the present application, the robot control training data of the original robot when performing the task to be learned is collected through the reinforcement learning environment interaction technology, thereby improving the robot's perception of the environment. The block temporal imitation learning model is used to learn the task to be learned, thereby solving the problem of increased compound errors of the traditional imitation learning model in long-term prediction, ensuring the quality and efficiency of the robot's skill learning of new tasks and improving control accuracy. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] Figure 1 A schematic diagram of an embodiment of the robot skill learning method of the present application;
[0017] Figure 2 A schematic diagram of another embodiment of the robot skill learning method of the present application;
[0018] Figure 3 This is a schematic diagram of the block timing imitation learning model framework for this application;
[0019] Figure 4 A schematic diagram of an embodiment of a robot skill learning device of the present application;
[0020] Figure 5 A schematic diagram of another embodiment of the robot skill learning device of the present application;
[0021] Figure 6This is a schematic diagram of an embodiment of a robotic arm robot of the present application. DETAILED DESCRIPTION
[0022] The present application provides a robot skill learning method, device, robotic arm robot and storage medium, which are used to solve the problems of increased compound error, poor skill learning effect and low control accuracy in traditional imitation learning models during long-term prediction.
[0023] The terms "first", "second", "third", "fourth", etc. (if any) in the specification and claims of this application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged where appropriate, so that the embodiments described herein can be implemented in an order other than that illustrated or described herein. In addition, the terms "including" or "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units that are clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0024] For ease of understanding, the specific process of the embodiment of the present application is described below. Figure 1 , an embodiment of the robot skill learning method in the embodiment of the present application includes:
[0025] 101. Robot control training data of the original robot when performing the task to be learned is collected through a preset robot control strategy.
[0026] It can be understood that the executor of the present application can be a robot skill learning device, or various intelligent entities, such as various types of robotic arm robots. In this embodiment, the robot control strategy is used to indicate the control strategy of the original robot, which can be any one of the master-slave control, drag teaching, code drive, etc., to control the original robot to perform the task to be learned, and obtain the robot control training data through the various control interfaces of the original robot, such as the motion control interface and the camera interface, for skill learning.
[0027] The above-mentioned robot control training data may include joint control data, which includes position information and speed information of each joint collected through each motion control interface on the robot, as well as the opening and closing status of each gripper, etc. The above-mentioned joint control data is a time series arranged according to time. The joint data at each moment can be converted into a matrix through reinforcement learning environment interaction technology to represent the state space corresponding to the robot at each moment.
[0028] The above-mentioned robot control training data may also include visual image sample data, which is real-time video data collected by a camera installed on the robot. The robot can perform environmental semantic analysis through the visual image sample data to improve the robot's perception of the environment. The number and installation positions of the cameras can be set according to actual conditions, and this embodiment does not impose any specific restrictions.
[0029] 102. Train a preset initial block timing imitation learning model according to the robot control training data to obtain a target block timing imitation learning model.
[0030] The present application provides a Block Temporal Imitation Learning (BTIL) model. The BTIL model is an improved imitation learning algorithm that divides a continuous time series (such as robot control training data) into blocks, extracts the comprehensive features of each time series block and the action prediction result of the previous time series block to predict the action of the next time series block.
[0031] Specifically, the robot control training data is input into the initial block timing imitation learning model, the action prediction results of each timing block are output, and the loss value of each timing block is determined according to the target loss function, the action prediction results of each timing block and the actual action sequence of each timing block. If the sum of the loss values is greater than the preset loss threshold, the model parameters of the initial block timing imitation learning model are adjusted until the target loss function converges to obtain the target block timing imitation learning model, and the target block timing imitation learning model is deployed on the original robot.
[0032] In a feasible implementation, the robot control training data is input into the initial block timing imitation learning model, and the action prediction result of each timing block is output, including: dividing the robot control training data according to a preset step size to obtain a robot data segment corresponding to each timing block; extracting features of the robot data segment corresponding to each timing block to obtain a comprehensive feature corresponding to each timing block; predicting the action of the next timing block according to the comprehensive feature corresponding to each timing block and the action prediction result of the previous timing block to obtain multiple initial timing action blocks; smoothing each initial timing action block according to the action weight value at each moment to obtain each action timing block.
[0033] The action weight value at each moment can be determined by the following formula:
[0034]
[0035] in, is the weight value corresponding to the i-th moment, is the natural logarithm, is a hyperparameter that determines the rate at which the influence of past actions on the current action decreases. The smaller it is, the slower the decay, and the impact of actions at past moments will have a greater impact on the current moment.
[0036] It can be understood that each timing block corresponds to a time segment, and the duration of the time segment is the preset step length. For example, the preset step length is 4 seconds, the time for the original robot to perform the task to be learned is 11 seconds, the first timing block corresponds to the 0th to 3rd seconds, the second timing block corresponds to the 4th to 7th seconds, and the third timing block corresponds to the 8th to 11th seconds. The robot control training data is the joint control data and visual image sample data collected by the original robot within the 0th to 11th seconds. The initial block timing imitation learning model of this embodiment can be used through the strategy prediction function express, Indicates Time to The action prediction result at the moment, That is, the preset step size, For the The robot control training data corresponding to the moment.
[0037] When performing action prediction for the first time sequence block, the action prediction result of the previous time sequence block can be set to the default value, and the action sequence of 4 seconds is inferred each time. According to the robot control training data corresponding to time t=0 and the preset action default values, the predicted action sequence of 0 to 3 seconds is output , which is the predicted first action timing block, According to the robot control training data corresponding to time t=4 and the first action timing block, the predicted action sequence of 4 to 7 seconds is output , which is the predicted second action timing block, According to the robot control training data corresponding to time t=8 and the second action timing block, the predicted action sequence of 8 to 11 seconds is output , that is, the predicted third action timing block; through the preset target loss function, the predicted action sequence of the above three predicted timing blocks and the real action sequence in the robot control training data, that is, the real action timing blocks divided according to the robot control training data, to determine the gap between the predicted value and the real value, and optimize the model through back propagation to obtain the target block timing imitation learning model.
[0038] In this embodiment, the comprehensive feature is a multi-dimensional feature vector extracted based on the robot control training data. The multi-dimensional feature vector may include position features, color features, texture features, etc. The comprehensive feature vector includes each feature information, as well as the correlation between features and global context information.
[0039] 103. Input the first robot control data of the original robot at the current moment into the target block timing imitation learning model, and perform robot control according to each output action timing block until the task to be learned is completed.
[0040] Specifically, after the model training is completed, the target block timing imitation learning model is loaded into the control program of the original robot, the first robot control data of the original robot at the current moment is used as input into the target block timing imitation learning model, and the first action timing block is output; the first action timing block is passed into the control module of the original robot, the original robot controls each joint to move according to the first action timing block, and executes the second robot control data of the first action timing block, repeats the above process, and the original robot completes the entire task according to the output of each action timing block.
[0041] 104. Each action sequence block is evaluated according to the preset reward function, and the model of the original robot is optimized according to the evaluation results to obtain the target robot.
[0042] Specifically, each action timing block is evaluated according to the preset reward function to obtain the sum of the reward values of the original task; if the sum of the reward values is less than the preset reward threshold, the hyperparameters of the target block timing imitation learning model are adjusted until the evaluation result is greater than or equal to the preset reward threshold, and the target robot is obtained.
[0043] In the embodiment of the present application, the robot control training data of the original robot when performing the task to be learned is collected through the reinforcement learning environment interaction technology, thereby improving the robot's perception of the environment. The block timing imitation learning model is used to learn the task to be learned to solve the problem of increased compound errors in traditional imitation learning models during long-term predictions, thereby ensuring the quality and efficiency of the robot's skill learning of new tasks and improving control accuracy.
[0044] See also Figure 2 Another embodiment of the robot skill learning method in the embodiment of the present application includes:
[0045] 201. Collect robot control training data when the original robot performs the task to be learned through a preset robot control strategy.
[0046] Specifically, if the preset robot control strategy is a master-slave control strategy, a master control robot is built, and the master control robot controls the original robot to perform the task to be learned through remote control; the joint control data and visual image sample data at each moment when the original robot performs the task to be learned are obtained to obtain the robot control training data.
[0047] Provide an example of a master-slave dual-arm robot: The dual-arm robot built includes two master control arms and two slave control arms. By manipulating the master control arms, the slave control arms can imitate the movements of the master control arms. The two slave control arms are of the same model, with 6 degrees of freedom, and both include a base, 6 joints, and 1 gripper. The angular motion limit range of joints 1, 4, and 6 is 360 degrees, while the angular motion limit range of joints 2, 3, and 5 is 180 degrees. The opening and closing limit of the grippers of the left and right arms is 180 degrees. Two wrist cameras are installed at the wrists of the two slave control arms, and one camera is installed at the top between the two slave control arms.
[0048] In the real scene, the task is to pick up the water cup on a rectangular table with the left arm, pick up the towel next to the table with the right arm, and then wipe the water stains under the water cup. Finally, the right arm puts the towel back to its original position, and the left arm puts the water cup back to its original position. During the operation, the operator operates the master control arm to perform the entire task, and the slave control arm imitates the master control arm to complete the corresponding task process. In the process of the slave control arm completing the task, the joint data information and video data information of the slave control arm are learned through the control interface of the robot arm.
[0049] 202. Divide the joint control sample data according to a preset step length to obtain timing block sample data, where the timing block sample data includes a real action sequence of each timing block.
[0050] The joint control sample data is divided according to the preset step length to obtain timing block sample data, and the timing block sample data includes a plurality of state space timing blocks and a plurality of real action timing blocks.
[0051] Among them, each real action timing block includes the action space corresponding to each moment in each time segment. For example, the real action sequence corresponding to the i-th timing block is , that is, the real action sequence of each time block, which is used as the real value and the predicted value of the model to calculate the loss value;
[0052] Among them, each state space time sequence block includes the state space components at each moment in each time segment. For example, the state space sequence corresponding to the i-th time sequence block is , as the input variable of the model for skill learning.
[0053] It can be understood that the action space at each moment includes the position information of each joint at each moment, and the position information of each jaw at each moment; the state space at each moment includes the position information of each joint and each jaw at each moment, the opening and closing state information of each jaw at each moment, and the angular velocity information of each joint at each moment.
[0054] The joint control data at each moment can be represented by the state space Observation matrix. For example, the state space corresponding to the 0th moment of the slave robot can be represented by The state matrix is expressed as:
[0055]
[0056] in, Indicates the left robot arm Moment The position information of each joint, Indicates the right robot arm The position information of the nth joint at the moment, Indicates the left gripper Location information at the moment, Indicates the right gripper Location information at the moment, Indicates the left robot arm The angular velocity information of the nth joint at the moment, Indicates the right robot arm The angular velocity information of the nth joint at the moment, Indicates the left gripper Opening and closing status information at all times, Indicates the right gripper Opening and closing status information at all times.
[0057] The joint control data at each moment mentioned above contains the action space Action matrix composed of the joint and gripper position information at each moment. For example, the action space corresponding to the zeroth moment of the arm robot can be obtained by The action matrix is:
[0058]
[0059] in, Indicates the left robot arm Moment The position information of each joint, Indicates the right robot arm The position information of the nth joint at the moment, Indicates the left gripper Location information at the moment, Indicates the right gripper Location information at the moment.
[0060] 203. Input the time series block sample data and the visual image sample data into the initial block time series imitation learning model to obtain the action prediction result of each time series block.
[0061] Specifically, the state space corresponding to each time block in the time block data is encoded by a first encoder to obtain the latent variables corresponding to each time block; the feature extraction network is used to extract features of the visual image sample data to obtain the multidimensional visual image features corresponding to each time block; the second encoder is used to convert the latent variables corresponding to each time block and the multidimensional visual image features corresponding to each time block to obtain the comprehensive features corresponding to each time block; the decoder is used to decode the comprehensive features corresponding to each time block and the action prediction results of the previous time block to obtain the action prediction results of each time block.
[0062] Reference Figure 3 The block temporal imitation learning model can be divided into an encoding stage and a decoding stage. In the encoding stage, the input of each joint of the robot arm, the position information of the gripper, the speed information, and the opening and closing state are spliced to form a complete vector, that is, the output latent variable z; in the decoding stage, the multi-dimensional visual image features of the visual image sample data and the latent variable z are extracted as the input of the second encoder. The second encoder outputs the comprehensive features, and the comprehensive features and the action prediction results of the previous time sequence block, that is, the predicted action space at the previous moment, are used as the input of the decoder. The decoder outputs the predicted action space of the current time sequence block, that is, the action sequence of each joint, that is, the action timing block.
[0063] Optionally, feature extraction is performed on the visual image sample data through a feature extraction network to obtain multi-dimensional visual image features corresponding to each time series block. The feature extraction formula is as follows:
[0064]
[0065] in, Represents the multi-dimensional visual image features extracted from the visual image, such as edge and texture features, local features, object spatial relationships and structures in the image, etc. Represents the feature extraction network as Figure 3 The input of the last layer of Inception-ResNet, It represents the mapping result of the residual module of the last layer through the forward operation, g is the number of residual layers, and g-1 represents the second to last layer of the residual layer. Used to indicate the width, height, and number of channels of the feature image.
[0066] In practical applications, the model will infer the action at k moments based on the comprehensive features of each moment during the decoding stage. That is, the input is the input variable at each moment, and the output is the predicted values of k moments. Multiple output values will cause the robot to move unsteadily. In order to alleviate the unsteady robot movement caused by the timing block prediction, this embodiment smoothes the action space at each moment in each predicted action timing block.
[0067] Optionally, the comprehensive features corresponding to each time sequence block and the action prediction results of the previous time sequence block are decoded by a decoder to obtain the action prediction results of each time sequence block, including: performing action prediction on the next time sequence block by a decoder based on the comprehensive features corresponding to each time sequence block and the action prediction results of the previous time sequence block to obtain an initial prediction result of each time sequence block; and smoothing the initial prediction results of each time sequence block to obtain the action prediction results of each time sequence block.
[0068] The following example uses exponential weighting to determine the weight value at each moment for smoothing: The smoothing formula in the decoding stage is:
[0069]
[0070] in, The output after smoothing Time to The predicted action sequence at the moment, For the The weight vector corresponding to the moment, is a hyperparameter that determines the rate at which the influence of past actions on the current action decreases. The smaller it is, the slower the decay, and the influence of actions at past moments has a greater impact on the current moment; strategy prediction function For the The unsmoothed output of the state space corresponding to the moment Time to The action prediction result at the moment.
[0071] Optionally, converting the latent variables corresponding to each time series block and the multidimensional visual image features corresponding to each time series block through a second encoder to obtain comprehensive features corresponding to each time series block includes: performing data conversion and feature extraction on the latent variables corresponding to each time series block and the multidimensional visual image features corresponding to each time series block through a second encoder to obtain comprehensive features corresponding to each time series block.
[0072] 204. Adjust the parameters of the initial block timing imitation learning model according to a preset target loss function, the action prediction result of each timing block and the real action sequence of each timing block until the target loss function converges to obtain the target block timing imitation learning model.
[0073] The target loss function of this embodiment uses the collaborative model learning target loss function to update the strategy. The complete target loss function formula of the model is as follows:
[0074]
[0075] Among them, the action prediction result of each time block and the actual action sequence of each timing block Determine the loss value of each time block; add the loss values of each time block to get the total loss value of this training; if the total loss value If the loss is greater than the preset loss threshold, the parameters of the initial block timing imitation learning model are adjusted and trained again until the total loss value is less than the preset loss threshold or the number of training times is greater than the number threshold, and the target block timing imitation learning model is obtained.
[0076] The above objective loss function can be divided into the first loss function in the encoding stage and the second loss function in the decoding stage:
[0077] The first loss function formula in the encoding stage is as follows:
[0078]
[0079] in, represents the loss function of the model encoding stage, Indicates Time to Real action sequences of moments, Indicates Time to The predicted action sequence at the moment, represents the KL divergence, which measures the difference between the probability distribution in the latent space and the prior distribution, q represents the probability distribution function of the latent space, is the real state space at the i-th moment, and β is the preset weight value. In this embodiment, the parameter value of the transformer encoder in the encoding stage is continuously adjusted through the first loss function formula to output a more accurate latent variable z.
[0080] The second loss function formula in the decoding stage is as follows:
[0081]
[0082] in, represents the loss function of the model decoding stage, is the cross entropy loss, which is used to measure the predicted action sequence output by the model With real action sequences The difference between Represents the input z and get The probability value of represents a gradient-based regularization term that encourages the model to be insensitive to small changes in input and parameters, where represents the gradient of the predicted action sequence block to the input, Represents the predicted action sequence block input and model parameters The gradient of Represents model parameters of Regularization term, used to prevent overfitting, where j is the number of model parameters; α and γ are both hyperparameters used to balance the contribution of different loss terms; is a hidden variable, is the real state space at the i-th moment.
[0083] This embodiment continuously adjusts the parameter values of the transformer encoder and the transformer decoder in the decoding stage through the second loss function formula, so that the output action prediction result is closer to the real action sequence.
[0084] 205. Input the first robot control data of the original robot at the current moment into the target block timing imitation learning model, and perform robot control according to each output action timing block until the task to be learned is completed.
[0085] The first robot control data of the original robot at the current moment is input into the target block timing imitation learning model to obtain the first action timing block; the joints of the original robot are controlled to move according to the first action timing block, and the second robot control data is obtained; the target block timing imitation learning model is used to predict according to the second robot control data and the first action timing block to obtain the second action timing block; the original robot is used to predict and execute each action timing block in turn until the task to be learned is completed.
[0086] 206. Each action sequence block is evaluated according to the preset reward function, and the model of the original robot is optimized according to the evaluation result to obtain the target robot.
[0087] The reward result corresponding to the task to be learned is determined according to the action reward corresponding to each moment; the hyperparameters of the target block timing imitation learning model are adjusted according to the reward result and the preset target reward threshold until the reward result is greater than or equal to the preset reward threshold, thereby obtaining the target robot.
[0088] Optionally, the reward result corresponding to completing the task to be learned is determined based on the action reward corresponding to each moment, including: if the next action timing block is a real action sequence, then the reward value corresponding to the next action timing block is determined to be a first preset value, otherwise it is a second preset value, the first preset value is a positive number, and the second preset value is a negative number; the sum of the reward values of each action timing block of the original robot performing the task to be learned is determined as the reward result.
[0089] For example, when the predicted action of the next time block is a real action, the reward is +1, otherwise the reward is -1. At the same time, the final reward is obtained by accumulating the rewards of the entire task process. The formula is as follows:
[0090]
[0091] in, is the reward result of the task to be learned, T is the total time of executing the task to be learned, t is the time of executing each action in the task to be learned, is the reward value corresponding to the tth moment.
[0092] In the embodiment of the present application, comprehensive feature extraction is performed by combining joint control data and visual image data with reinforcement learning environment interaction technology to further improve the robot's perception of the environment. Skill learning is performed through a block timing imitation learning model to alleviate the impact of the cumulative error of the traditional imitation learning algorithm on the robot's skill learning. The prediction results are optimized through smoothing processing to alleviate the unstable robot movement caused by the introduction of timing blocks, thereby ensuring the quality and efficiency of the robot's skill learning for new tasks and improving the accuracy and fluency of the robot's control.
[0093] The above describes the robot skill learning method in the embodiment of the present application. The following describes the robot skill learning device in the embodiment of the present application. Figure 4 In the embodiment of the present application, an embodiment of the robot skill learning device includes:
[0094] An acquisition module 401 is used to collect robot control training data when the original robot performs the task to be learned through a preset robot control strategy;
[0095] A training module 402 is used to train a preset initial block timing imitation learning model according to the robot control training data to obtain a target block timing imitation learning model;
[0096] An execution module 403 is used to input the first robot control data of the original robot at the current moment into the target block timing imitation learning model, and perform robot control according to each action timing block output until the task to be learned is completed;
[0097] The optimization module 404 is used to reward the prediction results of the learning task according to a preset reward function, and optimize the model of the original robot according to the reward result to obtain the target robot.
[0098] In the embodiment of the present application, the robot control training data of the original robot when performing the task to be learned is collected through the reinforcement learning environment interaction technology, thereby improving the robot's perception of the environment. The block timing imitation learning model is used to learn the task to be learned to solve the problem of increased compound errors in traditional imitation learning models during long-term predictions, thereby ensuring the quality and efficiency of the robot's skill learning of new tasks and improving control accuracy.
[0099] See also Figure 5 Another embodiment of the robot skill learning device in the embodiment of the present application includes:
[0100] An acquisition module 401 is used to collect robot control training data when the original robot performs the task to be learned through a preset robot control strategy;
[0101] A training module 402 is used to train a preset initial block timing imitation learning model according to the robot control training data to obtain a target block timing imitation learning model;
[0102] An execution module 403 is used to input the first robot control data of the original robot at the current moment into the target block timing imitation learning model, and perform robot control according to each action timing block output until the task to be learned is completed;
[0103] The optimization module 404 is used to reward the prediction results of the learning task according to a preset reward function, and optimize the model of the original robot according to the reward result to obtain the target robot.
[0104] Optionally, the training module 402 includes:
[0105] A division unit 4021 is used to divide the joint control sample data according to a preset step length to obtain timing block sample data, where the timing block sample data includes a real action sequence of each timing block;
[0106] Prediction unit 4022, used for inputting the time sequence block sample data and the visual image sample data into the initial block time sequence imitation learning model to obtain the action prediction result of each time sequence block;
[0107] The optimization unit 4023 is used to adjust the parameters of the initial block timing imitation learning model according to the preset target loss function, the action prediction result of each timing block and the real action sequence of each timing block until the target loss function converges to obtain the target block timing imitation learning model.
[0108] Optionally, the prediction unit 4022 includes:
[0109] The encoding subunit 40221 is used to encode the state space corresponding to each time sequence block in the time sequence block data through a first encoder to obtain a hidden variable corresponding to each time sequence block;
[0110] The feature extraction subunit 40222 is used to extract features from the visual image sample data through a feature extraction network to obtain multi-dimensional visual image features corresponding to each time series block;
[0111] The conversion subunit 40223 is used to convert the latent variables corresponding to each time sequence block and the multi-dimensional visual image features corresponding to each time sequence block through the second encoder to obtain the comprehensive features corresponding to each time sequence block;
[0112] The decoding subunit 40224 is used to decode the comprehensive features corresponding to each time sequence block and the action prediction result of the previous time sequence block through a decoder to obtain the action prediction result of each time sequence block.
[0113] Optionally, the decoding subunit 40224 is specifically used to: perform action prediction on the next time sequence block by using the decoder to perform action prediction on the comprehensive features corresponding to each time sequence block and the action prediction result of the previous time sequence block, so as to obtain an initial prediction result of each time sequence block;
[0114] The initial prediction result of each time series block is smoothed to obtain the action prediction result of each time series block.
[0115] Optionally, the execution module 403 is specifically used to: input the first robot control data of the original robot at the current moment into the target block timing imitation learning model to obtain a first action timing block;
[0116] Controlling each joint of the original robot to move according to the first action sequence block, and acquiring control data of the second robot;
[0117] The target block timing imitation learning model is used to predict the second robot control data and the first action timing block to obtain the second action timing block;
[0118] The original robot predicts and executes each action sequence block in sequence until the task to be learned is completed.
[0119] Optionally, the optimization module 404 is specifically used to: determine the reward result corresponding to completing the task to be learned according to the action reward corresponding to each moment;
[0120] According to the reward result and the preset target reward threshold, the hyperparameters of the target block timing imitation learning model are adjusted until the reward result is greater than or equal to the preset reward threshold, and the target robot is obtained.
[0121] Optionally, the acquisition module 401 is specifically used to build a master control robot if the preset robot control strategy is a master-slave control strategy, and the master control robot controls the original robot to perform the task to be learned through teleoperation;
[0122] The joint control data and visual image sample data at each moment when the original machine performs the task to be learned are obtained to obtain the robot control training data.
[0123] In the embodiments of the present application, comprehensive feature extraction is performed by combining joint control data and visual image data through reinforcement learning environment interaction technology, thereby further improving the robot's perception of the environment. Skill learning is performed through a block timing imitation learning model, thereby alleviating the impact of the cumulative error of the traditional imitation learning algorithm on the robot's skill learning. The prediction results are optimized through smoothing processing, thereby alleviating the unstable robot movement caused by the introduction of timing blocks, thereby ensuring the quality and efficiency of the robot's skill learning for new tasks, and improving the accuracy and fluency of the robot's control.
[0124] above Figure 4 and Figure 5 The robot skill learning device in the embodiment of the present application is described in detail from the perspective of modular functional entities, and the robotic arm robot in the embodiment of the present application is described in detail from the perspective of hardware processing.
[0125] See also Figure 6 As shown, the robotic arm robot includes a processor 600 and a memory 601, wherein the memory 601 stores machine executable instructions that can be executed by the processor 600, and the processor 600 executes the machine executable instructions to implement the above-mentioned robot skill learning method.
[0126] Further, Figure 6 The manipulator robot shown also includes a bus 602 and a communication interface 603 , and the processor 600 , the communication interface 603 and the memory 601 are connected via the bus 602 .
[0127] Among them, the memory 601 may include a high-speed random access memory (Random Access Memory, RAM), and may also include a non-volatile memory (non-volatile memory), for example, at least one disk storage. The communication connection between the system network element and at least one other network element is realized through at least one communication interface 603 (which can be wired or wireless), and the Internet, wide area network, local area network, metropolitan area network, etc. can be used. The bus 602 can be an ISA bus, a PCI bus or an EISA bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 6 Only one bidirectional arrow is used in the diagram, but this does not mean that there is only one bus or only one type of bus.
[0128] The processor 600 may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method can be completed by the hardware integrated logic circuit or software instructions in the processor 600. The above processor 600 can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components. The disclosed methods, steps and logic block diagrams in the embodiments of the present disclosure can be implemented or executed. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor. The steps of the method disclosed in conjunction with the embodiments of the present disclosure can be directly embodied as a hardware decoding processor to be executed, or a combination of hardware and software modules in the decoding processor can be executed. The software module can be located in a mature storage medium in the field such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory or an electrically erasable programmable memory, a register, etc. The storage medium is located in the memory 601 , and the processor 600 reads the information in the memory 601 and completes the method steps of the above-mentioned embodiment in combination with its hardware.
[0129] The present application also provides a computer-readable storage medium, which may be a non-volatile computer-readable storage medium or a volatile computer-readable storage medium. Instructions are stored in the computer-readable storage medium. When the instructions are executed on a computer, the computer executes the steps of the robot skill learning method.
[0130] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0131] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium, including several instructions to enable a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM), random access memory (RAM), disk or optical disk, etc., various media that can store program codes.
[0132] As described above, the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. A robot skill learning method, characterized in that: The robot skill learning method comprises: Collecting robot control training data when the original robot performs the task to be learned through a preset robot control strategy, wherein the robot control training data includes joint control sample data and visual image sample data; Training a preset initial block timing imitation learning model according to the robot control training data to obtain a target block timing imitation learning model, wherein the initial block timing imitation learning model includes a first encoder, a feature extraction network, a second encoder and a decoder; Inputting the first robot control data of the original robot at the current moment into the target block timing imitation learning model, and performing robot control according to each action timing block output until the task to be learned is completed; Each action timing block is evaluated according to a preset reward function, and a model of the original robot is optimized according to the evaluation result to obtain a target robot; The step of training a preset initial block timing imitation learning model according to the robot control training data to obtain a target block timing imitation learning model includes: Dividing the joint control sample data according to a preset step length to obtain timing block sample data, wherein the timing block sample data includes a real action sequence of each timing block; Inputting the time sequence block sample data and the visual image sample data into the initial block time sequence imitation learning model to obtain the action prediction result of each time sequence block; Adjusting parameters of the initial block timing imitation learning model according to a preset target loss function, an action prediction result of each timing block, and a real action sequence of each timing block until the target loss function converges to obtain a target block timing imitation learning model; The step of inputting the time sequence block sample data and the visual image sample data into the initial block time sequence imitation learning model to obtain the action prediction result of each time sequence block includes: Encoding the state space corresponding to each time sequence block in the time sequence block sample data by the first encoder to obtain a hidden variable corresponding to each time sequence block; Performing feature extraction on the visual image sample data through the feature extraction network to obtain multi-dimensional visual image features corresponding to each time sequence block; The second encoder converts the latent variables corresponding to each time sequence block and the multi-dimensional visual image features corresponding to each time sequence block to obtain the comprehensive features corresponding to each time sequence block; The decoder decodes the comprehensive features corresponding to each time sequence block and the action prediction result of the previous time sequence block to obtain the action prediction result of each time sequence block.
2. The robot skill learning method according to claim 1, characterized in that: The decoding of the comprehensive features corresponding to each time sequence block and the action prediction result of the previous time sequence block by the decoder to obtain the action prediction result of each time sequence block includes: The decoder performs action prediction on the next time sequence block based on the comprehensive features corresponding to each time sequence block and the action prediction result of the previous time sequence block, so as to obtain the initial prediction result of each time sequence block; The initial prediction result of each time series block is smoothed to obtain the action prediction result of each time series block.
3. The robot skill learning method according to claim 1, characterized in that: The first robot control data of the original robot at the current moment is input into the target block timing imitation learning model, and the robot is controlled according to each action timing block output until the task to be learned is completed, including: Inputting the first robot control data of the original robot at the current moment into the target block timing imitation learning model to obtain a first action timing block; Controlling each joint of the original robot to move according to the first action sequence block, and acquiring control data of the second robot; Predicting according to the second robot control data and the first action timing block by the target block timing imitation learning model to obtain a second action timing block; The original robot predicts and executes each action sequence block in sequence until the task to be learned is completed.
4. The robot skill learning method according to any one of claims 1 to 3, characterized in that: The step of evaluating each action timing block according to a preset reward function and optimizing the model of the original robot according to the evaluation result to obtain a target robot includes: Determine the reward result corresponding to completing the task to be learned according to the action reward corresponding to each moment; The hyperparameters of the target block timing imitation learning model are adjusted according to the reward result and the preset target reward threshold until the reward result is greater than or equal to the preset reward threshold, thereby obtaining the target robot.
5. The robot skill learning method according to claim 1, characterized in that: The robot control training data collected when the original robot performs the task to be learned by using a preset robot control strategy includes: If the preset robot control strategy is a master-slave control strategy, a master control robot is built, and the master control robot controls the original robot to perform the task to be learned through teleoperation; The joint control data and visual image sample data at each moment when the original robot performs the task to be learned are obtained to obtain robot control training data.
6. A robot skill learning device, characterized in that: The robot skill learning device comprises: An acquisition module, used for collecting robot control training data when the original robot performs the task to be learned through a preset robot control strategy, wherein the robot control training data includes joint control sample data and visual image sample data; A training module, used for training a preset initial block timing imitation learning model according to the robot control training data to obtain a target block timing imitation learning model, wherein the initial block timing imitation learning model includes a first encoder, a feature extraction network, a second encoder and a decoder; An execution module, used for inputting the first robot control data of the original robot at the current moment into the target block timing imitation learning model, and performing robot control according to each action timing block outputted until the task to be learned is completed; An optimization module, used to evaluate the execution result of the task to be learned according to a preset reward function, and optimize the model of the original robot according to the evaluation result to obtain a target robot; The training module includes: A division unit, used for dividing the joint control sample data according to a preset step length to obtain timing block sample data, wherein the timing block sample data includes a real action sequence of each timing block; A prediction unit, used for inputting the time sequence block sample data and the visual image sample data into the initial block time sequence imitation learning model to obtain an action prediction result for each time sequence block; An optimization unit, used for adjusting parameters of the initial block timing imitation learning model according to a preset target loss function, an action prediction result of each timing block, and a real action sequence of each timing block, until the target loss function converges to obtain a target block timing imitation learning model; The prediction unit includes: an encoding subunit, configured to encode the state space corresponding to each time sequence block in the time sequence block sample data by using the first encoder to obtain a hidden variable corresponding to each time sequence block; A feature extraction subunit, configured to extract features from the visual image sample data through the feature extraction network to obtain multi-dimensional visual image features corresponding to each time sequence block; a conversion subunit, configured to convert the latent variables corresponding to each time sequence block and the multidimensional visual image features corresponding to each time sequence block through the second encoder to obtain a comprehensive feature corresponding to each time sequence block; The decoding subunit is used to decode the comprehensive features corresponding to each time sequence block and the action prediction result of the previous time sequence block through the decoder to obtain the action prediction result of each time sequence block.
7. The robot skill learning device according to claim 6, characterized in that: The decoding subunit is specifically used for: The decoder performs action prediction on the next time sequence block based on the comprehensive features corresponding to each time sequence block and the action prediction result of the previous time sequence block, so as to obtain the initial prediction result of each time sequence block; The initial prediction result of each time series block is smoothed to obtain the action prediction result of each time series block.
8. The robot skill learning device according to claim 6, characterized in that: The execution module is specifically used to: input the first robot control data of the original robot at the current moment into the target block timing imitation learning model to obtain a first action timing block; Controlling each joint of the original robot to move according to the first action sequence block, and acquiring control data of the second robot; Predicting according to the second robot control data and the first action timing block by the target block timing imitation learning model to obtain a second action timing block; The original robot predicts and executes each action sequence block in sequence until the task to be learned is completed.
9. A mechanical arm robot, characterized in that: The mechanical arm robot comprises: a memory and at least one processor, wherein instructions are stored in the memory; The at least one processor calls the instructions in the memory so that the robotic arm robot executes the robot skill learning method as described in any one of claims 1-5.
10. A computer-readable storage medium having instructions stored thereon, characterized in that: When the instructions are read and executed, the robot skill learning method as described in any one of claims 1-5 is executed.
Citation Information
Patent Citations
Mechanical arm navigation obstacle avoidance method and system, computer equipment and storage medium
CN114603564A