Mechanical arm trajectory planning method and device based on generative adversarial imitation learning
Through the improved generative adversarial imitation learning model, combined with the fusion of dynamic reward function and global local information, the problem of difficulty in designing reward function in robotic arm trajectory planning is solved, the efficiency and accuracy of robotic arm operation are improved, and the adaptability and generalization ability of the model are enhanced.
Patent Information
- Application Number
- CN202510487518.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-18
- Publication Date
- 2025-08-01
- Estimated Expiration
- 2045-04-18
AI Technical Summary
In the prior art, the reward function design in robotic arm trajectory planning is difficult to effectively guide strategy optimization in complex tasks, resulting in insufficient operation efficiency and accuracy of robotic arm.
The improved generative adversarial imitation learning model is adopted, and the dynamic reward function is used to dynamically adjust the parameters of the reward function according to the task completion degree, and the global task objectives are combined with local state information to optimize the robotic arm trajectory planning.
It improves the learning efficiency and accuracy of robotic arm trajectory planning, enhances the generalization ability and adaptability of the model, can better deal with complex and mutated tasks, and achieves fine control and efficient operation.
Smart Images

Figure CN120395818A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of robotic arms, and in particular, to a robotic arm trajectory planning method and device based on generative adversarial imitation learning. Background Art
[0002] In many fields such as modern manufacturing, material synthesis, scientific research, and experiments, the efficient and accurate operation of robotic arms is very important. For example, in semiconductor manufacturing, a robotic arm needs to carry tiny chips with high precision; in material synthesis experiments, a robotic arm should accurately complete complex tasks such as reagent addition and sample processing. Trajectory planning, as the core link for a robotic arm to achieve precise operation, directly affects the quality and efficiency of task completion. Therefore, in related technologies, certain progress has been made in robotic arm trajectory planning based on generative adversarial imitation learning, but there are still some key problems. The main manifestation is that it is difficult to design a reward function. If faced with complex tasks, it is very difficult to design an effective reward function, which cannot fully guide policy optimization. Summary of the Invention
[0003] In order to solve the technical problem of the difficult design of the reward function in the prior art, embodiments of the present invention provide a robotic arm trajectory planning method and device based on generative adversarial imitation learning. The technical solutions are as follows:
[0004] On the one hand, a robotic arm trajectory planning method based on generative adversarial imitation learning is provided, and the method includes:
[0005] S1. Collect robotic arm joint state data;
[0006] S2. Input the robotic arm joint state data into a trained improved generative adversarial imitation learning model to obtain a new robotic arm action;
[0007] S3. Collect new robotic arm joint state data according to the new robotic arm action, and input the new robotic arm joint state data into a trained improved generative adversarial imitation learning model to obtain a new robotic arm action;
[0008] S4. Repeat S3 until a preset condition is reached, stop repeating, and obtain the final robotic arm trajectory;
[0009] Wherein, in the improved generative adversarial imitation learning model, the reward function is a dynamic reward function, and the parameters in the dynamic reward function are dynamically adjusted according to the task completion degree.
[0010] On the other hand, a robotic arm trajectory planning device based on generative adversarial imitation learning is provided. The device is applied to the robotic arm trajectory planning method based on generative adversarial imitation learning, and the device includes:
[0011] The acquisition module is used to acquire the mechanical arm joint state data;
[0012] The first processing module is used to input the mechanical arm joint state data into the trained improved generative adversarial imitation learning model to obtain a new mechanical arm motion;
[0013] The second processing module is used to acquire new mechanical arm joint state data according to the new mechanical arm motion, input the new mechanical arm joint state data into the trained improved generative adversarial imitation learning model, and obtain a new mechanical arm motion;
[0014] The loop processing module is used to repeatedly execute the second processing module until a preset condition is reached, stop repeating the execution, and obtain the final mechanical arm trajectory;
[0015] Among them, in the improved generative adversarial imitation learning model, the reward function is a dynamic reward function, and the parameters in the dynamic reward function are dynamically adjusted according to the task completion degree.
[0016] On the other hand, a mechanical arm trajectory planning device based on generative adversarial imitation learning is provided. The mechanical arm trajectory planning device based on generative adversarial imitation learning includes: a processor; a memory, and a computer-readable instruction is stored on the memory. When the computer-readable instruction is executed by the processor, any one of the methods in the above-mentioned mechanical arm trajectory planning method based on generative adversarial imitation learning is implemented.
[0017] On the other hand, a computer-readable storage medium is provided. At least one instruction is stored in the storage medium, and the at least one instruction is loaded and executed by the processor to implement any one of the methods in the above-mentioned mechanical arm trajectory planning method based on generative adversarial imitation learning.
[0018] The beneficial effects brought by the technical solutions provided in the embodiments of the present invention at least include:
[0019] The beneficial effects brought by the technical solutions provided in the embodiments of the present invention at least include: in the improved generative adversarial imitation learning model, the reward function is a dynamic reward function, and the parameters in the dynamic reward function are dynamically adjusted according to the task completion degree. It is beneficial to improve the learning efficiency. There is a significant improvement in stability, which can better imitate the expert trajectory, greatly improving the success rate and accuracy of the mechanical arm trajectory planning based on generative adversarial imitation learning, and laying a solid foundation for the practical application of the mechanical arm in fields such as material synthesis. Description of the Drawings
[0020] To more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the accompanying drawings required for the description of the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can be obtained based on these drawings.
[0021] Figure 1 is a flowchart of a robotic arm trajectory planning method based on generative adversarial imitation learning provided by an embodiment of the present invention;
[0022] Figure 2 is a schematic structural diagram of a robotic arm trajectory planning system based on generative adversarial imitation learning provided by an embodiment of the present invention;
[0023] Figure 3 is another schematic structural diagram of robotic arm trajectory planning based on generative adversarial imitation learning provided by an embodiment of the present invention;
[0024] Figure 4 is a block diagram of a robotic arm trajectory planning device based on generative adversarial imitation learning provided by an embodiment of the present invention;
[0025] Figure 5 is a schematic structural diagram of an electronic device provided by an embodiment of the present invention. Detailed implementation manners
[0026] The following describes the technical solutions in the present invention with reference to the accompanying drawings.
[0027] In the embodiments of the present invention, words such as "exemplarily" and "for example" are used to represent examples, illustrations or explanations. Any embodiment or design solution described as an "example" in the present invention should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Exactly, the use of the word "example" is intended to present concepts in a specific way. In addition, in the embodiments of the present invention, the meaning expressed by "and / or" can be both, or either one of the two.
[0028] In the embodiments of the present invention, "image" and "picture" can sometimes be used interchangeably. It should be noted that when the difference is not emphasized, the meanings they express are the same. "(of)", "corresponding", and "corresponding" can sometimes be used interchangeably. It should be noted that when the difference is not emphasized, the meanings they express are the same.
[0029] In the embodiments of the present invention, sometimes subscripts such as W1 may be written in a non-subscript form such as W1. When the difference is not emphasized, the meanings they express are the same.
[0030] To make the technical problems, technical solutions, and advantages to be solved by the present invention clearer, the following will be described in detail with reference to the accompanying drawings and specific embodiments.
[0031] An embodiment of the present invention provides a robotic arm trajectory planning method based on generative adversarial imitation learning. This method can be implemented by an electronic device, which can be a terminal or a server. As Figure 1 shown in the flowchart of the robotic arm trajectory planning method based on generative adversarial imitation learning, the processing flow of this method can include the following steps:
[0032] S1. Collect robotic arm joint state data.
[0033] In a feasible implementation, relying on rich experience and professional knowledge, an expert precisely controls the robotic arm through a teleoperation device to complete specific tasks, such as coating preparation, part assembly, etc. During the teleoperation process, the state of the robotic arm is sampled and recorded at a fixed interval of 0.05 seconds. By an expert teleoperating the robotic arm to perform tasks, the joint states q k and q k+1 are recorded.
[0034] Among them, the robotic arm joint state data includes but is not limited to joint angles, joint angular velocities, joint angular accelerations, joint torques or forces, and joint positions.
[0035] It can be collected using a data acquisition module, which is mainly responsible for obtaining the demonstration trajectory data of the expert. By connecting to the UR-UTDE 1.5.1 control interface, the state of the collaborative robot is sampled at a fixed interval of 0.05 seconds. It has data storage and management functions, and stores the collected state-action pair data in a local database or file system in a certain format for subsequent preprocessing and algorithm training.
[0036] The data acquisition also includes the following steps:
[0037] 1) Data cleaning, used to detect and process missing values and outliers.
[0038] 2) Data augmentation, expanding the dataset through transformations such as rotation and translation.
[0039] S2. Input the robotic arm joint state data into a trained improved generative adversarial imitation learning model to obtain new robotic arm actions.
[0040] S3. Collect new robotic arm joint state data according to the new robotic arm actions, and input the new robotic arm joint state data into a trained improved generative adversarial imitation learning model to obtain new robotic arm actions.
[0041] S4. Repeat the execution of S3 until a preset condition is reached, then stop the repeated execution to obtain the final trajectory of the robotic arm.
[0042] Among them, the preset condition can be that the robotic arm reaches the final target point.
[0043] Among them, in the improved generative adversarial imitation learning model, the reward function is a dynamic reward function, and the parameters in the dynamic reward function are dynamically adjusted according to the task completion degree.
[0044] The task completion degree is obtained by fusing the global task objective and local state information.
[0045] In a feasible implementation manner, the following is an embodiment of a robotic arm task to illustrate the global task objective, local state information, and how to fuse them to obtain the task completion degree:
[0046] The description of the global task objective is as follows:
[0047] The robotic arm needs to move an object from the initial position A to the target position B and maintain a stable posture after reaching the target position. This is the overall objective of the entire task, which clearly defines the final result that the robotic arm needs to achieve.
[0048] The local state information may include the following:
[0049] Joint angle information: The current angle values of each joint of the robotic arm, which reflect the current shape of the robotic arm.
[0050] End effector position: The actual coordinate position of the end effector of the robotic arm in space, which is used to judge its proximity to the target position.
[0051] Object grasping state: Indicates whether the robotic arm has successfully grasped the object. For example, it is determined through the feedback of the grasping sensor and is a boolean value (True for successful grasping, otherwise False).
[0052] In some embodiments, the fusion calculation of the task completion degree can adopt the following formula:
[0053] The task completion degree C is a value between 0 and 1, where 1 indicates that the task is completely completed and 0 indicates that the task is completely uncompleted. The task completion degree can be fused and calculated through the following formula:
[0054]
[0055] Among them: dmin is the actual distance between the current end effector of the robotic arm and the target position B, and dmax is a preset maximum distance threshold.
[0056] s is the object grasping state. If the object is successfully grasped, s = 1; otherwise, s = 0.
[0057] θ is the current attitude angle of the end effector of the robotic arm.
[0058] θ target is the target attitude angle, and θ max is the maximum allowable angle deviation.
[0059] (1 - |θ - θ target | / θ max ) represents the degree of proximity between the current attitude and the target attitude. When the attitudes are exactly the same, this value is 1.
[0060] w1, w2, and w3 are weight coefficients, which are set according to the key points and requirements of the task, and w1 + w2 + w3 = 1. For example, if more emphasis is placed on position accuracy, w1 can be set larger.
[0061] Through such a fusion method, it is possible to comprehensively consider the global task goal and local state information, obtain a value reflecting the degree of task completion, provide a basis for the dynamic reward function, so that the generative adversarial learning simulation model can learn and optimize according to the actual progress of the task.
[0062] In the technical solution of the present invention, in the improved generative adversarial imitation learning model, the reward function is a dynamic reward function, and the parameters in the dynamic reward function are dynamically adjusted according to the task completion degree. The task completion degree is obtained by fusing the global task goal and local state information.
[0063] The use of a dynamic reward function in the improved generative adversarial imitation learning model has the following effects:
[0064] First, it can improve the learning efficiency.
[0065] Guide rapid convergence. The dynamic reward function can adjust the parameters in real time according to the task completion degree, providing more accurate feedback for the model. For example, in a robotic arm task, when the robotic arm approaches the target position, the reward function will give a higher reward, guiding the model to converge towards the target faster, reducing the number of learning iterations, and improving the learning efficiency.
[0066] Adapt to complex tasks. For complex tasks, the fusion of the global task goal and local state information allows the model to comprehensively understand the task progress. For example, in a multi-step robotic arm operation task, not only considering the final position of the end effector but also fusing local information such as joint angles can enable the model to learn the correct execution method of each subtask faster and accelerate the learning of the overall task.
[0067] Second, it can enhance the generalization ability of the model.
[0068] In response to environmental changes, in different environmental conditions or task scenarios, by dynamically adjusting the reward function, the model can flexibly adapt to changes based on the current local state information and global task objectives. For example, when an obstacle appears in the working space of a robotic arm, the calculation of task completion based on local state information enables the model to promptly adjust its strategy, find a path to avoid the obstacle and complete the task, rather than being limited to the previously trained fixed pattern, thereby enhancing the model's generalization ability in different environments.
[0069] Handling task variations, the model can also handle different variations of the same type of task well. Since the reward function is dynamically adjusted according to the global objective of the specific task and the current state, even if some details of the task change, such as the change of the target position or the addition of new constraints to the task, the model can still adapt to the new task requirements through learning without the need for a large amount of retraining.
[0070] Thirdly, it can improve the quality of task completion.
[0071] Optimizing the intermediate process, the dynamic reward function focuses on the intermediate process of task completion. Through the feedback of local state information, the model can continuously optimize the execution of each intermediate step. Taking the robotic arm grasping an object as an example, the reward function not only gives a reward when the robotic arm successfully grasps and places the object at the target position, but also gives appropriate rewards during the grasping process according to local information such as the motion state of the robotic arm joints and the grasping stability of the object, prompting the model to learn a better grasping strategy and improve the quality of task completion.
[0072] Achieving fine control: The integration of the global task objective and local state information enables the model to achieve more fine control. In some tasks with high precision requirements, such as the robotic arm performing high-precision assembly work, the model can continuously adjust the motion of the robotic arm according to local state information such as the real-time joint angles and the position accuracy of the end effector to achieve a higher quality of task completion.
[0073] In some embodiments, the training of the improved generative adversarial imitation learning model may further include the following steps:
[0074] Collection and preprocessing of robotic arm joint state data.
[0075] The preprocessing includes: normalizing the collected robotic arm joint state data, using the min-max normalization method to map the data to the interval [0,1].
[0076] In some embodiments, the specific formula is as follows:
[0077]
[0078] where x is the original data, x min and x max are the minimum and maximum values of the data respectively, and x norm is the normalized data.
[0079] In this embodiment, during the teleoperation process, the state of the robotic arm is sampled and recorded at a fixed interval of 0.05 seconds. The collected data often has different dimensions and value ranges, which will have an adverse impact on the training of subsequent algorithms. Therefore, it is necessary to normalize the data. Common normalization methods such as min-max normalization can eliminate the dimensional differences by mapping the data to the interval [0,1], making the data comparable, thereby improving the training efficiency and stability of the algorithm.
[0080] In some embodiments, the improved generative adversarial imitation learning model includes: a generator and a discriminator.
[0081] The generator adopts a three-layer fully connected network, the hidden layer is set to (256×128), and the activation function is selected as the tanh function.
[0082] The discriminator adopts a three-layer fully connected network, and the hidden layer is (256×64).
[0083] In a feasible implementation, in the Ubuntu 20.04 and ROS (Robot Operating System) environment, a virtual laboratory is built using the Gazebo 11 simulator. The hidden layer of the generator is set to 256×128, and the hidden layer of the discriminator is 256×64. The activation function of both is the tanh function. The tanh function can map the input value to the interval [-1,1] and has good non-linear characteristics, which helps the generator learn complex trajectory patterns. The learning rate of the generator is set to 0.001, and the learning rate of the discriminator is set to 0.0001. The parameters of the generator and the discriminator are updated alternately. The main role of the discriminator is to distinguish the generated trajectory from the expert demonstration trajectory. Through a suitable network structure and activation function, its discrimination ability can be improved.
[0084] Different learning rate settings can achieve a better balance between the generator and the discriminator during the training process. Unrolling steps: The discriminator adopts multi-step iterative update, and the unrolling steps (K = 5). Through this unrolled training, the foresight of the generator is enhanced, enabling it to better meet complex task requirements.
[0085] The discriminator is updated according to the formula and the generator is updated according to the improved optimization objective Update. To enhance the foresight of the generator, the discriminator is updated through multiple-step iteration (K = 5). During the training process, the loss value of the discriminator, the reward of the generator, and the trajectory diversity metrics (such as trajectory coverage) are monitored in real time.
[0086] In some embodiments, the reward function of the improved generative adversarial imitation learning model is:
[0087] r new (s,a) = α(log(D(s,a)) - log(1 - D(s,a))) + β, where β is dynamically adjusted according to the task completion degree;
[0088] Among them, s is the state used as the input of the generator, that is, the joint state of the robotic arm; a is the output of the generator, that is, the joint action that the robotic arm is about to execute; α is the coefficient of the generator's pseudo-reward function, which is used to adjust the proportion of the reward; D represents the discriminator, which is used to distinguish between the expert trajectory and the false trajectory generated by the generator, and D(s,a) represents the output of the false trajectory through the discriminator; β represents the reward based on the task completion degree.
[0089] In this embodiment, during the training process, the global task goal and the local state information are fused through the real-time task completion monitoring module to obtain the real-time task completion degree, improving the decision-making ability of the policy.
[0090] Dynamic reward calculation: An improved reward function is adopted:
[0091] r new (s,a) = α(log(D(s,a)) - log(1 - D(s,a))) + β, where β is dynamically adjusted according to the task completion degree. During the training process, the task completion situation is monitored in real time. When the robot operation is approaching the task goal, the value of β increases, giving more rewards to the generator; otherwise, the value of β decreases.
[0092] Among them, (α > 1) is the gradient enhancement factor, which is used to enhance the intensity of the gradient signal. This dynamic reward mechanism can more effectively guide the generator to learn a better policy. By monitoring the task completion degree in real time, the value of β is dynamically adjusted to encourage the generator to learn a better policy.
[0093] In some embodiments, the value of β is calculated by the following formula:
[0094] β = β0 + γ·TaskProgress;
[0095] Among them, β0 is the initial reward offset, γ is the progress weight coefficient, and TaskProgress is the current task completion degree (0 - 1).
[0096] In some embodiments, the calculation formula of the TaskProgress is as follows:
[0097]
[0098] where m is the total number of segments of the complex multi-segment task, and i is the currently executed segment number; s goal is the target state of the current task stage; s cur is the current state, s0 is the initial state of the current task stage; s0 is the local state information.
[0099] In some embodiments, the labels of the discriminator are smoothed, the true label is adjusted from 1 to (1 - ∈), and the generated label is adjusted from 0 to ∈, where ∈ ∈ [0, 0.1];
[0100] The L2 norm of the discriminator gradient is clipped, and the threshold is 5.0;
[0101] where ∈ is a parameter used to adjust the label smoothing degree. When ∈ = 0, it means no label smoothing is performed. To prevent gradient explosion and ensure the stability of the training process.
[0102] In some embodiments, in the Generative Adversarial Imitation Learning (GAIL) algorithm, the policy plays the role of the generator, aiming to confuse the discriminator by generating samples indistinguishable from the state-action pairs samples generated by the expert policy, while the reward function acts as the discriminator, attempting to accurately distinguish these samples. The regularized minmax formula of GAIL is expressed as minθmaxωV(θ, ω), where V(θ, ω) is:
[0103]
[0104] where (s, a) represents the state-action pair sample, that is, the action a taken in the state s. The policy π θ plays the role of the generator, while the reward function D ω plays the role of the discriminator. They are both represented in the form of deep neural networks and have parameters θ and ω respectively. The causal entropy of the policy π θ , denoted as H(π θ ), is used as a regularization term together with the hyperparameter λH. When updating, the discriminator is updated first:
[0105]
[0106] Then the generator is updated:
[0107]
[0108] In the above-mentioned expansion generative adversarial imitation learning training strategy, the foresight of the generator is enhanced through certain improvements. Specifically, the objective function of the discriminator remains:
[0109]
[0110] The parameter update mode is to continuously update k times through gradient descent.
[0111]
[0112] The optimization objective of the generator is modified to:
[0113]
[0114] Specifically, during the update process of the generator, it not only considers the current state of the generator, but also considers the state of the discriminator after K updates, and combines these two parts of information to derive the optimal solution. The change of the gradient is as follows:
[0115]
[0116] Among them, the first term represents the gradient calculated in the standard GAIL formula, and the second term is an additional part that considers the state of the discriminator after K updates.
[0117] The Bayesian optimization algorithm adopts a Gaussian process regression model, and the objective function is: Objective(α) = SuccessRate(α), where SuccessRate is the success rate of the robotic arm task.
[0118] In some embodiments, during the training process, the loss value of the discriminator, the reward of the generator, and the Euclidean distance between the trajectory and the expert trajectory are monitored in real time. According to these monitoring metrics, the training parameters, such as the learning rate, the expansion step number, etc., are adjusted in a timely manner to ensure that the algorithm can converge stably.
[0119] In some embodiments, virtual-real migration and optimization. Domain randomization. The training is carried out in a simulation environment, and the domain randomization technology is adopted to enhance the adaptability of the model to the real environment. Specifically, it includes randomizing the friction coefficient (range: 0.05 - 0.2), the detection error (±0.005m), etc. By introducing these random factors in the simulation environment, the model can learn more robust strategies during the training process.
[0120] Model Deployment and Fine-tuning. Deploy the model trained in the simulation environment to the real robotic arm. Since there are certain differences between the simulation environment and the real environment, the model needs to be fine-tuned. The Bayesian optimization algorithm is used to adjust the parameter α of the reward function, with the task success rate (target ≥ 95%) as the optimization goal. The fine-tuning step size is set to 1 / 10 of the original training step size. By gradually adjusting the parameters, the model can better adapt to the real environment.
[0121] After the model is fine-tuned in the real environment, according to the current task requirements and environmental information, use the trained generator to generate the trajectory of the robotic arm. The generator will comprehensively consider the global task goal and local state information to generate an optimal trajectory. Then, send the generated trajectory to the control system of the robotic arm, and the robotic arm executes the task according to the planned trajectory. During the execution process, the state of the robotic arm and the task completion situation are monitored in real time, and adjustments are made in a timely manner if deviations are found. After the task is completed, feedback and optimization are performed on the model according to the actual execution results. If the task success rate is low or the trajectory has a large deviation, analyze the reasons and adjust the parameters of the model or retrain it to continuously improve the performance of the robotic arm trajectory planning.
[0122] The present invention proposes a robotic arm trajectory planning method based on an improved expanded Generative Adversarial Imitation Learning (GAIL) algorithm, which breaks through the bottleneck of the existing technology through the following innovative points. First, compared with the traditional GAIL method, combined with the idea of expanded Generative Adversarial Network (GAN), the expanded GAIL algorithm is used for training to alleviate the mode collapse problem in the GAIL algorithm. Second, the reward mechanism is improved, and a dynamically adjusted reward function is designed. First, optimize the discriminator feedback to alleviate the gradient vanishing problem, and then add a dynamic reward determined by the task completion degree to accelerate the training speed and improve the policy upper limit. Third, enhance the generalization ability. Through label smoothing and gradient clipping techniques, the risk of overfitting of the model to expert data is reduced, and the adaptability of the policy in unknown scenarios is improved. Fourth, virtual-real migration is achieved through domain randomization and Bayesian optimization to realize efficient virtual-real migration.
[0123] The present invention is based on an improved expanded Generative Adversarial Imitation Learning (GAIL) algorithm, aiming to provide an innovative solution for robotic arm trajectory planning based on generative adversarial imitation learning, improve the operating performance of the robotic arm in different scenarios, promote the development of automation and intelligence in related fields, and also lay a foundation for the expansion and application of robot technology in more emerging fields.
[0124] Please refer to Figures 2-3, this embodiment provides a robotic arm planning system for material synthesis. The system includes: a data acquisition module, which is responsible for collecting expert demonstration data, providing high-quality samples for algorithm training, and solving the problem of data sparsity. An algorithm engine module, which is the core module and generates trajectories through an improved expanded GAIL algorithm. A virtual experiment verification module, which verifies the algorithm performance in a simulation environment and enhances the model robustness through domain randomization methods. A human-computer interaction module, which provides a user interface, supports experts to remotely operate the robotic arm and displays the trajectory in real time. Visualized training metrics (such as loss curves and trajectory comparisons), which assist in parameter adjustment (such as Bayesian optimization) to achieve intuitive monitoring of algorithm training and task execution. Data acquisition drives algorithm training, virtual experiment verification optimizes the model generalization ability, and human-computer interaction realizes real-time monitoring and parameter tuning, forming a closed-loop system from data to decision-making, significantly improving the accuracy and robustness of robotic arm trajectory planning based on generative adversarial imitation learning. In a material synthesis laboratory, through the robotic arm trajectory planning method based on generative adversarial imitation learning proposed in the present invention, the leap from manual operation to intelligent and high-precision automation can be achieved, significantly improving the experimental efficiency and result reliability, and providing key technical support for the research and development of new materials.
[0125] The block diagram of a robotic arm trajectory planning device 4 provided by an embodiment of the present invention, which is used for a robotic arm trajectory planning method based on generative adversarial imitation learning. Refer to Figure 4 , the device includes:
[0126] An acquisition module 410, which is used to acquire robotic arm joint state data;
[0127] A first processing module 420, which is used to input the robotic arm joint state data into a trained improved generative adversarial imitation learning model to obtain new robotic arm actions;
[0128] A second processing module 430, which is used to acquire new robotic arm joint state data according to the new robotic arm actions, input the new robotic arm joint state data into a trained improved generative adversarial imitation learning model to obtain new robotic arm actions;
[0129] A loop processing module 440, which is used to repeatedly execute the second processing module until a preset condition is reached, stop repeating, and obtain the final robotic arm trajectory;
[0130] Among them, in the improved generative adversarial imitation learning model, the reward function is a dynamic reward function, and the parameters in the dynamic reward function are dynamically adjusted according to the task completion degree.
[0131] Figure 5 is the structural schematic diagram of an electronic device provided by an embodiment of the present invention, such as Figure 5As shown, the electronic device may include the above-mentioned robotic arm trajectory planning device based on generative adversarial imitation learning. Optionally, the electronic device 510 may include a first processor 2001.
[0132] Optionally, the electronic device 510 may further include a memory 2002 and a transceiver 2003.
[0133] Among them, the first processor 2001 is connected to the memory 2002 and the transceiver 2003, for example, through a communication bus.
[0134] Next, in combination with Figure 5 Specific introductions will be made to the various components of the electronic device 510:
[0135] Among them, the first processor 2001 is the control center of the electronic device 510, which can be a single processor or a collective term for multiple processing elements. For example, the first processor 2001 is one or more central processing units (CPUs), or can be an application specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present invention, for example: one or more digital signal processors (DSPs), or one or more field programmable gate arrays (FPGAs).
[0136] Optionally, the first processor 2001 can execute various functions of the electronic device 510 by running or executing software programs stored in the memory 2002 and calling data stored in the memory 2002.
[0137] In a specific implementation, as an embodiment, the first processor 2001 may include one or more CPUs, such as Figure 5 CPU0 and CPU1 shown in
[0138] In a specific implementation, as an embodiment, the electronic device 510 may also include multiple processors, such as Figure 5 the first processor 2001 and the second processor 2004 shown in
[0139] Among them, the memory 2002 is used to store the software program for implementing the solution of the present invention, and is controlled by the first processor 2001 to execute. The specific implementation manner can refer to the above method embodiment and will not be elaborated here.
[0140] Optionally, the memory 2002 may be a read-only memory (ROM) or other type of static storage device that can store static information and instructions, a random access memory (RAM) or other type of dynamic storage device that can store information and instructions, or may also be an electrically erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM) or other optical disc storage, optical disc storage (including compact discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), a magnetic disk storage medium or other magnetic storage device, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto. The memory 2002 may be integrated with the first processor 2001 or may exist independently and is coupled to the first processor 2001 through the interface circuit of the electronic device 510 ( Figure 5 not shown in the figure), and the embodiments of the present invention do not make specific limitations on this.
[0141] The transceiver 2003 is used to communicate with a network device or with a terminal device.
[0142] Optionally, the transceiver 2003 may include a receiver and a transmitter ( Figure 5 not shown separately in the figure). Among them, the receiver is used to implement the receiving function, and the transmitter is used to implement the sending function.
[0143] Optionally, the transceiver 2003 may be integrated with the first processor 2001 or may exist independently and is coupled to the first processor 2001 through the interface circuit of the electronic device 510 ( Figure 5 not shown in the figure), and the embodiments of the present invention do not make specific limitations on this.
[0144] It should be noted that Figure 5 the structure of the electronic device 510 shown in the figure does not constitute a limitation on the router. The actual knowledge structure recognition device may include more or fewer components than shown in the figure, or combine certain components, or have different component arrangements.
[0145] In addition, for the technical effects of the electronic device 510, reference may be made to the technical effects of the robotic arm trajectory planning method based on generative adversarial imitation learning described in the foregoing method embodiments, which will not be elaborated herein.
[0146] It should be understood that the first processor 2001 in the embodiments of the present invention may be a central processing unit (CPU), and this processor may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or this processor may also be any conventional processor, etc.
[0147] It should also be understood that the memory in the embodiments of the present invention may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable ROM (PROM), an erasable PROM (EPROM), an electrically erasable PROM (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of random access memory (RAM) are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced SDRAM (ESDRAM), synchlink DRAM (SLDRAM), and direct rambus RAM (DR RAM).
[0148] The above embodiments can be implemented in whole or in part by software, hardware (such as circuits), firmware, or any other combination. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, the processes or functions described in the embodiments of the present invention are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (such as infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or a data center that contains one or more collections of available media. The available medium can be a magnetic medium (such as a floppy disk, hard disk, or magnetic tape), an optical medium (such as a DVD), or a semiconductor medium. The semiconductor medium can be a solid-state drive.
[0149] It should be understood that the term "and / or" in this document is merely a description of the association relationship between associated objects, indicating that three relationships can exist. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. Here, A and B can be singular or plural. Additionally, the character " / " in this document generally represents an "or" relationship between the associated objects before and after, but it may also represent an "and / or" relationship, which can be specifically understood by referring to the context before and after.
[0150] In the present invention, "at least one" means one or more, and "a plurality" means two or more. "At least one of the following" or its similar expressions refer to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b, or c can represent: a, b, c, a - b, a - c, b - c, or a - b - c, where a, b, and c can be single or plural.
[0151] It should be understood that in various embodiments of the present invention, the magnitudes of the sequence numbers of the above processes do not imply the order of execution. The order of execution of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present invention.
[0152] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.
[0153] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the devices, apparatuses, and units described above can refer to the corresponding processes in the foregoing method embodiments, and will not be described herein again.
[0154] In several embodiments provided by the present invention, it should be understood that the disclosed devices, apparatuses, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of the devices or units can be electrical, mechanical, or other forms.
[0155] The units described as separate components may or may not be physically separated. The components shown as units may or may not be physical units, that is, they can be located in one place, or can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0156] In addition, in each embodiment of the present invention, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit.
[0157] When the above-mentioned function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs.
[0158] The above is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of changes or substitutions, which should all be covered by the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.
Claims
1. A robotic arm trajectory planning method based on generative adversarial imitation learning, characterized in that The method includes: S1. Collect the mechanical arm joint state data; S2. Input the mechanical arm joint state data into the trained improved generative adversarial imitation learning model to obtain a new mechanical arm motion; S3. Collect new mechanical arm joint state data according to the new mechanical arm motion, and input the new mechanical arm joint state data into the trained improved generative adversarial imitation learning model to obtain a new mechanical arm motion; S4. Repeat S3 until a preset condition is reached, stop repeating, and obtain the final mechanical arm trajectory; Among them, in the improved generative adversarial imitation learning model, the reward function is a dynamic reward function, and the parameters in the dynamic reward function are dynamically adjusted according to the task completion degree.
2. The robotic arm trajectory planning method based on generative adversarial imitation learning according to claim 1, wherein The training process of the improved generative adversarial imitation learning model includes: Collect and preprocess the mechanical arm joint state data; The preprocessing includes: normalizing the collected mechanical arm joint state data, using the min-max normalization method to map the data to the interval [0, 1].
3. The robotic arm trajectory planning method based on generative adversarial imitation learning according to claim 1, characterized in that The improved generative adversarial imitation learning model includes: a generator and a discriminator; The generator adopts a three-layer fully connected network, the hidden layer is set to (256×128), and the activation function selects the tanh function; The discriminator adopts a three-layer fully connected network, and the hidden layer is (256×64).
4. The robotic arm trajectory planning method based on generative adversarial imitation learning according to claim 3, wherein The reward function of the improved generative adversarial imitation learning model is as follows in formula (1): r new (s,a) = α(log(D(s,a)) - log(1 - D(s,a))) + β(1) Among them, β is dynamically adjusted according to the task completion degree; Among them, s is the state used as the input of the generator, that is, the joint state of the mechanical arm; a is the output of the generator, that is, the joint motion that the mechanical arm is about to execute; α is the coefficient of the generator's pseudo-reward function, used to adjust the proportion of the reward; D represents the discriminator, and the discriminator is used to distinguish the expert trajectory and the false trajectory generated by the generator, and D(s, a) represents the output of the false trajectory through the discriminator; β represents the reward according to the task completion degree.
5. The method for robotic arm trajectory planning based on generative adversarial imitation learning according to claim 4, wherein The β value is calculated by the following formula (2): β = β0 + γ·TaskProgress(2) Among them, β0 is the initial reward offset, γ is the progress weight coefficient, and TaskProgress is the current task completion degree.
6. The method for robotic arm trajectory planning based on generative adversarial imitation learning according to claim 5, wherein, The calculation formula of the TaskProgress is as follows in formula (3): Among them, m is the total number of segments of the complex multi-segment task, and i is the currently executed segment number; s goal is the target state of the current task stage; s cur is the current state, and s0 is the initial state of the current task stage. s0 is the local state information.
7. The method for robotic arm trajectory planning based on generative adversarial imitation learning according to claim 3, wherein Smooth the labels of the discriminator, adjust the true label from 1 to (1 - ∈), and adjust the generated label from 0 to ∈, ∈ ∈ [0, 0.1]; Clip the discriminator gradient by the L2 norm, and the threshold is 5.0; Among them, ∈ is a parameter used to adjust the label smoothing degree. When ∈ = 0, it means that no label smoothing processing is performed.
8. A robotic arm trajectory planning device based on generative adversarial imitation learning, the robotic arm trajectory planning device based on generative adversarial imitation learning is used to implement the robotic arm trajectory planning method based on generative adversarial imitation learning according to any one of claims 1-7, characterized in that, The device includes: A collection module, used to collect the mechanical arm joint state data; A first processing module, used to input the mechanical arm joint state data into the trained improved generative adversarial imitation learning model to obtain a new mechanical arm motion; A second processing module, used to collect new mechanical arm joint state data according to the new mechanical arm motion, and input the new mechanical arm joint state data into the trained improved generative adversarial imitation learning model to obtain a new mechanical arm motion; A loop processing module is used to repeatedly execute the second processing module until a preset condition is reached, and then stop the repeated execution to obtain the final manipulator trajectory; Among them, in the improved generative adversarial imitation learning model, the reward function is a dynamic reward function, and the parameters in the dynamic reward function are dynamically adjusted according to the task completion degree.
9. A robotic arm trajectory planning device based on generative adversarial imitation learning, characterized in that, The manipulator trajectory planning device based on generative adversarial imitation learning includes: A processor; A memory, on which computer-readable instructions are stored. When the computer-readable instructions are executed by the processor, the method described in any one of claims 1 to 7 is implemented.
10. A computer-readable storage medium, characterized in that, Program code is stored in the computer-readable storage medium, and the program code can be called by the processor to execute the method described in any one of claims 1 to 7.
Citation Information
Patent Citations
An industrial control system malicious sample generation method based on adversarial learning
CN109902709A
Adversarial imitation learning method and device based on state trajectory
CN111856925A
Robot learning method for model-based generative adversarial interactive imitation learning
CN116663651A
Industrial robot trajectory optimization control method based on intelligent algorithm
CN116901086A
Mechanical arm operation track optimization method based on generative adversarial network
CN119820569A
Cited By
Body intelligent decision-making control method and system based on joint state of mechanical arm
CN120620228A
A method and system for embodied intelligent decision control based on mechanical arm joint state
CN120620228B
Robot motion control strategy network training method and device based on imitation learning, robot motion control method and device, equipment, robot and storage medium
CN121403422A