A reinforcement learning driven trajectory adaptive optimization method for robotic arms

CN122807945APending Publication Date: 2026-09-25WUHAN INST OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611285262.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-08-24
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

[0003]针对现有智能轨迹优化技术难以适应工件实际负载变化、无法准确定位高风险局部轨迹段以及候选轨迹评价与当前运行状态匹配度不足的问题,本发明提出一种强化学习驱动的机械臂轨迹自适应优化方法

Benefits of technology

[0012]本发明通过采集机械臂执行抬升轨迹段时的关节位置、关节速度、驱动力矩和末端位姿,并根据关节控制指令与实际运动响应之间的偏差生成负载响应状态,实现了对机械臂与待转运工件组合动力学特性的在线表征,提升了对工件重量、重心位置及惯量参数变化的适应能力;同时,根据负载响应状态预测未执行轨迹段的关节位置余量、关节速度余量、驱动力矩余量和障碍物间距余量,实现了对高风险局部轨迹段的准确识别,解决了现有智能轨迹优化技术对全部剩余轨迹进行整体修正而导致计算量较大、优化针对性不足的问题,增强了机械臂轨迹优化的实时性和准确性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122807945A_ABST
    Figure CN122807945A_ABST
Patent Text Reader

Abstract

The application relates to the technical field of robot intelligent control, and discloses a trajectory self-adaptive optimization method of a mechanical arm driven by reinforcement learning. A reference joint trajectory is generated by acquiring a grabbing pose, a target placing pose and an obstacle position; a load response state is constructed according to joint positions, joint speeds, driving torques and an end pose in a mechanical arm execution process, and a target trajectory segment is determined in combination with various motion margins; an actual state of the mechanical arm is synchronized to a digital twin, a candidate correction trajectory is generated by using a trajectory correction strategy network, a target correction trajectory is determined through constraint screening and multidimensional evaluation, and subsequent trajectories are continuously optimized according to actual execution results. The application improves the trajectory safety, obstacle avoidance reliability and self-adaptive capacity of the mechanical arm under the condition of load change.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of robot intelligent control technology, and in particular to a reinforcement learning-driven adaptive optimization method for robotic arm trajectory. Background Technology

[0002] With the intelligent development of industrial robots, robotic arm trajectory optimization technologies based on reinforcement learning, digital twins, and state prediction have been widely applied in workpiece handling scenarios. Existing intelligent technologies typically generate or correct trajectories based on preset dynamic parameters. When the actual weight and center of gravity of the workpiece deviate from the preset parameters, problems such as increased driving torque, trajectory tracking deviation, and insufficient obstacle avoidance distance can easily occur. Furthermore, existing technologies usually optimize the remaining trajectory as a whole, making it difficult to accurately identify high-risk local trajectory segments based on the actual operating state of the robotic arm. Moreover, the evaluation methods for candidate trajectories are relatively fixed, making it difficult to simultaneously consider trajectory tracking, operational smoothness, driving torque safety, and obstacle avoidance safety. Therefore, a method is needed that can combine actual load response and digital twin simulation to perform local, adaptive optimization of the robotic arm trajectory. Summary of the Invention

[0003] To address the problems of existing intelligent trajectory optimization technologies, such as difficulty in adapting to actual workpiece load changes, inability to accurately locate high-risk local trajectory segments, and insufficient matching degree between candidate trajectory evaluation and current operating state, this invention proposes a reinforcement learning-driven adaptive optimization method for robotic arm trajectories. This method constructs a load response state based on the joint motion response deviation when the robotic arm executes a lifting trajectory segment. It then combines joint position margin, joint velocity margin, driving torque margin, and obstacle spacing margin to determine the target trajectory segment to be optimized from the unexecuted trajectory. A common simulation initial state consistent with the actual robotic arm is set in the robotic arm's digital twin. Multiple candidate trajectories with different correction directions are generated using a trajectory correction strategy network. The target corrected trajectory is determined through motion constraint screening and multi-dimensional trajectory evaluation. The weights of evaluation dimensions such as trajectory tracking, running smoothness, driving torque, and obstacle spacing are dynamically adjusted based on the load response state to ensure the trajectory selection result matches the current combined dynamic state of the robotic arm. By replacing the corresponding local trajectory segment with the target corrected trajectory and continuously updating the load response state using actual execution results, segmented iterative optimization of the robotic arm's trajectory is achieved. This improves the trajectory execution safety, obstacle avoidance reliability, and motion adaptability of the robotic arm when load parameters deviate.

[0004] This invention proposes a reinforcement learning-driven adaptive optimization method for robotic arm trajectories, which includes the following steps:

[0005] Step S1: Obtain the gripping pose, target placement pose, and obstacle position when the robotic arm is holding the workpiece to be transferred. Based on the gripping pose, target placement pose, obstacle position, and the kinematic constraints of the robotic arm, generate the current reference joint trajectory.

[0006] Step S2: Generate joint control commands based on the current reference joint trajectory, control the robotic arm to execute the lifting trajectory segment in the current reference joint trajectory, collect the joint position, joint speed, driving torque and end pose during the execution of the robotic arm, and generate a load response state characterizing the dynamic state of the combination of the robotic arm and the workpiece to be transferred based on the motion response deviation between the joint control commands and the collected results.

[0007] Step S3: Based on the load response status, determine the joint position margin, joint velocity margin, driving torque margin, and obstacle spacing margin corresponding to the unexecuted trajectory segment of the current reference joint trajectory, and determine the target trajectory segment from the unexecuted trajectory segment based on each margin.

[0008] Step S4: When the robotic arm moves to the starting position of the target trajectory segment, the actual motion state of the robotic arm, the pose of the workpiece to be transferred, and the position of the obstacle are synchronized to the digital twin of the robotic arm after load correction to obtain a common simulation initial state; the common simulation initial state, the load response state, and the target trajectory segment are input into the trajectory correction strategy network, and the trajectory correction strategy network generates multiple candidate correction trajectories under the same common simulation initial state to obtain a set of candidate correction trajectories;

[0009] Step S5: Using the digital twin of the robotic arm that has completed load correction, simulate and execute each candidate correction trajectory in the candidate correction trajectory set. Constraint and filter each candidate correction trajectory based on joint position limitations, joint speed limitations, driving torque limitations, and collision distance limitations to obtain an executable correction trajectory set. Perform multi-dimensional trajectory evaluation on each executable correction trajectory in the executable correction trajectory set, and determine the multi-dimensional representative credit result using a trajectory-by-trajectory elimination comparison method. Weight the multi-dimensional representative credit result according to the load response state to obtain the branch evaluation result. Update the trajectory correction strategy network using the branch evaluation result, calculate the selection probability corresponding to each executable correction trajectory using the updated trajectory correction strategy network, and determine the executable correction trajectory with the highest selection probability as the target correction trajectory.

[0010] Step S6: Replace the corresponding target trajectory segment in the current reference joint trajectory with the target correction trajectory, control the robotic arm to execute the replaced current reference joint trajectory, collect the actual execution results corresponding to the target correction trajectory, and update the load response status according to the actual execution results, so as to determine the target trajectory segment that needs to be further optimized from the subsequent unexecuted trajectory segments of the current reference joint trajectory, until the workpiece to be transferred is moved to the target placement pose.

[0011] By adopting the above solution, the beneficial effects achieved by the present invention are as follows:

[0012] This invention collects joint positions, joint velocities, driving torques, and end-effector poses of a robotic arm during the lifting trajectory segment. It then generates a load response state based on the deviation between joint control commands and the actual motion response. This enables online characterization of the combined dynamic characteristics of the robotic arm and the workpiece to be transferred, improving adaptability to changes in workpiece weight, center of gravity position, and inertia parameters. Simultaneously, based on the load response state, it predicts joint position margins, joint velocity margins, driving torque margins, and obstacle spacing margins for unexecuted trajectory segments. This allows for accurate identification of high-risk local trajectory segments, solving the problem of excessive computation and insufficient optimization specificity caused by the overall correction of all remaining trajectories in existing intelligent trajectory optimization technologies. This enhances the real-time performance and accuracy of robotic arm trajectory optimization.

[0013] This invention synchronizes the actual motion state of the robotic arm, the pose of the workpiece to be transferred, and the position of obstacles to a digital twin. A trajectory correction strategy network then generates multiple candidate correction trajectories under the same initial simulation state, enabling the generation and comparison of candidate trajectories on a unified basis. Through collaborative iterative correction of the candidate correction values, different candidate trajectories remain within the high-probability region output by the strategy network while exhibiting different trajectory correction directions, thus improving the effectiveness and diversity of candidate trajectories. Furthermore, the digital twin simulates and executes each candidate trajectory, and constraints are applied based on joint position, joint velocity, driving torque, and collision distance limitations for selection. This addresses the risks of joint overreach, excessive driving torque, and collisions that arise from directly applying unverified reinforcement learning-based trajectories to actual robotic arms, enhancing the executability and operational safety of the target correction trajectory.

[0014] This invention achieves a refined evaluation of the actual contribution of candidate trajectories by performing multi-dimensional evaluations of executable correction trajectories, including trajectory tracking, trajectory smoothing, driving torque, and obstacle spacing. It employs a trajectory-by-trajectory elimination comparison method to determine the representative credit of each trajectory across different evaluation dimensions. The weights of each evaluation dimension are dynamically adjusted based on the load response state and trajectory constraint margin, ensuring that the selection of the target correction trajectory matches the current operating state of the robotic arm. This improves trajectory tracking accuracy, motion smoothness, joint driving torque safety margin, and obstacle avoidance reliability. Simultaneously, by replacing the corresponding local trajectory segment with the target correction trajectory and continuously updating the load response state using actual execution results, it achieves closed-loop segmented iterative optimization of the robotic arm's subsequent trajectories. This solves the problem that a single trajectory correction cannot continuously adapt to changes in load state, enhancing the robotic arm's adaptive control capability and continuous operation stability in complex transport environments. Attached Figure Description

[0015] Figure 1This is a flowchart illustrating a reinforcement learning-driven adaptive optimization method for robotic arm trajectory proposed in this invention.

[0016] Figure 2 This is a schematic diagram illustrating the principle of generating branch evaluation results as proposed in Embodiment 5 of the present invention. Detailed Implementation

[0017] The technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort are within the scope of protection of the present invention.

[0018] Example 1, as Figure 1 As shown, this invention provides a reinforcement learning-driven adaptive optimization method for robotic arm trajectories, which includes the following steps:

[0019] Step S1: Obtain the gripping pose, target placement pose, and obstacle position when the robotic arm is holding the workpiece to be transferred; use the inverse kinematics of the robotic arm to solve for the starting joint configuration and target joint configuration corresponding to the gripping pose and target placement pose, respectively; based on the obstacle position, construct an obstacle-enclosed area in the robotic arm's workspace according to the obstacle's shape and size and a preset safety distance, and plan the joint path between the starting joint configuration and the target joint configuration while meeting the joint position constraints and obstacle-enclosed area avoidance conditions; parameterize the joint path in time according to the joint speed limit, and divide the time-parameterized joint path into multiple trajectory segments according to the trajectory function to obtain the current reference joint trajectory; the multiple trajectory segments include at least a lifting trajectory segment, a transfer transition trajectory segment, an obstacle avoidance trajectory segment, a target approach trajectory segment, and a placement trajectory segment connected in sequence;

[0020] Step S2: Generate joint control commands based on the current reference joint trajectory, control the robotic arm to execute the lifting trajectory segment in the current reference joint trajectory, collect the joint position, joint speed, driving torque and end pose during the execution of the robotic arm, and generate a load response state characterizing the dynamic state of the combination of the robotic arm and the workpiece to be transferred based on the motion response deviation between the joint control commands and the collected results.

[0021] Step S3: Identify the unexecuted trajectory segments after lifting, including the transfer transition trajectory segment, obstacle avoidance trajectory segment, target approach trajectory segment, and placement trajectory segment. Based on the load response status, predict the joint position, joint velocity, driving torque, and end-effector pose of each reference trajectory point in the unexecuted trajectory segment of the current reference joint trajectory. Based on the prediction results, determine the joint position margin, joint velocity margin, driving torque margin, and obstacle spacing margin corresponding to each reference trajectory point. Normalize each margin and determine the trajectory risk value corresponding to each reference trajectory point. Identify the continuous reference trajectory points whose trajectory risk values ​​reach the preset risk threshold, along with a preset number of reference trajectory points before and after them, as the target trajectory segment. Among these, the joint position margin is the difference between the joint position limit and the predicted joint position; the joint velocity margin is the difference between the joint velocity limit and the predicted joint velocity; the driving torque margin is the difference between the driving torque limit and the predicted driving torque; and the obstacle spacing margin is the difference between the predicted minimum obstacle spacing and the collision distance limit. The smaller any margin is, the greater its contribution to the trajectory risk value.

[0022] Step S4: When the robotic arm reaches the starting position of the target trajectory segment, the actual motion state of the robotic arm, the pose of the workpiece to be transferred, and the position of the obstacle are synchronized to the digital twin of the robotic arm after load correction, obtaining a common simulation initial state. The common simulation initial state, load response state, and target trajectory segment are input into the trajectory correction strategy network. The trajectory correction strategy network generates multiple candidate correction trajectories under the same common simulation initial state, obtaining a set of candidate correction trajectories. The robotic arm digital twin is a virtual simulation model established based on the linkage structure, joint dynamic parameters, parameters of the workpiece to be transferred, and obstacle positions of the robotic arm. It is used to simulate the joint position, joint velocity, driving torque, end effector pose, and obstacle spacing when the robotic arm executes the joint trajectory. The robotic arm digital twin is constructed using a multi-rigid-body dynamics simulation method. The trajectory correction strategy network includes a Gaussian Actor strategy network and a Critic value network.

[0023] Step S5: Using the digital twin of the robotic arm that has completed load correction, simulate and execute each candidate correction trajectory in the candidate correction trajectory set. Constraint and filter each candidate correction trajectory based on joint position limitations, joint speed limitations, driving torque limitations, and collision distance limitations to obtain an executable correction trajectory set. Perform multi-dimensional trajectory evaluation on each executable correction trajectory in the executable correction trajectory set, and determine the multi-dimensional representative credit result using a trajectory-by-trajectory elimination comparison method. Weight the multi-dimensional representative credit result according to the load response state to obtain the branch evaluation result. Update the trajectory correction strategy network using the branch evaluation result, calculate the selection probability corresponding to each executable correction trajectory using the updated trajectory correction strategy network, and determine the executable correction trajectory with the highest selection probability as the target correction trajectory.

[0024] Step S6: Replace the corresponding target trajectory segment in the current reference joint trajectory with the target correction trajectory, control the robotic arm to execute the replaced current reference joint trajectory, collect the actual execution results corresponding to the target correction trajectory, and update the load response status according to the actual execution results, so as to determine the target trajectory segment that needs to be further optimized from the subsequent unexecuted trajectory segments of the current reference joint trajectory, until the workpiece to be transferred is moved to the target placement pose.

[0025] In one specific embodiment, the robotic arm is a six-axis robotic arm, and the workpiece to be transferred is a motor housing. The robotic arm needs to transfer the motor housing from the picking station to the turnover box. An equipment column is set between the picking station and the turnover box. The nominal load of the robotic arm is 8kg, the actual weight of the motor housing is 10.5kg, and the actual center of gravity of the motor housing is offset relative to the nominal center of gravity.

[0026] Step S1: Obtain the gripping pose of the robotic arm when holding the motor housing, the target placement pose corresponding to the turnover box, and the position and external dimensions of the equipment column; use the inverse kinematics of the robotic arm to solve the initial joint configuration corresponding to the gripping pose and the target joint configuration corresponding to the target placement pose, and expand outward by 80mm according to the external dimensions of the equipment column to construct the obstacle surrounding area corresponding to the equipment column.

[0027] Under the conditions of satisfying the joint position constraints of the robotic arm and obstacle avoidance, the joint path between the initial joint configuration and the target joint configuration is planned, and time parameterization is performed according to the speed constraints of each joint to obtain the current reference joint trajectory. The current reference joint trajectory includes, in sequence:

[0028] The lifting trajectory section is used to lift the motor housing from the material picking height.

[0029] The transfer transition trajectory segment is used to transition the robotic arm from a material-picking posture to an obstacle-avoiding posture;

[0030] The obstacle avoidance trajectory segment is used to allow the motor housing to bypass the equipment column;

[0031] The target approach trajectory segment is used to gradually bring the motor housing closer to the turnover box;

[0032] The placement trajectory segment is used to lower the motor housing to the target placement position.

[0033] Step S2: The robotic arm first executes the lifting trajectory segment; during the lifting process, the joint position, joint velocity, driving torque and end pose of each joint are collected at a sampling frequency of 100Hz, and the collected results are compared with the expected motion response corresponding to the current reference joint trajectory.

[0034] The comparison results show that the driving torque of the second joint of the robotic arm increases by an average of 8.4 N·m compared to the expected driving torque, the maximum tracking deviation of the robotic arm end effector is 12 mm, the joint control response lag is 45 ms, and periodic vibrations exceeding the reference vibration level occur during the lifting process. The system combines the above torque increment, end effector tracking deviation, response lag, and vibration level to obtain the load response state when the robotic arm is holding the current motor housing. This load response state reflects the difference between the actual weight and center of gravity position of the motor housing and the nominal load parameters.

[0035] Step S3: Identify the unexecuted trajectory segments after lifting, including the transfer transition trajectory segment, obstacle avoidance trajectory segment, target approach trajectory segment, and placement trajectory segment; based on the load response status, predict the motion state of the unexecuted trajectory segments to obtain the margin corresponding to each unexecuted trajectory segment, as shown in Table 1:

[0036] Table 1

[0037] ;

[0038] Since the driving torque margin and obstacle spacing margin of the obstacle avoidance trajectory segment are the smallest, the system determines the part of the obstacle avoidance trajectory segment closest to the equipment column as the target trajectory segment; thus, the target trajectory segment can be either a complete functional trajectory segment or a higher-risk part of the functional trajectory segment.

[0039] Step S4: When the robotic arm moves to the starting position of the target trajectory segment, the robotic arm is briefly held at that position, and the joint position, joint speed, driving torque, end effector pose, motor housing pose, and equipment column position at that moment are synchronized to the digital twin of the robotic arm to obtain a common simulation initial state.

[0040] The trajectory correction strategy network generates 12 different sets of joint trajectory corrections based on the initial state of the common simulation, the load response state, and the target trajectory segment. Some of the joint trajectory corrections are used to reduce the running speed of the second joint, some are used to increase the distance between the motor housing and the equipment column, and some are used to reduce the change in joint acceleration at the trajectory turning position. The 12 sets of joint trajectory corrections are superimposed on the target trajectory segment to obtain 12 candidate correction trajectories.

[0041] Step S5: Restore the digital twin of the robotic arm to the initial state of the common simulation and simulate the execution of 12 candidate correction trajectories. Among them, 3 candidate correction trajectories are eliminated because the driving torque of the second joint exceeds the driving torque limit, 2 candidate correction trajectories are eliminated because the joint speed exceeds the joint speed limit, and another 2 candidate correction trajectories are eliminated because the minimum distance between the motor housing and the equipment column is less than the collision distance limit. Finally, an executable correction trajectory set consisting of 5 executable correction trajectories is obtained.

[0042] The system calculates the trajectory tracking score, trajectory smoothness score, driving torque score, and obstacle spacing score for each of the five executable correction trajectories, and uses a trajectory-by-trajectory elimination comparison method to determine the representative credit of each executable correction trajectory in each evaluation dimension.

[0043] Since the current load response indicates that the driving torque margin and obstacle spacing margin are small, the weights of the driving torque evaluation dimension and the obstacle spacing evaluation dimension are increased. Branch evaluation results are obtained based on the weighted multidimensional representative credit, and these results are used to update the trajectory correction strategy network. After the update, the fourth executable correction trajectory has the highest selection probability; therefore, the fourth executable correction trajectory is determined as the target correction trajectory.

[0044] The target correction trajectory reduces the peak velocity of the second joint relative to the original target trajectory segment and increases the lateral spacing of the motor housing when it passes around the equipment column by 35mm.

[0045] Step S6: Replace the target trajectory segment in the current reference joint trajectory with the target correction trajectory, and control the robotic arm to continue executing the replaced current reference joint trajectory. Actual execution results show that the driving torque margin of the second joint increased from 2.5 N·m to 6.8 N·m, the minimum distance between the motor housing and the equipment column increased from the originally predicted 18 mm to 53 mm, and the tracking deviation at the robotic arm end effector decreased from 12 mm to 5 mm.

[0046] The actual execution results are used to update the load response status and re-evaluate the subsequent target approach trajectory segment and placement trajectory segment. When the target approach trajectory segment still meets the optimization trigger condition, the target trajectory segment that needs to be optimized is re-determined from the target approach trajectory segment. When the subsequent trajectory segments all meet the corresponding margin requirements, the robotic arm continues to execute the target approach trajectory segment and placement trajectory segment until the motor housing is moved to the target placement pose in the turnover box.

[0047] Example 2, based on Example 1, further explains the process of generating the load response state in step S2. The process of generating the load response state specifically includes the following steps:

[0048] Step A1: According to the acquisition time, combine the acquired joint position, joint velocity, driving torque, and end effector pose into the actual motion response of the robotic arm corresponding to each acquisition time; compare the expected joint position, expected joint velocity, expected driving torque, and expected end effector pose corresponding to the current reference joint trajectory with the actual motion response of the robotic arm to obtain the joint position deviation, joint velocity deviation, driving torque deviation, and end effector pose deviation corresponding to each acquisition time, and combine the various deviations into a motion response deviation sample;

[0049] Step A2: Divide the motion response deviation samples according to the preset execution window, calculate the mean, peak value and rate of change of each deviation in each execution window, and determine the response lag based on the time difference between the joint control command issuance time and the corresponding robotic arm actual motion response. Combine the mean, peak value, rate of change and response lag into window load response characteristics.

[0050] Step A3: Compare the load response characteristics of each window with the baseline response characteristics of the robotic arm to determine the torque increment, end-effector tracking deviation, joint response hysteresis, and robotic arm vibration caused by the load, and combine the torque increment, end-effector tracking deviation, joint response hysteresis, and robotic arm vibration into a load response state.

[0051] Example 3, based on Example 1, further explains the process of obtaining the candidate corrected trajectory set in step S4. The load response state is generated using the method described in Example 2. The generation process of the candidate corrected trajectory set specifically includes the following steps:

[0052] Step B1: Determine the pose of the workpiece to be transferred based on the actual motion state of the robotic arm at the start of the target trajectory segment and the grasping relationship between the end effector of the robotic arm and the workpiece to be transferred; Set the actual motion state of the robotic arm, the pose of the workpiece to be transferred, and the position of the obstacle in the digital twin of the robotic arm, and save the corresponding digital twin state to obtain the common simulation initial state.

[0053] Step B2: Input the initial state, load response state and target trajectory segment of the common simulation into the Gaussian Actor policy network to obtain the joint trajectory correction distribution parameters corresponding to each reference trajectory point in the target trajectory segment, and construct the joint trajectory correction distribution corresponding to the target trajectory segment based on the joint trajectory correction distribution parameters.

[0054] Step B3: Independently sample multiple sets of initial joint trajectory corrections from the joint distribution of joint trajectory corrections, and combine each set of initial joint trajectory corrections according to the order of reference trajectory points and the order of robotic arm joints to obtain the set of initial trajectory corrections;

[0055] Step B4: Perform collaborative iterative correction on the initial trajectory correction set to obtain a distributed correction trajectory correction set; limit the amplitude of each group of joint trajectory corrections in the distributed correction trajectory correction set and superimpose them onto the corresponding reference trajectory points in the target trajectory segment; perform trajectory interpolation and trajectory segment boundary continuity processing on the superimposed reference trajectory points to obtain multiple candidate correction trajectories, and combine the candidate correction trajectories into a candidate correction trajectory set.

[0056] The process of performing cooperative iterative correction on the initial trajectory correction set to obtain a distributed correction trajectory correction set includes the following steps:

[0057] Step Q1: According to the order of reference trajectory points and robotic arm joints, convert each set of initial joint trajectory correction amounts into corresponding trajectory correction representation vectors, and normalize the vector components in each trajectory correction representation vector according to the amplitude limit of the joint trajectory correction amount corresponding to each joint; determine the trajectory risk weight corresponding to each reference trajectory point based on the joint position margin, joint velocity margin, driving torque margin, and obstacle spacing margin corresponding to each reference trajectory point in the target trajectory segment. Wherein, when any margin decreases, the trajectory risk weight of the corresponding reference trajectory point increases.

[0058] Step Q2: Based on the trajectory risk weights, the differences between the current trajectory correction representation vector and the corresponding vector components of other trajectory correction representation vectors are weighted to obtain the corresponding weighted distance. The weighted distance after adding a preset smoothing parameter is then exponentially calculated to obtain the trajectory difference kernel value between the current trajectory correction representation vector and other trajectory correction representation vectors. The trajectory difference kernel represents the degree of closeness between the two sets of joint trajectory correction values. The system first calculates the weighted distance between the two sets of trajectory correction representation vectors based on the trajectory risk weights, and then exponentially calculates the weighted distance after adding the smoothing parameter. The closer the two sets of trajectory correction values ​​are, the larger the trajectory difference kernel value, and the stronger the corresponding trajectory separation effect.

[0059] Step Q3: Based on the joint trajectory correction distribution parameters, determine the log probability gradient of each trajectory correction representation vector in the joint distribution of joint trajectory corrections. Excluding the current trajectory correction representation vector itself, use the corresponding trajectory difference kernel value to weightedly combine the log probability gradients of other trajectory correction representation vectors to obtain the distribution-preserving correction amount corresponding to the current trajectory correction representation vector. Based on the difference direction from other trajectory correction representation vectors to the current trajectory correction representation vector, and the rate of change of the corresponding trajectory difference kernel value with the weighted distance, determine the trajectory separation component corresponding to each difference direction, and combine these components to obtain the trajectory separation correction amount corresponding to the current trajectory correction representation vector. The distribution-preserving correction amount prevents the groups of joint trajectory corrections from deviating from the high-probability region of the Gaussian Actor policy network output due to mutual separation. The trajectory separation correction amount moves the closely approaching joint trajectory corrections in directions that move away from each other. Together, these two factors ensure that the corrected groups of joint trajectory corrections conform to the output distribution of the trajectory correction policy network while also having different trajectory change directions.

[0060] Step Q4: According to the preset distribution preservation coefficient and preset trajectory separation coefficient, the distribution preservation correction amount and trajectory separation correction amount corresponding to the current trajectory correction representation vector are weighted and combined to obtain the collaborative correction amount; the current trajectory correction representation vector is updated according to the preset correction step size and the collaborative correction amount, and the updated trajectory correction representation vectors are used for the next collaborative iteration correction; when the average vector change between two adjacent collaborative iteration corrections is not greater than the preset convergence threshold, or the number of collaborative iteration corrections reaches the preset number, the collaborative iteration correction is stopped; according to the reference trajectory point order and the robotic arm joint order, the trajectory correction representation vectors that have completed the collaborative iteration correction are denormalized and restored to the corresponding joint trajectory correction amount to obtain the set of dispersed correction trajectory correction amounts.

[0061] Example 4, based on Example 1, further explains the process of obtaining the executable correction trajectory set in step S5. The candidate correction trajectory set is generated using the method described in Example 3. The generation process of the executable correction trajectory set specifically includes the following steps:

[0062] Step T1: Repeatedly restore the digital twin of the robotic arm to the initial state of the common simulation, simulate and execute each candidate correction trajectory, and obtain the peak value of joint position, peak value of joint velocity, peak value of driving torque and minimum obstacle distance corresponding to each candidate correction trajectory;

[0063] Step T2: When the joint positions corresponding to a candidate correction trajectory are all between the lower limit and the upper limit of the corresponding joint position, the peak absolute value of the joint velocity and the peak absolute value of the driving torque do not exceed the corresponding limits, and the minimum obstacle spacing is not less than the collision distance limit, the candidate correction trajectory is retained; the retained candidate correction trajectories are combined into an executable correction trajectory set.

[0064] Example 5, as Figure 2 As shown, this embodiment, based on Embodiment 1, further explains the process of determining the branch evaluation result in step S5. The executable correction trajectory set is generated using the method described in Embodiment 4. The process of generating the branch evaluation result specifically includes the following steps:

[0065] Step C1: Calculate the trajectory tracking score, trajectory smoothing score, driving torque score, and obstacle spacing score for each executable corrected trajectory in the executable corrected trajectory set, and convert each score into an evaluation score where a larger value indicates better trajectory performance; combine the evaluation scores for each executable corrected trajectory to obtain the corresponding multi-dimensional trajectory evaluation results; wherein, the trajectory tracking score is determined based on the end pose deviation between the executable corrected trajectory and the target trajectory segment, the trajectory smoothing score is determined based on the joint acceleration and joint jerk corresponding to each simulation moment in the executable corrected trajectory, the driving torque score is determined based on the proportion of each joint driving torque relative to the corresponding driving torque limit, and the obstacle spacing score is determined based on the minimum distance between the robotic arm and the workpiece to be transferred and the obstacle during the execution of the executable corrected trajectory;

[0066] I. The trajectory tracking score formula is as follows:

[0067] ;

[0068] in, Indicates the first The trajectory tracking score of an executable correction trajectory; Indicates the number of simulation moments; and They represent the first Simulation end position and reference end position at each simulation moment; This represents the rotation angle deviation between the simulated end-effector attitude and the reference end-effector attitude; and These represent the normalized references for position deviation and attitude deviation, respectively; and Indicates weight;

[0069] II. The formula for trajectory smoothing score is as follows:

[0070] ;

[0071] in, Indicates the trajectory smoothness score; Indicates the number of joints in the robotic arm; and They represent the first The acceleration and jerk of each joint; and These represent the corresponding normalization benchmarks;

[0072] III. The formula for scoring the driving torque is as follows:

[0073] ;

[0074] in, Indicates the score for driving torque; Indicates the first The executable correction trajectory is in the first The simulation time corresponding to the first... Driven torque of each joint; Indicates the first The driving torque limit of each joint; the smaller the proportion of driving torque, the higher the driving torque score.

[0075] IV. The scoring formula for obstacle spacing is as follows:

[0076] ;

[0077] in, Indicates the score based on the distance between obstacles; This indicates the minimum distance between the robotic arm and the workpiece to be transferred, and the obstacle. Indicates the collision distance limit; Indicates the preset desired safe distance; This means that the calculation result will be restricted to the range [0, 1].

[0078] Step C2: For each executable correction trajectory, obtain the evaluation score for each evaluation dimension from the corresponding multidimensional trajectory evaluation results; when the set of executable correction trajectories includes at least two executable correction trajectories, for each evaluation dimension, determine the highest evaluation score of all executable correction trajectories in that evaluation dimension as the complete set representative score, determine the highest evaluation score after excluding the current executable correction trajectory as the excluded set representative score, and determine the difference between the complete set representative score and the excluded set representative score as the representative credit of the current executable correction trajectory in that evaluation dimension; when the set of executable correction trajectories includes only one executable correction trajectory, determine the difference between the evaluation scores of the executable correction trajectory and the uncorrected target trajectory segment in the corresponding evaluation dimension as the representative credit; combine the representative credits of the current executable correction trajectory in each evaluation dimension to obtain the corresponding multidimensional representative credit result;

[0079] Representative credit calculation formula:

[0080] ;

[0081] in, Indicates the first The executable correction trajectory is in the first Credit is represented on each evaluation dimension; This represents the set of executable correction trajectories; Indicates the number of trajectories in the set; Indicates the first The executable correction trajectory is in the first Evaluation scores on each evaluation dimension; This indicates the uncorrected target trajectory segment in the [number]th [year]. Evaluation scores on each evaluation dimension; This indicates that the first [track] is excluded from the set of executable modified tracks. A set of trajectories obtained after executing corrected trajectories;

[0082] Step C3: Based on the torque increment, end-effector tracking deviation, joint response hysteresis, and robotic arm vibration in the load response state, as well as the joint position margin, joint velocity margin, driving torque margin, and obstacle spacing margin corresponding to the target trajectory segment, determine the initial dimension weights for each evaluation dimension. When the end-effector tracking deviation or joint response hysteresis increases, or the joint position margin decreases, increase the initial dimension weights for the trajectory tracking evaluation dimension. When the robotic arm vibration increases or the joint velocity margin decreases, increase the initial dimension weights for the trajectory smoothing evaluation dimension. When the torque increment increases or the driving torque margin decreases, increase the initial dimension weights for the driving torque evaluation dimension. When the obstacle spacing margin decreases, increase the initial dimension weights for the obstacle spacing evaluation dimension. Normalize the adjusted initial dimension weights to obtain the dimension weights for each evaluation dimension. Obtain the representative credit of the current executable correction trajectory in each evaluation dimension from the multi-dimensional representative credit results, and perform a weighted sum according to the corresponding dimension weights to obtain the branch evaluation results corresponding to the current executable correction trajectory.

[0083] Example 6, based on Example 1, further explains the process of updating the trajectory correction strategy network in step S5. The branch evaluation results are generated using the method described in Example 5. The process of updating the trajectory correction strategy network specifically includes the following steps:

[0084] Step D1: Accumulate the task rewards corresponding to each simulation moment during the simulation of the current executable corrected trajectory to obtain the cumulative task reward; input the initial state, load response state, and load response state of the co-simulation into the Critic value network to obtain the state value estimation results; determine the task advantage based on the cumulative task reward and the state value estimation results; combine the task advantage with the branch evaluation results to obtain the comprehensive branch advantage corresponding to the current executable corrected trajectory; determine the policy probability ratio of the current Gaussian Actor policy network and the Gaussian Actor policy network before the update in generating the current executable corrected trajectory, and construct the truncation policy objective based on the policy probability ratio, the comprehensive branch advantage, and the preset truncation interval;

[0085] Step D2: Determine the policy loss corresponding to the Gaussian Actor policy network based on the truncation policy objective, determine the value estimation loss corresponding to the Critic value network based on the cumulative task reward and state value estimation results, and determine the policy entropy based on the joint trajectory correction distribution parameters output by the Gaussian Actor policy network; perform a weighted combination of the policy loss, value estimation loss, and policy entropy to obtain the total update loss, and update the network parameters of the Gaussian Actor policy network and the Critic value network based on the total update loss.

[0086] The present invention and its embodiments have been described above. This description is not restrictive, and the accompanying drawings are only one embodiment of the present invention. The actual structure is not limited to this. In short, if a person skilled in the art is inspired by this description and designs a similar structure and embodiment without departing from the spirit of the present invention, such design should fall within the protection scope of the present invention.

Claims

1. A reinforcement learning-driven adaptive optimization method for robotic arm trajectory, characterized in that, The method includes the following steps: Step S1: Obtain the gripping pose, target placement pose, and obstacle position of the robotic arm when gripping the workpiece to be transferred, and obtain the current reference joint trajectory; Step S2: Control the robotic arm to execute the lifting trajectory segment in the current reference joint trajectory, collect the joint position, joint speed, driving torque and end pose during the execution of the robotic arm, and generate the load response state; Step S3: Based on the load response status, determine the remaining amount of the unexecuted trajectory segment of the current reference joint trajectory, and determine the target trajectory segment based on the remaining amount; Step S4: When the robotic arm moves to the starting position of the target trajectory segment, the initial state of the common simulation is obtained using the digital twin of the robotic arm; the initial state of the common simulation, the load response state, and the target trajectory segment are input into the trajectory correction strategy network to obtain a set of candidate correction trajectories; Step S5: Using the digital twin of the robotic arm, simulate and constrain each candidate correction trajectory in the candidate correction trajectory set to obtain an executable correction trajectory set; perform multi-dimensional trajectory evaluation on each executable correction trajectory in the executable correction trajectory set to obtain branch evaluation results; use the branch evaluation results to update the trajectory correction strategy network and determine the target correction trajectory. Step S6: Replace the corresponding target trajectory segment in the current reference joint trajectory with the target correction trajectory, control the robotic arm to execute the replaced current reference joint trajectory, update the load response state, until the workpiece to be transferred is moved to the target placement pose.

2. The reinforcement learning-driven adaptive optimization method for robotic arm trajectory according to claim 1, characterized in that, The process of determining the target trajectory segment in step S3 includes: predicting the joint position, joint velocity, driving torque, and end pose corresponding to each reference trajectory point in the unexecuted trajectory segment based on the load response state; determining the joint position margin, joint velocity margin, driving torque margin, and obstacle spacing margin corresponding to each reference trajectory point based on the prediction results; normalizing each margin and determining the trajectory risk value corresponding to each reference trajectory point; and determining the target trajectory segment based on the trajectory risk value.

3. The reinforcement learning-driven adaptive optimization method for robotic arm trajectory as described in claim 1, characterized in that: Trajectory correction policy networks include Gaussian Actor policy networks and Critic value networks.

4. The reinforcement learning-driven adaptive optimization method for robotic arm trajectory according to claim 3, characterized in that: The process of generating the candidate corrected trajectory set includes the following steps: Step B1: Determine the pose of the workpiece to be transferred based on the actual motion state of the robotic arm at the start of the target trajectory segment and the grasping relationship between the end effector of the robotic arm and the workpiece to be transferred; Set the actual motion state of the robotic arm, the pose of the workpiece to be transferred, and the position of the obstacle in the digital twin of the robotic arm to obtain the common simulation initial state. Step B2: Input the initial state, load response state and target trajectory segment of the common simulation into the Gaussian Actor policy network to obtain the joint trajectory correction distribution parameters corresponding to each reference trajectory point in the target trajectory segment, and construct the joint trajectory correction distribution corresponding to the target trajectory segment based on the joint trajectory correction distribution parameters. Step B3: Independently sample multiple sets of initial joint trajectory corrections from the joint distribution of joint trajectory corrections, and combine each set of initial joint trajectory corrections according to the order of reference trajectory points and the order of robotic arm joints to obtain the set of initial trajectory corrections; Step B4: Perform collaborative iterative correction on the initial trajectory correction set to obtain a distributed correction trajectory correction set; limit the amplitude of each group of joint trajectory corrections in the distributed correction trajectory correction set and superimpose them onto the corresponding reference trajectory points in the target trajectory segment; perform trajectory interpolation and trajectory segment boundary continuity processing on the superimposed reference trajectory points to obtain multiple candidate correction trajectories, and combine the candidate correction trajectories into a candidate correction trajectory set.

5. The reinforcement learning-driven adaptive optimization method for robotic arm trajectory according to claim 4, characterized in that: The process of generating the set of dispersed correction trajectory adjustments includes the following steps: Step Q1: According to the order of reference trajectory points and the order of robotic arm joints, convert the initial joint trajectory correction amount of each group into the corresponding trajectory correction representation vector and determine the trajectory risk weight; Step Q2: Based on the trajectory risk weight, perform a weighted summation and negative exponentiation operation on the differences between the current trajectory correction representation vector and the corresponding vector components in other trajectory correction representation vectors to obtain the trajectory difference kernel value; Step Q3: Based on the joint trajectory correction distribution parameters, determine the log probability gradient of each trajectory correction representation vector in the joint distribution of joint trajectory corrections; use the trajectory difference kernel value to weight and combine the log probability gradients of other trajectory correction representation vectors to obtain the distribution preservation correction amount corresponding to the current trajectory correction representation vector; based on the difference direction from other trajectory correction representation vectors to the current trajectory correction representation vector, and the rate of change of the corresponding trajectory difference kernel value with the weighted distance, determine the trajectory separation component corresponding to each difference direction, and combine the trajectory separation components to obtain the trajectory separation correction amount corresponding to the current trajectory correction representation vector; Step Q4: Weight the distribution-preserving correction and trajectory separation correction corresponding to the current trajectory correction representation vector to obtain the collaborative correction amount; update the current trajectory correction representation vector according to the preset correction step size and collaborative correction amount, and use the updated trajectory correction representation vectors for the next collaborative iteration correction; when the preset number of times is reached, stop the collaborative iteration correction and obtain the set of distributed correction trajectory correction amounts.

6. The reinforcement learning-driven adaptive optimization method for robotic arm trajectory according to claim 1, characterized in that: The process of generating branch evaluation results includes the following steps: Step C1: Calculate the trajectory tracking score, trajectory smoothing score, driving torque score, and obstacle distance score for each executable correction trajectory in the executable correction trajectory set, and combine the scores to obtain the corresponding multi-dimensional trajectory evaluation results; Step C2: Obtain the evaluation scores for each evaluation dimension from the multidimensional trajectory evaluation results; when the executable correction trajectory set includes at least two executable correction trajectories, determine them as the complete set representative score and the excluded set representative score, and determine the difference between the complete set representative score and the excluded set representative score as the representative credit in that evaluation dimension; when the executable correction trajectory set includes only one executable correction trajectory, determine the difference between the evaluation scores of the executable correction trajectory and the uncorrected target trajectory segment in the corresponding evaluation dimension as the representative credit; combine the representative credits in each evaluation dimension to obtain the multidimensional representative credit result; Step C3: Obtain the representative credit of the current executable correction trajectory in each evaluation dimension from the multidimensional representative credit results, and perform weighted summation according to the corresponding dimension weights to obtain the branch evaluation results.