Reinforcement learning mechanical arm dynamic obstacle avoidance method and system based on past experience
Through the improved SAC algorithm and HER algorithm combined with trajectory interpolation constraints, the action space and reward function are constructed, real-time dynamic obstacle avoidance of high degree of freedom robotic arms is achieved, and the problems of high computational complexity and insufficient learning ability of traditional methods are solved, improving the flexibility and stability of obstacle avoidance.
Patent Information
- Application Number
- CN202510664740.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-22
- Publication Date
- 2025-08-29
AI Technical Summary
Traditional path planning algorithms have high computational complexity in the configuration space of high-degree of freedom robot arms, which is difficult to meet the real-time requirements of dynamic obstacle avoidance. In addition, the existing deep reinforcement learning methods lack learning ability in complex environments, are not flexible enough in obstacle avoidance, and have limited real-time and dynamic scenario adaptability.
The improved SAC algorithm is used to build a neural network, combining long and short-term memory networks and trajectory interpolation constraints, and introducing HER algorithms and experience playback pools. Through course learning methods, experience is accumulated step by step, action space, state space and reward functions are constructed, reinforcement learning training is performed, and action output is optimized through trajectory interpolation and smooth constraints.
Real-time dynamic obstacle avoidance of the robot arm in complex environments is realized, the flexibility of obstacle avoidance and dynamic scene adaptability is improved, the difficulty of obstacle description is reduced, the efficiency and stability of obstacle avoidance are improved, and the risk of jitter and collision of the robot arm movement is reduced.
Smart Images

Figure CN120552049A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of robot arm obstacle avoidance, and in particular to a method and system for dynamic obstacle avoidance of a robot arm using reinforcement learning based on past experience. Background Art
[0002] Traditional path planning algorithms, such as the Rapidly Exploring Random Tree (RRT*) or the A Star Algorithm (A*), face the curse of dimensionality in the configuration space of high-DOF manipulators. Their computational complexity increases exponentially with the number of degrees of freedom, making them unable to meet the real-time requirements of dynamic obstacle avoidance. Their real-time performance and computational efficiency are insufficient, and frequent path replanning can cause motion lag or trajectory jumps in the manipulator.
[0003] Obstacle avoidance for robotic arms requires processing a high-dimensional state space containing joint states, obstacle information, and sensor data. Traditional methods (such as the potential field method) are prone to a sharp drop in computational efficiency due to the curse of dimensionality, and their high-dimensional data processing efficiency is low.
[0004] The trajectories generated by traditional discretization planning methods (such as the grid method) have jumps or discontinuities, which affect the smoothness of the robot movement and are insufficient in the ability to generate continuous actions.
[0005] For example, the invention patent with publication number CN117873116A discloses a method for autonomous obstacle avoidance for multiple mobile robots based on deep reinforcement learning. This method updates the obstacle avoidance strategy through an improved deep reinforcement learning algorithm, employs a centralized learning and distributed training framework to complete multi-robot obstacle avoidance strategy training, and completes multi-robot advanced training from simple to complex environments through phased advanced training. This method can effectively address the problem of slow algorithm convergence in complex environments, improve the robustness of multi-robot obstacle avoidance, and avoid situations where the robot's reward function is sparse during obstacle avoidance tasks. However, this method pre-trains the model and then completes the obstacle avoidance task based on the trained obstacle avoidance model. This method also suffers from insufficient continuous action generation capabilities and reliance on pre-trained models, resulting in inflexible obstacle avoidance, insufficient real-time learning capabilities, and limited adaptability to dynamic scenarios. Summary of the Invention
[0006] The purpose of the present invention is to overcome the defects of the above-mentioned prior art and provide a reinforcement learning robot arm dynamic obstacle avoidance method and system based on past experience.
[0007] The purpose of the present invention can be achieved by the following technical solutions:
[0008] According to one aspect of the present invention, a method for dynamic obstacle avoidance of a robotic arm using reinforcement learning based on past experience is provided, the method comprising the following steps:
[0009] S1. Use the improved SAC algorithm to build a neural network. The improved SAC algorithm is embedded in the long short-term memory network and the trajectory interpolation constraint is introduced. The neural network includes a policy network and a value network.
[0010] S2. Construct action space, state space and reward function based on the environment, obstacle position and robot arm posture;
[0011] S3. Based on the constructed action space, state space, and reward function, the neural network is trained through reinforcement learning. The HER algorithm and experience replay pool are introduced for auxiliary training. During the training process, the curriculum learning method is used to enable the agent to accumulate experience step by step. Finally, the neural network converges and outputs actions.
[0012] S4, using the output action in S3 to control the robotic arm, and perform collision detection and error calculation to obtain the detection result;
[0013] S5. Adjust and update the neural network parameters based on the detection results and the experience replay pool, and return to execute step S2 until the robotic arm reaches the target point.
[0014] 2. A method for dynamic obstacle avoidance of a robotic arm using reinforcement learning based on past experience according to claim 1, wherein the trajectory interpolation constraint introduced in the improved SAC algorithm in S1 includes the first constraint or the second constraint;
[0015] The first constraint is to introduce the trajectory interpolation formula into the Actor network structure so that the actions generated by SAC have temporal continuity;
[0016] The second constraint is to add a trajectory smoothness constraint to the Actor's loss function and actively optimize to generate a smooth action sequence.
[0017] As a preferred technical solution, the first constraint is achieved by using a first-order low-pass filter, and its specific formula is:
[0018] a t =δ·a raw +(1-δ)·a t-1 ;
[0019] Among them, a t is the execution action of the current time step; a raw The original action directly output by the Actor; a t-1 is the execution action of the previous time step; δ is the smoothing coefficient.
[0020] As a preferred technical solution, the specific process of the second constraint is: first, define the first-order difference smoothing loss to suppress drastic changes in speed, then use the second-order difference smoothing loss to control the fluctuation of acceleration, and finally use the Actor loss function to control the trade-off between smoothness and policy optimization objectives. The specific formula is:
[0021] L smooth1 =∑||a t -a t-1 || 2 ;
[0022] L smooth2 =∑||(a t -a t-1 )-(a t-1 -a t-2 )|| 2 ;
[0023] L smooth =L smooth1 +L smooth2 ;
[0024] L actor =-Q(s,a)+λ smooth ·L smooth ;
[0025] Among them, L smooth1 is the first-order difference smoothing loss; L smooth2 is the second-order difference smoothing loss; L actor is the Actor loss function; a t is the action vector at time t; a t-1 is the action vector at time t-1; a t-2 is the action vector at time t-2; Q(s,a) is the input state s t and action a t Q function; λ smooth is the smoothing weight parameter; L smooth is the total smoothing loss.
[0026] As a preferred technical solution, the action space constructed in S2 is composed of a set of joint velocities. Indicates that the range of the action value is set within [-1,1]rad / s, and the maximum change of each control cycle is set to 0.1rad / s.
[0027] As a preferred technical solution, the state space constructed in S2 includes the state information of the robot arm, the position of the nearest obstacle, and the distance between the end effector and the target point. The specific formula is:
[0028] S=[S bot ,S obs ,Sgoal ];
[0029] S bot =[J s ,J v ,J p ];
[0030] Among them, J s Indicates the position of the end effector of the machine; J v represents the joint velocity; J p Indicates joint position; S bot is the status information of the robot arm; S obs is the position of the nearest obstacle; S goal is the distance between the end effector and the target point.
[0031] As a preferred technical solution, the reward function in S2 is constructed based on the APF algorithm. The specific process of constructing the reward function is: first, a repulsive field is constructed at the obstacle, and a gravitational field is constructed at the target point. Then, positive rewards are provided within the preset area of the gravitational field, and negative rewards are provided within the preset area of the repulsive field. Finally, the composite field formed by the gravitational field and the repulsive field constructs the reward for the current position of the robotic arm.
[0032] As a preferred technical solution, the reward function includes a gravitational field reward term, a repulsive field reward term, a collision penalty term, and a distance penalty term. The specific formula is:
[0033] R=R att +R collision -R rep -R dis ;
[0034] R att =α·d;
[0035]
[0036] R dis =η·d;
[0037] Among them, R att is the gravitational field bonus; R rep is the repulsive field reward item; R collision is the collision penalty term; R dis is the distance penalty term; α is the gain coefficient of the gravitational field; D(q) is the distance between the end effector and the nearest obstacle repulsion calculation point; β is the repulsion gain coefficient; is the action threshold for triggering repulsion; d is the distance between the end effector and the target point; η is the distance gain coefficient.
[0038] As a preferred technical solution, the specific process of using the curriculum learning method in S3 to enable the intelligent agent to accumulate experience step by step is: first, use the curriculum learning method to decompose complex tasks and store the tasks in the task space. Each task corresponds to a difficulty level. Then, a preset performance threshold is set for each task. When the performance of the robot arm on the current task reaches the preset performance threshold, it automatically switches to the next task.
[0039] According to another aspect of the present invention, a reinforcement learning robot arm dynamic obstacle avoidance system based on past experience is provided, the system comprising a network construction module, a state-action space modeling module, a reinforcement learning training module, and a control and feedback optimization module;
[0040] The network construction module is used to construct a neural network using the improved SAC algorithm;
[0041] The state-action space modeling module is used to construct the action space, state space, and reward function based on the environment, obstacle positions, and robot arm posture;
[0042] The reinforcement learning training module is used to combine the HER algorithm, experience replay pool and curriculum learning method to perform reinforcement learning training on the neural network until the neural network converges and outputs actions;
[0043] The control and feedback optimization module is used to achieve closed-loop control through collision detection, error calculation and parameter update, and iterates until the robotic arm reaches the target point.
[0044] Compared with the prior art, the present invention has the following beneficial effects:
[0045] 1. The present invention first constructs a neural network using an improved SAC algorithm. It then defines the action space, state space, and reward function, and uses the HER algorithm, experience replay pool, and curriculum learning to perform intensive training on the neural network, gradually converging the network. After training, the system outputs actions to control the robotic arm, detects collisions and motion errors in real time, and dynamically updates network parameters based on the detection results. Through repeated iterations of environmental interaction and parameter optimization, the robotic arm ultimately achieves precise obstacle avoidance and reaches the target position. Deep reinforcement learning technology enables real-time dynamic obstacle avoidance based on past experience, eliminating the need for pre-established mathematical models. By using a model-free approach and autonomous learning of task objectives with the neural network, the obstacle avoidance process possesses real-time learning capabilities, improving its flexibility and adaptability to dynamic scenarios.
[0046] 2. In a method of deep reinforcement learning based on past experience learning in the present invention, the improved SAC algorithm is embedded in a long short-term memory network. In the process of reinforcement learning training of the neural network, the HER algorithm and the experience replay pool are introduced for auxiliary training. In the training process, the curriculum learning method is used to enable the intelligent agent to accumulate experience step by step, accumulate experience by generating courses to learn simple tasks, update the neural network parameters through historical experience, and then explore more complex tasks with the updated parameters, effectively avoiding the situation where training is difficult to converge due to overly complex tasks, while improving the efficiency and stability of obstacle avoidance.
[0047] 3. In this invention, the reward function is constructed based on the APF algorithm. Its reward function includes a gravitational field reward term, a repulsive field reward term, a collision penalty term, and a distance penalty term. It eliminates the need to describe the number and shape of obstacles. Rewards are calculated by constructing repulsive forces near obstacles, significantly reducing the difficulty of describing obstacles in complex scenarios and designing reward functions. Reward calculation relies solely on relative position information, eliminating the need for complex geometric calculations. This results in lightweight computation, high real-time performance, and suitability for deployment on embedded devices.
[0048] 4. In the present invention, in order to solve the problem of jitter during the movement of the robotic arm caused by the uncertainty of the reinforcement learning output, the trajectory interpolation constraint introduced in the improved SAC algorithm includes a first constraint or a second constraint; wherein, the first constraint is implemented by adopting a first-order low-pass filter, so that the current robotic arm knows how to make a smooth transition with the past action. At the same time, the second constraint adds a trajectory smoothing constraint to the loss function of the strategy network, thereby suppressing the drastic changes in the action output by the reinforcement learning, ensuring the continuous and smooth movement of the robotic arm joints, and reducing the probability of sudden changes in the action, reducing the positioning deviation of the end effector or the risk of collision due to jitter; and improving the stability and safety of the robotic arm movement. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] Figure 1 A schematic diagram of the steps of a dynamic obstacle avoidance method for a robotic arm using reinforcement learning based on past experience in the present invention;
[0050] Figure 2 Schematic diagram of a simulation training scenario in the embodiment;
[0051] Figure 3 This is a flowchart of reinforcement learning training in the embodiment;
[0052] Figure 4a A schematic diagram of the construction of a non-spherical obstacle repulsive field for a single obstacle in an embodiment;
[0053] Figure 4b Schematic diagram of the construction of the repulsive force field of the non-spherical obstacle group in the embodiment;
[0054] Figure 5 HER workflow diagram in the embodiment;
[0055] Figure 6 Flowchart for model training and evaluation in the embodiment. DETAILED DESCRIPTION
[0056] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are part of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of the present invention.
[0057] Example 1
[0058] In this example, we address the numerous challenges of existing dynamic obstacle avoidance for robotic arms by implementing reinforcement learning based on past experience. Using a six-axis robotic arm as an example, we first construct a simulated training environment within the robot simulation software. Next, we design a deep reinforcement learning controller. By combining the actual working environment and tasks, we construct a state space, action space, and reward function, and then train the agent. Finally, we transfer the simulated training results to a real robotic arm for verification.
[0059] This embodiment adopts a reinforcement learning method for dynamic obstacle avoidance of a manipulator based on past experience. The steps of the method are as follows: Figure 1 As shown, specifically including:
[0060] S1. Build a neural network using an improved Soft Actor-Critic (SAC) algorithm. The improved SAC algorithm is embedded in a long short-term memory network and trajectory interpolation constraints are introduced. The neural network includes a policy network and a value network.
[0061] S2. Construct action space, state space and reward function based on the environment, obstacle position and robot arm posture;
[0062] S3: Based on the constructed action space, state space, and reward function, the neural network is trained through reinforcement learning. The Hindsight Experience Replay (HER) algorithm and experience replay pool are introduced to assist in training. During the training process, the curriculum learning method is used to enable the agent to accumulate experience step by step, and finally the neural network converges and outputs an action.
[0063] S4, using the output action in S3 to control the robotic arm, and perform collision detection and error calculation to obtain the detection result;
[0064] S5. Adjust and update the neural network parameters based on the detection results and the experience replay pool; return to execute step S2 until the robotic arm reaches the target point.
[0065] The specific implementation steps are as follows:
[0066] Step 1): Build a deep reinforcement learning and robotic arm simulation environment. The simulation environment is as follows Figure 2 shown
[0067] Step 2): Implement the Maximum Entropy Deep Reinforcement Learning (SAC) algorithm. SAC is an advanced algorithm based on the maximum entropy reinforcement learning framework. Unlike traditional reinforcement learning algorithms, SAC not only focuses on maximizing cumulative rewards but also enhances the agent's exploration capabilities by maximizing the entropy of the policy. This design enables SAC to excel in continuous control tasks, especially demonstrating strong adaptability in high-dimensional state spaces and complex dynamic environments.
[0068] In traditional reinforcement learning, the agent's goal is to maximize the expected value of the cumulative reward. In maximum entropy reinforcement learning, the objective function is expanded to maximize both the cumulative reward and the entropy of the policy. Entropy is a measure of the randomness of the policy, and a high-entropy policy means that the agent has greater diversity when exploring the environment. The maximum entropy objective function can be expressed as:
[0069]
[0070] in, is the policy π in state s t The entropy under , α is the temperature parameter used to adjust the relative importance between reward and entropy. Defined as:
[0071]
[0072] SAC combines the advantages of the Actor-Critic architecture and entropy regularization, aiming to maximize the entropy of the policy while maximizing the cumulative reward by learning a random policy. This design enables SAC to achieve a better balance between exploration and exploitation, and excels in continuous control tasks. The SAC algorithm is based on the Actor-Critic architecture and includes the following core components:
[0073] (1) Actor: responsible for learning strategy π φ (a t |s t ), where φ is the parameter of the policy network.
[0074] (2) Critic: Contains two Q networks and It is used to estimate the state-action value function and update the parameters by minimizing the Bellman error.
[0075] (3) Target Q network: Use the target network and to stabilize the training process.
[0076] The goal of SAC is to optimize the policy π φ and Q θ To maximize the cumulative reward of entropy regularization. Specifically, the update target of the Q function is:
[0077]
[0078] Among them, γ is the discount factor, and the loss function of the Q function is:
[0079]
[0080] The update target of the policy network is to maximize the entropy regularized Q value:
[0081]
[0082] In addition, SAC also introduces a mechanism to automatically adjust the entropy regularization coefficient α, which is dynamically adjusted by optimizing the following objectives:
[0083]
[0084] in, is the target entropy, usually set to the dimension of the action space.
[0085] By maximizing policy entropy, SAC can achieve a better balance between exploration and exploitation, avoiding premature convergence to suboptimal policies. In addition, using two Q networks and taking the minimum value as the goal reduces the overfitting problem of Q-value estimation.
[0086] Step 3): Action space, state space design and reward function design
[0087] In this embodiment, if Figure 3 The following is the specific strengthening training process.
[0088] (1) Action space: The action space is composed of a set of joint velocities. It should be noted that by selecting actions in the joint space as control inputs, singularity problems can be avoided. To achieve more reliable training, the range of action values is set to [-1, 1] rad / s, and the maximum change per control cycle is set to 0.1 rad / s.
[0089] (2) State space: The state space consists of three parts and is defined as [Sbot ,S obs ,S goal ]. bot =[J s ,J v ,J p ] is the state information of the robot arm, where J s Indicates the position of the end effector of the machine, J v represents the joint velocity, J p Indicates joint position; S obs is the position of the nearest obstacle; S goal Represents the distance between the end effector and the target point.
[0090] (3) Reward function design based on APF algorithm:
[0091] In complex scenarios, it's difficult to accurately describe the number, shape, and movement of obstacles. Therefore, constructing a repulsive field near obstacles can effectively address this problem. By setting the area around obstacles as a negative reward zone, the robot learns from the negative rewards that it's approaching an obstacle when it enters the repulsive field, allowing it to make decisions to avoid the obstacle.
[0092] Artificial Potential Field (APF) is applied to the design of reward function. First, we need to construct an attractive field and a repulsive field. In this case, the attractive field provides positive rewards, while the repulsive field provides negative rewards. The composite field formed by the attractive and repulsive forces determines the reward for the robot's current position. The reward target can be defined as:
[0093]
[0094] (1) Repulsive field construction: In practical application scenarios, it is difficult to accurately express the distance between the robot and the obstacle, especially when encountering non-spherical obstacles. To solve this problem, this paper adopts a conservative approach to define the repulsive field, and its formula is defined as:
[0095]
[0096] Among them, U p (i) is the calculation point of repulsive force; N is the number of calculation points.
[0097] The repulsive field of non-spherical obstacles is constructed as follows Figure 4a and Figure 4b As shown in the figure, repulsion calculation points are selected on the outer surface of the obstacle. These calculation points are mainly used to calculate the distance between it and the end effector. Each calculation point forms a repulsive field, and each repulsive field is ultimately combined into a repulsive field that can envelop the entire obstacle.
[0098] (2) Reward function design: The reward function includes positive rewards in the gravitational field, negative rewards in the repulsive field, collision rewards, and distance rewards. The artificial potential field repulsive reward does not require preset obstacle geometric features and is naturally adapted to dynamically increasing or decreasing obstacles or unstructured environments (such as obstacles with irregular shapes and randomly changing numbers), reducing the cost of scene modeling. The repulsive field provides a continuous gradient feedback signal, avoiding the problem of the agent having difficulty obtaining effective feedback under traditional sparse rewards, and accelerating strategy optimization. The overall definition of the reward function is:
[0099] R=R att +R collision -R rep -R dis ;
[0100] Gravitational Field Reward R att Defined as:
[0101] R att =α·d;
[0102] Where d is the distance between the end effector and the target point; α is the gain coefficient of the gravitational field.
[0103] Repulsion Field Reward R rep Defined as:
[0104]
[0105] Where D(q) is the distance between the end effector and the nearest obstacle repulsion calculation point; β is the repulsion gain coefficient; is the action threshold for triggering repulsion.
[0106] Collision Reward R collision definition:
[0107]
[0108] The collision reward is a piecewise function that gives a larger penalty when a collision occurs and a penalty of 0 in other cases.
[0109] The distance bonus is defined as:
[0110] R dis =η·d;
[0111] Where d is the distance between the end effector and the target point, and η is the distance gain coefficient. The distance reward is designed because the positive reward provided by the gravitational field decreases as the robot approaches the target point, potentially causing it to miss the target point. Therefore, the distance reward is designed to balance the gravitational field and encourage the robot to explore toward the target point.
[0112] Step 4): Design of past experience learning algorithm training framework
[0113] Training in complex dynamic scenarios is difficult, and direct training models have difficulty converging. This chapter designs a step-back experience learning (SEL) training method based on past experience. This section will describe the specific process of this method.
[0114] (1) Accumulating experience step by step through curriculum learning: Curriculum learning (CL) is a machine learning method that mimics the human learning process. Its core idea is to gradually increase the difficulty of tasks from simple to complex, so that the agent can gradually master complex skills during the learning process, thereby improving learning efficiency and final performance. By guiding the agent to transition from simple tasks to complex tasks in stages through a curriculum, the risk of strategy collapse caused by the sudden increase in task difficulty in the early stages of training is effectively reduced, and training stability is improved. Directed exploration based on updating parameters in the complex task stage can avoid the blindness of random exploration and accelerate strategy convergence.
[0115] CL first decomposes complex tasks and stores them in task space. Each task T i Corresponding to a difficulty level, by optimizing the task distribution, the agent can maximize the cumulative reward during the training process. Then, a performance threshold needs to be set for each task. When the robot's performance on the current task reaches the threshold, it automatically switches to the next task. The principle is as follows Figure 5 shown.
[0116] The robotic arm's dynamic obstacle avoidance research divides tasks into four phases. Phase 1: Only a target point is set in the training scenario, allowing the robotic arm to learn the reach task in a simple environment. Phase 2: A single static obstacle is added to the environment, allowing the robot to learn the simple obstacle avoidance task. Phase 3: A random number of static obstacles are added to the environment, allowing the robot to learn the reach task in a complex environment. Phase 4: Building on the third phase, a trajectory is created based on static obstacles in the environment, allowing the robot to learn dynamic obstacle avoidance and reach tasks in a complex environment. The performance threshold for each of these four phases is set as the success rate. When the success rate of each task remains stable above 95%, the robot switches to the next task.
[0117] (2) Implementation of SAC algorithm based on experience replay:
[0118] 1. In the field of deep reinforcement learning, reward function design has always been a difficult problem. An excellent reward function can help the model converge quickly, but in most tasks, the agent can only receive rewards when the task is completed, which leads to reward function problems. The same problem also exists in the dynamic obstacle avoidance task of the robotic arm. To solve this problem, this paper introduces the HER algorithm for auxiliary training when implementing the SAC algorithm. The HER experience storage mechanism perfectly matches the experience replay pool of the SAC algorithm.
[0119] The introduction of the HER algorithm can store a large amount of failed experiences in the early stages. For example, in the current training round, the target point is at point A, and the robot runs to point B. At this time, HER can store the trajectory of the robot reaching point B in the experience replay pool. When the subsequent task is refreshed to define the target point at point B, this experience can be directly utilized. The principle of the HER algorithm is shown in Figure 4, which is also the main technical point of the past experience learning framework. By utilizing the target information in the failed experience through the HER mechanism, the strategy convergence is significantly accelerated, and combined with the maximum entropy framework of SAC, SAC-HER can achieve a better balance between exploration and utilization. The parameterized update mechanism of historical experience enhances the cross-task knowledge transfer ability, avoids repeated exploration of the known state space, and significantly improves training efficiency;
[0120] 2. To fully leverage past experience, the SAC algorithm's network architecture utilizes a long short-term memory (LSTM) network. The unique mechanism of LSTM networks allows the network to dynamically determine information retention and forgetting, thereby better learning long-term dependencies. In the robotic arm's dynamic obstacle avoidance task, past environmental state information can be used to reason about the current strategy. In particular, combining the past motion trajectories of moving obstacles can help the model develop a better obstacle avoidance strategy.
[0121] Step 5): Design of smoothing method for robot arm motion trajectory
[0122] In the process of reinforcement learning to control a robotic arm, directly optimizing the Q value may result in an uneven action sequence, which in turn affects the stability of the robot's task execution. To address this problem, we introduce trajectory interpolation constraints into the SAC algorithm to make the output action smoother in the time dimension. Specifically, we use two methods:
[0123] First, the trajectory interpolation formula is directly introduced into the Actor network structure, so that the action generated by SAC has time continuity. t =δ·a raw +(1-δ)·a t-1 , where a raw The original action directly output by the Actor, at-1 is the action executed at the previous time step, and δ is the smoothing coefficient. This effectively reduces sudden changes in action, making the control signal smoother. Furthermore, we can use the LSTM structure to memorize historical information and combine it with a specific trajectory interpolation model to predict a more reasonable action sequence.
[0124] Secondly, we add a trajectory smoothness constraint to the Actor's loss function to actively optimize the generation of smooth action sequences. Specifically, we define the first-order difference smoothing loss:
[0125] L smooth =∑||a t -a t-1 || 2 ;
[0126] To suppress drastic changes in speed, use the second-order difference smoothing loss:
[0127] L smooth =∑||(a t -a t-1 )-(a t-1 -a t-2 )|| 2 ;
[0128] To control the fluctuation of acceleration. The final Actor loss function is:
[0129] L actor =-Q(s,a)+λ smooth ·L smooth ;
[0130] The trade-off between control smoothness and policy optimization objectives. Through the above method, we can simultaneously optimize task performance and action smoothness during reinforcement learning training, thereby improving the execution stability of the robotic arm.
[0131] Step 6): When tuning the SAC algorithm, focus on several key hyperparameters. First, select an appropriate network structure (such as the number of layers, number of neurons per layer, and activation function) to ensure the model has sufficient expressive power without overfitting. Second, optimize the learning rate and its scheduling strategy to avoid divergence or slow convergence during training. The specific parameters are shown in Table 1.
[0132] Table 1. Parameters of deep reinforcement learning controller
[0133]
[0134] Step 6): Training environment deployment and training testing
[0135] During the training process, the model is saved in real time. After the training is completed, the model is deployed to the environment for testing. The total number of training rounds is 2000, and the maximum step size of each round is 2000 steps. In order to obtain the best model during the training process, the average reward is calculated every 1000 training steps, and the model with the highest reward is saved in real time. The evaluation process is as follows Figure 6 shown.
[0136] Step 7): Migrate from simulation to real environment
[0137] To address the physical property differences between simulated and real-world environments, this paper employs a robustness enhancement method based on environmental noise injection. This method actively introduces state observation noise and action execution perturbations during simulation training to simulate sensor errors, actuator biases, and dynamic uncertainties present in real-world environments, thereby improving the generalization and transfer performance of reinforcement learning strategies. The specific implementation is as follows:
[0138] (1) State space noise injection: In the state observation phase of the simulation environment, Gaussian noise is added to the joint position, velocity, and obstacle distance information of the robot arm:
[0139]
[0140] To simulate the random error in real sensor measurement. The state observation value after noise injection It can be expressed as:
[0141]
[0142] Among them, s t is the original state value, is the noise standard deviation, and its value is set according to the real sensor calibration data. By introducing state noise, the policy network can learn a robust representation of observation deviations.
[0143] (2) Action output disturbance: During the action execution phase, a uniformly distributed disturbance is added to the robot arm joint velocity command. To simulate the dynamic response error of the real actuator. Action command after disturbance for:
[0144]
[0145] Among them, a t is the original action output by the policy network, and b is the perturbation amplitude. Through action perturbation, the policy network can adapt to the non-ideal response characteristics of the actuator.
[0146] (3) To further enhance the generalization ability, the noise parameters are randomly sampled within a preset range during training to avoid overfitting of the strategy to a fixed noise pattern.
[0147] In summary, this solution uses deep reinforcement learning technology to achieve real-time dynamic obstacle avoidance for the robotic arm based on past experience, without the need to pre-establish a mathematical model for solution. By using a model-free method and a neural network to autonomously learn task objectives, the obstacle avoidance process has real-time learning capabilities, improving the flexibility of obstacle avoidance and enhancing its adaptability to dynamic scenarios.
[0148] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in the present invention, and such modifications or substitutions are intended to be within the scope of protection of the present invention. Therefore, the scope of protection of the present invention shall be subject to the scope of protection of the claims.
Claims
1. A dynamic obstacle avoidance method for a robotic arm using reinforcement learning based on past experience, characterized in that: The method steps include: S1. Constructing a neural network using an improved SAC algorithm, wherein the improved SAC algorithm is embedded in a long short-term memory network and a trajectory interpolation constraint is introduced. The neural network includes a policy network and a value network; S2. Construct action space, state space and reward function based on the environment, obstacle position and robot arm posture; S3. Based on the constructed action space, state space, and reward function, the neural network is trained through reinforcement learning. The HER algorithm and experience replay pool are introduced for auxiliary training. During the training process, the curriculum learning method is used to enable the agent to accumulate experience step by step. Finally, the neural network converges and outputs actions. S4, using the output action in S3 to control the robotic arm, and perform collision detection and error calculation to obtain the detection result; S5. Adjust and update the neural network parameters based on the detection results and the experience replay pool, and return to execute step S2 until the robotic arm reaches the target point.
2. The method for dynamic obstacle avoidance of a robotic arm using reinforcement learning based on past experience according to claim 1, characterized in that: The trajectory interpolation constraint introduced in the improved SAC algorithm in S1 includes the first constraint or the second constraint; The first constraint is: introducing a trajectory interpolation formula into the Actor network structure so that the actions generated by SAC have temporal continuity; The second constraint is to add a trajectory smoothness constraint to the Actor's loss function and actively optimize to generate a smooth action sequence.
3. The method for dynamic obstacle avoidance of a robotic arm using reinforcement learning based on past experience according to claim 2, characterized in that: The first constraint is achieved by using a first-order low-pass filter, and its specific formula is: a t =δ·a raw +(1-δ)·a t-1 ; Among them, a t is the execution action of the current time step; a raw The original action directly output by the Actor; a t-1 is the execution action of the previous time step; δ is the smoothing coefficient.
4. The method for dynamic obstacle avoidance of a robotic arm using reinforcement learning based on past experience according to claim 2, characterized in that: The specific process of the second constraint is as follows: first, define the first-order difference smoothing loss to suppress drastic changes in speed, then use the second-order difference smoothing loss to control the fluctuation of acceleration, and finally use the Actor loss function to control the trade-off between smoothness and policy optimization objectives. The specific formula is: L smooth1 =∑||a t -a t-1 || 2 ; L smooth2 =∑||(a t -a t-1 )-(a t-1 -a t-2 )|| 2 ; L smooth =L smooth1 +L smooth2 ; L actor =-Q(s,a)+λ smooth ·L smooth ; Among them, L smooth1 is the first-order difference smoothing loss; L smooth2 is the second-order difference smoothing loss; L actor is the Actor loss function; a t is the action vector at time t; a t-1 is the action vector at time t-1; a t-2 is the action vector at time t-2; Q(s,a) is the input state s t and action a t Q function; λ smooth is the smoothing weight parameter; L smooth is the total smoothing loss.
5. The method for dynamic obstacle avoidance of a robotic arm using reinforcement learning based on past experience according to claim 1, characterized in that: The action space constructed in S2 is composed of a set of joint velocities, Indicates that the range of the action value is set within [-1,1]rad / s, and the maximum change of each control cycle is set to 0.1rad / s.
6. The method for dynamic obstacle avoidance of a robotic arm using reinforcement learning based on past experience according to claim 5, characterized in that: The state space constructed in S2 includes the state information of the manipulator, the position of the nearest obstacle, and the distance between the end effector and the target point. The specific formula is: S=[S bot ,S obs ,S goal ]; S bot =[J s ,J v ,J p ]: Among them, J s Indicates the position of the end effector of the machine; J v represents the joint velocity; J p Indicates joint position; S bot is the status information of the robot arm; S obs is the position of the nearest obstacle; S goal is the distance between the end effector and the target point.
7. The method for dynamic obstacle avoidance of a robotic arm using reinforcement learning based on past experience according to claim 6, characterized in that: The reward function in S2 is constructed based on the APF algorithm. The specific process of constructing the reward function is: first, a repulsive field is constructed at the obstacle, and a gravitational field is constructed at the target point. Then, positive rewards are provided within the preset area of the gravitational field, and negative rewards are provided within the preset area of the repulsive field. Finally, the composite field formed by the gravitational field and the repulsive field constructs a reward for the current position of the robotic arm.
8. The method for dynamic obstacle avoidance of a robotic arm using reinforcement learning based on past experience according to claim 7, characterized in that: The reward function includes a gravitational field reward term, a repulsive field reward term, a collision penalty term, and a distance penalty term. The specific formula is: R=R att +R collision -R rep -R dis ; R att =a·d; R dis =η·d; Among them, R att is the gravitational field bonus; R rep is the repulsive field reward item; R collision is the collision penalty term; R dis is the distance penalty term; α is the gain coefficient of the gravitational field; D(q) is the distance between the end effector and the nearest obstacle repulsion calculation point; β is the repulsion gain coefficient; is the action threshold for triggering repulsion; d is the distance between the end effector and the target point; η is the distance gain coefficient.
9. The method for dynamic obstacle avoidance of a robotic arm using reinforcement learning based on past experience according to claim 1, characterized in that: The specific process of using the curriculum learning method in S3 to enable the intelligent agent to accumulate experience step by step is: first, the complex tasks are decomposed using the curriculum learning method and the tasks are stored in the task space, where each task corresponds to a difficulty level. Then, a preset performance threshold is set for each task. When the performance of the robotic arm on the current task reaches the preset performance threshold, it automatically switches to the next task.
10. A dynamic obstacle avoidance system for a robotic arm using reinforcement learning based on past experience, characterized in that: The system applies a reinforcement learning manipulator dynamic obstacle avoidance method based on past experience as described in any one of claims 1 to 9, and the system includes a network construction module, a state-action space modeling module, a reinforcement learning training module, and a control and feedback optimization module; The network construction module is used to construct a neural network using an improved SAC algorithm; The state-action space modeling module is used to construct the action space, state space and reward function based on the environment, obstacle positions and robot arm posture; The reinforcement learning training module is used to combine the HER algorithm, the experience replay pool and the curriculum learning method to train the neural network until the neural network converges and outputs actions; The control and feedback optimization module is used to achieve closed-loop control through collision detection, error calculation and parameter update, and iterates until the robotic arm reaches the target point.
Citation Information
Patent Citations
Multi-mobile robot autonomous obstacle avoidance method based on deep reinforcement learning
CN117873116A
Cited By
Space debris intelligent capture and treatment method and system based on SAC algorithm
CN120941421A
Double-arm collaborative planning method, system and device based on reinforcement learning and medium
CN121004618A