Robot path planning method based on deep reinforcement learning

By applying a robot path planning method based on deep reinforcement learning in an agricultural environment, combining DDPG algorithm and differential game strategy, the challenge of dynamic obstacles to path planning is solved, and the efficiency of robot obstacle avoidance and path planning is improved.

CN120215511AActive Publication Date: 2025-06-27CHANGCHUN UNIV OF TECH

Patent Information

Application Number
CN202510681542.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-26
Publication Date
2025-06-27
Estimated Expiration
2045-05-26

AI Technical Summary

Technical Problem

The prior art is difficult to effectively deal with dynamic obstacles when planning robot paths in agricultural environments, resulting in low path planning efficiency, waste of computing resources and delayed response.

Method used

The robot path planning method based on deep reinforcement learning is adopted, combined with the improved deep deterministic strategy gradient (DDPG) algorithm and differential game strategy improvement, and path planning is optimized through the multimodal weighted combination reward mechanism.

Benefits of technology

It improves the obstacle avoidance ability and path planning efficiency of the robot in a dynamic environment, ensuring that the robot can effectively avoid obstacles and safely and efficiently reach the target point.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120215511A_ABST
    Figure CN120215511A_ABST
Patent Text Reader

Abstract

The invention discloses a robot path planning method based on deep reinforcement learning, and relates to the fields of intelligent agriculture, path planning, robots and the like. The method comprises the following steps: firstly, sensing a farm environment, defining a state space and an action space of a robot, and setting a multi-modal weighted combined reward mechanism and an experience playback buffer area; a learnable weight coefficient is introduced into a Critic network loss function in a traditional DDPG algorithm, an entropy regularization item is added into a target function of an Actor network, then a differential game is selected through an adaptive attenuation # imgabs0 # greedy strategy to generate a control strategy or a DDPG algorithm to generate an action, finally the action or the control strategy is executed, network parameters and target network parameters are updated, and a target network is obtained. And dynamically updating the experience playback buffer area. Compared with other path planning methods, the method improves the adaptability of path planning to a dynamic environment, and also has good efficiency and safety in a complex agricultural environment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of intelligent agriculture, path planning, and robotics, and particularly to a robot path planning method based on deep reinforcement learning. Background Art

[0002] With the development of global agriculture towards intelligence and automation, robots are increasingly widely used in agricultural operations. Robots can efficiently perform tasks such as sowing, weeding, fertilizing, spraying pesticides, and harvesting, greatly improving production efficiency and reducing labor costs. In complex agricultural environments such as orchards, greenhouses, and farms, robots need to have efficient path planning capabilities to adapt to dynamic, multi-obstacle, unstructured, and uneven ground agricultural environments and improve work efficiency.

[0003] The agricultural environment is more complex and changeable than other environments. The operation process of robots involves making decisions under uncertain and dynamic complex conditions, and path planning faces many challenges. Nowadays, multi-robot systems are widely used in the agricultural field. In agricultural environments such as farms and orchards, there are not only the interferences of static obstacles and uneven roads, but also dynamic obstacle avoidance problems and local path planning problems among robots.

[0004] Traditional local obstacle avoidance algorithms usually assume that obstacles are static or their movement is very slow, which limits their ability to handle dynamic obstacles. In a multi-robot agricultural operation system, dynamic obstacles are mainly other robots that are performing agricultural operations. Traditional methods often cannot effectively predict or avoid the movement of these obstacles. For commonly used path planning algorithms such as A* and Dijkstra algorithms, when facing a dynamically changing environment, they need to frequently perform replanning, resulting in a large amount of computational effort. If a robot needs to avoid obstacles in real time and consumes a large amount of computational resources each time it recalculates the path, this will affect the real-time performance of the robot in practical applications, especially in environments such as agriculture that require efficient operations, and may lead to response delays. The path planning method based on the artificial potential field method is prone to falling into local optimal solutions, which not only wastes time and computational resources but also may cause the robot to be unable to reach the target smoothly or to circle repeatedly and get lost during task execution. The traditional dynamic window method relies on local information and is prone to falling into local optimality. The dynamic window algorithm only considers the feasible speed within a short time and does not consider the global optimal path, which may cause the robot to fall into a local optimal solution, resulting in redundant paths or unnecessary stagnation. Summary of the Invention

[0005] In view of the above problems and actual requirements, the present invention proposes a robot path planning method based on deep reinforcement learning. The improved Deep Deterministic Policy Gradient (DDPG) algorithm is adopted as the core, and the differential game strategy is combined to enhance the obstacle avoidance ability of robots in the agricultural environment under dynamic conditions. This method balances the control strategy generated by the differential game and the action generated by the DDPG algorithm through an adaptive decay greedy strategy, improves the path planning efficiency, and ensures that the robot can effectively avoid obstacles in the dynamic environment and reach the target point safely and efficiently. At the same time, an optimization method based on a multi-modal weighted combination reward mechanism is adopted to ensure that the robot preferentially selects the optimal path.

[0006] The present invention proposes a robot path planning method based on deep reinforcement learning, which is realized through the following steps:

[0007] Step 1: The robot Performs environmental perception through a lidar in the farm, defines the current state according to the perception information, and defines the action space.

[0008] Step 1.1: Define the current state vector of the robot :

[0009] s t = [ d T , φ T , v t , ω t , d Oi , φ Oi ] ,

[0010] In the state vector, is the distance between the robot and the target, is the angle with the target, is the linear velocity, is the angular velocity, is the distance from the obstacle , is the angle with the obstacle , is the obstacle serial number;

[0011] Step 1.2: Based on the current state of the robot, define the position coordinates as , and perform kinematic modeling;

[0012] Step 1.3: Define the action space of the robot. The action space is represented by the set of actions that the robot can choose:

[0013] a t = [ v t , ω t ] .

[0014] Step 2: Design a multi-modal weighted combination reward mechanism, adjust the dynamic weight through the hyperbolic tangent function and the obstacle density function based on this reward mechanism, and set up an experience replay buffer;

[0015] Step 2.1: The reward mechanism consists of three sub - reward items, namely the path - guiding reward , the obstacle - avoidance constraint reward , and the smooth - constraint reward . By means of dynamic weights and , the priorities of the sub - rewards are adjusted, and the comprehensive reward function is as follows:

[0016] ,

[0017] Step 2.2: The path - guiding reward function is as follows:

[0018] ,

[0019] where is the distance between the robot and the target point , is the maximum reward of the path - guiding reward function, is the growth - rate constant, is the estimated maximum distance from the starting point to the ending point;

[0020] The obstacle - avoidance constraint reward function is as follows:

[0021] ,

[0022] where is the distance between the robot and the obstacle , is the maximum reward of the collision - constraint reward function, is the growth - rate constant;

[0023] The smooth - constraint reward function is as follows:

[0024] ,

[0025] where , is the reward - weight coefficient, is the curvature change rate, is the axial acceleration;

[0026] The adjustment methods of the weights and in the dynamic - weight adaptive mechanism are as follows:

[0027] ,

[0028] ,

[0029] wherein is the initial weight, is the final weight, is the time scaling coefficient, is the basic time coefficient, obstacle influence factor, is the obstacle density function in the farm environment, is the hyperbolic tangent function;

[0030] Step 2.3: Set up the experience replay buffer , the robot selects an action according to the current state , calculates the reward according to the preset reward mechanism , enters the next state , and stores the information in the experience replay buffer .

[0031] Step 3: Initialize the Actor-Critic network, introduce a learnable weight coefficient in the loss function of the Critic network , dynamically adjust the contribution of different states and actions to the loss, add an entropy regularization term to the objective function of the Actor network, and initialize the target network.

[0032] Step 3.1: The goal of the Critic network is to learn the Q-value function, that is, given the current state and action , calculate the long-term cumulative reward, and the expression is:

[0033] Q ( s t , a t ; θ Critic ) =  [ R t + γ ⋅ m a x ( a t + 1 ) Q ' ( s t + 1 , a t + 1 ; θ Critic − ) ] ,

[0034] wherein is the expected cumulative reward after executing the action in the state , is the average expected value, is the reward immediately obtained after executing the action in the state , is the discount factor, balancing the weights of immediate rewards and cumulative rewards, is the cumulative reward after executing the action in the next state , is the action generated by the Actor network in the next state, is the maximum Q-value estimate among all actions in the next state, are the parameters of the Critic network, are the parameters of the Critic target network;

[0035] The loss function of the Critic network is as follows:

[0036] L ( θ Critic ) =  s ~ D [ m ⋅ ( y t − Q ( s t , a t ; θ Critic ) ) 2 ] ,

[0037] where is a learnable weight coefficient that dynamically adjusts the contribution of the loss in the current state to the overall optimization according to the importance of the samples, is the target Q value. The parameters of the Critic network are optimized by minimizing the loss function to reduce the TD error, represents the expectation calculated for the state sampled from the experience replay buffer .

[0038] The goal of the Actor network is to optimize the policy by maximizing the Q value given by the Critic network. Entropy regularization is added to the objective function to improve exploration. The objective function is:

[0039] J Actor = −  s ~ D [ Q ( s t , a t ; θ Critic ) ] 2 + α ℋ ( a t ) ,

[0040] where is the entropy of the action , is the regularization coefficient;

[0041] Step 3.2: To determine whether to use differential game to generate control strategies or the deep deterministic policy gradient algorithm to generate actions in the current state, an adaptive decay greedy strategy is introduced, and the expression is:

[0042] ,

[0043] ,

[0044] where is the environmental complexity function of the state , is the decay adjustment parameter, is the exploration rate of the current state, is the minimum exploration rate, is the initial exploration rate, is the preset policy switching threshold, is a random number generated by uniform sampling, is the action with the largest Q value, is the control strategy generated by differential game.

[0045] Step 4: If the current state passes the adaptive attenuation The greedy strategy selects differential game for dynamic obstacle avoidance. By establishing a differential game model, a control strategy is generated.

[0046] Step 4.1: Establish a differential game model. Define the first robot as , and define the second robot as ;

[0047] Define and The control strategies are respectively u 0 = [ v 0 , ω 0 ] and u 1 = [ v 1 , ω 1 ] ;

[0048] Step 4.2: Set the cost function of which is expressed as:

[0049] J 0 = ∫ 0 T 1 [ w 1 d − 1 + w 2 ‖ u 0 ‖ 2 + w 3 ‖ d O i ‖ 2 ] d t ,

[0050] where , , are weight parameters to balance the importance of different terms in the cost function, to keep the two robots as far apart as possible, to balance the energy consumption, to make the robot approach the target point, is the distance between the two robots;

[0051] Set the cost function of robot which is expressed as:

[0052] J 1 = ∫ 0 T 1 [ w 4 d − 1 + w 5 ‖ u 1 ‖ 2 ] d t ,

[0053] where , are also weight parameters, to keep them as far apart as possible, to control the energy consumption;

[0054] Step 4.3: The feedback Nash equilibrium ensures that robot and robot optimize their strategies with each other in the differential game, and any unilateral change in strategy will not make the cost functions of the two smaller. Construct the Hamiltonian , :

[0055] ,​​

[0056] Among them , represents the kinematic equation of the robot, , is the co-state variable of the robot;

[0057] Step 4.4: Take the derivative of the constructed Hamiltonian with respect to the control strategy, set the derivative function to zero, and then solve the Hamiltonian equation to obtain the feedback control strategy;

[0058] Step 4.5: The control strategies of the robot and are updated by the gradient descent method:

[0059] ,

[0060] until the objective function satisfies the convergence condition. The feedback Nash equilibrium convergence condition is as follows:

[0061] ,

[0062] Among them , is the final control strategy, is the update parameter.

[0063] Step 5: The robot executes an action, obtains a reward and enters the next state, and stores the data in the experience replay buffer; calculates the target Q value, then generates the next action through the Actor network, and updates the Actor-Critic network and target network parameters, repeating the loop until the robot reaches the target point.

[0064] Step 5.1: Select and execute an action according to the current strategy, obtain a reward according to the current state and the reward mechanism after executing the action, and enter the next state;

[0065] Step 5.2: Execute the action and interact with the environment, obtain a reward according to the reward mechanism and observe the new state ;

[0066] Step 5.3: Store the experience of the current interaction in the experience pool for subsequent training. At the same time, extract a small batch of data samples from the experience pool and use this sample to update the Critic network parameters, and calculate the probability of each experience sample being sampled:

[0067] ,

[0068] Among them, is the normalization term to ensure that the sum of the sample sampling probabilities is 1, is a hyperparameter, is the TD error.

[0069] Step 5.4: Calculate the target Q value as follows:

[0070] ;

[0071] The goal is to make the Q value of the current state as close as possible to the target Q value. Therefore, training the Critic network requires minimizing the following mean squared error loss function:

[0072] L ( θ ) =  s ~ D [ m ⋅ ( y t − Q ( s t , a t ; θ Critic ) ) 2 ] ;

[0073] Step 5.5: Update the Actor-Critic parameters. The parameters of the Critic network are updated by gradient descent:

[0074] ,

[0075] where, is the learning rate, which controls the step size of each parameter update, is the gradient of the loss function with respect to the parameter ;

[0076] Use the gradient of the Critic network to update the parameters of the Actor network. The parameters of the Actor network are updated using the gradient ascent method, and the expression is as follows:

[0077] ,

[0078] where, is the gradient of the objective function with respect to the parameter ;

[0079] The target network adopts a soft update method, and the expression is as follows:

[0080] ,

[0081] ,

[0082] where, and are the network parameters of the Actor-Critic network, and are the target network parameters, is the network parameter update coefficient;

[0083] Step 5.6: Repeat the training, and the training ends when the termination condition is reached. Description of the Drawings

[0084] Figure 1 This is the overall flowchart of the embodiments of the present invention. Detailed implementation manners

[0085] In order to more clearly elaborate the purpose, technical solutions and their advantages of the present invention, the present invention will be described in detail below with reference to the accompanying drawings and specific embodiments.

[0086] Accompanying Figure 1 This is the overall flowchart of the embodiments of the present invention. This embodiment provides a robot path planning method based on deep reinforcement learning, which specifically includes the following processes: The farm environment is perceived, and the robot state space and action space are defined. A multi-modal weighted combination reward mechanism and an experience replay buffer are set up. The network parameters and target network parameters of the improved DDPG algorithm are initialized. Through adaptive attenuation The greedy strategy is used to select the differential game or the DDPG algorithm. According to the selection of the greedy strategy, the control strategy is generated through the differential game or the action is generated through the DDPG algorithm. The action or control strategy is executed, and the network parameters and target network parameters are updated, and the experience replay buffer is updated.

[0087] Step 1: The robot Perceives the environment in the farm through lidar, defines the current state according to the perception information, and defines the action space.

[0088] Step 1.1: Define the current state vector of the robot:

[0089] s t = [ d T , φ T , v t , ω t , d Oi , φ Oi ] ,

[0090] In the state vector, is the distance between the robot and the target, is the included angle with the target, is the linear velocity, is the angular velocity, is the distance from the obstacle, is the included angle with the obstacle, is the obstacle serial number;

[0091] Step 1.2: Based on the current state of the robot, define the position coordinates as , and perform kinematic modeling;

[0092] Step 1.3: Define the action space of the robot. The action space is represented by the set of actions that the robot can choose:

[0093] a t = [ v t , ω t ] .

[0094] Step 2: Design a multi-modal weighted combination reward mechanism. Based on this reward mechanism, adjust the dynamic weights through the hyperbolic tangent function and the obstacle density function, and set up an experience replay buffer;

[0095] Step 2.1: The reward mechanism includes three sub-reward items, namely the path guidance reward , the obstacle avoidance constraint reward and the smoothness constraint reward . Adjust the priorities of the sub-rewards through the dynamic weights and . The comprehensive reward function is:

[0096] ,

[0097] Step 2.2: The path guidance reward function is:

[0098] ,

[0099] where is the distance between the robot and the target point , is the maximum reward of the path guidance reward function, is the growth rate constant, is the estimated maximum distance from the starting point to the ending point;

[0100] The obstacle avoidance constraint reward function is:

[0101] ,

[0102] where is the distance between the robot and the obstacle , is the maximum reward of the collision constraint reward function, is the growth rate constant;

[0103] The smoothness constraint reward function is:

[0104] ,

[0105] where , is the reward weight coefficient, is the curvature change rate, is the axial acceleration;

[0106] In the dynamic weight adaptation mechanism, the weight and The adjustment methods are as follows:

[0107] ,

[0108] ,

[0109] where is the initial weight, is the final weight, is the time scaling factor, is the basic time coefficient, is the obstacle influence factor, is the obstacle density function in the farm environment, is the hyperbolic tangent function;

[0110] Step 2.3: Set up the experience replay buffer , the robot selects an action according to the current state , calculates the reward according to the preset reward mechanism , enters the next state , and stores the information in the experience replay buffer .

[0111] Step 3: Initialize the Actor-Critic network, introduce a learnable weight coefficient into the loss function of the Critic network , dynamically adjust the contribution of different states and actions to the loss, add an entropy regularization term to the objective function of the Actor network, and initialize the target network.

[0112] Step 3.1: The goal of the Critic network is to learn the Q-value function, that is, to calculate the long-term cumulative reward when given the current state and action , and the expression is:

[0113] Q ( s t , a t ; θ Critic ) =  [ R t + γ ⋅ m a x ( a t + 1 ) Q ' ( s t + 1 , a t + 1 ; θ Critic − ) ] ,

[0114] where is the expected cumulative reward after executing the action in the state , is the average expected value, is the reward immediately obtained after executing the action in the state , is the discount factor, balancing the weights of immediate rewards and cumulative rewards, is the cumulative reward after executing the action in the next state . is the action generated by the Actor network in the next state, is the maximum Q-value estimate among all actions in the next state, are the Critic network parameters, are the Critic target network parameters;

[0115] The loss function of the Critic network is as follows:

[0116] L ( θ Critic ) =  s ~ D [ m ⋅ ( y t − Q ( s t , a t ; θ Critic ) ) 2 ] ,

[0117] where is the learnable weight coefficient, which dynamically adjusts the contribution of the loss in the current state to the overall optimization according to the importance of the samples, is the target Q-value. By minimizing the loss function, the parameters of the Critic network are optimized to reduce the TD error, represents the expectation calculated for the state sampled from the experience replay buffer .

[0118] The goal of the Actor network is to optimize the policy by maximizing the Q-value given by the Critic network. Entropy regularization is added to the objective function to improve exploration. The objective function is:

[0119] J Actor = −  s ~ D [ Q ( s t , a t ; θ Critic ) ] 2 + α ℋ ( a t ) ,

[0120] where is the entropy of the action , is the regularization coefficient;

[0121] Step 3.2: To determine whether to use differential game to generate control strategies or use deep deterministic policy gradient algorithm to generate actions in the current state, an adaptive decay greedy policy is introduced, and the expression is:

[0122] ,

[0123] ,

[0124] where is the environmental complexity function of the state , is the decay adjustment parameter, is the exploration rate of the current state, is the minimum exploration rate, is the initial exploration rate, is the preset strategy switching threshold, is a random number generated by uniform sampling, is the action with the largest Q value, is the control strategy generated by the differential game.

[0125] Step 4: If the current state is adaptively decayed The greedy strategy selects differential game for dynamic obstacle avoidance, and generates control strategy by establishing differential game model.

[0126] Step 4.1: Establish a differential game model and define robot No. 1 as , define robot No. 2 as ;

[0127] definition and The control strategies are u 0 = [ v 0 , ω 0 ] and u 1 = [ v 1 , ω 1 ] ;

[0128] Step 4.2: Setup Cost function It is expressed as:

[0129] J 0 = ∫ 0 T 1 [ w 1 d − 1 + w 2 ‖ u 0 ‖ 2 + w 3 ‖ d O i ‖ 2 ] d t ,

[0130] in , , is a weight parameter that balances the importance of different terms in the cost function, Keep the two robots as far apart as possible. Balance energy consumption, Make the robot Close to the target point, is the distance between the two robots;

[0131] Setting up the robot Cost function It is expressed as:

[0132] J 1 = ∫ 0 T 1 [ w 4 d − 1 + w 5 ‖ u 1 ‖ 2 ] d t ,

[0133] in , Also the weight parameter, Keep the two as far apart as possible. Control energy consumption;

[0134] Step 4.3: Feedback Nash equilibrium ensures the robot and robots In differential games, optimize strategies mutually. Any unilateral change in strategy will not make the cost functions of both smaller. Construct the Hamiltonian , :

[0135] ,

[0136] where , represents the kinematic equation of the robot, , is the co-state variable of the robot;

[0137] Step 4.4: Take the derivative of the constructed Hamiltonian with respect to the control strategy, set the derivative function to zero, and then solve the Hamiltonian equation to find the feedback control strategy;

[0138] Step 4.5: The control strategies of the robot and are updated by the gradient descent method:

[0139] ,

[0140] until the objective function satisfies the convergence condition. The feedback Nash equilibrium convergence condition is as follows:

[0141] ,

[0142] where , is the final control strategy, is the update parameter.

[0143] Step 5: The robot executes an action, obtains a reward, and enters the next state, storing the data in the experience replay buffer; calculates the target Q value, then generates the next action through the Actor network, and updates the Actor-Critic network and target network parameters, repeating the loop until the robot reaches the target point.

[0144] Step 5.1: Select and execute an action according to the current strategy. After executing the action, obtain a reward according to the current state and the reward mechanism, and enter the next state;

[0145] Step 5.2: Execute the action , interact with the environment, obtain a reward according to the reward mechanism , and observe the new state ;

[0146] Step 5.3: Store the experience of the current interaction in the experience pool for subsequent training. At the same time, extract a small batch of data samples from the experience pool and use this sample to update the Critic network parameters, calculating the probability of each experience sample being sampled:

[0147] ,

[0148] Among them, is the normalization term, ensuring that the sampling probabilities of samples sum to 1, is a hyperparameter, is the TD error.

[0149] Step 5.4: Calculate the target Q value as follows:

[0150] ;

[0151] The goal is to make the Q value of the current state as close as possible to the target Q value. Therefore, training the Critic network requires minimizing the following mean squared error loss function:

[0152] L ( θ ) =  s ~ D [ m ⋅ ( y t − Q ( s t , a t ; θ Critic ) ) 2 ] ;

[0153] Step 5.5: Update the Actor-Critic parameters. The parameters of the Critic network are updated by gradient descent:

[0154] ,

[0155] Among them, is the learning rate, controlling the step size of each parameter update, is the gradient of the loss function with respect to the parameter ;

[0156] Use the gradient of the Critic network to update the parameters of the Actor network. The parameters of the Actor network are updated using the gradient ascent method, and the expression is as follows:

[0157] ,

[0158] Among them, is the gradient of the objective function with respect to the parameter ;

[0159] The target network adopts a soft update method, and the expression is as follows:

[0160] ,

[0161] ,

[0162] Among them, and are the network parameters of the Actor-Critic network, and are the parameters of the target network, is the network parameter update coefficient;

[0163] Step 5.6: Repeat the training, and the training ends when the termination condition is reached.

Claims

1. A robot path planning method based on deep reinforcement learning, characterized in that, Including the following steps: Step 1: Robot Perform environmental perception in the farm through lidar, define the current state according to the perception information, and define the action space; Step 2: Design a multi-modal weighted combination reward mechanism. On the basis of this reward mechanism, adjust the dynamic weights through the hyperbolic tangent function and the obstacle density function, and set up an experience replay buffer; Step 3: Initialize the Actor-Critic network, introduce a learnable weight coefficient into the loss function of the Critic network , dynamically adjust the contribution of different states and actions to the loss, add an entropy regularization term to the objective function of the Actor network, and initialize the target network; Step 4: If the current state passes the adaptive attenuation The greedy strategy selects differential game for dynamic obstacle avoidance, generates a control strategy by establishing a differential game model; Step 5: The robot executes an action, obtains a reward and enters the next state, and stores the data in the experience replay buffer; Calculate the target Q value, then generate the next action through the Actor network, and update the Actor-Critic network and the target network parameters, and repeat the loop until the robot reaches the target point.

2. The robot path planning method based on deep reinforcement learning according to claim 1, wherein, The design of the multi-modal weighted combination reward mechanism described in Step 2. On the basis of this reward mechanism, adjust the dynamic weights through the hyperbolic tangent function and the obstacle density function, and set up an experience replay buffer; Specifically, it is implemented according to the following steps: Step 2: Design a multi-modal weighted combination reward mechanism. On the basis of this reward mechanism, adjust the dynamic weights through the hyperbolic tangent function and the obstacle density function, and set up an experience replay buffer; Step 2.1: The reward mechanism includes three sub - reward items, namely the path - guiding reward , the obstacle - avoidance constraint reward , and the smooth - movement constraint reward . By using dynamic weights and to adjust the priorities of the sub - rewards, the comprehensive reward function is as follows: , Step 2.2: Path guidance reward function is as follows: , wherein is the robot and the target point distance, is the maximum reward of the path guidance reward function, is the growth rate constant, is the estimated maximum distance from the starting point to the ending point; Obstacle avoidance constraint reward function is as follows: , Among them is the robot and the obstacle distance is the maximum reward of the collision constraint reward function is the growth rate constant; Smoothing Constraint Reward Function is as follows: , Among them , is the reward weight coefficient, is the curvature change rate, is the axial acceleration; Weights in the dynamic weight adaptive mechanism and are adjusted as follows respectively: , , Among them is the initial weight, is the final weight, is the time scaling coefficient, is the basic time coefficient, obstacle influence factor, is the obstacle density function in the farm environment, is the hyperbolic tangent function; Step 2.3: Set up the experience replay buffer , the robot selects an action based on the current state and calculates the reward according to the preset reward mechanism , enters the next state , and stores the information in the experience replay buffer . ​ 3. A robot path planning method based on deep reinforcement learning according to claim 1, characterized in that, Initialize the Actor-Critic network described in step 3, introduce a learnable weight coefficient into the loss function of the Critic network , dynamically adjust the contribution of different states and actions to the loss, add an entropy regularization term to the objective function of the Actor network, and initialize the target network. The specific implementation steps are as follows: Step 3.1: The goal of the Critic network is to learn the Q-value function, that is, to calculate the long-term cumulative reward when given the current state and the action , and the expression is: , where is the expected cumulative reward after performing action in state ; is the average expected value; is the reward immediately obtained after performing action in state ; is the discount factor, which balances the weights of immediate rewards and cumulative rewards; is the cumulative reward after performing action in the next state ; is the action generated by the Actor network in the next state; is the maximum Q-value estimate among all actions in the next state; are the Critic network parameters; are the Critic target network parameters; Loss function of the Critic network As follows: , Among them is a learnable weight coefficient that dynamically adjusts the contribution of the loss in the current state to the overall optimization according to the importance of the samples is the target Q value, and the parameters of the Critic network are optimized by minimizing the loss function, thereby reducing the TD error represents calculating the expectation for the state sampled from the experience replay buffer ; The goal of the Actor network is to optimize the policy by maximizing the Q-value given by the Critic network. Entropy regularization is added to the objective function to improve exploration, and the objective function is as follows: , where is the entropy of the action , and is the regularization coefficient; Step 3.2: To determine whether to use differential game to generate control strategies or use the deep deterministic policy gradient algorithm to generate actions in the current state, an adaptive decay greedy strategy is introduced, and the expression is: , , where is the environmental complexity function of the state , is the decay adjustment parameter is the exploration rate of the current state is the minimum exploration rate is the initial exploration rate is the preset policy switching threshold is a random number generated by uniform sampling is the action with the largest Q value is the control policy generated by differential game 4. A robot path planning method based on deep reinforcement learning according to claim 1, characterized in that, If the current state passes through adaptive attenuation as described in step 4 The greedy strategy is used to select differential games for dynamic obstacle avoidance. By establishing a differential game model, a control strategy is generated. The specific implementation steps are as follows: Step 4.1: Establish a differential game model, define Robot 1 as , and define Robot 2 as ; Definition and the control strategies are respectively and ; Step 4.2: Set 's cost function is expressed as: , Among them , , are weight parameters that balance the importance of different terms in the cost function, making the two robots stay as far away as possible, balancing the energy consumption, making the robot approach the target point, is the distance between the two robots; Set the robot 's cost function is expressed as: , Among them , are also weight parameters, keep the two as far away as possible, control energy consumption; Step 4.3: The feedback Nash equilibrium ensures that robot and robot optimize their strategies with each other in the differential game, and any unilateral change in strategy will not make the cost functions of both smaller. Construct the Hamiltonian , : , Among them , represents the kinematic equation of the robot, , is the co-state variable of the robot; Step 4.4: Take the derivative of the constructed Hamiltonian with respect to the control strategy, set the derivative function to zero, and then solve the Hamiltonian equation to find the feedback control strategy; Step 4.5: The robot and control strategy is updated by the gradient descent method: , Until the objective function satisfies the convergence condition, the feedback Nash equilibrium convergence condition is as follows: , Among them , is the final control strategy, is the update parameter.

Citation Information

Patent Citations

  • Deep reinforcement learning energy management method for curiosity-driven hybrid power system

    CN112765723A

  • Intelligent household energy management system prediction and decision integrated scheduling method based on deep reinforcement learning

    CN116227883A

  • Multi-agent encircling control method based on reinforcement learning and auction algorithm

    CN117311356A

  • Deep reinforcement learning path planning method and system based on reward function improvement

    CN118760168A

  • Steel structure roof climbing robot path planning method, system, medium and equipment

    CN119374603A

Cited By

  • Intelligent unmanned aerial vehicle flight path planning method and system

    CN121612310A

  • Man-machine dynamic game control method driven by double reinforcement learning

    CN122035054A

  • Weeding robot operation fine control system based on reinforcement learning

    CN122195018A

  • Weeding robot operation fine control system based on reinforcement learning

    CN122195018B