Model training method for autonomous navigation of robot based on reinforcement learning in complex dynamic environment
By introducing the Stackelberg game structure and the Bayesian-RVO model, a two-layer strategy decision network was constructed, which solved the multi-objective conflict problem in robot navigation in complex dynamic environments, achieved a dynamic trade-off between safety, sociability and efficiency, and improved the flexibility and adaptability of the robot's autonomous navigation.
Patent Information
- Application Number
- CN202511063822.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-31
- Publication Date
- 2025-10-10
- Estimated Expiration
- 2045-07-31
AI Technical Summary
Existing multi-objective reinforcement learning methods find it difficult to effectively coordinate the conflicts among safety, sociability, and efficiency in complex dynamic environments, resulting in a lack of flexibility and adaptability in robot navigation strategies and an inability to make reasonable decisions in multi-objective conflict situations.
By combining game theory with reinforcement learning, a two-layer decision-making architecture based on the Stackelberg game structure is constructed. The Stackelberg game is combined with the advantage function policy gradient algorithm (GT-A2C) and the Bayesian reciprocal velocity obstacle method (Bayesian-RVO) to model pedestrian behavior. Safety, sociability, and efficiency reward functions are designed to construct a two-layer policy decision-making network.
It improves the robot's autonomous navigation safety and navigation task efficiency in complex dynamic environments, achieves an effective trade-off between social interaction with people and task completion, and enhances the robustness and generalization ability of the model.
Smart Images

Figure CN120558244B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of reinforcement learning and robot navigation, and specifically to a robot autonomous navigation method based on reinforcement learning in a complex dynamic environment. Background Art
[0002] In recent years, autonomous navigation robots have been widely used in real life, such as automatic returning robots for supermarket shopping carts, handling robots in warehousing environments, robots used for guiding or delivering food in hotels, and many other complex environments. This type of autonomous navigation robots will be more widely used in the future.
[0003] Autonomous navigation robots still face complex challenges in real-world social navigation tasks. These tasks often involve dynamic trade-offs between multiple objectives. For example, when avoiding pedestrians, a robot may have to deviate from the shortest path. In such cases, improperly designed reward functions can make it difficult for navigation strategies to balance the conflict between safety and efficiency, preventing them from making flexible and reasonable decisions. This ultimately leads to decreased reward accumulation and reduced task completion rates.
[0004] Existing multi-objective reinforcement learning methods generally compress multiple objectives into a single scalar reward function, a weighted sum, for unified optimization. However, this approach has significant limitations in practice: the research process requires numerous experiments to continuously try different weight combinations, but with limited time and computing resources, it is impossible to exhaust all possible combinations. The resulting strategy is generally only acceptable, not optimal. Furthermore, this scalar merging approach fails to truly reflect the dynamically changing priorities and dependencies between objectives in complex environments, resulting in a lack of adaptability in the strategy. Although hierarchical reinforcement learning (HRL) alleviates the multi-objective conflict problem to some extent by structurally separating high-level goal planning from low-level action execution, its learning process often relies on a manually defined subtask structure, resulting in a robot's strong passivity during training and difficulty in actively optimizing high-level strategies through environmental feedback.
[0005] In summary, while current mainstream approaches each have their own strengths, they still lack a mature and scalable policy structure that can effectively model and resolve conflicts between multiple efficiency rewards, especially in dynamic, complex, and socially intensive navigation environments. Therefore, designing a reinforcement learning framework that can dynamically balance, proactively respond, and possess structured decision-making capabilities is a key direction for further development of social robot navigation methods.
[0006] To solve the problem of multi-objective conflict difficult to coordinate, the application proposes a robot autonomous navigation method based on reinforcement learning suitable for complex dynamic environment. The method innovatively combines game theory (GT) and reinforcement learning (RL), introduces Stackelberg game structure in strategy design, and realizes dynamic trade-off in multi-objective conflict situation by constructing a double-layer decision mechanism. When the robot conflicts between safety, sociality and efficiency, etc., the Stackelberg equilibrium can be solved to prioritize safety while considering social etiquette and path efficiency, thereby generating more reasonable and stable behavior strategy.
[0007] In addition, to improve the authenticity and diversity of crowd behavior in simulation experiments, the system uses the Bayesian reciprocal velocity obstacle method (Bayesian Reciprocal Velocity Obstacles, Bayesian-RVO) to model pedestrian behavior. This method describes the speed intention distribution of pedestrians through Gaussian mixture model, and realizes dynamic collision avoidance prediction by combining velocity obstacle model, effectively improving the complexity and interaction authenticity of pedestrian modeling in simulation environment. With this modeling framework, the application constructs an adaptive, generalizable and robust dense crowd navigation decision system suitable for robot autonomous navigation tasks in various complex scenarios. SUMMARY
[0008] The application aims to solve the problem that robots cannot make flexible decisions due to conflicts between multiple objectives in the multi-objective reinforcement learning process in complex environments. To this end, the application introduces the game theory into the reward mechanism design of reinforcement learning, and combines the Stackelberg game model with the advantage function strategy gradient algorithm (Advantage Actor-Critic, A2C) in reinforcement learning (Game Theory-Advantage Actor-Critic, GT-A2C) to construct a double-layer strategy decision architecture. The architecture includes Leader decision network and Follower decision network, which replaces the traditional single Actor network with sequential game decision mechanism, so that the robot can prioritize safety when facing multi-objective conflicts, and optimize sociality and task efficiency on this basis. This method not only improves the safety of robot autonomous navigation in complex dynamic environment, but also realizes effective trade-off between crowd social interaction and navigation task completion.
[0009] The specific implementation steps are as follows:
[0010] Step 1: Establish a simulation environment. The algorithm in this invention needs to be trained and tested in a crowded and dynamic human-machine co-existing environment. Use Webots simulation software to customize the environment, add robots and pedestrians, and fixed obstacles, build a Bayesian-RVO pedestrian simulation framework, and introduce data-driven prediction to enhance the authenticity of the simulation. Use probability distribution to represent the pedestrian's intention. Suppose the pedestrian's speed v is . i Obeying the Gaussian mixture model, the following formula is obtained:
[0011]
[0012] where μ k is the expected speed of the kth intention, ∑ k Depends on the density of pedestrians, v i is the velocity vector of pedestrian i, ω k is the weight coefficient of the kth Gaussian distribution, and ω k ≥0, Ν(μ k ,∑ k ) is a multivariate Gaussian distribution, which represents the probability density function of the kth Gaussian component.
[0013] Get the speed v of pedestrian i i Then, calculate the expected collision probability with surrounding pedestrians or obstacles:
[0014]
[0015] in Choose a speed v for pedestrian i i The total collision probability with all surrounding pedestrians or obstacles j, P(v j ) is the probability density function of the speed of pedestrian or obstacle j, v j is the speed of all pedestrians or obstacles j around, is the speed obstacle area, indicating v i and v j The collection of velocities that will cause a collision.
[0016] Optimize the pedestrian speed according to the collision probability, and finally obtain the optimal speed of pedestrian i:
[0017]
[0018] in is the optimal speed of the pedestrian, v pref is the pedestrian’s preferred speed towards the target, v currentRepresents the current speed of pedestrian i. λ1 is used to adjust the pedestrian's tendency to reach the target. A larger λ1 means the more likely the pedestrian is to reach the target, while a smaller λ1 means the pedestrian tends to move conservatively, prioritizing safety. λ2 is used to adjust the pedestrian's speed. A larger λ2 means the pedestrian tends to maintain the current speed, while a smaller λ2 means the pedestrian will adjust the speed more flexibly.
[0019] Step 2: Design the state and action space, breaking down the environmental state into the states of the robot, pedestrians, and fixed obstacles. The robot state includes its position, velocity, and direction of motion; the pedestrian state includes its position, velocity, and direction of motion; and the fixed obstacle state includes its position. This foundation introduces the concept of social relationships, such as the distance between the robot and pedestrian. The robot can perform actions such as acceleration, deceleration, and turning. To ensure the safety of both the robot and pedestrian, the robot's maximum speed and acceleration must be limited, and the pedestrian's motion state must be fully considered in the decision-making process.
[0020] Assume s t , a t , s i,t , represents the state, action of the robot and the state of pedestrian i at time t. The state vector has two parts, namely the observable part s t y , and the unobservable s t n The observable part consists of the robot's position coordinates p(x,y) and velocity vector v a , where the robot velocity vector is divided into a component v towards the nearest pedestrian ah , and towards the target point v ag Direction, acceleration a, distance d between the robot and the target point ag , the distance d between the robot and the nearest pedestrian ap , the distance d between the machine and the nearest obstacle obs , where observable refers to the state part visible to adjacent pedestrians, including the position coordinates p'(x,y) of the pedestrian i, the speed v i , is the optimal speed of pedestrian i. Target position P, optimal speed v * and direction θ * , where θ * is the direction of the two-dimensional environment, The direction of the three-dimensional environment and other hidden pedestrians constitute the unobservable part of the state. The robot's action space includes: going straight, turning left, turning right, accelerating, decelerating, etc.
[0021] Step 3: Introduce game theory to design the reward function. The reward function includes a safety reward function, a social reward function, and an efficiency reward function, which respectively ensure the safety, sociality, and work efficiency of the robot during autonomous navigation.
[0022] Step 3.1: Design a reward mechanism that takes safety into consideration. The robot receives a negative reward when it is too close to pedestrians or other fixed obstacles, or when the velocity component toward the nearest pedestrian is too large. Otherwise, a positive reward is given.
[0023]
[0024] where R safe is the safety reward function, d ap is the distance between the robot and the nearest pedestrian, v is the speed of the robot, and v max is the maximum velocity of the robot, v ah is the velocity component of the robot toward the nearest pedestrian, d obs is the distance between the robot and the nearest obstacle, τ obs The minimum distance to activate the obstacle penalty. original Safety reward function that penalizes inactivating obstacles:
[0025]
[0026] Step 3.2: Design a reward mechanism that takes sociality into consideration. The robot needs to consider whether its speed is facing pedestrians, causing pedestrians to feel uneasy or other negative social emotions. A negative reward is given when the speed component facing pedestrians is too large, and a positive reward is given when the robot moves away from pedestrians.
[0027]
[0028] R social is the social reward function, v ah is the component of the robot's velocity towards the nearest pedestrian.
[0029] Step 3.3: Consider the reward mechanism for task completion efficiency. To prevent the robot from stagnating or taking too long to reach the target point, a positive reward is given when the robot gets closer to the target point, and a negative reward is given when it moves closer.
[0030]
[0031] R goal is the efficiency reward function, v ag is the velocity component of the robot towards the target point.
[0032] In the reward function designed by the present invention, the behavioral decision-making needs of robots in complex environments are fully considered. The design emphasizes that the robot must give priority to ensuring the safety of pedestrians during navigation, and uses the safety reward function to model the penalty relationship between the distance relationship with the nearest pedestrian and the distance relationship with fixed obstacles. The social reward measures the robot's respect for the behavior of surrounding individuals by calculating the projection of the robot's speed in the direction of the nearest pedestrian, thereby reflecting its social coordination ability. The efficiency reward measures the efficiency of the robot's task completion based on the component of its speed in the direction pointing to the target point. The three-part reward design checks and balances each other, which is in line with the basic idea of strategic confrontation and equilibrium response in game theory, and provides a structured theoretical basis for multi-objective decision-making.
[0033] Step 4: Construct a game theory reward function framework. The three reward functions involved in the game are safety reward function, social reward function, and efficiency reward function. Set the safety reward function as the leader, and set the efficiency reward function and social reward function as followers. The optimization goal of the safety reward function is:
[0034] a safe* =argmax[R safe (d ap ,v,d obs )],
[0035] where a safe* is the set of actions that can maximize the cumulative safety reward, d ap is the distance between the robot and the nearest pedestrian, v is the speed of the robot, d obs is the distance between the robot and the nearest obstacle. The follower includes two rewards, namely social reward and efficiency reward. When the leader makes a decision, the social reward is social* ∈a safe* To maximize the efficiency reward under the condition of , the optimization goal of the social reward function is:
[0036] a social* =argmaxR social ,
[0037]
[0038] in represents the set of actions that conform to the social reward function under the Stackelberg game framework, Action strategy for social rewards, is the global action set, a safe The set of actions representing the safety reward function, a social The action set representing the social reward function, a social* Indicates the largest set of actions that can accumulate social rewards.
[0039] Step 4.1: Similarly, the safety reward function and efficiency reward function are also built in the same game theory framework, then, in a goal* ∈a safe* In the case of , the optimization objective of the efficiency reward function is:
[0040] a goal* =argmax R goal ,
[0041]
[0042] in represents the set of actions that conform to the efficiency reward function under the Stackelberg game framework, A strategy that rewards efficiency, is a global action collection, a goal The set of actions representing the efficiency reward function, a goal* Indicates the largest set of actions that can accumulate efficiency rewards.
[0043] Step 5: Design the GT-A2C network, using BNN-Critic as the critic network in GT-A2C, and perform model training. A2C is an improved version of the actor-critic algorithm framework, adding parallel computing to the traditional actor-critic algorithm. A global actor and global critic are maintained on the master node, while the reinforcement learning robot on each slave node communicates with the global actor and global critic through a coordinator. In this design, the coordinator is responsible for collecting experience from each slave node interacting with the environment and then further updating it based on the collected trajectories. After the update, the global actor is synchronized to each slave node to continue interacting with the environment. In this design, the slave nodes are solely responsible for interacting with the environment, while all computation and updates occur in the master node, facilitating data processing and computation. Actor networks are typically used to guide robots to perform actions. Their input typically includes observations of the robot, such as the state of the surrounding environment and the position and velocity of pedestrians. The output is the robot's action. The goal is to maximize reward accumulation, that is, to make the robot perform actions that are more consistent with expectations.
[0044] Step 5.1: Build the GT-A2C actor network. Based on the traditional A2C framework, we introduce the Stackelberg game structure. The input is the state environment s. The leader sub-policy network first outputs the robot's velocity v. The follower sub-policy network then generates a direction combination weight (a·β·γ) based on this. This two-layer structure replaces the original single strategy, retaining the differentiability of the parameters in the actor network. The output action a is:
[0045]
[0046] in They are respectively the vector towards the target, the vector towards the nearest pedestrian, and the obstacle avoidance vector, satisfying α≥0, β≥, γ≥0, and α+β+γ=1. The strategy distribution of speed control and direction fusion is learned through the Leader and Follower networks respectively.
[0047] In order to introduce environmental uncertainty, adaptive randomness can still be achieved by adjusting the Leader output speed. Introduce Gaussian perturbation to model:
[0048] π θ (a|s)=π θF (α,β,γ|s,v)·π θL (v|s),
[0049] where σ a =f(σ z ,s),σ z To quantify the uncertainty of the state value of the Critic network, the Actor can use σ z Adaptively adjust the randomness of the action to adapt to the new environment, π θF is the Follower strategy, π θL (v|s) is the Leader strategy, μ a is the speed mean, which can be output through the Leader strategy network.
[0050] Step 5.2: Construct a BNN-Critic network. This critic network is implemented by a Bayesian neural network. Its input layer is the environment state s. Unlike the value function of the traditional critic network, the reward of its output layer is defined as the probability distribution of the value function to reflect the uncertainty of the value estimation. Its formula is as follows:
[0051]
[0052] Among them H π (s) is the value function V π (s), and H π (s) and V π (s). r(s) is the reward of the current state, δ is the discount factor, and H π (s') is the reward of the next state.
[0053] Step 5.3: Update the BNN-Critic network. Since the BNN-Critic network outputs a probability distribution rather than a definite value, variational inference is used to maximize the evidence lower bound (ELBO) while minimizing the TD error of reinforcement learning. ELBO is essentially an indirect approach to the real evidence by optimizing a computable lower bound. First, the posterior distribution of the parameter Z (network weight) is calculated. The formula is as follows:
[0054]
[0055] Where P(D|Z) is the likelihood function used to simulate the ability of fitting data, P(Z) is the prior distribution of network weights, P(D) is the evidence (marginal likelihood), and Z is the network weight. Since it is not feasible to directly calculate the posterior P(Z|D), an adjustable distribution Q(Z|λ) is introduced, and by finding the optimal parameter η, the difference (KL divergence) between Q(Z|λ) and P(Z|D) is minimized. The formula is as follows:
[0056] η * =argminKLQ(Z|λ)||P(Z|D),
[0057] The formula for KL divergence is:
[0058] KL(Q(Z|λ)||P(Z|D))=E[logQ(Z|λ)]-E[logP(Z|D)]+logP(D),
[0059] Q(Z|λ) is the adjustable distribution, P(Z|D) is the posterior distribution, λ is the variational parameter, P(D) is the evidence, i.e., marginal likelihood, and E is the expectation. Given a data set D(X,Y), where X is the input data and Y is the target value corresponding to the input, as the observation value, the data set is used to measure whether the model is correct and guide the optimization of the network structure. The above KL divergence formula can be transformed into:
[0060] logP(D)=KL(Q(Z|λ)||P(Z|D))+E[logP(Z,D)]-E[logQ(Z|λ)],
[0061] P(Z,D) is the joint distribution, Q(Z|λ) is the adjustable distribution, and E is the expectation. Since the data set D(X,Y) is fixed, minimizing the KL divergence is equivalent to maximizing the evidence lower bound ELBO:
[0062] ELBO(η)=E[logP(Z,D)]-E[logQ(Z|λ)],
[0063] E[logP(Z,D)] is the logarithmic expectation of the joint distribution P(Z,D), and E[logQ(Z|λ)] is the logarithmic expectation of the adjustable distribution Q(Z|λ);
[0064] The loss function of BNN-Critic is defined as:
[0065]
[0066] Where (y m -V(s m ; Z m )) 2 is the square error of the mth sample, used to fit the target return, KL(Q(Z|λ)||P(Z)) is the divergence, used to prevent BNN overfitting, N is the sample size, β is the regularization strength coefficient, which controls the trade-off between data fitting and distribution conservatism, y m is the actual return of the mth sample, s m is the state input of the mth sample, Z m is the network weight when training the mth sample;
[0067] Step 5.4: Actor network update. The Actor network update goal is the maximum advantage strategy, and its formula is as follows:
[0068] L A-loss =-E s,a [logπ θ (a|s)A(s,a)],
[0069] where π θ (a|s) is the strategy output by the Actor network, A(s,a) is the advantage function, and E is the expectation.
[0070] Step 5.5: Model training. In each training step, the pedestrian will be trained according to the optimal speed in the Bayesian-RVO framework. To choose the speed, on the basis of reaching the target point, avoid other pedestrians as much as possible, so that the pedestrian's actions are more in line with the behavioral laws in the real scene.
[0071] During training, the Actor network, which incorporates game theory, enables machines to make decisions more flexibly and improves the safety of robots in complex environments. At the same time, the BNN-Critic network can solve the overfitting problem of traditional Critic networks, allowing robots to obtain feedback information from the environment more accurately.
[0072] The specific steps are as follows:
[0073] Step 1: Initialize the GT-A2C network;
[0074] Step 2: Integrate the Bayesian Reciprocal Velocity Obstacles framework into the simulation environment to simulate the motion trajectory of pedestrians.
[0075] Step 3: Use the master node to control the training process and use the Actor decision network that integrates the Stackelberg game and the BNN-Criric network model to generate the value function of the robot's actions and evaluation status;
[0076] Step 4: Use multiple synchronized slave nodes to continue interacting with the environment. Each slave node uses the Actor model to generate the robot's actions, executes the actions, observes the environment, and finally updates the parameters of the BNN-Critic model.
[0077] Step 5: Use gradient descent to update the parameters of the Actor decision network and the BNN-Criric network that integrate the Stackelberg game, and update the parameters to the master node in a synchronous manner;
[0078] Step 6: End the training if the training requirements are met. If the training requirements are not met, repeat Step 3.
[0079] This invention provides a robot autonomous navigation method based on the GT-A2C algorithm. By introducing the Stackelberg game into the decision network, the robot prioritizes safety when facing multi-objective conflicts, and on this basis optimizes sociability and task efficiency. This method not only improves the safety of the robot's autonomous navigation in complex dynamic environments, but also achieves an effective trade-off between social interaction with people and the completion of navigation tasks. This method has high application value and can be used for the autonomous navigation of robots in various complex environments, such as airports, production workshops, hospitals, and other scenarios. The technical solution of the present invention has the following beneficial technical effects:
[0080] 1. In the field of autonomous navigation of social robots, a realistic, dynamic, and challenging pedestrian interaction environment is crucial for model training. The Bayersian-Rvo pedestrian simulation framework constructed in this paper utilizes Bayesian updating to dynamically predict pedestrian intentions and adapt to individual behavioral differences. It supports multimodal interactive decision-making and optimizes avoidance strategies in dense crowds. It also provides interpretable confidence assessments to enhance safety. Using this framework for pedestrian simulation in simulation training can effectively improve the robustness and generalization of the model.
[0081] 2. This paper proposes a two-layer Actor decision network that integrates the Stackelberg game structure. By dividing a single strategy into a two-level decision system, the leader prioritizes safety goals, while the follower optimizes sociability and efficiency. This explicitly decomposes the multi-objective optimization process in reinforcement learning, effectively avoiding goal conflicts and strategy oscillations. This not only improves training stability and convergence speed, but also enhances the strategy's expressiveness and navigation rationality in complex and dynamic crowd environments.
[0082] 3. The BNN-Critic network used in this invention effectively addresses the overfitting problem often seen in traditional critic networks using artificial neural networks, enabling robots to better adapt to new or complex environments and improving the model's generalization capabilities. This approach has high application value and can be used for social movement planning in various complex environments, such as restaurants, airports, hotels, and hospitals.
[0083] The present invention will be further described in detail below with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0084] Figure 1 It is the overall flow chart of the present invention;
[0085] Figure 2 This is the schematic diagram of the GT-A2C algorithm. DETAILED DESCRIPTION
[0086] In order to make the purpose, technical solutions and advantages of the present invention more clear, the present invention will be described in detail below with reference to the accompanying drawings and specific embodiments.
[0087] Attachment Figure 1 This embodiment provides a robot autonomous navigation method based on reinforcement learning in a complex dynamic environment, including using Webots simulation software to establish a simulation environment and build a Bayesian-RVO pedestrian simulation framework, designing the environment state and action space, designing the reward function, building a reward function game theory framework, designing the GT-A2C network, and model training.
[0088] Attachment Figure 2This is an improved GT-A2C algorithm provided in this implementation case. Its principle is to use a decision network that integrates Stackelberg game and a BNN-Critic network to replace the actor network and critic network in the traditional A2C network based on the parallel computing of the A2C algorithm. Each slave node interacts with the environment and feeds the feedback given by the environment back to the coordinator of the main node. The main node will use gradient descent to update the model parameters based on the data transmitted back by each slave node, and then pass the updated model parameters back to the slave node to continue interacting with the environment.
[0089] To facilitate understanding, a specific example of a complex environment and a robot is given. Taking the airport environment as an example, the complex environment is composed of robots, staff, pedestrians and fixed obstacles in the airport, and the robot is a luggage transport robot.
[0090] The specific implementation steps are as follows:
[0091] Implementation step 1: Establish a simulation environment. The algorithm in this invention needs to be trained and tested in a crowded, dynamic human-machine co-existence environment. Use Webots simulation software to customize the environment, add robots, pedestrians, and fixed obstacles, build a Bayesian-RVO pedestrian simulation framework, and introduce data-driven prediction to enhance the authenticity of the simulation. Use probability distribution to represent the pedestrian's intention. Suppose the pedestrian's speed v is 0. i Obeying the Gaussian mixture model, the following formula is obtained:
[0092]
[0093] where μ k is the expected speed of the kth intention, Σ k Depends on the density of pedestrians, v i is the velocity vector of pedestrian i, ω k is the weight coefficient of the kth Gaussian distribution, and ω k ≥0, Ν(μ k ,Σ k ) is a multivariate Gaussian distribution, which represents the probability density function of the kth Gaussian component.
[0094] Get the speed v of pedestrian i i Then, calculate the expected collision probability with surrounding pedestrians or obstacles:
[0095]
[0096] in Choose a speed v for pedestrian i i The total collision probability with all surrounding pedestrians or obstacles j, P(v j ) is the probability density function of the speed of pedestrian or obstacle j, vj is the speed of all pedestrians or obstacles j around, is the speed obstacle area, indicating v i and v j The collection of velocities that will cause a collision.
[0097] Optimize the pedestrian speed according to the collision probability, and finally obtain the optimal speed of pedestrian i:
[0098]
[0099] in is the optimal speed of the pedestrian, v pref is the pedestrian’s preferred speed towards the target, v current Represents the current speed of pedestrian i. λ1 is used to adjust the pedestrian's tendency to reach the target. A larger λ1 means the more likely the pedestrian is to reach the target, while a smaller λ1 means the pedestrian tends to move conservatively, prioritizing safety. λ2 is used to adjust the pedestrian's speed. A larger λ2 means the pedestrian tends to maintain the current speed, while a smaller λ2 means the pedestrian will adjust the speed more flexibly.
[0100] Implementation Step 2: Design the state and action space, breaking down the environmental state into the states of the robot, pedestrians, and fixed obstacles. The robot state includes its position, velocity, and direction of motion; the pedestrian state includes its position, velocity, and direction of motion; and the fixed obstacle state includes its position. This foundation introduces the concept of social relationships, such as the distance between the robot and pedestrian. The robot can perform actions such as acceleration, deceleration, and turning. To ensure the safety of both the robot and pedestrian, the robot's maximum speed and acceleration must be limited, and the pedestrian's motion state must be fully considered in the decision-making process.
[0101] Assume s t , a t , s i,t , represents the state, action of the robot and the state of pedestrian i at time t. The state vector has two parts, namely the observable part s t y , and the unobservable s t n The observable part consists of the robot's position coordinates p(x,y) and velocity vector v a , where the robot velocity vector is divided into a component v towards the nearest pedestrian ah , and towards the target point v ag Direction, acceleration a, distance d between the robot and the target point ag , the distance d between the robot and the nearest pedestrian ap , the distance d between the machine and the nearest obstacle obs, where observable refers to the state part visible to adjacent pedestrians, including the position coordinates p'(x,y) of the pedestrian i, the speed v i , is the optimal speed of pedestrian i. Target position P, optimal speed v * and direction θ * , where θ * is the direction of the two-dimensional environment, The direction of the three-dimensional environment and other hidden pedestrians constitute the unobservable part of the state. The robot's action space includes: going straight, turning left, turning right, accelerating, decelerating, etc.
[0102] Implementation step 3: Introduce game theory to design the reward function. The reward function includes a safety reward function, a social reward function, and an efficiency reward function, which respectively ensure the safety, sociality, and work efficiency of the robot during autonomous navigation.
[0103] Step 3.1: Design a reward mechanism that takes safety into consideration. The robot receives a negative reward when it is too close to pedestrians or other fixed obstacles, or when the velocity component toward the nearest pedestrian is too large. Otherwise, a positive reward is given.
[0104]
[0105] where R safe is the safety reward function, d ap is the distance between the robot and the nearest pedestrian, v is the speed of the robot, and v max is the maximum velocity of the robot, v ah is the velocity component of the robot toward the nearest pedestrian, d obs is the distance between the robot and the nearest obstacle, τ obs The minimum distance to activate the obstacle penalty. original Safety reward function that penalizes inactivating obstacles:
[0106]
[0107] Step 3.2: Design a reward mechanism that takes sociality into consideration. The robot needs to consider whether its speed is facing pedestrians, causing pedestrians to feel uneasy or other negative social emotions. A negative reward is given when the speed component facing pedestrians is too large, and a positive reward is given when the robot moves away from pedestrians.
[0108]
[0109] R social is the social reward function, v ah is the component of the robot's velocity towards the nearest pedestrian.
[0110] Step 3.3: Consider the reward mechanism for task completion efficiency. To prevent the robot from stagnating or taking too long to reach the target point, a positive reward is given when the robot gets closer to the target point, and a negative reward is given when it moves closer.
[0111]
[0112] R goal is the efficiency reward function, v ag is the velocity component of the robot towards the target point.
[0113] In the reward function designed by the present invention, the behavioral decision-making needs of robots in complex environments are fully considered. The design emphasizes that the robot must give priority to ensuring the safety of pedestrians during navigation, and uses the safety reward function to model the penalty relationship between the distance relationship with the nearest pedestrian and the distance relationship with fixed obstacles. The social reward measures the robot's respect for the behavior of surrounding individuals by calculating the projection of the robot's speed in the direction of the nearest pedestrian, thereby reflecting its social coordination ability. The efficiency reward measures the efficiency of the robot's task completion based on the component of its speed in the direction pointing to the target point. The three-part reward design checks and balances each other, which is in line with the basic idea of strategic confrontation and equilibrium response in game theory, and provides a structured theoretical basis for multi-objective decision-making.
[0114] Implementation step 4: Construct a game theory reward function framework. The three reward functions involved in the game are safety reward function, social reward function, and efficiency reward function. Set the safety reward function as the leader, and set the efficiency reward function and social reward function as followers. The optimization goal of the safety reward function is:
[0115] a safe* =argmax[R safe (d ap ,v,d obs )],
[0116] where a safe* is the set of actions that can maximize the cumulative safety reward, d ap is the distance between the robot and the nearest pedestrian, v is the speed of the robot, d obs is the distance between the robot and the nearest obstacle. The follower includes two rewards, namely social reward and efficiency reward. When the leader makes a decision, the social reward is social* ∈a safe* To maximize the efficiency reward under the condition of , the optimization goal of the social reward function is:
[0117] a social* =argmaxR social ,
[0118]
[0119] in represents the set of actions that conform to the social reward function under the Stackelberg game framework, Action strategy for social rewards, is the global action set, a safe The set of actions representing the safety reward function, a social The action set representing the social reward function, a social* Indicates the largest set of actions that can accumulate social rewards.
[0120] Step 4.1: Similarly, the safety reward function and efficiency reward function are also built in the same game theory framework, then, in a goal* ∈a safe* In the case of , the optimization objective of the efficiency reward function is:
[0121] a goal* =argmaxR goal ,
[0122]
[0123] in represents the set of actions that conform to the efficiency reward function under the Stackelberg game framework, A strategy that rewards efficiency, is a global action collection, a goal The set of actions representing the efficiency reward function, a goal* Indicates the largest set of actions that can accumulate efficiency rewards.
[0124] Step 5: Design GT-A2C network, use BNN-Critic as Critic network in GT-A2C, and train the model. A2C is an improved form of Actor-Critic algorithm framework, which adds parallel computing design based on traditional Actor-Critic algorithm. Global actor and global critic are maintained in the master node. Each secondary node's reinforcement learning robot communicates with the global actor and global critic through the coordinator. In this design, the coordinator is responsible for collecting the experience of each secondary node interacting with the environment, and then updating the collected trajectory. After updating, the global actor is synchronized to each secondary node to continue interacting with the environment. In this design, the secondary nodes are only responsible for interacting with the environment, and all calculations and updates are performed in the master node, which helps to process and calculate data. The Actor network is usually used to guide the robot to make actions. Its input usually includes the robot's observation information, such as the environment state around the robot, the position and speed of pedestrians, etc. The output is the robot's action, and the goal is to maximize the cumulative reward, that is, to make the robot make actions more in line with expectations.
[0125] Step 5.1: Construct the Actor network of GT-A2C. Introduce Stackelberg game structure based on traditional A2C framework. The input is the state environment s. First, the Leader sub-policy network outputs the robot speed v. Then the Follower sub-policy network generates the direction combination weight (a·β·γ) based on this. This double-layer structure replaces the original single policy, preserves the differentiability of the parameters in the Actor network, and outputs the action a:
[0126]
[0127] where are the vectors towards the target, the nearest pedestrian, and the obstacle avoidance, respectively. They satisfy α≥0, β≥, γ≥0, and α+β+γ=1. The Leader and Follower networks learn the speed control and direction fusion policy distribution, respectively.
[0128] To introduce environmental uncertainty, adaptive randomness can still be introduced by introducing Gaussian disturbance to model:
[0129] π θ (a|s)=π θF (α,β,γ|s,v)·π θL (v|s),
[0130] where σ a =f(σ z ,s), σ zTo quantify the uncertainty of the state value for the Critic network, the Actor can pass σ z Adaptively adjust the randomness of the action to adapt to the new environment, π θF For the Follower policy, π θL (v|s) is the Leader policy, μ a The average speed, which can be output by the Leader policy network.
[0131] Step 5.2: Construct the BNN-Critic network, which is implemented by a Bayesian neural network, with the input layer being the environment state s. Unlike the value function of the traditional Critic network, the output layer of the return is defined as the probability distribution of the value function, which reflects the uncertainty of the value estimate, and its formula is as follows:
[0132]
[0133] Where H π (s) is the probability distribution of the value function V π (s), and H π (s) has the same expectation as V π (s). r(s) is the reward of the current state, δ is the discount factor, and H π (s') is the return of the next state.
[0134] Step 5.3: Update the BNN-Critic network. Since the BNN-Critic network outputs a probability distribution rather than a deterministic value, the evidence lower bound (ELBO) is maximized using variational inference, while the TD error of reinforcement learning is minimized. ELBO is essentially an indirect approximation of the true evidence by optimizing a computable lower bound. First, calculate the posterior distribution of the parameter Z (network weight), whose formula is as follows:
[0135]
[0136] Where P(D|Z) is the likelihood function used to simulate the ability to fit the data, P(Z) is the prior distribution of the network weight, P(D) is the evidence (marginal likelihood), and Z is the network weight. Since it is not feasible to directly calculate the posterior P(Z|D), a adjustable distribution Q(Z|λ) is introduced, and by finding the optimal parameter η, the difference (KL divergence) between Q(Z|λ) and P(Z|D) is minimized, as follows:
[0137] η * = argmin KL Q(Z|λ)‖P(Z|D),
[0138] Where the formula of the KL divergence is:
[0139] KL(Q(Z|λ)||P(Z|D))=E[logQ(Z|λ)]-E[logP(Z|D)]+logP(D),
[0140] Q(Z|λ) is the adjustable distribution, P(Z|D) is the posterior distribution, λ is the variational parameter, P(D) is the evidence, i.e., marginal likelihood, and E is the expectation. Given a data set D(X,Y), where X is the input data and Y is the target value corresponding to the input, as the observation value, the data set is used to measure whether the model is correct and guide the optimization of the network structure. The above KL divergence formula can be transformed into:
[0141] logP(D)=KL(Q(Z|λ)||P(Z||D))+E[logP(Z,D)]-E[logQ(Z|λ)],
[0142] P(Z,D) is the joint distribution, Q(Z|λ) is the adjustable distribution, and E is the expectation. Since the data set D(X,Y) is fixed, minimizing the KL divergence is equivalent to maximizing the evidence lower bound ELBO:
[0143] ELBO(η)=E[logP(Z,D)]-E[logQ(Z|λ)],
[0144] E[logP(Z,D)] is the logarithmic expectation of the joint distribution P(Z,D), and E[logQ(Z|λ)] is the logarithmic expectation of the adjustable distribution Q(Z|λ);
[0145] The loss function of BNN-Critic is defined as:
[0146]
[0147] Where (y m -V(s m ; Z m )) 2 is the square error of the mth sample, used to fit the target return, KL(Q(Z|λ)||P(Z)) is the divergence, used to prevent BNN overfitting, N is the sample size, β is the regularization strength coefficient, which controls the trade-off between data fitting and distribution conservatism, y m is the actual return of the mth sample, s m is the state input of the mth sample, Z m is the network weight when training the mth sample;
[0148] Step 5.4: Actor network update. The Actor network update goal is the maximum advantage strategy, and its formula is as follows:
[0149] L A-loss =-E s,a [logπθ (a|s)A(s,a)],
[0150] where π θ (a|s) is the strategy output by the Actor network, A(s,a) is the advantage function, and E is the expectation.
[0151] Step 5.5: Model training. In each training step, the pedestrian will be trained according to the optimal speed in the Bayesian-RVO framework. To choose the speed, on the basis of reaching the target point, avoid other pedestrians as much as possible, so that the pedestrian's actions are more in line with the behavioral laws in the real scene.
[0152] During training, the Actor network, which incorporates game theory, enables machines to make decisions more flexibly and improves the safety of robots in complex environments. At the same time, the BNN-Critic network can solve the overfitting problem of traditional Critic networks, allowing robots to obtain feedback information from the environment more accurately.
[0153] The specific steps are as follows:
[0154] Step 1: Initialize the GT-A2C network;
[0155] Step 2: Integrate the Bayesian Reciprocal Velocity Obstacles framework into the simulation environment to simulate the motion trajectory of pedestrians.
[0156] Step 3: Use the master node to control the training process and use the Actor decision network that integrates the Stackelberg game and the BNN-Criric network model to generate the value function of the robot's actions and evaluation status;
[0157] Step 4: Use multiple synchronized slave nodes to continue interacting with the environment. Each slave node uses the Actor model to generate the robot's actions, executes the actions, observes the environment, and finally updates the parameters of the BNN-Critic model.
[0158] Step 5: Use gradient descent to update the parameters of the Actor decision network and the BNN-Criric network that integrate the Stackelberg game, and update the parameters to the master node in a synchronous manner;
[0159] Step 6: End the training if the training requirements are met. If the training requirements are not met, repeat Step 3.
Claims
1. A model training method for robot autonomous navigation based on reinforcement learning in complex dynamic environments, characterized by: The following steps are involved: Step 1: Establish a simulation environment. Use Webots simulation software to customize the environment, add robots, pedestrians, and fixed obstacles, and build a Bayesian-RVO pedestrian simulation framework. Step 2: Design the state and action space. The environmental state is subdivided into the states of the robot, pedestrians, and fixed obstacles. The robot state includes the robot's position, speed, and direction of motion; the pedestrian state includes its position, speed, and direction of motion; and the fixed obstacle state includes its position information. On this basis, the concept of social relationships and the distance relationship between the robot and pedestrian are introduced. The robot's executable actions include acceleration, deceleration, and turning. The robot's maximum speed and acceleration are restricted, and the pedestrian's motion state is fully considered in the decision-making process. Step 3: Introduce game theory to design reward functions, which include safety reward function, social reward function, and efficiency reward function; Step 4: Construct a game theory reward function framework. The three reward functions involved in the game are safety reward function, social reward function, and efficiency reward function. Set the safety reward function as the leader, and set the efficiency reward function and social reward function as followers. Step 5: Design the GT-A2C network. Use BNN-Critic as the Critic network in GT-A2C and design an Actor network that integrates the Stackelberg game structure. The Leader sub-policy network outputs the robot's velocity v, and the Follower sub-policy network generates a directional combination weight (a·β·γ) based on this output to form the final action. Then, perform model training and optimize the model parameters: Step 1: Initialize the GT-A2C network; Step 2: Integrate the Bayesian Reciprocal Velocity Obstacles framework into the simulation environment to simulate the motion trajectory of pedestrians. Step 3: Use the master node to control the training process and use the Actor decision network that integrates the Stackelberg game and the BNN-Criric network model to generate the value function of the robot's actions and evaluation status; Step 4: Use multiple synchronized slave nodes to continue interacting with the environment. Each slave node uses the Actor model to generate the robot's actions, executes the actions, observes the environment, and finally updates the parameters of the BNN-Critic model. Step 5: Use gradient descent to update the parameters of the Actor decision network and the BNN-Criric network that integrate the Stackelberg game, and update the parameters to the master node in a synchronous manner; Step 6: End the training if the training requirements are met. If the training requirements are not met, repeat Step 3.
2. The model training method for robot autonomous navigation based on reinforcement learning in a complex dynamic environment according to claim 1 is characterized in that: As described in step 1, we set up a simulation environment, customized the environment using Webots simulation software, added robots, pedestrians, and fixed obstacles, and constructed a Bayesian-RVO pedestrian simulation framework. The implementation is as follows: Use probability distribution to represent the pedestrian's intention, set the pedestrian's speed v i Obeying the Gaussian mixture model, there is the following formula: where μ k is the expected speed of the kth intention, ∑ k Depends on the density of pedestrians, v i is the velocity vector of pedestrian i, ω k is the weight coefficient of the kth Gaussian distribution, and ω k ≥0, Ν(μ k ,∑ k ) is a multivariate Gaussian distribution, which represents the probability density function of the kth Gaussian component; Get the speed v of pedestrian i i Then, calculate the expected collision probability with surrounding pedestrians or obstacles: Among them, P collision(vi) Choose a speed v for pedestrian i i The total collision probability with all surrounding pedestrians or obstacles j, P(v j ) is the probability density function of the speed of pedestrian or obstacle j, v j is the speed of all pedestrians or obstacles j around, is the speed obstacle area, indicating v i and v j A collection of velocities that will cause a collision; Optimize the pedestrian speed according to the collision probability, and finally obtain the optimal speed of pedestrian i: in is the optimal speed of the pedestrian, v pref is the pedestrian’s preferred speed towards the target, v current represents the current speed of pedestrian i, λ1 is the pedestrian's approach parameter to the target, and λ2 is the pedestrian's speed parameter.
3. The model training method for robot autonomous navigation based on reinforcement learning in a complex dynamic environment according to claim 1 is characterized in that: In step 3, game theory is introduced to design the reward function. The reward function includes a safety reward function, a social reward function, and an efficiency reward function. The specific implementation steps are as follows: The security reward function based on the game theory structure is set as: where R safe is the safety reward function, d ap is the distance between the robot and the nearest pedestrian, v is the speed of the robot, and v max is the maximum velocity of the robot, v ah is the velocity component of the robot toward the nearest pedestrian, d obs is the distance between the robot and the nearest obstacle, τ obs The minimum distance to activate the obstacle penalty, where R original Safety reward function that penalizes inactivating obstacles: A social reward mechanism based on game theory is designed to encourage robots to move away from pedestrians. Specifically: R social is the social reward function, v ah is the component of the robot's velocity toward the nearest pedestrian; The efficiency reward mechanism designed based on game theory encourages the robot to increase the velocity component towards the target point. Specifically: R goal is the efficiency reward function, v ag is the velocity component of the robot towards the target point.
4. The model training method for robot autonomous navigation based on reinforcement learning in a complex dynamic environment according to claim 1 is characterized in that: In step 4, we construct a game theory reward function framework. The three reward functions involved in the game are safety reward function, social reward function, and efficiency reward function. Set the safety reward function as the leader, and set the efficiency reward function and social reward function as followers. Specific implementation steps are as follows: First, the goal of safety reward optimization is: a safe* =argmax[R safe (d ap ,v,d obs )], where a safe* is the set of actions that maximizes the safety reward, d ap is the distance between the robot and the nearest pedestrian, v is the speed of the robot, d obs is the distance between the robot and the nearest obstacle; The optimization goal of social rewards is: a social* =argmaxR social , in represents the set of actions that conform to the social reward function under the Stackelberg game framework, Strategies for social rewards, is the global action set, a safe The set of actions representing the safety reward function, a social The action set representing the social reward function, a social* It is the set of actions that maximizes the accumulation of social rewards; The optimization goal of efficiency reward is: a goal* =argmaxR goal , in represents the set of actions that conform to the efficiency reward function under the Stackelberg game framework, A strategy that rewards efficiency, is a global action collection, a goal The set of actions representing the efficiency reward function, a goal* It is the collection of actions that maximizes the efficiency bonus.
5. The model training method for robot autonomous navigation based on reinforcement learning in a complex dynamic environment according to claim 1 is characterized in that: Design the GT-A2C network as described in step 5, use BNN-Critic as the critic network in GT-A2C, and train the model. Specific implementation steps are as follows: The Actor network that integrates the Stackelberg game structure first uses the Leader sub-strategy network to output the robot's speed v. Then, the Follower sub-strategy network generates a direction combination weight (a·β·γ) based on this output to form the final action. This two-layer structure replaces the original single strategy. The output action a is: in They are respectively the target vector, the pedestrian vector, and the obstacle avoidance vector, satisfying α≥0, β≥, γ≥0, and α+β+γ=1. The strategy distribution of speed control and direction fusion is learned through the Leader and Follower networks respectively; The output strategy of the Actor network is: p θ (a|s)=π θF (a,b,c|s,v)·π θL (v|s), π θF Follower decision network output strategy, π θL (v|s) is the Leader decision network output strategy; Constructing BNN-Critic, unlike the value function output by traditional critic networks, BNN-Critic outputs a probability distribution, and its estimated return is defined as: Among them H π (s) is the value function V π (s), and H π (s) and V π (s), r(s) is the reward of the current state, δ is the discount factor, and H π (s') is the reward of the next state; In order to update the BNN-Critic network, variational inference is introduced to approximate its posterior distribution. The posterior distribution of the network weight Z is approximately expressed by the Bayesian formula: Where P(D|Z) is the likelihood function, P(Z) is the prior distribution of network weights, P(D) is the marginal likelihood, and Z is the network weight; Then introduce the adjustable distribution Q(Z|λ) and calculate the posterior distribution by minimizing the KL divergence between Q(Z|λ) and P(Z|D). The formula is: η * =arg min KLQ(Z|λ)||P(Z|D), The KL divergence formula is: KL(Q(Z|λ)||P(Z|D))=E[log Q(Z|λ)]-E[log P(Z|D)]+log P(D), Q(Z|λ) is the adjustable distribution, P(Z|D) is the posterior distribution, λ is the variational parameter, P(D) is the marginal likelihood, and E is the expectation. Given a data set D(X,Y), where X is the input data and Y is the target value corresponding to the input, the KL divergence formula is transformed into: log P(D)=KL(Q(Z|λ)||P(Z|D))+E[log P(Z,D)]-E[log Q(Z|λ)], P(Z,D) is the joint distribution, Q(Z|λ) is the adjustable distribution, E is the expectation, the data set D(X,Y) is fixed, and minimizing the KL divergence is equivalent to maximizing the evidence lower bound ELBO, and its formula is: ELBO(η)=E[log P(Z,D)]-E[log Q(Z|λ)], E[log P(Z,D)] is the logarithmic expectation of the joint distribution P(Z,D), and E[log Q(Z|λ)] is the logarithmic expectation of the adjustable distribution Q(Z|λ); The loss function of the BNN-Critic network is defined as: Where (y m -V(s m ; Z m )) 2 is the square error of the mth sample, KL(Q(Z|λ)||P(Z)) is the divergence, N is the sample size, β is the regularization strength coefficient, y m is the actual return of the mth sample, s m is the state input of the mth sample, Z m is the network weight when training the mth sample; The loss function of the Actor network is defined as: THE A-loss =-E s,a [logπ θ (a|s)A(s,a)], where π θ (a|s) is the strategy output by the Actor network, A(s,a) is the advantage function, and E is the expectation.
Citation Information
Patent Citations
Voltage reactive power optimization method based on double-layer reinforcement learning power grid-user cooperation
CN115313407A
Robot social motion planning method in man-machine co-fusion environment
CN116562332A