An agent obstacle avoidance path planning method and system for dynamic obstacles
Patent Information
- Application Number
- CN202611092903.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-22
- Publication Date
- 2026-09-18
- Estimated Expiration
- 2046-07-22
AI Technical Summary
[0005]本发明实施例提供一种面向动态障碍物的智能体避障路径规划方法、系统、电子设备和存储介质,以解决现有技术中智能体面对高速动态障碍物时因缺乏未来状态建模能力而导致的避障反应滞后、安全性差的技术问题
(1)通过策略网络直接输出纵向加速度和转向角速度指令,结合包含线性前进奖励、航向偏差惩罚、动态避障惩罚及风险感知奖励的综合奖励函数,使智能体能够主动预测动态障碍物运动趋势并提前调整航向,避免被动应激式避障,从而减少路径冗余与航行震荡,获得平滑且高效的避障路径。
Smart Images

Figure CN122613982B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of intelligent agent autonomous navigation technology, and in particular to an intelligent agent obstacle avoidance path planning method, system, electronic device and storage medium for dynamic obstacles. Background Technology
[0002] Intelligent agents have been widely applied in autonomous navigation missions such as marine operations, maritime patrols, and search and rescue. Path planning is the core technology for autonomous navigation of intelligent agents, aiming to plan a feasible path from the starting point to the target point in complex marine environments that meets the requirements of safe obstacle avoidance and navigation efficiency. Currently, research on obstacle avoidance for dynamic obstacles both domestically and internationally focuses on low-speed scenarios, or simply approximates dynamic obstacles as static obstacles. Related algorithms (such as the traditional Soft Actor-Critic (SAC) algorithm) can adapt to low-speed obstacles and steady-state ocean current environments to a certain extent. In addition, existing technologies typically use reward functions based on distance or collision detection to guide intelligent agents in learning obstacle avoidance strategies.
[0003] However, the aforementioned existing technologies have significant technical defects and limitations in practical applications. First, when an agent faces high-speed dynamic obstacles common in real-world navigation environments, the obstacles' high randomness, relative speed, and unpredictable trajectories greatly compress the agent's decision-making response time. Traditional SAC algorithms can only learn statistical correlations from historical state information, lacking the ability to explicitly model and reason about highly uncertain future states. This results in a "passive-reactive" obstacle avoidance strategy, where actions are only triggered when the obstacle approaches a dangerous distance, causing obstacle avoidance lag and making it difficult to ensure navigation safety. Second, existing reward functions are designed with coarse granularity, typically considering only macroscopic factors such as collision probability and target distance. They cannot quantify the real-time danger levels corresponding to different spatial distances, leading to a high failure rate in the early stages of algorithm training. Furthermore, the planned paths contain unnecessary redundant detours, reducing path smoothness and task arrival efficiency. Furthermore, existing technologies generally lack continuous constraints on the global navigation direction of intelligent agents, which can easily lead to frequent or large-scale heading deviations during obstacle avoidance, deviating from the optimal target direction, thereby increasing navigation energy consumption and path length, and failing to effectively balance obstacle avoidance safety and heading stability.
[0004] In summary, existing intelligent agent obstacle avoidance path planning methods for high-speed dynamic obstacle environments have significant shortcomings in terms of decision response timeliness, risk quantification fineness, and heading constraint coordination, making it difficult to meet the application requirements of actual high-speed navigation scenarios. Summary of the Invention
[0005] This invention provides an intelligent agent obstacle avoidance path planning method, system, electronic device, and storage medium for dynamic obstacles, to solve the technical problems of delayed obstacle avoidance response and poor safety caused by the lack of future state modeling ability when intelligent agents face high-speed dynamic obstacles.
[0006] In a first aspect, embodiments of the present invention provide an intelligent agent obstacle avoidance path planning method for dynamic obstacles, comprising: S1. Obtain the environmental state information of the intelligent agent; wherein, the environmental state information includes the state information of the intelligent agent, the relative navigation state of the target point relative to the intelligent agent, and the motion state of at least one dynamic obstacle; S2. Input the environmental state information into a pre-trained policy network, and output the action instructions of the agent according to the policy network; the action instructions include longitudinal acceleration instructions and steering angular velocity instructions. S3. Based on the kinematic model of the agent, update the position and heading of the agent according to the action command, and determine the immediate reward obtained by executing the action command according to the comprehensive reward function; the comprehensive reward function includes a target approach progress reward, a linear forward progress reward, a heading deviation penalty, a dynamic obstacle avoidance penalty, a target arrival reward, and a risk perception reward. S4. Store the current environmental state information, the action command, the instant reward, and the updated state information of the agent as a state transition sample into the experience replay pool, and sample from the experience replay pool to update the parameters of the policy network and the evaluation network. S5. Repeat S1 to S4 until the agent reaches the target point or triggers the preset task failure condition.
[0007] Preferably, in step S1, the state information includes the position, heading, and velocity of the agent; the relative navigation state includes the first relative distance and the first relative azimuth angle between the target point and the agent; and the motion state includes the position, velocity, and heading of the dynamic obstacle. The environmental state information also includes a relative velocity vector calculated based on the velocity vector difference between the agent and the dynamic obstacle, and a collision risk index calculated based on the second relative distance between the agent and the dynamic obstacle, the relative velocity vector, and the second relative azimuth angle of the dynamic obstacle relative to the agent.
[0008] Preferably, in step S3, the risk perception reward is configured as follows: a reward value is calculated using an exponential function based on the real-time distance between the agent and the dynamic obstacle, wherein the reward value decreases exponentially as the real-time distance decreases.
[0009] Preferably, in step S3, the dynamic obstacle avoidance penalty term adopts a three-level progressive structure: When the second relative distance is less than the first distance threshold, a constant collision termination penalty value is assigned, and the task failure condition is triggered. When the second relative distance is greater than or equal to the first distance threshold and less than the second distance threshold, a progressive penalty value that is linearly negatively correlated with the second relative distance is assigned. The penalty is zero when the second relative distance is greater than or equal to the second distance threshold; The calculation of the dynamic obstacle avoidance penalty also includes the extrapolation of the obstacle's motion state: based on the current speed and heading angle of the dynamic obstacle, a uniform linear motion model is used to linearly extrapolate and predict the position of the dynamic obstacle at the next moment. The predicted distance between the agent and the dynamic obstacle is used as the second relative distance, and the progressive penalty value is calculated.
[0010] Preferably, in step S3, the linear forward reward is determined based on the product of the agent's velocity and the cosine of the angle between the target direction and the target direction, where the target direction is the angle between the agent's heading and the direction from the agent to the target point. The heading deviation penalty term is determined based on the square of the angle between the target directions; The target approach progress reward is determined based on the difference between the distance from the agent to the target point at the previous time step and the current time step; The target arrival reward is configured as follows: when the first relative distance between the agent and the target point is less than the third distance threshold, a preset positive reward value is triggered.
[0011] Preferably, in step S4, updating the parameters of the policy network and the evaluation network includes: A dual-evaluation network structure is adopted, and the parameters of the two evaluation networks are updated by minimizing the loss function. When calculating the target Q value, the minimum Q value output by the two target evaluation networks is selected. The policy network is updated by minimizing the policy loss function, which is constructed based on the Q-value and temperature coefficient output by the evaluation network. The temperature coefficient is an adaptive temperature coefficient, which is optimized based on a preset target entropy and is limited to a preset lower bound.
[0012] Preferably, in step S4, when the number of samples in the experience replay pool reaches a preset batch size, the preset batch size of samples is randomly sampled for network updates. The task failure condition includes the distance between the agent and the dynamic obstacle being less than a preset collision distance threshold.
[0013] Secondly, embodiments of the present invention provide an intelligent agent obstacle avoidance path planning system for dynamic obstacles, comprising: A state acquisition module is used to acquire environmental state information of the intelligent agent; wherein, the environmental state information includes the state information of the intelligent agent, the relative navigation state of the target point relative to the intelligent agent, and the motion state of at least one dynamic obstacle; An action decision module is used to input the environmental state information into a pre-trained policy network and output action commands for the agent according to the policy network; the action commands include longitudinal acceleration commands and steering angular velocity commands. The state update and reward calculation module is used to update the position and heading of the agent according to the action command based on the agent's kinematic model, and to determine the immediate reward obtained by executing the action command according to the comprehensive reward function; the comprehensive reward function includes a target approach progress reward, a linear forward progress reward, a heading deviation penalty, a dynamic obstacle avoidance penalty, a target arrival reward, and a risk perception reward. The experience storage and network update module is used to store the current environmental state information, the action instruction, the instant reward, and the updated state information of the agent as a state transition sample into the experience replay pool, and to sample from the experience replay pool to update the parameters of the policy network and the evaluation network. The iterative control module is used to repeatedly call the state acquisition module, the action decision module, the state update and reward calculation module, and the experience storage and network update module until the agent reaches the target point or triggers the preset task failure condition.
[0014] Thirdly, embodiments of the present invention provide an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, it implements the steps of the intelligent agent obstacle avoidance path planning method for dynamic obstacles as described in the first aspect of the present invention.
[0015] Fourthly, embodiments of the present invention provide a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the intelligent agent obstacle avoidance path planning method for dynamic obstacles as described in the first aspect of the present invention.
[0016] This invention provides an intelligent agent obstacle avoidance path planning method, system, electronic device, and storage medium for dynamic obstacles. By acquiring multi-dimensional environmental state information such as the agent's own state, the target point's relative navigation state, and the motion state of dynamic obstacles, a comprehensive perception of the navigation scenario is constructed. Subsequently, this environmental state information is input into a pre-trained policy network, directly outputting two types of low-level action commands: longitudinal acceleration and steering angular velocity, achieving end-to-end decision-making from perception to control. After executing an action, the agent's pose is updated based on a kinematic model, and a comprehensive reward function containing six reward and penalty items—target approach progress, linear progress, heading deviation penalty, dynamic obstacle avoidance penalty, target arrival, and risk perception—quantifies the immediate reward of the action. This reward function considers multiple objectives such as approaching the target, maintaining heading, avoiding obstacles, and risk warning. Finally, interaction samples are stored in an experience replay pool and sampled for updating the policy network and evaluation network. Through iterative optimization, the agent gradually learns a safe and efficient obstacle avoidance strategy. Compared with existing technologies, it has the following beneficial effects: (1) The longitudinal acceleration and steering angular velocity commands are directly output through the policy network. Combined with the comprehensive reward function which includes linear forward reward, heading deviation penalty, dynamic obstacle avoidance penalty and risk perception reward, the agent can actively predict the dynamic obstacle movement trend and adjust the heading in advance to avoid passive stress-based obstacle avoidance, thereby reducing path redundancy and navigation oscillation, and obtaining a smooth and efficient obstacle avoidance path.
[0017] (2) The experience replay pool is used to store state transition samples and random sampling is performed to break the temporal correlation of continuous samples and suppress oscillations during training. At the same time, the comprehensive reward function provides multiple fine-grained feedback signals such as target approach progress, heading deviation, and dynamic obstacle avoidance, which reduces the blind exploration and high failure rate caused by sparse rewards in the early stage of training and accelerates the stable convergence of the policy network.
[0018] (3) The comprehensive reward function includes target approach progress reward, linear progress reward, and target arrival reward to incentivize the agent to drive towards the target point efficiently, as well as heading deviation penalty, dynamic obstacle avoidance penalty, and risk perception reward to constrain the agent to maintain a safe attitude and distance. This enables the agent to effectively avoid collision risks in complex dynamic environments and avoid path redundancy or delays caused by excessive obstacle avoidance, thus achieving coordinated optimization of safety and efficiency. Attached Figure Description
[0019] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0020] Figure 1 This is a flowchart of an intelligent agent obstacle avoidance path planning method for dynamic obstacles provided in an embodiment of the present invention; Figure 2 This is a schematic diagram of an experiment on high-speed dynamic obstacle avoidance path planning for an intelligent agent provided in an embodiment of the present invention; Figure 3 This invention provides an algorithmic framework for a path planning method in an embodiment of the invention. Figure 4 This is a schematic diagram of the structure of an intelligent agent obstacle avoidance path planning system for dynamic obstacles provided in an embodiment of the present invention; Figure 5 This is a schematic diagram of a physical structure provided for an embodiment of the present invention. Detailed Implementation
[0021] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0022] Figure 1 This is a flowchart of an intelligent agent obstacle avoidance path planning method for dynamic obstacles according to an embodiment of the present invention, with reference to... Figure 1 The method includes: S1. Obtain the environmental state information of the intelligent agent; wherein, the environmental state information includes the state information of the intelligent agent, the relative navigation state of the target point relative to the intelligent agent, and the motion state of at least one dynamic obstacle.
[0023] Among them, the intelligent agent is an autonomous mobile platform that can autonomously perceive the environment, plan paths and execute movements, including unmanned ships, vehicles and machines; environmental state information refers to the full-dimensional scene data perceived by the intelligent agent during navigation; the state information of the intelligent agent is its own position, heading and speed, which is the basis for control decisions; the relative navigation state is the distance and azimuth of the target point relative to the intelligent agent, which is used to guide the navigation direction; the motion state of dynamic obstacles is the position, speed and heading of the obstacles, which reflects their rapid movement trend in high-speed scenarios.
[0024] It should be noted that in this embodiment, the dynamic obstacle is a high-speed dynamic obstacle with a speed greater than 20 knots. Due to the high speed of the obstacle, its state needs to be updated in real time.
[0025] Existing technologies often rely on single static or low-speed obstacle information, typically only acquiring the instantaneous position of the obstacle while ignoring its dynamic characteristics such as speed and heading. This makes it impossible to predict the movement trend of high-speed obstacles, and difficult to characterize the randomness and high relative speed of high-speed dynamic obstacles, resulting in incomplete perception and delayed decision-making. In this embodiment, S1, through multi-dimensional state fusion, comprehensively captures the real-time motion characteristics of high-speed dynamic obstacles, providing accurate data support for subsequent early prediction and proactive obstacle avoidance, significantly improving the completeness and timeliness of environmental perception.
[0026] S2. Input the environmental state information into the pre-trained policy network, and output the action instructions of the agent according to the policy network; the action instructions include longitudinal acceleration instructions and steering angular velocity instructions.
[0027] Among them, the policy network is a neural network used for decision-making in deep reinforcement learning, which learns the optimal obstacle avoidance policy through offline training; the action command is a direct command to control the movement of the agent, the longitudinal acceleration command controls the navigation speed, and the steering angular velocity command controls the navigation direction.
[0028] By using a trained policy network to achieve end-to-end decision-making, it can quickly respond to the dynamic changes of high-speed obstacles and output precise movement commands, transforming passive, reactive obstacle avoidance into proactive predictive planning, thus significantly improving decision-making response speed and the rationality of actions.
[0029] S3. Based on the kinematic model of the agent, update the position and heading of the agent according to the action command, and determine the immediate reward obtained by executing the action command according to the comprehensive reward function; the comprehensive reward function includes a target approach progress reward, a linear forward progress reward, a heading deviation penalty, a dynamic obstacle avoidance penalty, a target arrival reward, and a risk perception reward.
[0030] The kinematic model is a mathematical model describing the changes in the agent's position and heading with speed and turn, used to simulate navigation. The comprehensive reward function is the core guiding mechanism of reinforcement learning, with six reward items quantifying the merits of actions from multiple dimensions, including goal approach, navigation efficiency, heading stability, obstacle avoidance safety, task completion, and risk warning.
[0031] Traditional SAC (Soft Actor-Critic) reward functions are coarse-grained, focusing only on a single objective, which easily leads to obstacle avoidance lag, path redundancy, high training failure rates, and an inability to quantify the collision risk of high-speed obstacles. This embodiment, through a finely designed set of six rewards, balances navigation efficiency, heading stability, and high-speed obstacle risk control, quantifies hazard levels in real time, and suppresses high-risk behaviors, effectively balancing obstacle avoidance safety and navigation efficiency, reducing redundant paths, and improving training stability and obstacle avoidance safety.
[0032] S4. Store the current environmental state information, the action command, the instant reward, and the updated state information of the agent as a state transition sample into the experience replay pool, and sample from the experience replay pool to update the parameters of the policy network and the evaluation network.
[0033] The state transition samples are time-series data of (current state, action, reward, next state), recording a decision-making and feedback process; the experience replay pool is used to store historical samples, breaking the temporal correlation of samples; the policy network and the evaluation network are a dual network structure of reinforcement learning. The evaluation network evaluates the value of actions, and the policy network optimizes decisions based on value. The evaluation network, also known as the Critic network, evaluates the value of state-action pairs to guide the policy network update.
[0034] Traditional algorithms are susceptible to the influence of temporal correlations among samples, resulting in unstable training, slow convergence, and an inability to effectively learn complex decision-making patterns in high-speed dynamic obstacle scenarios. This embodiment disrupts the temporal sequence of samples through an experience replay mechanism, and, in conjunction with a dual-network structure, accurately evaluates the value of actions, stabilizes the network training process, and rapidly converges to the optimal obstacle avoidance strategy, continuously improving the model's adaptability and decision-making accuracy in high-speed dynamic obstacle scenarios.
[0035] S5. Repeat S1 to S4 until the agent reaches the target point or triggers the preset task failure condition.
[0036] Through continuous iteration, the agent gradually optimizes its strategy through trial and error, ultimately learning the complete obstacle avoidance path from the starting point to the target point. Addressing the characteristics of short decision-making response times and the need for rapid iterative learning in high-speed dynamic obstacle environments, this embodiment ensures that the entire reinforcement learning framework can run continuously until the task is completed or a clear failure occurs. Its technical effect is to achieve end-to-end autonomous obstacle avoidance navigation for the agent in complex dynamic environments, efficiently reaching the target point while ensuring safety. This provides a convergent training framework and reliable decision-making strategies for practical deployment.
[0037] like Figure 2 As shown, in a 600m × 600m two-dimensional simulated water area, one intelligent agent, three high-speed dynamic obstacle boats, and a fixed target point are deployed. The initial speeds of the three obstacle boats are set to 20m / s, 21m / s, and 22m / s, respectively, and their course is random but all face or intersect the planned route of the intelligent agent, simulating high-speed ships or floating bodies moving through a real marine environment. The starting point of the intelligent agent is located on one side of the water area, and the target point is located on the other side, with high-speed moving obstacles blocking the way.
[0038] The agent acquires its own state (position, heading, velocity), the relative state of the target point (relative distance, azimuth), and the motion state of each obstacle (position, velocity vector, heading angle) in real time through shipboard sensors (such as radar and AIS). This environmental information is fused into a current state vector stst, which is then input into a pre-trained RA-SAC (Risk Awareness Soft Actor-Commentator) policy network.
[0039] Policy network according to s t Output continuous action instructions a This includes longitudinal acceleration (controlling throttle / brake) and steering angular velocity (controlling rudder angle). The agent updates its position and heading based on a three-degree-of-freedom kinematic model to reach the next state. s t+1 Simultaneously, the immediate reward is calculated based on the comprehensive reward function. r t The reward function includes six components: target approach progress reward, linear progress reward, heading deviation penalty, dynamic obstacle avoidance penalty (three-level progressive structure + extrapolation), target arrival reward, and risk perception reward. The dynamic obstacle avoidance penalty uses the obstacle's speed and heading to perform linear extrapolation to predict the relative distance at the next moment, thus triggering the obstacle avoidance penalty in advance before the obstacle approaches.
[0040] The experience samples generated by each interaction ( s t , a t , r t , s t+1 The samples are stored in the experience replay buffer. When the number of samples in the buffer reaches a preset batch size (e.g., 256), a random batch of samples is sampled to update the policy network and the evaluation network (dual-mode). Q (Network, adaptive temperature coefficient, soft update) to continuously optimize obstacle avoidance strategies.
[0041] The entire simulation process lasted for 800 training rounds. In each round, the agent started from the starting point and repeated the above "perception-decision-execution-store-update" cycle until it successfully reached the target point (distance to the target point less than 20 meters) or triggered the failure condition (distance to any obstacle less than 25 meters). After sufficient training, the RA-SAC algorithm enabled the agent to learn to proactively predict risks and adjust its course in advance under high-speed dynamic obstacle interference, realizing the transformation from passive reactive obstacle avoidance to proactive predictive planning, and finally safely reaching the target point with a smooth and efficient path.
[0042] Based on the above embodiments, as a preferred implementation, in step S1, the state information includes the position, heading, and speed of the agent; the relative navigation state includes the first relative distance and the first relative azimuth angle between the target point and the agent; and the motion state includes the position, speed, and heading of the dynamic obstacle.
[0043] The intelligent agent collects and fuses multi-dimensional environmental data in real time through shipborne sensors (such as radar, AIS, and visual sensors) to construct complete environmental state information, including the intelligent agent's own state, the target point's relative navigation state, the dynamic obstacle's motion state, and derived risk characteristics.
[0044] Specifically, the agent's state information includes its two-dimensional position coordinates in the geodetic coordinate system, heading angle (the angle between the agent's longitudinal axis and true north), and longitudinal velocity (a scalar value, the velocity component along the heading direction). These parameters collectively describe the agent's current motion state. The target point's relative navigation state to the agent includes the first relative distance (the Euclidean distance from the agent's current position to the target point) and the first relative azimuth angle (the angle between the line connecting the agent to the target point and true north), used to guide the agent towards the target point. The motion state of dynamic obstacles includes the obstacle's position, velocity vector (including magnitude and direction), and heading angle. Due to the obstacle's high-speed motion characteristics (e.g., speeds exceeding 20 m / s), its state needs to be updated every frame to capture its motion trend.
[0045] The environmental state information also includes a relative velocity vector calculated based on the velocity vector difference between the agent and the dynamic obstacle, and a collision risk index calculated based on the second relative distance between the agent and the dynamic obstacle, the relative velocity vector, and the second relative azimuth angle of the dynamic obstacle relative to the agent.
[0046] Building upon the aforementioned basic state, this preferred embodiment further constructs two derived risk characteristic parameters to enhance the quantitative perception capability of environmental threats. First, the relative velocity vector, calculated based on the velocity vector difference between the agent and the dynamic obstacle, reflects the magnitude and direction of the obstacle's approaching speed relative to the agent, and is a key physical quantity for judging the urgency of a collision. Second, the collision risk index, a comprehensive scalar indicator, is calculated based on the second relative distance between the agent and the dynamic obstacle (i.e., the real-time Euclidean distance between them), the aforementioned relative velocity vector, and the second relative azimuth angle of the dynamic obstacle relative to the agent (i.e., the angle between the line connecting the agent to the obstacle and the agent's current heading). This index can be fused using a weighted function or a lookup table, incorporating factors such as the reciprocal of the distance, the component of the relative velocity in the direction of the connecting line, and the sine of the azimuth angle; a higher value indicates a greater collision risk. By incorporating the relative velocity vector and the collision risk index into the environmental state information, the agent not only knows where the obstacle is, but can also quantitatively assess its approach speed and the magnitude of the threat.
[0047] This embodiment addresses the shortcomings of existing technologies, such as limited environmental perception dimensions and a lack of dynamic threat trend modeling. Traditional methods often only acquire the instantaneous position of obstacles, treating them as static or low-speed targets, failing to predict whether high-speed approaching obstacles will cut into the flight path, leading to passive obstacle avoidance decisions. This embodiment quantifies the movement trend and potential threat of obstacles into feature values that can be input into the network by explicitly calculating the relative velocity vector and collision risk index, enabling the policy network to learn the ability to perceive risks in advance. The technical effects are twofold: firstly, it provides direct quantitative basis for the risk perception reward term and dynamic obstacle avoidance penalty term in the subsequent reward function, allowing the agent to proactively adjust its course or speed before obstacles enter a dangerous distance; secondly, it improves the richness and information density of the deep reinforcement learning state input, accelerating the policy network's identification and convergence of high-risk patterns, ultimately achieving a paradigm shift from passive reaction to proactive prediction in obstacle avoidance.
[0048] Based on the above embodiments, as a preferred implementation, in step S3, the risk perception reward item is configured as follows: the reward value is calculated using an exponential function based on the real-time distance between the agent and the dynamic obstacle, wherein the reward value decreases exponentially as the real-time distance decreases.
[0049] Real-time distance refers to the distance at the current moment. t The Euclidean distance between the agent and any obstacle is denoted as . d obs The unit is meters. This real-time distance can be calculated in real time from sensor data. An exponential function refers to a function that uses natural numbers as the unit of measurement. eAn exponential function with base 0, where the reward value refers to the output value of the risk perception reward item, denoted as . r risk In this embodiment, the specific formula for the reward value is as follows: r risk =-exp(- d obs ) In the above formula, exp(·) represents the expression in terms of the natural constant. e It is an exponential function with base 0. The mathematical property of this function is: when 0. d obs When the value is large (e.g., greater than 50 meters), exp(- d obs The value approaches 0, therefore r risk Approaching 0 (weak negative incentive or no punishment); when d obs As it gradually decreases, exp(- d obs It gradually increases, but because it has a negative sign in front of it, r risk The value becomes negative and its absolute value increases exponentially with decreasing distance. In other words, the smaller the distance, the smaller the perceived risk reward (i.e., the heavier the punishment), and the rate of increase in punishment intensity accelerates sharply as the distance approaches. This exponential design means that the agent receives only slight negative incentives when far from obstacles, but once it enters the close range (e.g., within 25 meters), the punishment will rapidly amplify, creating strong avoidance pressure.
[0050] Existing reward functions typically employ only linear or piecewise constant obstacle avoidance penalties, failing to differentiate threat levels at different distances. This results in agents only receiving significant negative feedback when very close to obstacles, leading to frequent collisions in the early stages of training and making it difficult for policy networks to learn to avoid obstacles in advance. Furthermore, traditional methods lack non-linear risk quantification, causing the agent's reaction to obstacles to be threshold-triggered—action only occurs when the distance is below a certain value, and is ignored when it exceeds that value, neglecting the continuity and urgency of risk changes.
[0051] This preferred embodiment designs an exponential risk perception reward based on real-time distance, implicitly incorporating the danger thresholds of the 25m warning zone and the 50m early warning zone into the reward change pattern. It quantifies the spatial danger level in real time. When the agent approaches an obstacle and the real-time distance decreases, the reward value decreases exponentially and rapidly, creating a strong negative incentive and effectively inhibiting the agent from entering high-risk areas. This embodiment transforms the traditional "passive reactive obstacle avoidance" into "active predictive avoidance": the agent perceives the exponentially increasing risk even when the obstacle is still within the warning distance, thus reacting in advance. In practical applications, for example, when a high-speed obstacle boat approaches from the side at 20m / s, the exponential risk reward will cause the agent to feel a significant negative incentive at a distance of 50 meters, leading to a small turn in advance and avoiding the need for subsequent sharp turns or sudden stops. This significantly reduces the collision rate in the early training phase, accelerates policy network convergence, improves the smoothness and lead time of obstacle avoidance actions, and enables the agent to be sensitive to risk trends, enhancing its survivability in high-speed dynamic obstacle environments.
[0052] Based on the above embodiments, as a preferred implementation, in step S3, the dynamic obstacle avoidance penalty term adopts a three-level progressive structure: When the second relative distance d obs Less than the first distance threshold d When the value is 1, a constant collision termination penalty is assigned, and the mission failure condition is triggered.
[0053] When the second relative distance d obs Greater than or equal to the first distance threshold d 1 and less than the second distance threshold d At time 2, assign a relative distance to the second. d obs A progressive penalty value that exhibits a linear negative correlation.
[0054] When the second relative distance d obs Greater than or equal to the second distance threshold d At time 2, the penalty is zero.
[0055] The calculation of the dynamic obstacle avoidance penalty also includes the extrapolation of the obstacle's motion state: based on the current speed and heading angle of the dynamic obstacle, a uniform linear motion model is used to linearly extrapolate and predict the position of the dynamic obstacle at the next moment. The predicted distance between the agent and the dynamic obstacle is used as the second relative distance, and the progressive penalty value is calculated.
[0056] Specifically, in this embodiment, the formula for the dynamic obstacle avoidance penalty term is:
[0057] in, K 1 represents the collision termination penalty constant, which is negative. k 3 is a negative linear penalty coefficient, which determines the slope of the penalty as a function of distance within the warning zone.
[0058] Furthermore, in this embodiment, d 1. The value is 25 meters. d 2 is taken as 50 meters.
[0059] Traditional obstacle avoidance methods only penalize based on the distance at the current moment. When an obstacle rushes towards the agent at high speed, due to decision lag, the agent may fail to react in time when the distance is still in the warning zone (e.g., 40 meters). By the time the action is executed, the distance has plummeted to the collision zone (<25 meters), causing an irreversible collision. Existing technologies lack advance modeling of obstacle movement trends, resulting in passive, reactive obstacle avoidance that cannot proactively avoid obstacles before they approach. Furthermore, a single linear penalty or a simple threshold penalty cannot reflect the non-uniform changes in risk across different distance ranges. This leaves the agent without sufficient avoidance pressure in the warning zone, while the sudden, large penalty in the collision zone makes it difficult to learn effective avoidance strategies.
[0060] In this embodiment, the space is first divided into a safe zone (≥50 meters), a warning zone (25-50 meters), and a danger zone (<25 meters) through a three-level progressive structure. No penalty is imposed in the safe zone, allowing free navigation. In the warning zone, a linear penalty negatively correlated with distance is applied, with the penalty increasing as the distance increases, forming a continuous pressure gradient that guides the agent to actively adjust its course or decelerate before entering the danger zone. Once the agent enters the danger zone, a termination penalty is triggered, ending the mission and avoiding ineffective exploration.
[0061] Furthermore, this embodiment introduces obstacle motion state extrapolation. Based on the current position, speed, and heading angle of the dynamic obstacle, its predicted position after a preset time step is calculated. Then, the Euclidean distance between the agent's current position and the predicted position is calculated as the predicted distance. This allows the agent to see the obstacle's movement trend: even if the current distance is still in the warning zone (e.g., 45 meters), if the obstacle approaches at high speed, the predicted distance may have already entered a smaller value (e.g., 35 meters), thus obtaining a higher penalty in advance and driving the agent to take obstacle avoidance actions earlier. In practical applications, when facing an obstacle boat approaching head-on at a speed of 22 m / s, the agent, at a distance of 50 meters, extrapolates and discovers that the distance will drop to 39 meters in 0.5 seconds, and will immediately obtain the corresponding progressive penalty, thus turning or decelerating in time. This transforms passive reactive obstacle avoidance into active predictive planning, significantly improving the response lead to high-speed dynamic obstacles; reducing the collision rate in the early stages of training through three-level gradient penalties, accelerating strategy convergence; and improving the smoothness and safety of the obstacle avoidance path, avoiding emergency sharp turns or collisions due to insufficient reaction.
[0062] Based on the above embodiments, as a preferred implementation, in step S3, the linear forward reward is determined based on the product of the agent's velocity and the cosine of the angle between the target direction and the target direction, where the angle between the agent's heading and the direction from the agent to the target point is the angle between the agent's heading and the direction from the agent to the target point.
[0063] Specifically, velocity refers to the longitudinal velocity of the agent, denoted as . u The unit is m / s; the included angle of the target direction is denoted as Δ. ψ , is the angle between the agent's current heading and the target direction (the direction of the line connecting the agent to the target point), and its value range is usually [ π,π];cosine value is cos(Δ ψ ).
[0064] In this embodiment, the formula for calculating the linear forward reward is: r forward = k 1· u ·cos(Δ ψ ) in, k A positive weighting coefficient of 1 is used for this linear forward reward term, which applies when the agent is traveling at a high speed and its heading is aligned with the target direction (Δ). ψ =0, cos(Δ ψ When )=1), the maximum positive reward is obtained; when the course deviation is large (such as Δ), the maximum positive reward is obtained. ψ When =90°), cos(Δ ψWhen the heading deviates by more than 90° (i.e., away from the target direction), the reward is zero; when the heading deviates by more than 90° (i.e., away from the target direction), the reward is zero. ψ A negative value for the reward function generates a negative reward (penalty), encouraging the agent to turn around as quickly as possible. In existing technologies, the reward function lacks guidance on navigation efficiency, causing the agent to easily travel along circuitous paths at low speeds during obstacle avoidance, resulting in excessively long task times or high energy consumption. This invention, by multiplying speed by heading alignment, incentivizes the agent to travel as directly as possible to the target at high speed while ensuring safety, avoiding ineffective detours. This significantly improves the task arrival efficiency of path planning, reduces travel time and energy consumption, and provides a positive complementary signal for subsequent heading deviation penalties.
[0065] The heading deviation penalty is determined based on the square of the included angle of the target direction; the specific calculation formula is as follows: r angle = k 2·(Δ ψ ) 2 in, k 2 represents a negative weighting coefficient. This heading deviation penalty term results in a smaller penalty value when the heading deviation is small (the square function makes the impact of small deviations much smaller than that of a linear function); as the deviation increases, the penalty value increases quadratically, effectively suppressing large-angle yawing. Unlike the linear forward reward, this term is always non-negative (zero or positive). Since it is included as a penalty term in the comprehensive reward function, a larger value indicates a heavier penalty for heading deviation, effectively providing a negative incentive. Existing methods lack heading stability constraints during obstacle avoidance, leading to frequent or large-amplitude turns by the agent, resulting in path oscillations, poor smoothness, and even serpentine navigation. This invention uses a square penalty, whose value increases with the heading deviation, to suppress excessive heading deviations in the comprehensive reward function. This improves the smoothness and heading stability of the planned path, reduces unnecessary turning actions, lowers the burden on the control system, and works synergistically with the linear forward reward term to achieve a balance between obstacle avoidance safety and navigation stability.
[0066] The target proximity progress reward is determined based on the difference between the distance from the agent to the target point at the previous time step and the current time step; the specific calculation formula is as follows: r progress = d prev - d curr in, d prev This represents the distance from the target point at the previous moment. d currThis represents the distance to the target point at the current moment. The progress reward for getting closer to the target point is awarded if the current moment is closer to the target point than the previous moment (…). d curr < d prev If the distance to the target point is far, the difference is positive, and a positive reward is given; conversely, if the distance is far from the target point, the difference is negative, and a penalty is given. This is a dense, step-by-step progress-driven reward that complements the sparse final-arrival reward. Traditional methods only set sparse arrival rewards (rewards are given only upon reaching the target point), which makes it difficult for the agent to obtain effective feedback in the early stages of training, resulting in low exploration efficiency and a tendency to get stuck in local minima or wander aimlessly. This invention provides instant feedback through the distance change at each time step, so that the agent clearly knows that moving closer to the target is the correct behavior, and can continue to receive positive incentives even if the target is still far away. This significantly reduces the difficulty of reinforcement learning training, accelerates the convergence of the policy network to goal-oriented behavior, avoids ineffective exploration, and improves the task success rate.
[0067] The target arrival reward is configured as follows: when the first relative distance between the agent and the target point is less than a third distance threshold, a preset positive reward value is triggered. The specific calculation formula is:
[0068] in, K 2 represents a preset positive reward value. The third distance threshold in this embodiment is 20 meters. This reward is triggered only once when the agent enters a circular area with a radius of 20 meters centered on the target point, providing a high positive reward and signifying task completion. Relying solely on progress rewards (the distance difference per step) may cause the agent to deviate slightly due to obstacle avoidance when approaching the target, preventing it from ever entering the final threshold range; simultaneously, the lack of a clear success signal makes it difficult to determine whether the task is finished. This embodiment of the invention sets a terminal reward, providing high positive feedback when the agent is sufficiently close to the target point, reinforcing the value of reaching the target as a key behavior, and serving as a termination condition. Its technical effect is: providing a clear task success signal, ensuring that the agent can ultimately converge to the vicinity of the target point after safely avoiding obstacles, avoiding wandering around the target point or missing the target, and improving the task completion rate and practicality of path planning.
[0069] Furthermore, the comprehensive reward function consists of the following components: target approach progress reward. r progress Linear progress reward r forward Heading deviation penalty r angle Dynamic obstacle avoidance penalty items r obstacle Target achievement reward rgoal Risk perception reward r risk Specifically: R= r progress + r forward + r angle + r obstacle + r goal + r risk Based on the above embodiments, as a preferred implementation, in step S4, as follows: Figure 3 As shown, updating the parameters of the policy network and the evaluation network includes: A dual-evaluation network structure is adopted, and the parameters of the two evaluation networks are updated by minimizing the loss function. When calculating the target Q value, the minimum Q value output by the two target evaluation networks is selected.
[0070] The policy network is updated by minimizing the policy loss function, which is constructed based on the Q-value and temperature coefficient output by the evaluation network.
[0071] The temperature coefficient is an adaptive temperature coefficient, which is optimized based on a preset target entropy and is limited to a preset lower bound.
[0072] Furthermore, the dual-evaluation network structure refers to the present invention employing two target evaluation networks with identical structures but independent parameters, and their corresponding online evaluation networks, denoted as... Q target,θ1 , Q target,θ2 and Q 1. Q 2. Used to suppress common problems in traditional single-evaluation networks. Q The problem is overestimation; the loss function is the optimization objective when evaluating network updates, i.e., minimizing the prediction. Q Values and Objectives Q Mean square error between values; target Q Value is denoted as y t The sum of the current reward and the expected reward discounted to the next state is used to evaluate the network's regression objective; the policy loss function is used to update the policy network (Actor network, with parameters denoted as...). Its goal is to maximize the evaluation of the network output. Q The value, while minimizing the entropy of the strategy, encourages exploration; the temperature coefficient is denoted as... αThis is used to control the weight of policy entropy in the loss function, balancing exploration and exploitation; the preset target entropy is denoted as... H target is a hyperparameter that is expected to have an entropy value close to the target value to ensure that the policy has sufficient randomness; the lower bound is denoted as . α min To prevent insufficient exploration due to an excessively small temperature coefficient; the experience replay pool is recorded as... D Used to store historical state transition samples ( s t , a t , r t , s t+1 ),in, s t Given the current environmental state, a t For the action instructions to be executed, r t For instant rewards, s t+1 The state at the next moment; the discount factor is denoted as c It is used to weigh the importance of current rewards against future rewards; the soft update coefficient is denoted as... t This is used to smooth the updates of the target network parameters, and its value is 1×10. 3 The action probability distribution output by the policy network is denoted as... π ( a t | s t ), indicating the state s t The probability density of taking each action.
[0073] Furthermore, the evaluation network loss function (dual network) is as follows:
[0074] in, i 1, i 2 represents the parameters of two online evaluation networks. Q 1( s t , a t )and Q 2( s t , a t ) represent the two evaluation networks in state. s t Next action at Prediction Q value; y t For the goal Q value; E ( s t , a t , r t , s t+1 ) indicates from the experience replay pool D The mathematical expectation of random sampling.
[0075] Target Q The value (using dual-network minimum suppression to suppress overestimation) is:
[0076] in, This represents the output of the evaluation network for two objectives. Q The minimum value among the values is used to suppress overestimation; a t+1 The action for the next moment is determined by the policy network based on the next state. s t+1 Obtained by sampling; π ( a t+1 | s t+1 ) represents the action probability density output by the policy network; α The temperature coefficient controls the entropy term log. π The weight.
[0077] The policy network loss function is:
[0078] in, π ( a t | s t ) indicates that the current policy network is in state. s t Down Output Action a t The probability density; This is an entropy regularization term that encourages strategies to maintain randomness. The output of two online evaluation networks Q The minimum value among the values is used to evaluate the quality of the action.
[0079] The adaptive temperature coefficient loss function and its update are as follows:
[0080] in, H target The target entropy is a preset value (take the negative value of the action dimension, such as -dim(A)). α min This is the lower bound of the temperature coefficient (to prevent the exploration from disappearing completely). or α The learning rate for the temperature coefficient; L α ( α ) represents the loss function pair α The gradient; during update, gradient descent is used first, then compared with... α min Take the maximum value and then crop it.
[0081] The target network uses a soft update method:
[0082]
[0083] in, t This is the soft update coefficient, with a value of 1×10. 3 This allows the target network parameters to slowly track the online network parameters, improving training stability.
[0084] This invention employs a dual-evaluation network to minimize overestimation, utilizes an adaptive temperature coefficient to automatically adjust the exploration level based on the current policy entropy, breaks temporal correlations between samples through random sampling in an experience replay pool, and combines this with soft updates to smooth changes in target network parameters, thus constructing a stable deep reinforcement learning training framework. This significantly improves the stability and convergence speed of the training process, avoiding policy degradation caused by overestimation. The adaptive temperature coefficient allows the agent to fully explore high-risk areas in the early stages of training and stably utilize learned policies in the later stages. Soft updates avoid oscillations caused by parameter mutations. Overall, this enables the policy network to learn robust and efficient obstacle avoidance path planning strategies, making it particularly suitable for continuous control tasks in high-speed dynamic obstacle environments.
[0085] Based on the above embodiments, as a preferred implementation, in step S4, when the number of samples in the experience replay pool reaches a preset batch size, the preset batch size of samples is randomly sampled for network updates.
[0086] The task failure condition includes the distance between the agent and the dynamic obstacle being less than a preset collision distance threshold.
[0087] Specifically, the experience replay pool is a cache for storing historical interaction samples, with each sample being a state transition quadruple. s t , a t , r t , s t+1 The preset batch size refers to the number of samples randomly drawn from the experience replay pool each time. In this embodiment, it is set to 256. This value ensures the stability and computational efficiency of gradient estimation. Random sampling means that samples are not drawn in chronological order, but are selected uniformly and randomly from the pool to break the temporal correlation between consecutive samples. Network update uses this batch of samples to calculate the loss function and backpropagates to update the parameters of the policy network and the evaluation network. The preset collision distance threshold refers to the distance limit for determining whether the agent collides with a dynamic obstacle. In this embodiment, it is set to 25 meters. When the second relative distance is less than this threshold, the task fails and the current round is terminated.
[0088] In deep reinforcement learning, if consecutive samples are used for updates sequentially in chronological order, the strong correlation between adjacent samples can lead to unstable network training and a tendency to get trapped in local optima. Furthermore, without effective task failure conditions, the agent will continue ineffective exploration after collisions, wasting computational resources and hindering the learning of obstacle avoidance strategies. This embodiment addresses this by using a random sampling mechanism in the experience replay pool, ensuring that the samples used for each update approximately satisfy the independent and identically distributed assumption, significantly improving training stability and sample utilization. By setting a collision distance threshold, the task is immediately terminated when the agent enters a danger zone (less than 25 meters in this embodiment). This avoids prolonging training time with invalid trajectories and provides the network with a clear failure signal, reinforcing the behavior criterion of maintaining a safe distance.
[0089] Secondly, embodiments of the present invention provide an intelligent agent obstacle avoidance path planning system for dynamic obstacles, based on the methods described in the above embodiments, such as... Figure 4 As shown, the system 400 includes: The state acquisition module 410 is used to acquire the environmental state information of the intelligent agent; wherein, the environmental state information includes the state information of the intelligent agent, the relative navigation state of the target point relative to the intelligent agent, and the motion state of at least one dynamic obstacle.
[0090] The action decision module 420 is used to input the environmental state information into a pre-trained policy network and output the action instructions of the agent according to the policy network; the action instructions include longitudinal acceleration instructions and steering angular velocity instructions.
[0091] The state update and reward calculation module 430 is used to update the position and heading of the agent according to the action command based on the agent's kinematic model, and to determine the immediate reward obtained by executing the action command according to the comprehensive reward function; the comprehensive reward function includes a target approach progress reward, a linear forward progress reward, a heading deviation penalty, a dynamic obstacle avoidance penalty, a target arrival reward, and a risk perception reward.
[0092] The experience storage and network update module 440 is used to store the current environmental state information, the action instruction, the instant reward, and the updated state information of the agent as a state transition sample into the experience replay pool, and to sample from the experience replay pool to update the parameters of the policy network and the evaluation network.
[0093] The iterative control module 450 is used to repeatedly call the state acquisition module, the action decision module, the state update and reward calculation module, and the experience storage and network update module until the agent reaches the target point or triggers the preset task failure condition.
[0094] Based on the same concept, this invention also provides a schematic diagram of a physical structure, such as... Figure 5 As shown, the server may include a processor 510, a communications interface 520, a memory 530, and a communication bus 540, wherein the processor 510, communications interface 520, and memory 530 communicate with each other via the communication bus 540. The processor 510 can call logical instructions in the memory 530 to execute the steps of the intelligent agent obstacle avoidance path planning method for dynamic obstacles as described in the above embodiments.
[0095] Furthermore, the logical instructions in the aforementioned memory 530 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0096] Based on the same concept, embodiments of the present invention also provide a non-transitory computer-readable storage medium storing a computer program containing at least one piece of code that can be executed by a master control device to control the master control device to implement the steps of the intelligent agent obstacle avoidance path planning method for dynamic obstacles as described in the above embodiments.
[0097] Based on the same technical concept, this application also provides a computer program, which, when executed by a main control device, is used to implement the above-described method embodiments.
[0098] The program may be stored, in whole or in part, on a storage medium packaged with the processor, or in part or in whole on a memory not packaged with the processor.
[0099] The embodiments described above are merely illustrative of several implementations of the present invention, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of the invention. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of the present invention, and these modifications and improvements all fall within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the appended claims.
[0100] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for intelligent agent obstacle avoidance path planning for dynamic obstacles, characterized in that, include: S1. Obtain the environmental state information of the intelligent agent; wherein, the environmental state information includes the state information of the intelligent agent, the relative navigation state of the target point relative to the intelligent agent, and the motion state of at least one dynamic obstacle; S2. Input the environmental state information into a pre-trained policy network, and output the action instructions of the agent according to the policy network; the action instructions include longitudinal acceleration instructions and steering angular velocity instructions; the environmental state information includes a relative velocity vector calculated based on the velocity vector difference between the agent and the dynamic obstacle, and a collision risk index calculated based on the second relative distance between the agent and the dynamic obstacle, the relative velocity vector, and the second relative azimuth angle of the dynamic obstacle relative to the agent; S3. Based on the kinematic model of the agent, update the position and heading of the agent according to the action command, and determine the immediate reward obtained by executing the action command according to the comprehensive reward function; the comprehensive reward function includes a target approach progress reward, a linear forward progress reward, a heading deviation penalty, a dynamic obstacle avoidance penalty, a target arrival reward, and a risk perception reward; the risk perception reward is configured as follows: based on the real-time distance between the agent and the dynamic obstacle, the reward value is calculated through an exponential function, wherein the reward value decreases exponentially as the real-time distance decreases; S4. Store the current environmental state information, the action command, the instant reward, and the updated state information of the agent as a state transition sample into the experience replay pool, and sample from the experience replay pool to update the parameters of the policy network and the evaluation network. S5. Repeat S1 to S4 until the agent reaches the target point or triggers the preset task failure condition. The dynamic obstacle avoidance penalty adopts a three-level progressive structure: When the second relative distance is less than the first distance threshold, a constant collision termination penalty value is assigned, and the task failure condition is triggered. When the second relative distance is greater than or equal to the first distance threshold and less than the second distance threshold, a progressive penalty value that is linearly negatively correlated with the second relative distance is assigned. The penalty is zero when the second relative distance is greater than or equal to the second distance threshold; The calculation of the dynamic obstacle avoidance penalty also includes the extrapolation of the obstacle's motion state: based on the current speed and heading angle of the dynamic obstacle, a uniform linear motion model is used to linearly extrapolate and predict the position of the dynamic obstacle at the next moment. The predicted distance between the agent and the dynamic obstacle is used as the second relative distance, and the progressive penalty value is calculated.
2. The intelligent agent obstacle avoidance path planning method for dynamic obstacles according to claim 1, characterized in that, In step S1, the state information includes the position, heading, and velocity of the agent; the relative navigation state includes the first relative distance and the first relative azimuth angle between the target point and the agent; and the motion state includes the position, velocity, and heading of the dynamic obstacle.
3. The intelligent agent obstacle avoidance path planning method for dynamic obstacles according to claim 2, characterized in that, In S3, the linear forward reward is determined based on the product of the agent's velocity and the cosine of the angle between the target direction and the target direction. The angle between the target direction and the heading of the agent is the angle between the direction from the agent to the target point. The heading deviation penalty term is determined based on the square of the angle between the target directions; The target approach progress reward is determined based on the difference between the distance from the agent to the target point at the previous time step and the current time step; The target arrival reward is configured as follows: when the first relative distance between the agent and the target point is less than the third distance threshold, a preset positive reward value is triggered.
4. The intelligent agent obstacle avoidance path planning method for dynamic obstacles according to claim 1, characterized in that, In step S4, updating the parameters of the policy network and the evaluation network includes: A dual-evaluation network structure is adopted, and the parameters of the two evaluation networks are updated by minimizing the loss function. When calculating the target Q value, the minimum Q value output by the two target evaluation networks is selected. The policy network is updated by minimizing the policy loss function, which is constructed based on the Q-value and temperature coefficient output by the evaluation network. The temperature coefficient is an adaptive temperature coefficient, which is optimized based on a preset target entropy and is limited to a preset lower bound.
5. The intelligent agent obstacle avoidance path planning method for dynamic obstacles according to claim 1, characterized in that, In step S4, when the number of samples in the experience replay pool reaches a preset batch size, the preset batch size of samples is randomly sampled for network updates. The task failure condition includes the distance between the agent and the dynamic obstacle being less than a preset collision distance threshold.
6. A smart agent obstacle avoidance path planning system for dynamic obstacles, used to execute the smart agent obstacle avoidance path planning method for dynamic obstacles as described in any one of claims 1 to 5, characterized in that, include: A state acquisition module is used to acquire environmental state information of the intelligent agent; wherein, the environmental state information includes the state information of the intelligent agent, the relative navigation state of the target point relative to the intelligent agent, and the motion state of at least one dynamic obstacle; An action decision module is used to input the environmental state information into a pre-trained policy network and output action commands for the agent according to the policy network; the action commands include longitudinal acceleration commands and steering angular velocity commands. The state update and reward calculation module is used to update the position and heading of the agent according to the action command based on the agent's kinematic model, and to determine the immediate reward obtained by executing the action command according to the comprehensive reward function; the comprehensive reward function includes a target approach progress reward, a linear forward progress reward, a heading deviation penalty, a dynamic obstacle avoidance penalty, a target arrival reward, and a risk perception reward. The experience storage and network update module is used to store the current environmental state information, the action instruction, the instant reward, and the updated state information of the agent as a state transition sample into the experience replay pool, and to sample from the experience replay pool to update the parameters of the policy network and the evaluation network. The iterative control module is used to repeatedly call the state acquisition module, the action decision module, the state update and reward calculation module, and the experience storage and network update module until the agent reaches the target point or triggers the preset task failure condition.
7. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the steps of the intelligent agent obstacle avoidance path planning method for dynamic obstacles as described in any one of claims 1 to 5.
8. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When executed by a processor, the computer program implements the steps of the intelligent agent obstacle avoidance path planning method for dynamic obstacles as described in any one of claims 1 to 5.