Mobile robot dynamic path planning method based on improved DDQN algorithm

Through the integration of the improved DDQN algorithm and the BOAE mechanism, the problem of insufficient adaptability and autonomous learning capabilities of traditional path planning algorithms in dynamic environments is solved, and the efficient path planning and obstacle avoidance capabilities of mobile robots in complex environments is achieved.

CN119952697AActive Publication Date: 2025-05-09JIANGSU UNIV OF SCI & TECH

Patent Information

Application Number
CN202510049295.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-13
Publication Date
2025-05-09
Estimated Expiration
2045-01-13

AI Technical Summary

Technical Problem

Traditional path planning algorithms lack adaptability and independent learning ability in dynamic environments, making it difficult to effectively avoid obstacles, increasing collision risks and reducing the efficiency and accuracy of path planning.

Method used

The improved DDQN algorithm and BOAE mechanism are used to integrate the improved DDQN algorithm and the BOAE mechanism, and the synergistic effect of the neural network module, experience playback module and the agent control module can realize efficient path planning of the robot in a dynamic environment.

Benefits of technology

It significantly improves the obstacle avoidance ability and path planning of mobile robots in dynamic environments, and reduces training time and resource consumption.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119952697A_ABST
    Figure CN119952697A_ABST
Patent Text Reader

Abstract

The invention discloses a mobile robot dynamic path planning method based on an improved DDQN algorithm. The method comprises the following steps: establishing a dynamic simulation environment; constructing a path planning model which comprises a neural network module, an experience playback module and an agent control module; initializing a dynamic simulation environment and a path planning model; training the path planning model based on a fusion strategy of an improved DDQN and BOAE mechanism; the trained mobile robot path planning model is loaded, the intelligent agent drives the robot to operate in the dynamic environment according to the optimal action decision output by the strategy network, the environment simulation module feeds back the position and state of the robot in real time, and the visualization module displays the optimal path of the robot. Through an innovative mode of combining BOAE and an improved DDQN algorithm, the mobile robot can perform path planning more efficiently and accurately in a dynamic environment, the obstacle avoidance capability and the success rate of reaching a target are remarkably improved, and the method has extremely wide application prospects and popularization values.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of mobile robots and relates to a dynamic path planning technology, and in particular to a dynamic path planning method for a mobile robot based on an improved DDQN algorithm. Background Art

[0002] In the development history of mobile robots, path planning has always been the core link to realize their autonomous navigation function. Traditional path planning algorithms, represented by the A* algorithm, show certain planning efficiency in static environments and can plan relatively reasonable paths based on preset environmental maps and target information. However, with the increasing expansion and complexity of robot application scenarios, especially in scenarios with dynamic obstacles and real-time changes in environmental information, traditional algorithms have exposed serious limitations.

[0003] Traditional algorithms lack adaptability and autonomous learning capabilities to dynamic changes in the environment, and are unable to adjust path planning strategies in a timely manner based on real-time changes in the environment. This defect makes it difficult for robots to make effective obstacle avoidance decisions when facing dynamic obstacles, increasing the risk of collision and reducing the efficiency and accuracy of path planning.

[0004] The emergence of the DDQN (Deep Double Q Network) algorithm in the field of deep reinforcement learning has brought hope for solving the problem of path planning in dynamic environments. The DDQN algorithm approximates the Q-value function through a deep neural network, enabling the robot to gradually learn the optimal action strategy in the process of interacting with the environment. Despite this, the DDQN algorithm still has many shortcomings under the test of complex dynamic scenes. Its sample efficiency is low, and a large number of interaction samples are required to achieve a relatively ideal training effect, which undoubtedly increases training time and resource consumption. In addition, it has poor adaptability to dynamic obstacles, and it is difficult to accurately cope with the complex and changeable movement patterns of obstacles, resulting in a significant reduction in the accuracy and stability of path planning in practical applications. Summary of the invention

[0005] Purpose of the invention: In order to overcome the deficiencies in the prior art, a dynamic path planning method for a mobile robot based on an improved DDQN algorithm is provided, which innovatively introduces the BOAE (Bidirectional Obstacle Avoidance Enhancement) mechanism and organically integrates it with the improved DDQN algorithm. The BOAE mechanism enhances the robot's obstacle avoidance capability from multiple dimensions, and works synergistically with the improved DDQN algorithm to form an efficient and adaptable dynamic path planning solution, which opens up a new way to improve the path planning performance of mobile robots in complex dynamic environments.

[0006] Technical solution: To achieve the above purpose, the present invention provides a mobile robot dynamic path planning method based on an improved DDQN algorithm, comprising the following steps:

[0007] S1: Build a dynamic simulation environment;

[0008] S2: Build a path planning model, including a neural network module, an experience playback module, and an agent control module;

[0009] S3: Initialize the dynamic simulation environment and path planning model;

[0010] S4: Train the path planning model based on the fusion strategy of improved DDQN and BOAE mechanism;

[0011] S5: Load the trained mobile robot path planning model. The intelligent agent drives the robot to operate in a dynamic environment based on the optimal action decision output by the policy network. The environmental simulation module provides real-time feedback on the robot's position and status. The visualization module dynamically records the robot's path. After the test, the visualization module displays the robot's optimal path.

[0012] Furthermore, the construction of the dynamic simulation environment in step S1 includes:

[0013] Accurately set the starting coordinates (such as np.array([-2.2,-3.3])) and target coordinates (such as np.array([2.88,3.2])) of the mobile robot to clarify the starting and ending points of the path planning, providing a basic target orientation for subsequent path planning and navigation;

[0014] For dynamic obstacles, a detailed list of their attributes is defined, covering geometric information (such as shape and size) and dynamic motion rules (including motion parameters such as initial position, speed, acceleration, etc.); this rich obstacle information provides a key basis for the robot in the obstacle avoidance decision-making process, enabling it to predict the movement trajectory and trend of obstacles in advance, thereby formulating a more accurate and efficient obstacle avoidance strategy;

[0015] For various sensors, the sampling period of GPS sensor is set to 1ms, and the touch sensor and distance sensor collect data synchronously based on this. The distance sensor is obtained through strict screening and sorting. The screening rule is that the name contains 'so' and contains numeric characters, and is sorted in ascending order according to the numeric part. The collected sensor data is normalized and integrated into the observation vector. According to the maximum and minimum values ​​of each sensor's own range, the data range is standardized to [0,1] to meet the neural network input requirements, ensure the consistency and validity of the data, and facilitate the neural network to perform efficient processing and analysis, thereby providing accurate data support for the robot's decision-making.

[0016] The action space is defined as spaces.Discrete(3), which corresponds to the three discrete actions of moving forward, turning left, and turning right. This concise and effective action definition can meet the basic motion requirements of the robot in most scenarios. The shape of the observation space is set to ([X]), which fully integrates key elements such as target distance, sensor data, current position, and obstacle avoidance related information provided by the BOAE mechanism. [X] is the dimension of the observation space after considering the BOAE information, which is determined according to the specific output of BOAE. This enables the robot to make comprehensive and accurate path planning decisions based on multi-faceted information, thereby improving its perception and response capabilities to complex environments.

[0017] Furthermore, the neural network module in step S2 includes a constructed policy network and a target network, and a multi-layer fully connected layer structure is built based on the PyTorch framework; in the multi-layer fully connected layer structure, the first layer fc1 maps the input dimension to 256 dimensions, the middle layer fc2 maintains 256-dimensional processing, and the output dimension of the last layer fc3 matches the dimension of the action space; ReLU is selected as the activation function between each layer, and the policy network generates the Q value corresponding to each discrete action based on the observation vector that integrates the BOAE information, providing an accurate basis for the robot's action decision; the target network is used to assist the policy network training, initially copying the policy network parameters, and subsequently adjusting the parameters according to specific update rules; through this dual network structure and reasonable network layer design, the network's learning ability and decision-making accuracy can be effectively improved. At the same time, with the help of the auxiliary role of the target network, the training process can be stabilized and the occurrence of problems such as overestimation of the Q value can be reduced.

[0018] Furthermore, the experience replay module in step S2 uses a priority experience replay mechanism based on the SumTree data structure to store rich experience data generated by the interaction between the robot and the dynamic environment; each experience data unit covers core information such as the current state, execution action, reward, next state, and whether the task is completed. This module can efficiently sample experience batches according to the set priority, generate importance sampling weights in combination with the experience priority when sampling, effectively balance the deviation caused by priority sampling, and dynamically update the priority of the stored experience according to the training process. In particular, for experience related to dynamic obstacles and in which the BOAE mechanism plays an important role in obstacle avoidance, its priority can be appropriately adjusted to make high-value experience easier to sample. This experience replay mechanism can make full use of the valuable experience generated in the process of interaction between the robot and the environment, improve sample utilization, accelerate the convergence speed of network training, and further improve the adaptability and coping ability to dynamic obstacle scenes through reasonable priority adjustment.

[0019] Furthermore, in step S2, the intelligent agent control module cooperates with other modules to reasonably configure key training parameters, including discount factor, learning rate, exploration rate and its decay rate, minimum exploration rate, target network update frequency, memory capacity of experience replay, etc., and introduces n-step parameter, soft update parameter τ and regularization coefficient λ to implement specific improvement strategies.

[0020] Furthermore, in step S4, the path planning model is trained by combining the n-step reward calculation mechanism, the soft update and regularization combination strategy and the BOAE mechanism fusion strategy, wherein:

[0021] n-step reward calculation mechanism: The existing single-step reward calculation often fails to fully consider the long-term impact of the robot's actions, while the n-step reward calculation mechanism of the present invention can comprehensively consider the cumulative reward effect brought by the robot's actions in n consecutive steps. In this way, the robot's action strategy can be evaluated more comprehensively, and the Q value overestimation problem caused by the single-step reward calculation in the DDQN algorithm can be alleviated, so that the robot can make more forward-looking and long-term decisions in the path planning process, thereby improving the accuracy and stability of path planning.

[0022] Combination of soft update and regularization strategy: In the target network update link, the traditional hard update method is abandoned and the soft update strategy is adopted. The traditional hard update method may cause sudden changes in the target network parameters, affecting the stability of training. The soft update strategy introduces the soft update parameter τ, so that the target network parameters can gradually and smoothly approach the policy network parameters, avoiding drastic fluctuations in the parameter update process; at the same time, a regularization term is added to the loss function calculation, and the loss function is adjusted by the regularization coefficient λ; the regularization term can effectively prevent the network from overfitting, so that the network can not only adapt to the training data during the learning process, but also maintain the generalization ability of unknown data. Through the combination of soft update and regularization, the training process is stabilized, the phenomenon of Q value overestimation is reduced, and overfitting is prevented, thereby improving the accuracy and efficiency of the robot's path planning in complex dynamic environments.

[0023] BOAE mechanism fusion strategy: During the training phase, the agent selects actions to drive the robot according to the exploration rate rule combined with the obstacle avoidance suggested actions or action adjustment information of the BOAE mechanism. The BOAE mechanism analyzes obstacle avoidance reactions from the state and action, and uses cross-attention and duel networks and generates auxiliary reward factors to enhance the robot's obstacle avoidance ability in multiple dimensions. Combining it with the improved DDQN algorithm, the robot can perceive the environment more accurately and adjust the action strategy more effectively when facing dynamic obstacles, further improving the obstacle avoidance ability and the adaptability of path planning.

[0024] Furthermore, the training of the path planning model in step S4 includes the following process:

[0025] The intelligent agent carefully selects actions based on the current observation vector, exploration rate, and obstacle avoidance recommended actions or action adjustment information provided by the BOAE mechanism. After the robot executes the selected action in a dynamic environment, the environment simulation module provides real-time feedback on rewards, next state, and whether it ends. The intelligent agent stores the interaction experience in the experience replay module, which samples experience according to priority to provide data batches for training. During the training process, the n-step reward, time difference error, etc. are calculated, and the policy network and target network parameters are updated using soft update strategies and regularization terms. At the same time, the experience priority is adjusted according to the evaluation of the obstacle avoidance effect by the BOAE mechanism to accelerate the convergence of the policy network. In this process, the intelligent agent continuously interacts with the environment and gradually improves its path planning ability through learning and optimization. The close collaboration and parameter adjustment between modules are the key to achieving efficient training.

[0026] In combination with the above content, the innovations of the present invention are mainly reflected in the following three aspects: First, an n-step reward calculation mechanism is introduced to comprehensively consider the long-term cumulative rewards of the robot's actions, effectively alleviating the problem of overestimation of Q values; second, soft updates are combined with regularization to smoothly update the target network parameters and prevent network overfitting, thereby improving the stability and accuracy of training; third, DDQN and BOAE mechanisms are organically integrated to enhance the robot's obstacle avoidance capability and path planning adaptability from multiple dimensions, providing strong technical support for the autonomous navigation of mobile robots in complex dynamic environments.

[0027] Beneficial effects: Compared with the prior art, the present invention, through the above three innovative aspects, by combining BOAE with the improved DDQN algorithm, can enable the mobile robot to plan paths more efficiently and accurately in a dynamic environment, significantly improve the obstacle avoidance capability and the success rate of reaching the target, and has extremely broad application prospects and promotion value. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] Figure 1 It is a schematic diagram of the process of the present invention;

[0029] Figure 2 This is the work flow diagram of the BOAE mechanism;

[0030] Figure 3 It is a multi-layer fully connected layer structure diagram;

[0031] Figure 4 It is a smooth curve graph of the rewards of the three DRL algorithms;

[0032] Figure 5 The trajectory planning effect diagram of three DRL algorithms;

[0033] Figure 6 Trajectory length graphs planned for the three DRL algorithms; DETAILED DESCRIPTION

[0034] The present invention is further explained below in conjunction with the accompanying drawings and specific embodiments. It should be understood that these embodiments are only used to illustrate the present invention and are not used to limit the scope of the present invention. After reading the present invention, various equivalent forms of modifications to the present invention by those skilled in the art all fall within the scope defined by the claims attached to this application.

[0035] like Figure 1 As shown, the present invention provides a mobile robot dynamic path planning method based on an improved DDQN algorithm, comprising the following steps:

[0036] S1: Build a dynamic simulation environment;

[0037] Accurately set the starting coordinates (such as np.array([-2.2,-3.3])) and target coordinates (such as np.array([2.88,3.2])) of the mobile robot to clarify the starting and ending points of the path planning, providing a basic target orientation for subsequent path planning and navigation;

[0038] For dynamic obstacles, a detailed list of their attributes is defined, covering geometric information (such as shape and size) and dynamic motion rules (including motion parameters such as initial position, speed, acceleration, etc.); this rich obstacle information provides a key basis for the robot in the obstacle avoidance decision-making process, enabling it to predict the movement trajectory and trend of obstacles in advance, thereby formulating a more accurate and efficient obstacle avoidance strategy;

[0039] For various sensors, the sampling period of GPS sensor is set to 1ms, and the touch sensor and distance sensor collect data synchronously based on this. The distance sensor is obtained through strict screening and sorting. The screening rule is that the name contains 'so' and contains numeric characters, and is sorted in ascending order according to the numeric part. The collected sensor data is normalized and integrated into the observation vector. According to the maximum and minimum values ​​of each sensor's own range, the data range is standardized to [0,1] to meet the neural network input requirements, ensure the consistency and validity of the data, and facilitate the neural network to perform efficient processing and analysis, thereby providing accurate data support for the robot's decision-making.

[0040] The action space is defined as spaces.Discrete(3), which corresponds to the three discrete actions of moving forward, turning left, and turning right. This concise and effective action definition can meet the basic motion requirements of the robot in most scenarios. The shape of the observation space is set to ([X]), which fully integrates key elements such as target distance, sensor data, current position, and obstacle avoidance related information provided by the BOAE mechanism. [X] is the dimension of the observation space after considering the BOAE information, which is determined according to the specific output of BOAE. This enables the robot to make comprehensive and accurate path planning decisions based on multi-faceted information, thereby improving its perception and response capabilities to complex environments.

[0041] S2: Build a path planning model, including a neural network module, an experience playback module, and an agent control module;

[0042] The neural network module includes the constructed policy network and target network, and builds a multi-layer fully connected layer structure based on the PyTorch framework; Figure 3 As shown in the figure, in the multi-layer fully connected layer structure, the first layer fc1 maps the input dimension to 256 dimensions, the middle layer fc2 maintains 256 dimensions, and the output dimension of the last layer fc3 matches the dimension of the action space; ReLU is used as the activation function between each layer, and the policy network generates the Q value corresponding to each discrete action based on the observation vector that integrates the BOAE information, providing an accurate basis for the robot's action decision; the target network is used to assist the policy network training, initially copying the policy network parameters, and then adjusting the parameters according to specific update rules; through this dual network structure and reasonable network layer design, the network's learning ability and decision-making accuracy can be effectively improved. At the same time, with the help of the auxiliary role of the target network, the training process can be stabilized and the occurrence of problems such as overestimation of Q values ​​can be reduced;

[0043] The multi-layer fully connected layer extracts features and learns representations of input data through each layer, realizing feature combination, transformation, and mapping from low-level to high-level. In classification and regression tasks, the classification or numerical prediction of data is realized based on the matching of output dimensions and tasks. The data is fitted with powerful function fitting capabilities, and generalization capabilities are guaranteed through regularization and other means, learning the general laws of data for accurate prediction and classification.

[0044] The policy network generates the Q value corresponding to each discrete action based on the observation vector that integrates the BOAE information. The calculation process of the Q value includes:

[0045] 1) Input processing

[0046] First, the observation vector x fused with BOAE information is taken as input. Assume that the observation vector x = {x1, x2, ..., x n}, where n is the dimension of the observation vector;

[0047] 2) First fully connected layer calculation

[0048] The observation vector x is input to the first fully connected layer fc1. Each neuron j in fc1 performs a weighted summation of the input and adds a bias term b. 1j ,Right now: where w 1ij is the connection weight from the ith input to the jth neuron in fc1;

[0049] Then apply the ReLU activation function: 1j =ReLU(z 1j )=max(0,z 1j ), and get the output of fc1: a1 = [a 11 ,a 12 ,...,a 1m ], where m = 256, which is the output dimension of fc1;

[0050] 3) Middle layer fully connected layer calculation

[0051] The calculation of the intermediate layer fc2 is similar to that of fc1. For each neuron k in fc2, w 2jk is the connection weight from the jth input to the kth neuron in fc2, b 2k is the bias term;

[0052] Apply the ReLU activation function again: a 2k =ReLU(z 2k )=max(0,z 2k ), and get the output of fc2: a2 = [a 21 ,a 22 ,...,a 2m ]

[0053] 4) Calculation of the last fully connected layer

[0054] The calculation of the last layer fc3, for each output action corresponding to the dimension l, w 3kl is the connection weight from the kth input to the lth output in fc3, b 3l is the bias term. Here, the value range of l corresponds to the dimension of the action space;

[0055] At this time, z 3l It is the Q value corresponding to each discrete action, that is, Q l =z 3l , the final Q value vector Q=[Q1,Q2,...,Q s ], s is the dimension of the action space, which represents the value of each discrete action taken by the robot under the current observation.

[0056] The experience replay module uses a priority experience replay mechanism based on the SumTree data structure to store rich experience data generated by the interaction between the robot and the dynamic environment; each experience data unit covers core information such as the current state, execution action, reward, next state, and whether the task is completed. This module can efficiently sample experience batches according to the set priority, generate importance sampling weights based on the experience priority when sampling, effectively balance the deviation caused by priority sampling, and dynamically update the priority of stored experience according to the training process. In particular, for experience related to dynamic obstacles and where the BOAE mechanism plays an important role in obstacle avoidance, its priority can be appropriately adjusted to make high-value experience easier to sample. This experience replay mechanism can make full use of the valuable experience generated in the process of interaction between the robot and the environment, improve sample utilization, and accelerate the convergence speed of network training. At the same time, through reasonable priority adjustment, it can further improve the adaptability and coping ability to dynamic obstacle scenes;

[0057] The intelligent agent control module coordinates the operation of various modules and reasonably configures key training parameters, including discount factor, learning rate, exploration rate and its decay rate, minimum exploration rate, target network update frequency, memory capacity of experience replay, etc. It also introduces the n-step parameter, soft update parameter τ and regularization coefficient λ to implement specific improvement strategies.

[0058] S3: Initialize the dynamic simulation environment and path planning model;

[0059] 1. Dynamic simulation environment initialization

[0060] 1) Coordinate setting: Accurately set the starting coordinates and target coordinates of the mobile robot, such as setting the starting coordinates to np.array([-2.2,-3.3]) and the target coordinates to np.array([2.88,3.2]) to clearly define the starting and ending points of the path planning.

[0061] 2) Obstacle definition: A detailed definition of the attribute list of dynamic obstacles, covering geometric information (such as shape and size) and dynamic motion laws (including initial position, velocity, acceleration and other motion parameters).

[0062] 3) Sensor settings: Set the sampling period of the GPS sensor to 1ms, and the touch sensor and distance sensor collect data synchronously based on this. Strictly filter and sort the distance sensors. The filter rules are that the names contain "so" and contain numeric characters, and are sorted in ascending order according to the numeric part. The collected sensor data is normalized and integrated into the observation vector.

[0063] 4) Space definition: The action space is defined as spaces.Discrete(3), corresponding to the three discrete actions of moving forward, turning left, and turning right. The shape of the observation space is set to ([X]), which fully integrates key elements such as target distance, sensor data, current position, and obstacle avoidance related information provided by the BOAE mechanism.

[0064] 2. Path planning model initialization

[0065] Neural network module: The policy network and target network are built with a multi-layer fully connected layer structure based on the PyTorch framework. The first layer fc1 maps the input dimension to 256 dimensions, the middle layer fc2 maintains 256-dimensional processing, and the output dimension of the last layer fc3 matches the dimension of the action space. ReLU is used as the activation function between each layer. The target network initially copies the policy network parameters.

[0066] Experience replay module: Initialize the priority experience replay mechanism based on the SumTree data structure, and set parameters such as the memory capacity of the experience replay.

[0067] Agent control module: Rationally configure key training parameters, including discount factor, learning rate, mining rate and its decay rate, minimum exploration rate, target network update frequency, etc., and initialize n-step parameter, soft update parameter τ and regularization coefficient λ.

[0068] S4: Train the path planning model based on the fusion strategy of improved DDQN and BOAE mechanism, including:

[0069] The path planning model is trained by combining the n-step reward calculation mechanism, the soft update and regularization combination strategy, and the BOAE mechanism fusion strategy, where:

[0070] n-step reward calculation mechanism: The existing single-step reward calculation often fails to fully consider the long-term impact of the robot's actions, while the n-step reward calculation mechanism of the present invention can comprehensively consider the cumulative reward effect brought by the robot's actions in n consecutive steps. In this way, the robot's action strategy can be evaluated more comprehensively, and the Q value overestimation problem caused by the single-step reward calculation in the DDQN algorithm can be alleviated, so that the robot can make more forward-looking and long-term decisions in the path planning process, thereby improving the accuracy and stability of path planning.

[0071] The n-step reward calculation mechanism is as follows: Assume that the single-step reward of step t is r t , discount factor γ, then the reward for n steps is R t It can be expressed as

[0072] Combination of soft update and regularization strategy: In the target network update link, the traditional hard update method is abandoned and the soft update strategy is adopted. The traditional hard update method may cause sudden changes in the target network parameters, affecting the stability of training. The soft update strategy introduces the soft update parameter τ, so that the target network parameters can gradually and smoothly approach the policy network parameters, avoiding drastic fluctuations in the parameter update process; at the same time, a regularization term is added to the loss function calculation, and the loss function is adjusted by the regularization coefficient λ; the regularization term can effectively prevent the network from overfitting, so that the network can not only adapt to the training data during the learning process, but also maintain the generalization ability of unknown data. Through the combination of soft update and regularization, the training process is stabilized, the phenomenon of Q value overestimation is reduced, and overfitting is prevented, thereby improving the accuracy and efficiency of the robot's path planning in complex dynamic environments.

[0073] Soft update strategy: The update formula of the target network parameter θ' is θ'←(1-τ)θ'+τθ, where θ' is the policy network parameter and τ is the soft update parameter, which is usually small, so that the target network parameter can gradually and smoothly approach the policy network parameter.

[0074] Regularization term: Add a regularization term to the calculation of the loss function L. Assuming that the original loss function is L0, the regularization term is Ω(θ), and the regularization coefficient is λ, the loss function after adding regularization is L=L0+λΩ(θ). By adjusting λ, the strength of regularization can be controlled to prevent network overfitting.

[0075] Reference Figure 2 , BOAE mechanism fusion strategy: During the training phase, the agent selects actions to drive the robot according to the exploration rate rule combined with the obstacle avoidance suggested actions or action adjustment information of the BOAE mechanism. The BOAE mechanism analyzes the obstacle avoidance response from the state and action, and uses cross-attention and duel networks and generates auxiliary reward factors to enhance the robot's obstacle avoidance ability in multiple dimensions. Combining it with the improved DDQN algorithm, the robot can perceive the environment more accurately and adjust the action strategy more effectively when facing dynamic obstacles, further improving the obstacle avoidance ability and the adaptability of path planning.

[0076] The improvements in the improved DDQN algorithm of the present invention are as follows:

[0077] First: introduce n-step rewards and combine multi-step rewards with priority experience playback. (Multi-step rewards are an important way to calculate rewards in reinforcement learning. By considering the accumulation of rewards in multiple time steps in the future, the objective function can be optimized more comprehensively.)

[0078] Second: A soft update strategy is adopted when updating the network and regularization is combined.

[0079] The above improvements are all aimed at solving the problem of overestimation of Q value in the traditional DDQN algorithm.

[0080] Combination of BOAE mechanism and improved DDQN algorithm:

[0081] During the training phase, the agent selects actions to drive the robot based on the exploration rate rule combined with the obstacle avoidance recommended actions or action adjustment information of the BOAE mechanism. The BOAE mechanism analyzes the obstacle avoidance response from the state and action, and the information it provides is integrated into the observation vector. The policy network generates the Q value corresponding to each discrete action based on the observation vector that integrates the BOAE information, thereby driving the robot's action decision.

[0082] The following is a detailed description of the workflow of the BOAE mechanism:

[0083] 1. Analysis of obstacle avoidance response unit

[0084] 1.1 Status Assessment

[0085] The robot collects information about the environment and its own state through a variety of sensors (such as GPS, touch, distance sensors, etc.). This information is normalized and integrated into an observation vector s. Taking the distance sensor data d as an example, the normalization formula is where d min and d max are the minimum and maximum values ​​of the sensor range, ensuring that the data is in the range of [0,1] for easy processing by the neural network.

[0086] The system determines whether the robot is close to an obstacle based on the observation vector s. Set the distance threshold D thresh , when the normalized distance d between the robot and the obstacle norm <D thresh When the robot is judged to be in a potential danger area, the obstacle avoidance program needs to be started.

[0087] 1.2 Action Evaluation

[0088] The action space is defined as a discrete set A = {forward, turn left, turn right}, corresponding to a∈{0,1,2}. This unit analyzes whether the current action selection will cause a collision risk. For example, if the robot chooses to move forward and an obstacle is detected in front, this action has a high risk. Combining the current state s and action a, it is determined whether the robot is in a dangerous state. If it is in a dangerous state, it is necessary to adjust the action according to the real-time environment to avoid collision.

[0089] 2. Obstacle avoidance direction suggestions

[0090] 2.1 Direction determination based on sensor data

[0091] Determine the position of the obstacle relative to the robot based on the distance sensor data in multiple directions. Assume that there are n distance sensors around the robot, and the data are d1, d2, ..., d n . Compare the data size of each sensor. If d i If it is the smallest and significantly smaller than other sensor data, it indicates that the obstacle is in the direction corresponding to sensor i.

[0092] Plan a moving path for the robot to move away from this direction. For example, if the obstacle is in front (assuming the front sensor corresponds to d j is the smallest), the recommended action is to turn left or right to avoid direct collision.

[0093] 2.2 Optimization based on environmental information

[0094] Considering the robot's target position information, select an obstacle avoidance direction that can avoid obstacles and get as close to the target direction as possible. Suppose the robot's current position is P = (x, y), and the target position is T = (x t ,y t ), the obstacle position is O = (x0, y0). Calculate the vector from the robot's current position to the target position and the vector from the robot's current position to the obstacle's position

[0095] Through vector operations, such as calculating the angle between vectors If the angle is small, it means that the obstacle is close to the target direction. At this time, the obstacle avoidance direction should be selected to stay close to the target direction.

[0096] 3. Obstacle avoidance urgency assessment

[0097] 3.1 Distance-Speed ​​Evaluation Model

[0098] According to the distance d between the robot and the obstacle and the relative speed v of the obstacle rel Evaluate the urgency of obstacle avoidance. Define the urgency index U, the formula is where v rel It is the relative speed after considering the movement of obstacles and the movement of the robot itself. When U exceeds the set threshold U thresh When U is high, it indicates that the obstacle avoidance situation is urgent and immediate action is required. When U is low, the obstacle avoidance path can be planned relatively calmly.

[0099] 3.2 Time-Collision Assessment

[0100] Predict the collision time T between the robot and the obstacle in the current motion state collisionAssume that the motion trajectories of the robot and the obstacle are straight lines (a simplified model, a more accurate trajectory can actually be obtained through a motion prediction model), and calculate the collision time based on their current position, speed, and motion direction. If T collision Less than the set safety time threshold T safe , then the urgency of obstacle avoidance is high and quick decision-making is required; otherwise, obstacle avoidance can be planned according to conventional strategies.

[0101] 4. Optimize action value assessment

[0102] 4.1 Introducing auxiliary features

[0103] When calculating the action value Q(s,a), in addition to the state s and action a, auxiliary features related to obstacle avoidance are introduced. For example, the distance d between the robot and the nearest obstacle is min , relative speed v rel The angle θ between the obstacle avoidance direction and the target direction is taken as additional input features. These features are integrated into the input layer of the neural network and are calculated together with the original input features through an additional weight matrix to more comprehensively evaluate the value of the action in the current obstacle avoidance scenario.

[0104] 4.2 Value Adjustment Based on Priority

[0105] Adjust the action value based on the obstacle avoidance urgency assessment results. For situations with high urgency, pay more attention to improving the action value that can immediately avoid the collision; for situations with low urgency, you can optimize the action value while taking into account long-term goals.

[0106] Assume that the urgency adjustment factor is α (0<α<1, the higher the urgency, the closer α is to 1), the adjusted action value Q'(s,a)=(1-α)Q(s,a)+αQ urgent (s,a), where Q urgent (s,a) is the action value calculated for emergency obstacle avoidance situations.

[0107] 5. Reward factor generation unit

[0108] 5.1 Distance Reward

[0109] Design distance reward function r d , encouraging the robot to stay away from obstacles. When the distance d between the robot and the obstacle is greater than the safety distance d safe When d =k1(dd safe ) where k1 is a positive coefficient, indicating the reward the robot gets for staying away from obstacles; when d≤d safe When d =-k2(d safe -d), where k2 is a positive coefficient representing the penalty the robot receives for approaching an obstacle.

[0110] 5.2 Direction Rewards

[0111] Considering the consistency between the obstacle avoidance direction and the target direction, define the direction reward function r θ Suppose the angle between the robot's current direction and the target direction is θ. When the direction chosen by the robot during obstacle avoidance reduces θ, a positive reward is given; otherwise, a negative reward is given. θ =k3cosθ, where k3 is a positive coefficient that measures the degree of directional proximity through cosθ.

[0112] 5.3 Comprehensive Reward Factor

[0113] The distance reward r d , Direction Reward θ And other related rewards (such as action smoothness rewards, etc.) are combined to obtain the final auxiliary reward factor r aux =r d +r θ +r other This auxiliary reward factor is used as an additional reward signal, combined with the reward r of environmental feedback, to train the agent, prompting the agent to pay more attention to obstacle avoidance and goal-oriented behavior during the learning process.

[0114] The training of the path planning model includes the following steps:

[0115] The intelligent agent carefully selects actions based on the current observation vector, exploration rate, and obstacle avoidance recommended actions or action adjustment information provided by the BOAE mechanism. After the robot executes the selected action in a dynamic environment, the environment simulation module provides real-time feedback on rewards, next state, and whether it ends. The intelligent agent stores the interaction experience in the experience replay module, which samples experience according to priority to provide data batches for training. During the training process, the n-step reward, time difference error, etc. are calculated, and the policy network and target network parameters are updated using soft update strategies and regularization terms. At the same time, the experience priority is adjusted according to the evaluation of the obstacle avoidance effect by the BOAE mechanism to accelerate the convergence of the policy network. In this process, the intelligent agent continuously interacts with the environment and gradually improves its path planning ability through learning and optimization. The close collaboration and parameter adjustment between modules are the key to achieving efficient training.

[0116] The specific operation process of the environmental simulation module includes:

[0117] 1) Environmental status feedback

[0118] Position feedback:

[0119] The environment simulation module can track the robot's position information in a dynamic environment in real time. For example, in a two-dimensional plane environment, it will record the robot's coordinates (x, y) and feed this information back to the agent. This position information is crucial for the agent to understand the robot's current location and plan the next action.

[0120] Status feedback:

[0121] In addition to the position, other states of the robot are also fed back, such as speed, direction, etc. If the robot is in an environment with complex terrain or special conditions, the environment simulation module will also feedback state information such as whether the robot is on a slope, whether it is disturbed by a magnetic field, etc. These state information together with the robot's position information constitute a complete state feedback, helping the intelligent agent to fully understand the current state of the robot.

[0122] 2) Reward feedback

[0123] Reward generation principle:

[0124] Generate reward signals based on the robot's actions and environmental changes. For example, if the robot successfully avoids a dynamic obstacle, the environment simulation module will generate a positive reward r (such as r = +1), and the size and positive and negative of the reward value are determined according to the preset reward mechanism.

[0125] The reward mechanism is usually related to factors such as whether the robot is close to the target, whether it successfully avoids obstacles, and whether the action is efficient. If the robot moves towards the target, it may get a small positive reward; if the robot collides with an obstacle or moves away from the target, it may get a negative reward (such as r = -1).

[0126] Reward feedback function:

[0127] The reward signal is fed back to the agent to update the agent's policy network. The agent continuously tries different actions, receives rewards from the environment simulation module, and learns action strategies that can get more rewards, thereby optimizing path planning.

[0128] 3) Next state provision

[0129] State transfer simulation:

[0130] When the agent selects an action a (such as moving forward, turning left or right), the environment simulation module simulates the changes in the environment and the update of the robot's state according to the robot's current state and the selected action to obtain the robot's next state s'.

[0131] This process involves complex environment models and robot motion models. For example, for a simple two-dimensional robot motion, if the robot's current position is (x, y) and its velocity is (v x ,v y ), after selecting the forward action, the next position may be updated to (x+v x Δt,y+v y Δt), where Δt is the time step. At the same time, the environment simulation module also updates the position of dynamic obstacles in the environment, taking into account their own motion laws, such as updating the position according to speed and acceleration. The importance of providing the next state:

[0132] The next state s' together with the current state s, the executed action a and the obtained reward r constitute the experience data unit (s, a, r, s', done), which is stored in the experience replay module. These experience data are an important basis for training the intelligent agent, helping the intelligent agent learn the relationship between state-action-reward, so as to continuously optimize the path planning strategy.

[0133] 4) Task completion judgment

[0134] Judgment basis:

[0135] The environment simulation module determines whether the task is completed based on whether the robot reaches the target position. For example, if the target position is (x t ,y t ), when the robot's position (x, y) meets certain distance conditions (such as Euclidean distance When the value is less than a certain threshold, the task is considered completed and the done flag is set to True. Impact on the training process:

[0136] The done flag is also part of the experience unit. When done = True, the experience unit may be handled differently during training. For example, when calculating cumulative rewards, the reward after the task is completed may be treated specially, or the priority of this experience unit in the experience replay module may be adjusted based on whether the task is successfully completed, thus affecting the training process of the agent.

[0137] Based on the above content, the training implementation of the path planning model specifically includes the following steps:

[0138] Step 1: Build a dynamic simulation environment

[0139] 1) Set the robot coordinates:

[0140] Accurately set the starting coordinates S of the mobile robot start =np.array([x s ,y s]) and the target coordinates S start =np.array([x t ,y t ]) to determine the starting and ending points for path planning.

[0141] 2) Define obstacle properties:

[0142] Define the property list of dynamic obstacles in detail. For each obstacle, let its shape be shape, its size be size, and its initial position be O. init =np.array([x o0 ,y o0 ]), the speed is The acceleration is These parameters can be used to predict the position of the obstacle at time t.

[0143] 3) Processing sensor data:

[0144] The sampling period of GPS sensor is set to T GPS =1ms, the touch sensor and distance sensor collect data synchronously.

[0145] Filter and sort the distance sensor data. Suppose the distance sensor set is {S d1 ,S d2 , ...}, filter out sensor data whose names contain 'so' and numeric characters. Suppose the filtered data is And arrange them in ascending order according to the numerical part. Normalize the collected sensor data. For the distance sensor data d, the normalization formula is: where d min and d max are the minimum and maximum values ​​of the sensor range respectively. The normalized data is integrated into the observation vector s.

[0146] 4) Define action and observation space:

[0147] The action space is defined as A = spaces.Discrete(3), which corresponds to three discrete actions: forward (a = 0), left turn (a = 1), and right turn (a = 2).

[0148] The shape of the observation space is set to ([X]), and the observation vector s is integrated into the target distance Sensor data, current position (x s ,y s ) and key elements such as obstacle avoidance related information provided by the BOAE mechanism.

[0149] Step 2: Train the path planning model

[0150] 1) Constructing a neural network module:

[0151] The strategy network policy_net and the target network target_net are constructed, both relying on the PyTorch framework to build a multi-layer fully connected layer structure.

[0152] Assume the input dimension is I, the first layer of the policy network fc1:h1=ReLU(W1s+b1), where W is the weight matrix and b1 is the bias vector, mapping the input dimension I to 256 dimensions; the middle layer fc2:h2=ReLU(W2h1+b2), maintaining 256-dimensional processing; the last layer fc3:Q(s,a)=W3h2+b3, the output dimension matches the action space dimension, and generates the Q value corresponding to each discrete action.

[0153] The target network initially replicates the policy network parameters, and then adjusts the parameters according to specific update rules.

[0154] 2) Set up the experience playback module:

[0155] The priority experience playback mechanism based on the SumTree data structure is used. Let the experience data unit be (s, a, r, s', done), where s is the current state, a is the action to be performed, r is the reward, s' is the next state, and done indicates whether the task is completed.

[0156] Sample the experience batch according to the set priority p, and generate the importance sampling weight ω based on the experience priority when sampling i ,like Where N is the number of samples, p i is the priority of the ith sample. At the same time, the priority of the stored experience can be dynamically updated according to the training process. For the experience related to dynamic obstacles and where the BOAE mechanism plays an important role in obstacle avoidance, its priority can be appropriately increased.

[0157] 3) Configure the intelligent agent control module:

[0158] Configure key training parameters, including discount factor γ, learning rate lr, exploration rate ε and its decay rate ε decay , minimum exploration rate ε min , target network update frequency target_update, memory capacity memory_size for experience replay, etc. Introduce step size parameter n, soft update parameter τ and regularization coefficient λ.

[0159] 4) Adopting n-step reward calculation mechanism:

[0160] The current single-step reward calculation is R1=r, and under the n-step reward calculation mechanism, the n-step reward R n for where r t+iis the reward obtained from the i-th step starting from the current time t. In this way, the robot's action strategy is evaluated more comprehensively.

[0161] 5) Implement a soft update and regularization combination strategy:

[0162] In the target network update phase, a soft update strategy is adopted, and the target network parameters θ target The update formula is θ target ←(1-τ)θ target +τθ policy , where θ policy It is the policy network parameter, which avoids drastic fluctuations in the target network parameter update process.

[0163] Add a regularization term to the loss function calculation, set the loss function to L, and the original loss function to L original , the regularization term is λ||θ|| 2 (θ is the network parameter), then the adjusted loss function L = L original +λ||θ|| 2 , preventing the network from overfitting, so that the network can not only adapt to the training data during the learning process, but also maintain the ability to generalize to unknown data.

[0164] 6) Integration of BOAE mechanism strategy:

[0165] During the training phase, the agent selects actions to drive the robot to act according to the exploration rate rule combined with the obstacle avoidance recommended actions or action adjustment information of the BOAE mechanism. The BOAE mechanism analyzes obstacle avoidance reactions from the state and action, and uses cross-attention and duel networks and generates auxiliary reward factors to enhance the robot's obstacle avoidance ability in multiple dimensions. For example, through the cross-attention mechanism, the attention score between the observation vector s and the action preference information is calculated, and then the value function is decomposed into the state value function V(s) and the advantage function A(s,a) through the duel network, that is, Q(s,a)=V(s)+A(s,a), and the advantage function is adjusted so that its mean is 0, that is, The final Q value is Q(s,a)=V(s)+A'(s,a). At the same time, the auxiliary reward factor r is generated aux , such as combining the distance reward r d (When the distance d between the robot and the obstacle is greater than the safety distance d safe When d =k1(dd safe ); when d≤d safe When d =-k2(d safe -d)) and direction bonus r θ (such as r θ = k3cosθ, θ is the angle between the robot's current direction and the target direction) and so on to get raux =r d +r θ +r other , combined with the reward r from the environment feedback for training.

[0166] Step 3: Get the path planning results

[0167] 1) Run the robot:

[0168] Load the trained mobile robot path planning model, and the agent makes the optimal action decision based on the output of the policy network Drive robots to operate in dynamic environments.

[0169] 2) Feedback and Records:

[0170] The environmental simulation module provides real-time feedback on the robot's position and status, and the visualization module dynamically records the robot's path.

[0171] Performance evaluation and preservation:

[0172] After the test, the visualization module displays the robot's optimal path and statistical performance indicators, such as path length. ((x i ,y i ) is a point on the path), and it takes time t to reach the target total , number of collisions n collision , minimum safe distance d from dynamic obstacles min-safe And other information, and integrate these indicators in the form of charts and save them locally, providing detailed data support for in-depth analysis of algorithm performance and subsequent improvements.

[0173] Based on the above, the update of policy network and target network parameters can be summarized as follows:

[0174] Calculate n-step rewards: Calculate the robot's cumulative reward R in n consecutive steps according to the n-step reward calculation mechanism t .

[0175] Calculate time difference error: Calculate time difference error Where Q θ (s t ,a t ) is the policy network in state s t Next, perform action a t Q value, Q θ' (s t+n ,a') is the target network in state s t+n The Q value below.

[0176] Update network parameters: According to the time difference error δ, the strategy network parameters θ are updated using methods such as gradient descent, and the target network parameters θ' are updated according to the soft update strategy θ'←(1-τ)θ'+τθ.

[0177] Evaluation of obstacle avoidance effect by BOAE mechanism: BOAE mechanism analyzes obstacle avoidance response from state and action, and evaluates obstacle avoidance effect in multiple dimensions such as cross-attention and duel network and generation of auxiliary reward factors. For example, if the robot successfully avoids dynamic obstacles according to the recommended actions of BOAE mechanism and maintains a safe distance from the obstacles, it can be considered that the obstacle avoidance effect is good, and the priority of experience data related to the obstacle avoidance experience in the experience playback module is increased accordingly; if the robot fails to avoid obstacles effectively or is too close to obstacles, the priority of the experience is lowered so that more high-priority obstacle avoidance experiences can be sampled in subsequent training to accelerate the convergence of the policy network.

[0178] S5: Load the trained mobile robot path planning model. The intelligent agent drives the robot to run in a dynamic environment based on the optimal action decision output by the policy network. The environmental simulation module feeds back the robot's position and status in real time. The visualization module dynamically records the robot's path. After the test, the visualization module displays the robot's optimal path and counts the operating performance indicators, such as path length, time to reach the target, number of collisions, and minimum safe distance from dynamic obstacles. These indicators are integrated and saved locally in the form of charts to provide detailed data support for in-depth analysis of algorithm performance and subsequent improvements. Through the display and data recording functions of the visualization module, not only can we intuitively understand the path planning effect of the robot in a specific environment, but we can also deeply analyze the advantages and disadvantages of the algorithm based on statistical data, provide a strong basis for further optimization of the algorithm, and promote the continuous development and improvement of mobile robot path planning technology.

[0179] In order to verify the effectiveness and practical effect of the method of the present invention, this embodiment conducts a simulation comparison experiment between the method of the present invention and the dynamic path planning method of the existing DDQN algorithm, as follows:

[0180] This example applies n-step rewards combined with priority experience replay to the DDQN algorithm (DDQN-NstepPRE) and conducts simulation experiments for comparison. The experimental results are as follows:

[0181] In the same scenario, DDQN with basically the same parameters and network structure, DDQN with added priority replay, and the DDQN-NstepPER algorithm in this paper are used for path planning. The training reward value curve of the algorithm is as follows: Figure 4As shown, the horizontal axis is the number of training rounds, and the vertical axis is the average reward value accumulated every 200 rounds. The reward curves of the three algorithms eventually tend to be similar, but there are slight differences in convergence speed and reward value. It can be seen that the DDQN-NstepPER algorithm finally obtains a higher average reward and converges faster than the DDQN algorithm with increased priority replay. After 4000 rounds of training, the average reward fluctuation trend of the DDQN-NstepPER algorithm gradually decreases, while DDQN and the DDQN algorithm with increased priority replay still have a larger fluctuation trend. Moreover, after 6000 rounds of training, the average reward value is always higher than that of DDQN and the DDQN algorithm with increased priority replay under the same conditions. This shows that in terms of path planning, the DDQN-NstepPER algorithm has better performance.

[0182] And by comparing the path length and path smoothness planned by the Pioneer 3-DX robot from the starting point to the target point, such as Figure 5 and Figure 6 As shown. It can be seen that although the paths planned by the three algorithms have certain similarities, the path planned by the DDQN-NstepPER algorithm shows better smoothness and coherence, and the planned path is shorter. Therefore, it can also be concluded that the performance of the DDQN-NstepPER algorithm provided by the present invention is better than the other two algorithms.

[0183] According to the above experiments, the advantages of the method of the present invention can be summarized as follows:

[0184] (1) Introduction of multi-step return priority experience replay: This paper combines multi-step returns (N-step Returns) with priority experience replay (PER) to propose the DDQN-NstepPER algorithm. In traditional Q learning, the calculation method of single-step returns often cannot effectively capture long-term rewards, while multi-step returns effectively reduce the variance caused by single-step returns by aggregating reward information from multiple time steps. After combining with priority experience replay, it is possible to give priority to those experiences that have made a high contribution to strategy improvement, thereby accelerating the learning process and improving training efficiency.

[0185] (2) Combination with DDQN: By combining the DDQN algorithm, the present invention effectively solves the problem of overestimation of Q values ​​and further improves the stability of training. DDDQN estimates Q values ​​using the behavior network and the target network respectively, avoiding the overestimation problem caused by the traditional DQN because the same network is used for both action selection and Q value estimation. This improvement enables the algorithm to update strategies more stably and accurately when dealing with problems such as long-term dependencies and delayed rewards.

[0186] (3) Efficient experience replay mechanism: In the design of multi-step reward priority experience replay, the SumTree data structure is used to efficiently implement the storage and sampling of priority experience. Through this structure, the sampling and updating of experience replay can be completed within the time complexity of O(logN), greatly improving the efficiency of the training process.

[0187] Through these improvements, the DDQN-NstepPER algorithm proposed in this paper not only improves the training efficiency of the reinforcement learning model in complex environments, but also significantly improves the model's perception of long-term rewards through the combination of priority replay and multi-step rewards, enabling the intelligent agent to better cope with complex decision-making tasks. It provides an effective method and tool for the behavior learning of mobile robots in complex environments.

Claims

1. A mobile robot dynamic path planning method based on an improved DDQN algorithm, characterized in that: The steps include: S1: Build a dynamic simulation environment; S2: Build a path planning model, including a neural network module, an experience playback module, and an agent control module; S3: Initialize the dynamic simulation environment and path planning model; S4: Train the path planning model based on the fusion strategy of improved DDQN and BOAE mechanism; S5: Load the trained mobile robot path planning model. The intelligent agent drives the robot to operate in a dynamic environment based on the optimal action decision output by the policy network. The environmental simulation module provides real-time feedback on the robot's position and status. The visualization module dynamically records the robot's path. After the test, the visualization module displays the robot's optimal path.

2. According to claim 1, a mobile robot dynamic path planning method based on an improved DDQN algorithm is characterized in that: The construction of the dynamic simulation environment in step S1 includes: Accurately set the starting and target coordinates of the mobile robot; For dynamic obstacles, define their attribute list in detail, covering geometric information and dynamic motion rules; For various sensors, the sampling period of GPS sensor is set to 1ms, and the touch sensor and distance sensor collect data synchronously based on this; distance sensors are obtained through strict screening and sorting, and the screening rules are that the name contains 'so' and contains numeric characters, and is sorted in ascending order according to the numeric part; the collected sensor data is normalized and integrated into the observation vector, and the data range is normalized to [0,1] according to the maximum and minimum values ​​of each sensor's own range to meet the neural network input requirements; The action space is defined as spaces.Discrete(3), corresponding to the three discrete actions of moving forward, turning left, and turning right. The shape of the observation space is set to ([X]), which fully integrates the target distance, sensor data, current position, and obstacle avoidance related information provided by the BOAE mechanism, where [X] is the dimension of the observation space after considering the BOAE information, which is determined according to the specific output of BOAE.

3. According to claim 1, a mobile robot dynamic path planning method based on an improved DDQN algorithm is characterized in that: The neural network module in step S2 includes a constructed policy network and a target network, and a multi-layer fully connected layer structure is built based on the PyTorch framework; in the multi-layer fully connected layer structure, the first layer fc1 maps the input dimension to 256 dimensions, the middle layer fc2 maintains 256-dimensional processing, and the output dimension of the last layer fc3 matches the action space dimension; ReLU is used as the activation function between each layer, and the policy network generates the Q value corresponding to each discrete action based on the observation vector that integrates the BOAE information, providing an accurate basis for the robot's action decision; the target network is used to assist the policy network training, initially copying the policy network parameters, and subsequently adjusting the parameters according to specific update rules.

4. According to claim 1, a mobile robot dynamic path planning method based on an improved DDQN algorithm is characterized in that: In step S2, the policy network generates the Q value corresponding to each discrete action according to the observation vector fused with BOAE information. The calculation process of the Q value includes: 1) Input processing First, the observation vector x fused with BOAE information is taken as input; assuming that the observation vector x = {x1, x2, …, x n }, where n is the dimension of the observation vector; 2) First fully connected layer calculation The observation vector x is input to the first fully connected layer fc1. Each neuron j in fc1 performs a weighted summation of the input and adds a bias term b. 1j ,Right now: where w 1ij is the connection weight from the ith input to the jth neuron in fc1; Then apply the ReLU activation function: 1j =ReLU(z 1j )=max(0,z 1j ), and get the output of fc1: a1 = [a 11 ,a 12 ,…,a 1m ], where m = 256, which is the output dimension of fc1; 3) Middle layer fully connected layer calculation For each neuron k in fc2, w 2jk is the connection weight from the jth input to the kth neuron in fc2, b 2k is the bias term; Apply the ReLU activation function again: a 2k =ReLU(z 2k )=max(0,z 2k ), and get the output of fc2: a2 = [a 21 ,a 22 ,…,a 2m ] 4) Calculation of the last fully connected layer The calculation of the last layer fc3, for each output action corresponding to the dimension l, w 3kl is the connection weight from the kth input to the lth output in fc3, b 3l is the bias term; At this time, z 3l It is the Q value corresponding to each discrete action, that is, Q l =z 3l , the final Q value vector Q=[Q1,Q2,...,Q s ], s is the dimension of the action space, which represents the value of each discrete action taken by the robot under the current observation.

5. The method for dynamic path planning of a mobile robot based on an improved DDQN algorithm according to claim 1, characterized in that: The experience replay module in step S2 uses a priority experience replay mechanism based on the SumTree data structure to store rich experience data generated by the interaction between the robot and the dynamic environment; each experience data unit covers the current state, execution action, reward, next state, and whether the task is completed.

6. The method for dynamic path planning of a mobile robot based on an improved DDQN algorithm according to claim 1, characterized in that: In step S2, the intelligent agent control module cooperates with other modules to reasonably configure key training parameters, including discount factor, learning rate, exploration rate and its decay rate, minimum exploration rate, target network update frequency, memory capacity of experience replay, and introduces n-step parameter, soft update parameter τ and regularization coefficient λ to implement specific improvement strategies.

7. A mobile robot dynamic path planning method based on an improved DDQN algorithm according to claim 6, characterized in that: In step S4, the path planning model is trained by combining the n-step reward calculation mechanism, the soft update and regularization combination strategy and the BOAE mechanism fusion strategy, wherein: n-step reward calculation mechanism: Soft update and regularization combination strategy: The soft update strategy introduces a soft update parameter τ, which allows the target network parameters to gradually and steadily approach the policy network parameters, avoiding drastic fluctuations during the parameter update process. At the same time, a regularization term is added to the loss function calculation, and the loss function is adjusted by the regularization coefficient λ. BOAE mechanism fusion strategy: The BOAE mechanism analyzes obstacle avoidance reactions from the state and action, and enhances the robot's obstacle avoidance ability in multiple dimensions by using cross-attention and duel networks and generating auxiliary reward factors.

8. The method for dynamic path planning of a mobile robot based on an improved DDQN algorithm according to claim 7, characterized in that: The training of the path planning model in step S4 includes the following process: The agent carefully selects actions based on the current observation vector, exploration rate, and obstacle avoidance recommended actions or action adjustment information provided by the BOAE mechanism. After the robot performs the selected action in a dynamic environment, the environment simulation module provides real-time feedback on rewards, next state, and whether it ends. The agent stores the interaction experience in the experience playback module, which samples experience according to priority and provides data batches for training. During the training process, the n-step reward and time difference error are calculated, and the policy network and target network parameters are updated using the soft update strategy and regularization term. At the same time, the experience priority is adjusted according to the evaluation of the obstacle avoidance effect by the BOAE mechanism.

9. The method for dynamic path planning of a mobile robot based on an improved DDQN algorithm according to claim 8, characterized in that: In the path planning model training of step S4: Calculate n-step rewards: Calculate the robot's cumulative reward R in n consecutive steps according to the n-step reward calculation mechanism t ; Calculate time difference error: Calculate time difference error Where Q θ (s t ,a t ) is the policy network in state s t Next, perform action a t Q value, Q θ' (s t+n ,a') is the target network in state s t+n The Q value under Update network parameters: According to the time difference error δ, the strategy network parameters θ are updated using methods such as gradient descent, and the target network parameters θ' are updated according to the soft update strategy θ'←(1-τ)θ'+τθ.

Citation Information

Patent Citations

  • Unmanned aerial vehicle path planning method based on transfer learning strategy deep Q-network

    CN110703766A

  • Method, device and equipment for planning flight path of unmanned aerial vehicle in dense city

    CN118605558A

  • Autonomous robot path planning method based on deep reinforcement learning

    CN119105512A

  • Method of Route Construction of UAV Network, UAV and Storage Medium thereof

    US20200359297A1

Cited By

  • DDQN unmanned aerial vehicle interruption scene path planning method based on attention mechanism

    CN120576777A

  • Path planning method for drone interruption scenarios based on DDQN with attention mechanism

    CN120576777B

  • VLA control method and system fusing world model and uncertainty quantitative decision

    CN121245857A

  • DDDPG-based autonomous path planning and obstacle avoidance multi-target continuous control method

    CN121541679A