A Deep Reinforcement Learning Robot Navigation Method Based on Attention Mechanism
By adopting a deep reinforcement learning method based on attention mechanism in mobile robots, integrating camera and lidar data, autonomous navigation in unknown environments is achieved, the shortcomings of traditional navigation technology in unknown environments are solved, and the adaptability and efficiency of navigation are improved.
Patent Information
- Application Number
- CN202211397557.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-09
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2042-11-09
AI Technical Summary
Existing mobile robot navigation technology relies on map information and high-precision sensors, making it difficult to achieve effective navigation in unknown environments, and does not have cross-environment adaptability.
A deep reinforcement learning method based on attention mechanism is adopted, and a SAC algorithm Actor network and a critic network are constructed by fusing camera image data and single-line lidar data to realize the robot's independent learning and navigation in the absence of map information.
It realizes safe and fast navigation in unfamiliar and complex environments, reduces dependence on map information and high-precision sensors, and improves the adaptability and generalization capabilities of navigation strategies.
Smart Images

Figure CN115585813B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of mobile robots, and relates to a deep reinforcement learning robot navigation method based on an attention mechanism. Background Art
[0002] Mobile robots play an increasingly important role in industrial fields, space exploration fields such as lunar exploration, military fields such as mine-sweeping robots, medical fields such as body temperature detection, logistics fields, service fields, etc. And navigation technology is the most important part in the research of mobile robots. An excellent navigation strategy can greatly improve the performance of mobile robots.
[0003] Currently, the adopted navigation strategies are mostly based on map information. This method can usually obtain the optimal or sub-optimal path, but the disadvantages are also obvious, mainly in the following aspects:
[0004] 1. Traditional navigation methods need to draw an environmental map, while service robots, driverless vehicles, planetary rovers, etc. often need to work in unknown environments, and map drawing is difficult.
[0005] 2. The navigation effect of traditional navigation systems is highly dependent on the accuracy of sensors, and the more accurate the sensor, the more expensive it is, which also increases the cost of the navigation system.
[0006] 3. Traditional navigation systems do not have the ability of generalization, that is, when the robot switches from one working environment to another, it needs to redraw the map.
[0007] Therefore, there is an urgent need for a navigation method that can achieve navigation in strange and complex environments when the mobile robot lacks map information. Summary of the Invention
[0008] To solve the above technical problems, the present invention provides a deep reinforcement learning robot navigation method based on an attention mechanism, which can achieve safe and fast navigation of a mobile robot in a strange and complex environment when lacking map information.
[0009] The present invention provides a deep reinforcement learning robot navigation method based on an attention mechanism, including the following steps:
[0010] Step 1: Collect image data of the robot's on-vehicle camera, on-vehicle single-line lidar data, and the linear velocity and angular velocity of the robot;
[0011] Step 2: Fuse the image data and the single-line lidar data to obtain fused image information with a large field of view and depth values, as the state parameter s t of the robot, and use the linear velocity and angular velocity as the action parameter a t ;
[0012] Step 3: Construct an experience pool for the SAC algorithm to store the combined data of action-state-reward R(s t , a t , s t+1 , r t+1 );
[0013] Step 4: Construct an Actor network for the SAC algorithm with spatial and temporal attention mechanisms, input the initial state parameters and initial action parameters into the Actor network to obtain action space probability parameters, and sample the action parameters according to the probability;
[0014] Step 5: The robot moves according to the action parameters, and fuses the newly acquired image data and single-line lidar data to obtain the current state parameters;
[0015] Step 6: Construct a reward function model, calculate the reward value according to the current state parameters of the robot, form a set of action-state-reward combined data with the current action parameters, state parameters and reward value, and put them into the experience pool. Repeat the above process until the number of combined data in the experience pool reaches the requirement;
[0016] Step 7: Construct a Q critic network and a V critic network for the SAC algorithm. The output of the Q critic network is the predicted value of the action-state value, and the output of the V critic network is the predicted value of the state value;
[0017] Step 8: Sample multiple groups of combined data from the experience pool and input them into the Actor network, Q critic network and V critic network to train each network until the network converges;
[0018] Step 9: Deploy the converged network to the robot system. The robot collects data through the camera and single-line lidar sensor as input, and outputs the linear velocity and angular velocity of the robot to achieve mapless navigation of the robot.
[0019] A deep reinforcement learning robot navigation method based on attention mechanism of the present invention has at least the following beneficial effects:
[0020] (1) By learning end-to-end, no map information is required and it does not depend on the accuracy of sensors, reducing system costs.
[0021] (2) The navigation strategy implemented by the deep reinforcement learning algorithm has strong adaptability to unfamiliar and complex environments and can be directly deployed from one environment to another completely unfamiliar and complex environment. Description of the Drawings
[0022] Figure 1It is a flowchart of a deep reinforcement learning robot navigation method based on the attention mechanism of the present invention;
[0023] Figure 2 It is the network structure of the Actor network in the present invention. Detailed implementation manner
[0024] The present invention provides a deep reinforcement learning robot navigation method based on the attention mechanism. Through this method, a mobile robot can achieve navigation in a strange and complex environment under the prerequisite of missing map information. Compared with the traditional robot navigation method based on map information, this method uses sparse lidar sensor data, camera image data, robot linear velocity, and angular velocity as input parameters, and adopts the Attention-SAC algorithm to realize the autonomous learning of the optimal or sub-optimal navigation strategy by the robot.
[0025] As Figure 1 shown, a deep reinforcement learning robot navigation method based on the attention mechanism of the present invention includes the following steps:
[0026] Step 1: Collect the image data of the robot's on-vehicle camera, the on-vehicle single-line lidar data, and the linear velocity and angular velocity of the robot;
[0027] Specifically in implementation, the single-line lidar emits laser light and provides points in the surrounding environment based on the light reflected and returned around. Selecting a single-line lidar can provide a 360° horizontal field of view (FOV) and a limited vertical field of view of approximately 15°. The main advantage of the single-line lidar is that it can provide highly accurate depth values. However, their output is sparse, that is, they do not provide a very high-resolution output.
[0028] Cameras are very common in life and are 2D sensors. The advantage is that they can give a clear and optionally high-resolution image, the disadvantage is that the field of view is limited, focused on a limited field of view, and there is no depth value information.
[0029] Step 2: Fuse the image data and the single-line lidar data to obtain fused image information with a large field of view and depth values, as the state parameter s of the robot t , and use the linear velocity and angular velocity as the action parameter a of the robot t ;
[0030] The fusion of the image data and the single-line lidar data is specifically as follows:
[0031] Step 2.1: Obtain the lidar extrinsic data and the camera extrinsic data, and publish the lidar extrinsic data and the camera extrinsic data to the robot operating system;
[0032] Obtaining the structural parameters of the lidar on the overall robot from the hardware structure, which is the coordinate information relative to the overall reference system of the robot, and publishing the pose of the lidar to the coordinate transformation tool TF when the appropriate launch file is executed. The same principle applies to obtaining the external parameters of the camera.
[0033] Then, through the coordinate transformation of the static_transfrom_publisher in the robot operating system, the coordinates of the lidar and the camera are published to TF. Transform is a tool in the robot operating system for managing 3D coordinate system transformations. As long as you tell TF the coordinate transformation information of two related coordinate systems, TF will help you continuously record the coordinate transformation of these two coordinate systems, even if the two coordinate systems are in motion.
[0034] Step 2.2: Calibrate the internal parameters of the camera to obtain the internal parameter data of the camera, and project the 3D coordinate points in the camera coordinate system onto the pixel plane of the camera without distortion. The specific steps of Step 2.2 are as follows:
[0035] Step 2.2.1: Print a checkerboard and paste it on a plane as a calibration object;
[0036] Step 2.2.2: Take some photos of the calibration object from different directions by adjusting the direction of the calibration object or the camera;
[0037] Step 2.2.3: Extract the checkerboard corner points from the photos;
[0038] Step 2.2.4: Estimate 5 internal parameter data and 6 external parameter data in the ideal non-distorted case;
[0039] Step 2.2.5: According to the 5 internal parameter data and 6 external parameter data, use the least squares method to calculate the distortion coefficients under the actual existing radial distortion;
[0040] Step 2.2.6: According to the distortion coefficients under the radial distortion, use the maximum likelihood method for optimization estimation to obtain the internal parameter matrix, radial distortion, and tangential distortion of the camera.
[0041] Step 2.3: Jointly calibrate the lidar and the camera to establish the correspondence between the coordinate points of the 3D point cloud data in the lidar coordinate system and the 3D points in the camera coordinate system. The specific steps of Step 2.3 are as follows:
[0042] Considering that this is a single-line lidar and it is not convenient to use the existing joint calibration toolkit for joint calibration, we can only work on the measurement and calculation accuracy of the single-line lidar and the camera to ensure that the external parameters published to TF in the first and second steps are as accurate and reliable as possible. When specifically using it, only by using the TF function of the robot operating system can we perform the coordinate transformation of the 3D point cloud data in the single-line lidar coordinate system and the 3D points in the camera coordinate system. The specific steps are as follows:
[0043] Determine the conversion relationship between the camera coordinate system and the robot coordinate system, and the conversion relationship between the lidar coordinate system and the robot coordinate system through TF transformation according to the installation positions of the camera and the lidar on the robot.
[0044] Step 2.4: Obtain the camera internal parameter data and align the time lines of the single-line lidar data and the image data.
[0045] When specifically implementing, since the information fusion of the single-line lidar and the camera depends on the camera internal parameters, after the program starts, it is necessary to wait for the subscribed camera internal parameter data to be obtained. Also, because the publishing frequencies of the single-line lidar topic message and the camera image topic message are inconsistent, it is necessary to compare the timestamps of the two types of topic messages to ensure that the lidar scan data and the camera image data both reflect the state at the same moment.
[0046] When specifically implementing, the data type of the single-line lidar is actually in polar coordinate form, which is different from the point cloud data type of the multi-line lidar. It is necessary to convert the data from polar coordinate representation to point cloud data through steps 2.5 and 2.6.
[0047] Step 2.5: Convert the single-line lidar data from polar coordinate representation to Cartesian coordinate representation, generate the initial point cloud data in the Cartesian coordinate system, perform horizontal interpolation on the initial point cloud data to ensure the point cloud density, and generate new point cloud data.
[0048] Step 2.6: Expand each row of the new point cloud data in the Cartesian coordinate system vertically, expand downward to the intersection of the X-axis plane and the Y-axis plane, and expand upward by j values to generate the final point cloud data in the Cartesian coordinate system.
[0049] Step 2.7: Use TF to transform the final point cloud data into the camera coordinate system to generate the radar point cloud data in the camera coordinate system. At the same time, use the obtained camera internal parameter data to map the radar point cloud data in the camera coordinate system to the pixel plane coordinate system of the two-dimensional image data collected by the camera to obtain the fusion data, and obtain the corresponding color RGB values from the pixel plane.
[0050] Step 2.8: According to the conversion relationship between the camera coordinate system and the robot coordinate system, the fused data is converted into the robot coordinate system to obtain fused image information with a large field of view and depth values.
[0051] Step 3: Construct an experience pool for the SAC algorithm to store the combined data R(s t , a t , s t+1 , r t+1 );
[0052] Step 4: Construct an Actor network for the SAC algorithm with spatial and temporal attention mechanisms. Input the initial state parameters and initial action parameters into the Actor network to obtain action space probability parameters, and sample the action parameters according to the probability;
[0053] Specifically, when implemented, the Actor network introduces a spatial attention mechanism, including the following steps:
[0054] Step 4.1: A set of m×n original feature maps are obtained through the convolutional layer neural network from the obtained fused image information, and these feature maps are regarded as a set of region vectors of length d Expressed by the following formula:
[0055]
[0056] Step 4.2: The feature maps are respectively passed through max pooling and average pooling to form two [1, H, W] weight vectors. The number of channels changes from [C, H, W] to [1, H, W], and pooling is performed on all channels of the same feature point;
[0057] Step 4.3: Stack the two obtained weight vectors to form a [2, H, W] feature space weight vector;
[0058] Step 4.4: The feature space weight vector passes through a convolutional layer and a sigmoid activation function, and the feature dimension changes from [2, H, W] to [1, H, W]. This [1, H, W] weight vector represents the importance of each point on the feature map, and the larger the value, the more important;
[0059] Step 4.5: Add a long short-term memory (LSTM) neural network module at the output end of the spatial attention mechanism. Softmax is performed on the hidden vector generated by the LSTM and the output of the activation function to obtain the spatial weight
[0060]
[0061] In the above formula The weight vector of [1, H, W] output by the sigmoid activation function in step 4.4, w g and w h are coefficients, h t-1 is the hidden vector generated by the LSTM recurrent neural network;
[0062] Step 4.6: Multiply the spatial weight of [1, H, W] by the original feature map of [C, H, W], that is, each point of [H, W] on the feature map is assigned a weight;
[0063]
[0064] In the formula, z t is the context vector of the spatial attention mechanism and is the weighted sum of all region vectors .
[0065] Specifically in implementation, the Actor network introduces a temporal attention mechanism specifically as follows:
[0066] Process the temporal information in the network through the LSTM module, not only taking the stacked historical prediction values as input; learn through the LSTM which frames are the most important for the current action selection in past predictions at different time nodes. The output w of each LSTM t is the inner product of the feature vector v t and the hidden vector h generated by the LSTM i , and then perform normalization processing through the softmax function:
[0067] w t = Softmax(v t ·h t )
[0068] Calculate the context vector c of the temporal attention mechanism through the output of each LSTM t :
[0069]
[0070] In the above formula, the output w of the LSTM t is regarded as the importance in a certain frame of the LSTM input. Therefore, by selecting the frames relatively important for the current selected action for learning, the training time and the navigation effect are optimized. Figure 2 is the network structure of the Actor network in the present invention.
[0071] Step 5: The robot moves according to the action parameters, and fuses the newly acquired image data and the single-line lidar data to obtain the current state parameters;
[0072] Step 6: Construct a reward function model, calculate the reward value according to the reward function model, and form a set of action-state-reward combination data with the current action parameters, state parameters and reward value, and put it into the experience pool. Repeat the above process until the number of combination data in the experience pool reaches the requirement;
[0073] Specifically, the established reward function model is as follows:
[0074]
[0075] In the above formula, r success represents the reward obtained by the robot when it reaches the destination without colliding with obstacles, which is a positive constant value, used to encourage the robot to choose behaviors that do not collide with obstacles and can reach the destination; r collisicon represents the reward obtained by the robot after colliding with an obstacle, which is a negative constant value, used to punish the action chosen by the robot that causes the collision and guide the robot to choose other actions; r process represents the reward obtained by the robot when it does not reach the destination and does not collide with the robot, r process is designed as follows:
[0076]
[0077] In the formula, c1, c2, c3 are hyperparameters, d t -d t-1 represents the distance difference between the current time step and the previous time step from the target position. If the difference is less than 0, it means the robot is moving away from the target point. At this time, this term is negative and a penalty is given; otherwise, a positive value is given for encouragement; θ t -θ t-1 represents the angle difference between the current time step and the previous time step of the robot from the target position. If the angle deviates, it is negative, otherwise it is positive; represents the change rate of the angular velocity of the robot, which is to prevent the angular velocity of the robot from changing greatly between adjacent time steps.
[0078] Step 7: Construct the Q critic network and V critic network of the SAC algorithm. The output of the Q critic network is the predicted value of the action-state value, and the output of the V critic network is the predicted value of the state value;
[0079] Step 8: Sample multiple groups of combination data from the experience pool and input them into the Actor network, Q critic network and V critic network to train each network until the network converges;
[0080] Specifically, the update process of the V critic network is as follows:
[0081] Select data (st , a t , s t+1 , r t+1 ) Update the V critic network. Calculate the theoretical value of the state value output by the V critic network through the following formula:
[0082]
[0083] In the formula, represents taking the expectation of the subsequent content; π(a t |s t ) is the policy, which is a mapping from state to action, defining all possible behaviors and probabilities of the robot in each state; a t ~π(·|s t ) means sampling a t from the policy π, and π(·|s t ) represents all actions and probabilities in state s t ; i is the number of Q critic networks; A t represents the action space composed of multiple action parameters; α is the reward coefficient of entropy.
[0084] Use the mean square error between the predicted value and the theoretical value of the state value output by the V critic network as the loss function to train the V critic network. The loss function of the V critic network is:
[0085]
[0086] In the formula, B represents the experience pool, that is, when calculating Loss1, it is necessary to take the average of the samples taken from the experience pool to reflect the quality of the samples taken in the average sense; v(s t ) is the predicted value of the state value output by the V critic network, is the theoretical value of the state value output by the V critic network.
[0087] Specifically, the update process of the Q critic network is as follows:
[0088] Select data (s t , a t , s t+1 , r t+1 ) from the experience pool to update the Q critic network. Based on the optimal Bellman equation, use as the theoretical value of the action-state value output by the Q critic network, and use the mean square error between the predicted value and the theoretical value of the action-state value output by the Q critic network as the loss function to train the Q critic network. The loss function of the Q critic network is:
[0089]
[0090] Wherein, q i (s t , a t ) is the predicted value of the action-state value output by the i-th Q critic network.
[0091] Specifically, the loss function of the Actor is:
[0092]
[0093]
[0094] Wherein, α is the reward coefficient of entropy, which determines the importance of the entropy lnπ(a t |s t ). The larger it is, the more important it is.
[0095] Step 9: Deploy the converged network into the robot system. The robot collects data through a camera and a single-line lidar sensor as input, and outputs the linear velocity and angular velocity of the robot to achieve mapless navigation of the robot.
[0096] The above are only the preferred embodiments of the present invention and are not intended to limit the idea of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A deep reinforcement learning robot navigation method based on an attention mechanism, characterized in that, It includes the following steps: Step 1: Collect the image data of the robot's on-vehicle camera, the on-vehicle single-line lidar data, as well as the linear velocity and angular velocity of the robot; Step 2: Fuse the image data and the single-line lidar data to obtain fused image information with a large field of view and depth values, which serves as the state parameter s of the robot t , and use the linear velocity and angular velocity as the action parameter a of the robot t ; Step 3: Construct an experience pool for the SAC algorithm to store the combined data of action-state-reward R(s t , a t , s t+1 , r t+1 ); Step 4: Construct an Actor network of the SAC algorithm with spatial attention mechanism and temporal attention mechanism, input the initial state parameters and initial action parameters into the Actor network to obtain action probability parameters, and sample the action parameters according to the probability; Step 5: The robot moves according to the action parameters, and fuses the newly obtained image data and single-line lidar data to obtain the current state parameters; Step 6: Construct a reward function model, calculate the reward value according to the reward function model, form a set of action-state-reward combination data with the current action parameters, state parameters and reward value, and put them into the experience pool. Repeat the above process until the number of combination data in the experience pool reaches the requirement; Step 7: Construct a Q critic network and a V critic network of the SAC algorithm. The output of the Q critic network is the predicted value of the action-state value, and the output of the V critic network is the predicted value of the state value; Step 8: Sample multiple groups of combination data from the experience pool and input them into the Actor network, Q critic network and V critic network to train each network until the network converges; Step 9: Deploy the converged network to the robot system. The robot collects data through the camera and lidar sensors as input, and outputs the linear velocity and angular velocity of the robot to achieve mapless navigation of the robot.
2. The method for deep reinforcement learning robot navigation based on the attention mechanism according to claim 1, wherein The specific fusion of the image data and the single-line lidar data in Step 2 is as follows: Step 2.1: Obtain the lidar extrinsic data and camera extrinsic data, and publish the lidar extrinsic data and camera extrinsic data to the robot operating system; Step 2.2: Calibrate the camera intrinsic parameters, obtain the camera intrinsic data, and project the 3D coordinate points in the camera coordinate system onto the pixel plane of the camera without distortion; Step 2.3: Jointly calibrate the lidar and the camera to establish the correspondence between the coordinate points of the 3D point cloud data in the lidar coordinate system and the 3D points in the camera coordinate system; Step 2.4: Obtain the camera intrinsic data, and perform time alignment of the single-line lidar data and the image data; Step 2.5: Convert the single-line lidar data from polar coordinate representation to Cartesian coordinate representation, generate the initial point cloud data in the Cartesian coordinate system, and perform horizontal interpolation on the initial point cloud data to ensure the point cloud density and generate new point cloud data; Step 2.6: Perform vertical expansion on each row of the new point cloud data in the Cartesian coordinate system, expand downward to the intersection of the X-axis plane and the Y-axis plane, and expand upward by j values to generate the final point cloud data in the Cartesian coordinate system; Step 2.7: Use TF to transform the final point cloud data into the camera coordinate system to generate the radar point cloud data in the camera coordinate system. At the same time, use the obtained camera intrinsic data to map the radar point cloud data in the camera coordinate system to the pixel plane coordinate system of the two-dimensional image data collected by the camera to obtain the fusion data, and obtain the corresponding color RGB value from the pixel plane; Step 2.8: According to the conversion relationship between the camera coordinate system and the robot coordinate system, the fused data is converted into the robot coordinate system to obtain the fused image information with a large field of view and depth values.
3. The method for deep reinforcement learning robot navigation based on the attention mechanism according to claim 2, characterized in that, The specific content of step 2.2 is as follows: Step 2.2.1: Print a checkerboard and paste it on a plane as a calibration object. Step 2.2.2: Take some photos of the calibration object in different directions by adjusting the direction of the calibration object or the camera. Step 2.2.3: Extract the checkerboard corner points from the photos. Step 2.2.4: Estimate 5 internal parameter data and 6 external parameter data in the case of ideal undistortion. Step 2.2.5: According to the 5 internal parameter data and 6 external parameter data, use the least squares method to calculate the distortion coefficients under actual radial distortion. Step 2.2.6: According to the distortion coefficients under radial distortion, use the maximum likelihood method for optimization estimation to obtain the internal parameter matrix, radial distortion, and tangential distortion of the camera.
4. The method for deep reinforcement learning robot navigation based on the attention mechanism according to claim 2, wherein, The specific content of step 2.3 is as follows: According to the installation positions of the camera and the lidar on the robot through TF transformation, determine the conversion relationship between the camera coordinate system and the robot coordinate system, and the conversion relationship between the lidar coordinate system and the robot coordinate system.
5. The method for deep reinforcement learning robot navigation based on the attention mechanism according to claim 1, wherein The Actor network in step 4 introduces a spatial attention mechanism, including the following steps: Step 4.1: A set of original feature maps of m×n is obtained through the convolutional layer neural network using the acquired fused image information, and these feature maps are regarded as a set of regional vectors of length d which is expressed by the following formula: Step 4.2: The feature map is respectively passed through max pooling and average pooling to form two weight vectors of [1, H, W], and the number of channels changes from [C, H, W] to [1, H, W], pooling all channels of the same feature point. Step 4.3: Stack the two obtained weight vectors to form a feature space weight vector of [2, H, W]. Step 4.4: The feature space weight vector passes through a convolutional layer and a sigmoid activation function, and the feature dimension changes from [2, H, W] to [1, H, W]. This weight vector of [1, H, W] characterizes the importance of each point on the feature map, and the larger the value, the more important. Step 4.5: Add a long short-term memory (LSTM) module at the output end of the spatial attention mechanism, perform softmax on the hidden vector generated by the LSTM and the output of the activation function to obtain the spatial weights In the above formula is the weight vector of [1, H, W] output by the sigmoid activation function in step 4.4, w g and w h are coefficients, h t-1 is the hidden vector generated by the LSTM recurrent neural network; Step 4.6: Multiply the spatial weight of [1, H, W] by the original feature map of [C, H, W], that is, each point of [H, W] on the feature map is assigned a weight; where z t is the context vector of the spatial attention mechanism and is the weighted sum of all region vectors .
6. The method for deep reinforcement learning robot navigation based on the attention mechanism according to claim 1, wherein, The specific content of the time attention mechanism introduced by the Actor network in step 4 is as follows: Process the temporal information in the network through the LSTM module, not just taking the accumulated historical prediction values as input; learn through the LSTM which frames are most important for the current action selection in past predictions at different time nodes, and the output w of each LSTM t is the feature vector v t and the hidden vector h generated by the LSTM i is the inner product of, and then is normalized through the softmax function: w t = Softmax(v t ·h t ) The context vector c of the temporal attention mechanism is calculated from the output of each LSTM t : The output w of the LSTM in the above formula t is regarded as the importance in a certain frame of the LSTM input. Therefore, by selecting the frames that are relatively important for the currently selected action for learning, the training time and the navigation effect can be optimized.
7. The method for robot navigation based on attention mechanism in claim 1, the update process of the V critic network in step 8 is as follows: Select data (s from the experience pool t , a t , s t+1 , r t+1 ) to update the V-critic network, and calculate the theoretical value of the state value output by the V-critic network through the following formula: wherein, denotes calculating the expectation of the subsequent content; π(a t |s t ) is the policy, which is a mapping from state to action, defining various possible behaviors and probabilities of the robot in each state; a t :π(·|s t ) means that a t is sampled in the policy π, and π(·|s t ) represents all actions and probabilities in the state s t ; i is the number of Q critic networks; A t represents the action space composed of multiple action parameters; α is the reward coefficient of entropy; Use the mean square error between the predicted value of the state value output by the V critic network and the theoretical value of the state value as the loss function to train the V critic network. The loss function of the V critic network is: In the formula, B represents the experience pool, that is, when calculating Loss1, the samples taken from the experience pool need to be averaged to reflect the quality of the samples taken in the average sense; v(s t ) is the predicted value of the state value output by the V critic network, and is the theoretical value of the state value output by the V critic network.
8. The method for deep reinforcement learning robot navigation based on the attention mechanism according to claim 7, wherein, The update process of the Q critic network in step 8 is: Select data (s from the experience pool t , a t , s t+1 , r t+1 ) to perform Q-critic network update. Based on the optimal Bellman equation, use as the theoretical value of the action-state value output by the Q-critic network. Use the mean square error between the predicted value of the action-state value output by the Q-critic network and the theoretical value of the action-state value as the loss function to train the Q-critic network. The loss function of the Q-critic network is: where q i (s t , a t ) is the predicted value of the action-state value output by the i-th Q critic network.
9. The method for deep reinforcement learning robot navigation based on the attention mechanism according to claim 8, characterized in that, The loss function of the Actor in step 8 is: where α is the reward coefficient of entropy, which determines the importance of the entropy lnπ(a t |s t ), the larger it is, the more important it is.
10. The method for deep reinforcement learning robot navigation based on the attention mechanism according to claim 1, characterized in that, The reward function model established in step 6 is as follows: In the above formula, r success represents the reward obtained by the robot to reach the destination without colliding with obstacles, which is a positive constant value used to encourage the robot to choose behaviors that do not collide with obstacles and can reach the destination; r collisicon represents the reward obtained by the robot after colliding with an obstacle, which is a negative constant value used to punish the action chosen by the robot that causes the collision and guide the robot to choose other actions; r process represents the reward obtained when the robot does not reach the destination and does not collide with the robot, r process is designed as follows: where c1, c2, c3 are hyperparameters, and d t -d t-1 represents the distance difference between the current time step and the previous time step from the target position. If the difference is less than 0, it means the robot is moving away from the target point. In this case, this term is negative and a penalty is given; otherwise, a positive value is given as encouragement. θ t -θ t-1 Indicates the angular difference between the robot and the target position at the current time step and the previous time step. If the angle deviates, it is negative; otherwise, it is positive. Represents the change rate of the robot's angular velocity, which is to prevent the angular velocity of the robot from changing significantly between adjacent time steps.