Lunar rover path planning method based on depth deterministic gradient strategy
Through a path planning method based on a depth deterministic gradient strategy, combined with a multi-dimensional reward function and an adaptive step size mechanism, the problem of insufficient local obstacle avoidance capabilities in complex terrain is solved, and efficient and safe path planning is achieved.
Patent Information
- Application Number
- CN202510466417.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-15
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2045-04-15
AI Technical Summary
The existing lunar rover path planning method has insufficient local obstacle avoidance capabilities under complex terrain and communication delay conditions, making it difficult to achieve efficient extraterrestrial detection tasks.
A path planning method based on depth deterministic gradient strategy is adopted to construct a kinematic model by introducing terrain slope angle and wheel subsidence; a multi-dimensional reward function is defined, combining distance, progress, direction and smoothness rewards; an adaptive step size mechanism is introduced to dynamically adjust the step size to adapt to different terrains.
The improvement of obstacle avoidance accuracy and autonomous exploration capabilities in complex lunar terrain has been achieved, local optimization is avoided, path continuity is improved, path planning efficiency and motion safety are enhanced.
Smart Images

Figure CN120010495A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of unmanned vehicle control and deep reinforcement learning technology, and in particular to a lunar rover path planning method based on a deep deterministic gradient strategy. Background Art
[0002] With the advancement of science and technology, especially the continuous development of space exploration technology, celestial exploration vehicles (such as lunar rovers, Mars rovers, etc.) are increasingly used in unknown environments. These exploration vehicles are used to perform various tasks, such as geological exploration, sampling, data collection, etc., so research on path planning is particularly important. Traditional ground robot path planning methods usually rely on pre-built environmental maps. However, in celestial exploration missions, due to the complexity, unknown and dynamic nature of the terrain, traditional path planning methods face many challenges. Rule-based algorithms (such as Dijkstra's Dijkstra method) need to rely on high-precision pre-built maps, and the complex terrain of the moon leads to large computational complexity in real-time map updates, and local obstacle avoidance algorithms (such as the dynamic window method DWA) are prone to fall into local optimality in continuous obstacle areas because they only optimize single-step paths. The state space design of existing deep reinforcement learning methods does not fully integrate multi-source sensor data (such as slope angle, roughness and terrain complexity), resulting in the inability of the policy network to accurately perceive the mechanical properties of the terrain and path safety assessment errors. The traditional deep deterministic gradient strategy (DDPG) algorithm's simplified reward function only optimizes distance or time, ignoring terrain features, vehicle stability, and motion smoothness, which makes it difficult for the strategy to converge during training and it is difficult to generate the optimal path. Summary of the invention
[0003] Purpose of the invention: In view of the problem that the existing lunar rover path planning methods have insufficient local obstacle avoidance capabilities under the complex lunar terrain and communication delay conditions, thus being unable to achieve efficient extraterrestrial exploration missions, the present invention aims to provide a lunar rover path planning method based on a deep deterministic gradient strategy.
[0004] Technical solution: The lunar rover path planning method based on deep deterministic gradient strategy described in the present invention comprises the following steps: (1) Introducing terrain slope angle , based on the Ackerman steering principle, determine the heading angle and steering angle The differential equations of the lunar rover are used to obtain the kinematic model of the lunar rover. (2) Based on the Baker pressure sinking theory, calculate the wheel sinking amount Z; according to the terrain slope angle and wheel sinking amount Z to calculate the maximum steering angle of the vehicle ; According to the maximum allowable steering angle of the vehicle and wheel sinking amount Z, calculate the minimum wheel turning radius ; Based on the steering angle, minimum turning radius, maximum speed and slope angle, the constraints of the lunar rover kinematic model based on the deep deterministic gradient strategy are constructed; (3) According to the lunar rover kinematic state vector and the lunar terrain feature vector , define the state space S; according to the velocity increment and heading angle increment , define the action space ; Based on distance reward, dynamic instantaneous progress reward, dynamic instantaneous direction reward and smoothness reward, define the multi-dimensional reward function r of path planning; (4) Constructing adaptive step size based on terrain roughness, slope gradient and terrain complexity ; Expand the state space to , to the action space Introducing Adaptive Step Size , get the basic action space , input the critic network and output the step evaluation value , used to update the actor Actor network parameters; Introducing Ornstein-Uhlenbeck OU noise, we get the final motion , perform the final action , if satisfied and , then go to step (5); otherwise, recalculate the final action; (5) Calculate the multi-dimensional reward function for path planning; (6) Calculate the priority weight based on the terrain risk function, multi-dimensional reward function and step loss , filter the experience samples into the experience pool; perform model training, and select the best path based on the path accessibility rate.
[0005] Furthermore, the kinematic model of the lunar rover is ; in, is the lunar rover position information, is the vehicle linear velocity, is the vehicle wheelbase.
[0006] Furthermore, the wheel sinking amount Z is ; in, is the vertical soil pressure, is the soil cohesion modulus, is the vehicle wheel width, is the soil friction modulus, is the subsidence index; Vehicle maximum steering angle for ; in, It indicates the maximum turning angle on flat ground when considering the subsidence and slope. is the vehicle radius; Minimum wheel turning radius for ; Maximum speed for ; in, is the friction coefficient, is the lunar gravity coefficient.
[0007] Furthermore, the state space S is ; Among them, the lunar rover kinematic state vector Including lunar rover location information , vehicle linear speed and heading angle , lunar terrain feature vector Include elevation standard deviation , Terrain slope angle and roughness R; Roughness R is ; Where N represents the total number of sampling points. Represents the elevation value of the i-th sampling point in the grid cell. Represents the average elevation of all sampling points in the grid cell. When the lunar rover encounters rough terrain, the lunar rover's speed will be limited. ; Action Space for ; in, and .
[0008] Furthermore, the multi-dimensional reward function of path planning for ; in, , , and is the weight, Distance Rewards, For dynamic instant progress rewards, is the dynamic instantaneous direction reward, Reward for smoothness.
[0009] Further, distance reward for ; Dynamic Instant Progress Rewards for ; Dynamic instantaneous direction reward for ; Smoothness Bonus for ; Weight , , and As shown below: ; ; in, Represents the Euclidean distance from the current position to the target, is the initial distance, is the current speed, is the target speed, is the remaining distance from the current position to the target point, is the total mission distance, is the distance from the previous state to the target point, is the basic direction reward weight, , is the adjustment coefficient, Indicates the angle between the current heading and the target direction. is the maximum tolerance threshold angle, is the heading angle of the target direction, is the heading angle at the current moment, is the basic smoothness reward weight, is the rate of change of acceleration between adjacent time steps, is the maximum allowable rate of change, is the heading angle at the current moment, is the maximum allowed heading angle.
[0010] Furthermore, adaptive step size for ; ; ; in, is the base step length, is the terrain roughness, which is determined by the elevation standard deviation Normalized to get; is the slope gradient, which is obtained by calculating the slope change rate of adjacent grids; It is the terrain complexity, which is a composite index based on the fusion of terrain roughness and terrain slope; , , is the dynamic weight coefficient; when When the value is not less than 0.7, i.e. when encountering high-density terrain, the step length is shortened to 30%~50% of the baseline value; when When it is less than 0.3, that is, when encountering low-density terrain, the step length is expanded to 120%~150% of the baseline value.
[0011] Furthermore, the basic action space for ; Step length evaluation value for ; in, is the weight coefficient, is the terrain feature embedding vector, is the dynamic state vector, is the bias term; Final Action for ; in, represents the mean reversion rate, ; represents the noise mean, is the maximum slope threshold, represents the noise value generated by the OU process, is the noise value at time t, is the noise disturbance term, is the terrain complexity threshold.
[0012] Furthermore, the priority weight for ; in, is the time difference error, is the step loss, represents the terrain risk function, is the state vector, Indicates i The absolute value of the multi-dimensional rewards of experience. is the step size decision error term coefficient, is the terrain risk weight coefficient, is the reward coefficient.
[0013] Furthermore, the time difference error for ; in, is the discount factor, is the target Q value of the target network for the next state and action, The predicted Q value of the critic network for the current state and action; Step loss for ; in, is the recommended value for the step size, is the actual execution step value; Terrain risk function for ; Among them, 1 represents high-risk experience and 0.2 represents low-risk experience.
[0014] Beneficial effects: Compared with the prior art, the present invention has the following significant advantages: 1. The present invention integrates four types of goals, namely, distance reward, instantaneous progress reward, instantaneous heading reward and motion smoothness reward, to form a gradient signal that can adapt to the complexity of the terrain, guide the strategy network to achieve a dynamic balance between exploration and safety, construct a dynamic multi-dimensional reward function mechanism, dynamically adjust the weight coefficient through terrain features, and achieve multi-target adaptive coordination; 2. The present invention proposes an adaptive step length mechanism based on multi-scale terrain features, integrates terrain features, and dynamically adjusts the step length in different risk terrains to improve the obstacle avoidance accuracy and autonomous exploration capability of the lunar rover in different terrains; 3. The present invention uses dynamic multi-dimensional reward function mechanism to dynamically adjust the weight coefficient through terrain features to achieve multi-target adaptive coordination; 4. The multi-level adaptive mechanism of dimensional reward function, lightweight network structure and adaptive step-size strategy avoids local optimality and improves path continuity, thereby completing the lunar rover's autonomous decision-making and detection mission, achieving a dual improvement in path planning efficiency and motion safety, and solving the core problems of traditional reinforcement learning in complex lunar environments, such as slow convergence, frequent path oscillations, poor terrain adaptability and low model deployment efficiency, and reducing model complexity; 4. The present invention integrates terrain feature perception with an improved reward function and a collaborative multi-level optimization framework of the network structure, and proposes a dynamically adaptive step-size control strategy based on terrain features, and dynamically adjusts the strategy through the terrain feature fusion amount. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] Figure 1 is a flow chart of the present invention; Figure 2 It is a structural schematic diagram of the adaptive step size strategy of the present invention. DETAILED DESCRIPTION
[0016] The present invention will be further described below in conjunction with the accompanying drawings.
[0017] The lunar rover path planning method based on the deep deterministic gradient strategy described in the present invention comprises the following steps: (1) Introducing terrain slope angle , based on the Ackerman steering principle, determine the heading angle and steering angle The differential equations of the lunar rover are used to obtain the kinematic model of the lunar rover.
[0018] Establish the heading angle-steering angle differential equation: ; in, is the vehicle heading angle, in rad, is the linear velocity of the vehicle, in m / s, is the front wheel steering angle, in rad, Indicates the vehicle wheelbase, Represents the terrain slope angle in rad.
[0019] The kinematic model of the lunar rover considering the influence of terrain slope angle is: ; in, is the lunar rover position information, is the vehicle linear velocity, is the vehicle wheelbase.
[0020] (2) Based on the Baker pressure sinking theory, calculate the wheel sinking amount Z; according to the terrain slope angle and wheel sinking amount Z to calculate the maximum steering angle of the vehicle ; According to the maximum allowable steering angle of the vehicle and wheel sinking amount Z, calculate the minimum wheel turning radius ,; Based on the steering angle, minimum turning radius, maximum speed and slope angle, the constraints of the lunar rover kinematic model based on the deep deterministic gradient strategy are constructed.
[0021] Based on Bekker pressure-sinking theory, the wheel sinking amount Z is: ; in, is the vertical soil pressure in kPa, is the soil cohesion modulus, in kPa / m n+1 , is the wheel width of the vehicle, in m, is the soil friction modulus, in kPa / m n+1 , is the subsidence index.
[0022] Vehicle maximum steering angle for ; in, It indicates the maximum turning angle on flat ground when considering the subsidence and slope. is the vehicle radius. Front wheel steering angle .
[0023] Minimum wheel turning radius for ; Maximum speed for ; in, is the friction coefficient, is the lunar gravity coefficient.
[0024] (3) Calculate terrain characteristics.
[0025] Import the elevation DEM data file information that needs to be extracted, and perform downsampling feature extraction through Gaussian pyramid. Apply a 5×5 Gaussian filter kernel (standard deviation σ=1.0 pixel) to the original DEM data, and eliminate high-frequency noise and small terrain undulations through weighted averaging. Downsample the filtered DEM with a step size of 2, that is, retain the center point of each 2×2 pixel area as the new resolution grid node. Repeat 3 downsampling iterations (k=3) to finally generate a hierarchical terrain model with gradually decreasing resolution.
[0026] According to the generated multi-scale terrain model (hierarchical terrain model), the terrain features of each grid are quantitatively calculated. The feature calculations of the high-resolution terrain model and the low-resolution terrain model are retained separately, and the feature resolution is dynamically selected according to the path planning requirements. The high-resolution grid feature calculation supports global path planning, and the low-resolution grid feature calculation ensures the obstacle avoidance task and dynamic feasibility.
[0027] Define the elevation standard deviation (intensity of terrain relief): ; in, Indicates the elevation value of the i-th sampling point in the grid cell (m). Indicates the average elevation of all sampling points in the grid cell (m), N indicates the total number of sampling points, and the global layer (5x5 neighborhood), local layer (10×10 neighborhood). This indicates that significant terrain undulations (such as the edge of a crater) have been encountered and the obstacle avoidance strategy needs to be triggered.
[0028] Define the slope (inclination of the terrain): ; in, , represents the grid cell size, Indicates location The elevation value at . When considering the lunar soil friction coefficient ( ) Determine the risk of skidding and constrain the path slope.
[0029] Define roughness (terrain irregularity): ; When the lunar rover encounters rough terrain (such as a gravel-dense area), the lunar rover's speed is limited ( ).
[0030] (4) According to the lunar rover kinematic state vector and the lunar terrain feature vector , define the state space S; according to the velocity increment and heading angle increment , define the action space ; Based on distance reward, dynamic instantaneous progress reward, dynamic instantaneous direction reward and smoothness reward, define the multi-dimensional reward function r for path planning.
[0031] The feature maps of the global layer and the local layer are spatially aligned through bilinear interpolation, feature fusion is performed, a fused feature matrix is generated, and a three-layer pyramid is constructed to store feature data.
[0032] Design state space: ; in, represents the lunar rover kinematic state vector, represents the lunar terrain feature vector, Represents vector concatenation, taking into account the position information of the lunar rover , vehicle linear speed and the vehicle heading angle These dynamic states and the standard deviation of elevation ,slope The terrain characteristic states such as and roughness R are used as state space inputs.
[0033] Designing the action space: ; Considering speed increment and heading angle increment As the input of the action space. Constraints: , .
[0034] Design the Actor-Critic network structure, the input layer is the state vector S, the dimension is 7. In the hidden layer, the Actor network is set to a 256×128 fully connected layer and a ReLU activation function, and the Critic network is set to a 64-dimensional CNN to process elevation information and a 128-dimensional fully connected layer. The output layer Actor network is set to an action vector , the Critic network is designed as the Q-value estimation function.
[0035] Define the multi-dimensional reward function r for path planning: ; in, , , and is the weight, Distance Rewards, For dynamic instant progress rewards, is the dynamic instantaneous direction reward, Reward for smoothness.
[0036] Further, distance reward for ; Dynamic Instant Progress Rewards for ; Dynamic instantaneous direction reward for ; Smoothness Bonus for ; Weight , , and As shown below: ; ; in, Represents the Euclidean distance from the current position to the target, is the initial distance, is the current speed, is the target speed, is the remaining distance from the current position to the target point, is the total mission distance, is the distance from the previous state to the target point, is the basic direction reward weight, , is the adjustment coefficient, Indicates the angle between the current heading and the target direction. is the maximum tolerance threshold angle, is the heading angle of the target direction, is the heading angle at the current moment, is the basic smoothness reward weight, is the rate of change of acceleration between adjacent time steps, is the maximum allowable rate of change, is the heading angle at the current moment, is the maximum allowed heading angle.
[0037] Select represents the basic weight, driving global convergence, Indicates dynamic instantaneous progress rewards, abandons the traditional [0-1] reward setting, sets the reward function based on the distance between the lunar rover and the target point, and adjusts the sparseness of the reward distribution. represents the dynamic instantaneous direction reward, which encourages the lunar rover to align its heading angle with the target direction during exploration, thereby increasing the reward and adjusting the sparsity of the reward distribution. represents the smoothness reward, , when the heading angle increment exceeds , giving appropriate punishment for sudden changes in heading angle.
[0038] According to the state space Dynamically adjust weights as input , strengthen distance guidance when the terrain is undulating. is the progress reward weight, is the current speed, is the target speed, is the remaining distance from the current position to the target point, is the total mission distance. In the early stage of mission execution, when the remaining distance of the lunar rover to the target position is large, the progress incentive is increased; when the remaining distance is small, the weight is reduced to avoid overly aggressive actions. is the direction reward weight, is the basic direction reward weight, Indicates the angle between the current heading and the target direction. is the maximum tolerance threshold angle, represents the smoothness reward weight, represents the basic smoothness reward weight, is the rate of change of acceleration between adjacent time steps, is the maximum allowable rate of change.
[0039] (5) Construct an adaptive step size based on terrain roughness, slope gradient, and terrain complexity ; Expand the state space to , to the action space Introducing Adaptive Step Size , get the basic action space , input the critic network and output the step evaluation value , used to update the actor Actor network parameters; Introducing Ornstein-Uhlenbeck OU noise, we get the final motion , perform the final action , if satisfied and , then go to step (6); otherwise, recalculate the final action.
[0040] Adaptive step size for ; ; in, As the benchmark step length, it can be set to 0.3m according to the maximum speed of the lunar rover. is the terrain roughness, which is determined by the elevation standard deviation Normalized to get; is the slope gradient, which is obtained by calculating the slope change rate of adjacent grids; It is the terrain complexity, which is a composite index based on the fusion of terrain roughness and terrain slope; , , is the dynamic weight coefficient. Step size constraint .
[0041] When encountering high-density terrain ( ), appropriately shorten the step length to 30%~50% of the baseline value to improve obstacle avoidance accuracy; When encountering low-density terrain ( ), appropriately expand the step size to 120%~150% of the baseline value to speed up exploration.
[0042] Dynamic weight coefficient , , Adjust according to the sports status, as follows:
[0043] ; Constraints: .
[0044] Among them, the higher the speed, the more the terrain roughness affects the weight Reduce; when the heading deviation is large, the slope gradient weight Enhanced; weight when terrain complexity exceeds threshold Double.
[0045] Expand the state space to , to the action space Introducing Adaptive Step Size Get the basic action space , Critic network output step evaluation value : ; in, is the weight coefficient, is the terrain feature embedding vector, is the dynamic state vector, is the bias term.
[0046] ; Among them, the terrain feature embedding vector By elevation standard deviation ,slope , roughness R and terrain complexity constitute, , , , is the weight coefficient. The dynamic state vector By heading angle , vehicle linear speed , Maximum steering angle and safe speed constitute, , , is the weight coefficient.
[0047] The output of the Actor network is the recommended step size, which is compared with the actual execution step size to calculate the loss. , and a weight of 0.3 is added to the policy gradient update, where is the recommended value for the step size, is the actual execution step value.
[0048] Towards Introducing Ornstein-Uhlenbeck OU noise, we get the final motion : ; in, represents the mean reversion rate, ; represents the noise mean, is the maximum slope threshold, , so that The mean reversion rate is controlled at ; represents the noise value generated by the OU process, is the noise value at time t, is the noise disturbance term, is the terrain complexity threshold. Final action Increases with the complexity of the terrain.
[0049] The final action With adaptive step size Send it to the underlying controller to execute the action and verify the feasibility of the lunar rover dynamics model. , , then go to step (6); otherwise, verification is not feasible, otherwise the Actor network output , Zoom to 70% of the feasible interval and clear the noise coefficient, and recalculate the final action.
[0050] (6) Calculate the multi-dimensional reward function for path planning.
[0051] Based on the current state and target position, the multi-dimensional reward function r is calculated to instantly reward, and the weights are dynamically adjusted to balance path safety and efficiency. When , the smoothness reward weight increases to 0.5, when , the dynamic instantaneous progress reward coefficient is reduced to 2.0.
[0052] (7) Calculate the priority weight based on the terrain risk function, multi-dimensional reward function and step loss , filter the experience samples into the experience pool; perform model training, and select the best path based on the path accessibility rate.
[0053] The priority experience replay mechanism calculates the priority weight based on the terrain risk level and the absolute value of the reward in the experience for ; in, is the time difference error, which is used to measure the prediction deviation of the network Q value; is the step loss, represents the terrain risk function, is the state vector, Indicates i The absolute value of the multi-dimensional rewards of experience. is the step size decision error term coefficient, is the terrain risk weight coefficient, is the reward coefficient. , .
[0054] Time difference error for ; in, is the discount factor, is the target Q value of the target network for the next state and action, It is the predicted Q value of the critic network for the current state and action.
[0055] Terrain risk function for ; Among them, 1 represents high-risk experience and 0.2 represents low-risk experience.
[0056] Dynamically choose batch size based on training stage : ; If the proportion of high-risk experience exceeds 30%, adjust the batch size to 1.5 . Enter batch data ; Calculate the target Q value and minimize the timing difference error . Calculate the policy gradient , Update strategy network parameters , update the target network .
[0057] When the average reward for 500 consecutive iterations And standard deviation When , convergence is determined and the optimal strategy network parameters are saved , output training indicators.
[0058] Output and analyze training indicators based on training results, such as path accessibility, path length, path smoothness evaluation, training time, etc. Use path accessibility to select the best path. Path accessibility is defined as the probability of successfully reaching the target point within a given time. Path accessibility evaluation is specifically: ; in, is the Euclidean distance between the position and the target at the end of the i-th training episode, is the task success threshold, The function represents an indicator function, which takes 1 when the condition is met, otherwise it takes 0. The path with the optimal value of path reachability is the optimal path.
Claims
1. A lunar rover path planning method based on a deep deterministic gradient strategy, characterized in that: The following steps are involved: (1) Introducing terrain slope angle , based on the Ackerman steering principle, determine the heading angle and steering angle The differential equations of the lunar rover are used to obtain the kinematic model of the lunar rover. (2) Based on the Baker pressure sinking theory, calculate the wheel sinking amount Z; according to the terrain slope angle and wheel sinking amount Z to calculate the maximum steering angle of the vehicle ; According to the maximum allowable steering angle of the vehicle and wheel sinking amount Z, calculate the minimum wheel turning radius ; Based on the steering angle, minimum turning radius, maximum speed and slope angle, the constraints of the lunar rover kinematic model based on the deep deterministic gradient strategy are constructed; (3) According to the lunar rover kinematic state vector and the lunar terrain feature vector , define the state space S; according to the velocity increment and heading angle increment , define the action space ; Based on distance reward, dynamic instantaneous progress reward, dynamic instantaneous direction reward and smoothness reward, define the multi-dimensional reward function r of path planning; (4) Constructing adaptive step size based on terrain roughness, slope gradient and terrain complexity ; Expand the state space to , to the action space Introducing Adaptive Step Size , get the basic action space , input the critic network and output the step evaluation value , used to update the actor Actor network parameters; Introducing Ornstein-Uhlenbeck OU noise, we get the final motion , perform the final action , if satisfied and , then go to step (5); otherwise, recalculate the final action; (5) Calculate the multi-dimensional reward function for path planning; (6) Calculate the priority weight based on the terrain risk function, multi-dimensional reward function and step loss , filter the experience samples into the experience pool; Perform model training and select the best path based on the path accessibility rate.
2. The lunar rover path planning method based on the deep deterministic gradient strategy according to claim 1 is characterized in that: The kinematic model of the lunar rover is ; in, is the lunar rover position information, is the vehicle linear velocity, is the vehicle wheelbase.
3. The lunar rover path planning method based on the deep deterministic gradient strategy according to claim 2 is characterized in that: The wheel sinking amount Z is ; in, is the vertical soil pressure, is the soil cohesion modulus, is the vehicle wheel width, is the soil friction modulus, is the subsidence index; Vehicle maximum steering angle for ; in, It indicates the maximum turning angle on flat ground when considering the subsidence and slope. is the vehicle radius; Minimum wheel turning radius for ; Maximum speed for ; in, is the friction coefficient, is the lunar gravity coefficient.
4. The lunar rover path planning method based on the deep deterministic gradient strategy according to claim 3 is characterized in that: The state space S is ; Among them, the lunar rover kinematic state vector Including lunar rover location information , vehicle linear speed and heading angle , lunar terrain feature vector Include elevation standard deviation , Terrain slope angle and roughness R; Roughness R is ; Where N represents the total number of sampling points. Represents the elevation value of the i-th sampling point in the grid cell. Represents the average elevation of all sampling points in the grid cell. When the lunar rover encounters rough terrain, the lunar rover's speed will be limited. ; Action Space for ; in, and .
5. The lunar rover path planning method based on deep deterministic gradient strategy according to claim 4 is characterized in that: Multi-dimensional reward function for path planning for ; in, , , and is the weight, Distance Rewards, For dynamic instant progress rewards, is the dynamic instantaneous direction reward, Reward for smoothness.
6. The lunar rover path planning method based on deep deterministic gradient strategy according to claim 5 is characterized in that: Distance Rewards for ; Dynamic Instant Progress Rewards for ; Dynamic instantaneous direction reward for ; Smoothness Bonus for ; Weight , , and As shown below: ; ; in, Represents the Euclidean distance from the current position to the target. is the initial distance, is the current speed, is the target speed, is the remaining distance from the current position to the target point, is the total mission distance, is the distance from the previous state to the target point, is the basic direction reward weight, , is the adjustment coefficient, Indicates the angle between the current heading and the target direction. is the maximum tolerance threshold angle, is the heading angle of the target direction, is the heading angle at the current moment, is the basic smoothness reward weight, is the rate of change of acceleration between adjacent time steps, is the maximum allowable rate of change, is the heading angle at the current moment, is the maximum allowed heading angle.
7. The lunar rover path planning method based on the deep deterministic gradient strategy according to claim 6 is characterized in that: Adaptive step size for ; ; ; in, is the base step length, is the terrain roughness, which is determined by the elevation standard deviation Normalized to get; is the slope gradient, which is obtained by calculating the slope change rate of adjacent grids; It is the terrain complexity, which is a composite index based on the fusion of terrain roughness and terrain slope; , , is the dynamic weight coefficient; when When the value is not less than 0.7, i.e. when encountering high-density terrain, the step length is shortened to 30%~50% of the baseline value; when When it is less than 0.3, that is, when encountering low-density terrain, the step length is expanded to 120%~150% of the baseline value.
8. The lunar rover path planning method based on deep deterministic gradient strategy according to claim 7 is characterized in that: Basic action space for ; Step length evaluation value for ; in, is the weight coefficient, is the terrain feature embedding vector, is the dynamic state vector, is the bias term; Final Action for ; in, represents the mean reversion rate, ; represents the noise mean, is the maximum slope threshold, represents the noise value generated by the OU process, is the noise value at time t, is the noise disturbance term, is the terrain complexity threshold.
9. The lunar rover path planning method based on the deep deterministic gradient strategy according to claim 8 is characterized in that: Priority Weight for ; in, is the time difference error, is the step loss, represents the terrain risk function, is the state vector, Indicates i The absolute value of the multi-dimensional rewards of experience. is the step size decision error term coefficient, is the terrain risk weight coefficient, is the reward coefficient.
10. The lunar rover path planning method based on deep deterministic gradient strategy according to claim 9, characterized in that: Time difference error for ; in, is the discount factor, is the target Q value of the target network for the next state and action, The predicted Q value of the critic network for the current state and action; Step loss for ; in, is the recommended value for the step size, is the actual execution step value; Terrain risk function for ; Among them, 1 represents high-risk experience and 0.2 represents low-risk experience.
Citation Information
Patent Citations
Lunar surface safe landing area-based inspector path planning method and system
CN109597415A
Ground robot path planning method based on deep reinforcement learning
CN116625369A
Lunar surface path planning method based on reinforcement learning
CN117516572A
Planetary probe vehicle multi-target path planning method and system, storage medium and computing equipment
CN117606501A
Autonomous mobile robot path planning method based on deep reinforcement learning
CN118259669A
Cited By
Method and device for generating safety buffer area for movement of lunar rover
CN121115742A
Intelligent unmanned aerial vehicle flight path planning method and system
CN121612310A