A lunar rover path planning method based on deep deterministic gradient policy
Through the path planning method based on the depth deterministic gradient strategy, combined with the multi-dimensional reward function and the adaptive step size mechanism, the problem of insufficient local obstacle avoidance capabilities of the lunar rover under complex terrain is solved, the path planning efficiency and motion safety is improved, and the problems of slow convergence speed and poor terrain adaptability in traditional methods are solved.
Patent Information
- Application Number
- CN202510466417.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-15
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2045-04-15
AI Technical Summary
Traditional lunar rover path planning methods lack local obstacle avoidance capabilities under complex terrain and communication delay conditions, making it difficult to achieve efficient extraterrestrial detection tasks. The existing deep reinforcement learning methods do not fully integrate multi-source sensor data, resulting in difficulty in strategy convergence and path safety evaluation errors.
The path planning method based on the depth deterministic gradient strategy is adopted, and a multi-dimensional reward function is constructed by introducing the terrain slope angle and wheel subsidence. Combined with the adaptive step size mechanism, the policy network is dynamically adjusted to improve obstacle avoidance accuracy and autonomous exploration capabilities, avoid local optimization, and achieve path continuity and security.
It improves the obstacle avoidance accuracy and independent exploration capabilities of the lunar rover in complex terrain, solves the problems of slow convergence speed, frequent path oscillation and poor terrain adaptability in traditional methods, and achieves the dual improvement of path planning efficiency and motion safety, reducing the complexity of the model.
Smart Images

Figure CN120010495B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of unmanned vehicle control and deep reinforcement learning, and particularly relates to a lunar rover path planning method based on a deep deterministic gradient policy. Background Art
[0002] With the progress of technology, especially the continuous development of space exploration technology, the application of celestial exploration vehicles (such as lunar rovers, Mars rovers, etc.) in unknown environments is becoming increasingly widespread. These exploration vehicles are used to perform various tasks, such as geological exploration, sampling, data collection, etc. Therefore, research on path planning is particularly important. Traditional path planning methods for ground robots usually rely on pre-constructed environmental maps. However, in celestial exploration tasks, due to the complexity, unknownness, and dynamics of the terrain, traditional path planning methods face many challenges. Rule-based algorithms (such as Dijkstra's algorithm) need to rely on high-precision pre-built maps, and the complex terrain on the moon leads to a large amount of real-time map update calculations. Moreover, local obstacle avoidance algorithms (such as the Dynamic Window Approach - DWA) are prone to falling into local optima in continuous obstacle regions because they only optimize single-step paths. The state space design of existing deep reinforcement learning methods does not fully integrate multi-source sensor data (such as slope angle, roughness, and terrain complexity), resulting in the policy network being unable to accurately perceive terrain mechanical characteristics and path safety evaluation errors. The single reward function of the traditional Deep Deterministic Policy Gradient - DDPG algorithm only optimizes the objective of distance or time, ignoring the consideration of terrain features, vehicle body stability, and motion smoothness, thus leading to difficulties in policy convergence during training and making it difficult to generate the optimal path. Summary of the Invention
[0003] Object of the Invention: Aiming at the problem that the existing lunar rover path planning methods have insufficient local obstacle avoidance ability under the conditions of complex lunar terrain and communication delay, resulting in the inability to achieve efficient extraterrestrial exploration tasks, the present invention aims to provide a lunar rover path planning method based on a deep deterministic gradient policy.
[0004] Technical Solution: The lunar rover path planning method based on a deep deterministic gradient policy according to the present invention includes the following steps:
[0005] (1) Introduce the terrain slope angle , and based on the Ackermann steering principle, determine the differential equations of the heading angle and the steering angle to obtain the kinematic model of the lunar rover;
[0006] (2) Based on the Bekker pressure - settlement theory, calculate the wheel settlement amount Z; according to the terrain slope angle and the wheel settlement amount Z, calculate the maximum steering angle of the vehicle; according to the maximum allowable steering angle Calculate the minimum wheel turning radius based on the wheel subsidence amount Z ; Based on the steering angle, minimum turning radius, maximum speed, and slope angle, construct the constraint conditions of the lunar rover kinematic model based on the depth deterministic gradient strategy;
[0007] (3)According to the lunar rover kinematic state vector and the lunar terrain feature vector , define the state space S; According to the speed increment and the heading angle increment , define the action space ; Based on the distance reward, dynamic instantaneous progress reward, dynamic instantaneous direction reward, and smoothness reward, define the multi-dimensional reward function r of path planning;
[0008] (4)Construct an adaptive step size according to the terrain roughness, slope gradient, and terrain complexity ; Expand the state space to , introduce the adaptive step size into the action space , obtain the basic action space , input it into the critic network, and output the step size evaluation value , which is used to update the actor network parameters; Introduce Ornstein-Uhlenbeck OU noise into , obtain the final action , execute the final action , if it satisfies and , then enter step (5); Otherwise, recalculate the final action;
[0009] (5)Calculate the multi-dimensional reward function of path planning;
[0010] (6)Calculate the priority weight according to the terrain risk function, multi-dimensional reward function, and step size loss , screen the experience samples into the experience pool; Perform model training, and select the best path according to the path reach rate.
[0011] Furthermore, the lunar rover kinematic model is
[0012] ;
[0013] Among them, is the lunar rover position information, is the vehicle body linear velocity, is the vehicle wheelbase.
[0014] Furthermore, the wheel subsidence amount Z is
[0015] ;
[0016] Among them, is the vertical soil pressure, is the soil cohesion modulus, is the width of the vehicle wheel, is the soil friction modulus, is the settlement index;
[0017] The maximum steering angle of the vehicle is
[0018] ;
[0019] Among them, represents the maximum steering angle on flat ground considering the settlement and slope, is the vehicle radius;
[0020] The minimum turning radius of the wheel is
[0021] ;
[0022] The maximum speed is
[0023] ;
[0024] Among them, is the friction coefficient, is the lunar gravity coefficient.
[0025] Furthermore, the state space S is
[0026] ;
[0027] Among them, the kinematic state vector of the lunar rover includes the position information of the lunar rover , the linear velocity of the vehicle body and the heading angle , and the lunar terrain feature vector includes the elevation standard deviation , the terrain slope angle and the roughness R;
[0028] The roughness R is
[0029] ;
[0030] Among them, N represents the total number of sampling points, represents the elevation value of the i-th sampling point in the grid cell, represents the average elevation of all sampling points in the grid cell. If When it indicates that the lunar rover encounters rough terrain, the driving speed of the lunar rover is restricted ;
[0031] Action space is
[0032] ;
[0033] Among them, and .
[0034] Furthermore, the multi-dimensional reward function of path planning is
[0035] ;
[0036] Among them, 、 、 and are weights, Distance reward, is the dynamic instantaneous progress reward, is the dynamic instantaneous direction reward, is the smoothness reward.
[0037] Furthermore, the distance reward is
[0038] ;
[0039] The dynamic instantaneous progress reward is
[0040] ;
[0041] The dynamic instantaneous direction reward is
[0042] ;
[0043] The smoothness reward is
[0044] ;
[0045] Weights 、 、 and are shown as follows:
[0046] ;
[0047] ;
[0048] Among them, represents the Euclidean distance from the current position to the target, is the initial distance, is the current speed, is the target speed, is the remaining distance from the current position to the target point, is the total distance of the task, is the distance from the position of the previous state to the target point, is the basic direction reward weight, 、 is the adjustment coefficient, represents the angle between the current heading and the target direction, is the maximum tolerance threshold angle, is the target direction heading angle, is the current moment heading angle, is the basic smoothness reward weight, is the acceleration change rate of adjacent time steps, is the maximum allowable change rate, is the current moment heading angle, is the maximum allowable heading angle.
[0049] Furthermore, the adaptive step size is
[0050] ;
[0051] ;
[0052] ;
[0053] Among them, is the reference step size, is the terrain roughness, which is obtained by normalizing the elevation standard deviation ; is the slope gradient, which is obtained by calculating the slope change rate of adjacent grids; is the terrain complexity, which is a composite index based on the fusion of terrain roughness and terrain slope; 、 、 are dynamic weight coefficients;
[0054] When is not less than 0.7, that is, when encountering high-density terrain, the step size is shortened to 30% - 50% of the reference value;
[0055] When is less than 0.3, that is, when encountering low-density terrain, the step size is expanded to 120% - 150% of the reference value.
[0056] Furthermore, the basic action space is
[0057] ;
[0058] The step evaluation value is
[0059] ;
[0060] Among them, is the weight coefficient, is the terrain feature embedding vector, is the dynamic state vector, is the bias term;
[0061] The final action is
[0062] ;
[0063] Among them, represents the mean reversion rate, ; represents the noise mean, is the maximum slope threshold, represents the noise value generated by the OU process, is the noise value at time t, is the noise perturbation term, is the terrain complexity threshold.
[0064] Furthermore, the priority weight is
[0065] ;
[0066] Among them, is the time difference error, is the step loss, represents the terrain risk function, is the state vector, represents the absolute value of the multi-dimensional reward of the i th experience, is the step decision error term coefficient, is the terrain risk weight coefficient, is the reward coefficient.
[0067] Furthermore, the time difference error is
[0068] ;
[0069] Among them, is the discount factor, is the target Q-value of the target network for the next state and action, is the predicted Q-value of the Critic network for the current state and action;
[0070] Step size loss is
[0071] ;
[0072] wherein, is the step size suggestion value, is the actual executed step size value;
[0073] Terrain risk function is
[0074] ;
[0075] wherein, 1 represents high-risk experience and 0.2 represents low-risk experience.
[0076] Beneficial effects: Compared with the prior art, the present invention has the following remarkable advantages: 1. By fusing four types of objectives, namely distance reward, instantaneous progress reward, instantaneous heading reward, and motion smoothness reward, the present invention forms a gradient signal that can adapt to terrain complexity, guides the policy network to achieve a dynamic balance between exploration and safety, constructs a dynamic multi-dimensional reward function mechanism, and dynamically adjusts the weight coefficient through terrain features to achieve multi-objective adaptive collaboration; 2. The present invention proposes an adaptive step size mechanism based on multi-scale terrain features, fuses terrain features, and dynamically adjusts the step size in different risk terrains to improve the obstacle avoidance accuracy and autonomous exploration ability of the lunar rover in different terrains; 3. Through a multi-level adaptive mechanism of a dynamic multi-dimensional reward function, a lightweight network structure, and an adaptive step size adjustment strategy, the present invention avoids local optima, improves path continuity, and further completes the autonomous decision-making detection task of the lunar rover, achieving a double improvement in path planning efficiency and motion safety, solving the core problems of traditional reinforcement learning in complex lunar surface environments, such as slow convergence speed, frequent path oscillations, poor terrain adaptability, and low model deployment efficiency, and reducing the model complexity; 4. The present invention fuses a collaborative multi-level optimization framework of terrain feature perception, improved reward function, and network structure, and at the same time proposes a dynamic adaptive step size control strategy according to terrain features, and dynamically adjusts the strategy through the terrain feature fusion amount. Brief description of the drawings
[0077] Figure 1 is the flow chart of the present invention;
[0078] Figure 2 is the structural schematic diagram of the adaptive step size strategy of the present invention. Detailed implementation manners
[0079] The present invention will be further described below in conjunction with the accompanying drawings.
[0080] The lunar rover path planning method based on the deep deterministic gradient strategy according to the present invention includes the following steps:
[0081] (1) Introduce the terrain slope angle , and based on the Ackermann steering principle, determine the differential equations of the heading angle and the steering angle to obtain the kinematic model of the lunar rover.
[0082] Establish the heading angle - steering angle differential equation:
[0083] ;
[0084] Among them, is the vehicle heading angle, in rad, is the vehicle body linear velocity, in m / s, is the front wheel steering angle, in rad, represents the vehicle wheelbase, represents the terrain slope angle, in rad.
[0085] The kinematic model of the lunar rover considering the influence of the terrain slope angle is
[0086] ;
[0087] Among them, is the lunar rover position information, is the vehicle body linear velocity, is the vehicle wheelbase.
[0088] (2) Based on Bekker's pressure - settlement theory, calculate the wheel settlement amount Z; according to the terrain slope angle and the wheel settlement amount Z, calculate the maximum steering angle of the vehicle; according to the maximum allowable steering angle of the vehicle and the wheel settlement amount Z, calculate the minimum wheel turning radius , ; Based on the steering angle, minimum turning radius, maximum speed and slope angle, construct the constraint conditions of the kinematic model of the lunar rover based on the deep deterministic gradient strategy.
[0089] Based on Bekker's pressure - settlement theory, the wheel settlement amount Z is
[0090] ;
[0091] Among them, is the soil vertical pressure, in kPa, is the soil cohesion modulus, in kPa / m n+1, is the vehicle wheel width, in m is the soil friction modulus, in kPa / m n+1 , is the settlement index
[0092] Maximum vehicle steering angle is
[0093] ;
[0094] Among them, represents the maximum steering angle on flat ground considering the settlement amount and slope is the vehicle radius. Front wheel steering angle .
[0095] Minimum wheel turning radius is
[0096] ;
[0097] Maximum speed is
[0098] ;
[0099] Among them, is the friction coefficient is the lunar gravity coefficient
[0100] (3) Calculate the terrain feature quantity
[0101] Import the elevation DEM data file information that needs to extract the body shape features, and perform downsampling feature extraction through the Gaussian pyramid. Apply a 5×5 Gaussian filter kernel (standard deviation σ = 1.0 pixel) to the original DEM data, and eliminate high-frequency noise and small terrain undulations through weighted averaging. Perform downsampling with a step size of 2 on the filtered DEM, that is, retain the center point of each 2×2 pixel area as the new resolution grid node. Repeat the downsampling iteration 3 times (k = 3), and finally generate a hierarchical terrain model with gradually decreasing resolution
[0102] According to the generated multi-scale terrain model (hierarchical terrain model), perform terrain feature quantization calculation on each grid. Retain the feature calculations of the high-resolution terrain model and the low-resolution terrain model respectively, and dynamically select the feature resolution according to the path planning requirements. The high-resolution grid feature calculation supports global path planning, and the low-resolution grid feature calculation ensures obstacle avoidance tasks and dynamic feasibility
[0103] Define the elevation standard deviation (terrain undulation intensity):
[0104] ;
[0105] Among them, represents the elevation value (m) of the i-th sampling point in the grid cell, represents the average elevation of all sampling points in the grid cell (m), N represents the total number of sampling points, global layer (5x5 neighborhood), local layer (10×10 neighborhood). When it indicates encountering significant terrain undulations (such as the edge of a crater), the obstacle avoidance strategy needs to be triggered.
[0106] Define the slope (terrain inclination):
[0107] ;
[0108] Among them, and represent the grid cell size, represents the position of the elevation value. When, combined with the lunar soil friction coefficient ( ), the slipping risk is judged and the path slope is restricted.
[0109] Define the roughness (terrain irregularity):
[0110] ;
[0111] When it means that the lunar rover encounters rough terrain (such as a densely gravel area), the driving speed of the lunar rover is restricted ( ).
[0112] (4) According to the kinematic state vector of the lunar rover and the lunar terrain feature vector , define the state space S; according to the speed increment and the heading angle increment , define the action space ; based on the distance reward, dynamic instantaneous progress reward, dynamic instantaneous direction reward and smoothness reward, define the multi-dimensional reward function r for path planning.
[0113] Align the global layer and local layer features through bilinear interpolation for spatial alignment, perform feature fusion, generate a fused feature matrix, and construct a three-layer pyramid to store feature data.
[0114] Design the state space:
[0115] ;
[0116] Among them, represents the kinematic state vector of the lunar rover, Represents the lunar terrain feature vector, denotes vector concatenation, considering the lunar rover's position information , the linear velocity of the vehicle body and the vehicle heading angle These dynamic states and the elevation standard deviation , slope and the roughness R, these terrain feature quantity states are used as the state space input.
[0117] Design the action space:
[0118] ;
[0119] Consider the velocity increment and the heading angle increment as the input of the action space. Constraints: , .
[0120] Design the Actor-Critic network structure. The input layer is the state vector S with a dimension of 7. In the hidden layer, the Actor network is set with a 256×128 fully connected layer and a ReLU activation function, and the Critic network is set as a 64-dimensional CNN to process the elevation information and a 128-dimensional fully connected layer. The output layer of the Actor network is set as the action vector , and the Critic network is designed as a Q-value estimation function.
[0121] Define the multi-dimensional reward function r for path planning:
[0122] ;
[0123] Among them, , , and are weights, distance reward, is the dynamic instantaneous progress reward, is the dynamic instantaneous direction reward, is the smoothness reward.
[0124] Furthermore, the distance reward is
[0125] ;
[0126] The dynamic instantaneous progress reward is
[0127] ;
[0128] The dynamic instantaneous direction reward For
[0129] ;
[0130] Smoothness reward For
[0131] ;
[0132] Weight , , And Are as shown in the following formula:
[0133] ;
[0134] ;
[0135] Wherein, Represents the Euclidean distance from the current position to the target, Is the initial distance, Is the current speed, Is the target speed, Is the remaining distance from the current position to the target point, Is the total task distance, Is the distance from the position of the previous state to the target point, Is the basic direction reward weight, , Are adjustment coefficients, Represents the angle between the current heading and the target direction, Is the maximum tolerance threshold angle, Is the target direction heading angle, Is the current moment heading angle, Is the basic smoothness reward weight, Is the acceleration change rate between adjacent time steps, Is the maximum allowable change rate, Is the current moment heading angle, Is the maximum allowable heading angle.
[0136] Select Represents the basic weight, driving global convergence, Represents the dynamic instantaneous progress reward, abandoning the traditional [0 - 1] reward setting, setting the reward function based on the distance between the lunar rover and the target point, and adjusting the sparsity of the reward distribution, Represents the dynamic instantaneous direction reward, encouraging the lunar rover to align the heading angle with the target direction during exploration, thereby increasing the reward and adjusting the sparsity of the reward distribution, Represents the smoothness reward, , when the heading angle increment exceeds Give appropriate punishment for sudden changes in the heading angle.
[0137] According to the dynamically adjust the weight as the input Strengthen distance guidance when the terrain undulates violently. is the progress reward weight, is the current speed, is the target speed, is the remaining distance from the current position to the target point, is the total distance of the mission. At the initial stage of mission execution, when the remaining distance of the lunar rover to the target position is large, enhance the progress incentive; when the remaining distance is small, reduce the weight to avoid overly aggressive actions. is the direction reward weight, is the basic direction reward weight, represents the angle between the current heading and the target direction, is the maximum tolerance threshold angle, represents the smoothness reward weight, represents the basic smoothness reward weight, is the acceleration change rate between adjacent time steps, is the maximum allowable change rate.
[0138] (5) Construct an adaptive step size according to terrain roughness, slope gradient, and terrain complexity ; Expand the state space to , introduce the adaptive step size into the action space , obtain the basic action space , input it into the critic network, and output the step size evaluation value , which is used to update the parameters of the actor network; introduce Ornstein-Uhlenbeck OU noise into , obtain the final action , execute the final action , if it satisfies and and , then enter step (6); otherwise, recalculate the final action.
[0139] Adaptive step size is
[0140] ;
[0141] ;
[0142] Among them, is the reference step size, which can be set to 0.3m according to the maximum speed of the lunar rover, is the terrain roughness, which is obtained by normalizing the standard deviation of elevation ; is the slope gradient, which is obtained by calculating the change rate of slopes of adjacent grids; is the terrain complexity, which is a composite index based on the fusion of terrain roughness and terrain slope; 、 、 are dynamic weight coefficients. Step size constraint .
[0143] When encountering high-density terrain ( ), the step size is appropriately shortened to 30% - 50% of the reference value to improve the obstacle avoidance accuracy;
[0144] When encountering low-density terrain ( ), the step size is appropriately enlarged to 120% - 150% of the reference value to accelerate exploration.
[0145] The dynamic weight coefficients 、 、 are adjusted according to the motion state, specifically as follows:
[0146] ;
[0147] Constraint conditions: .
[0148] Among them, the higher the speed, the lower the influence weight of terrain roughness ; when the heading deviation is large, the weight of slope gradient is enhanced; when the terrain complexity exceeds the threshold, the weight doubles.
[0149] The state space is expanded to , and an adaptive step size is introduced into the action space to obtain the basic action space , and the Critic network outputs the step size evaluation value :
[0150] ;
[0151] Among them, is the weight coefficient, is the terrain feature embedding vector, is the dynamic state vector, is the bias term.
[0152] ;
[0153] Among them, the terrain feature embedding vector Composed of elevation standard deviation , slope , roughness R, and terrain complexity . , , , are weight coefficients. The dynamic state vector is composed of heading angle , vehicle body linear velocity , maximum steering angle , and safety speed . , , are weight coefficients.
[0154] The output of the Actor network is the step size recommendation value, which calculates the loss with the actual executed step size and adds it to the policy gradient update with a weight of 0.3. Among them, is the step size recommendation value, is the actual executed step size value.
[0155] Introduce Ornstein-Uhlenbeck OU noise to obtain the final action :
[0156] ;
[0157] Among them, represents the mean reversion rate, ; represents the noise mean, is the maximum slope threshold, , so that the mean reversion rate is controlled within ; represents the noise value generated by the OU process, is the noise value at time t, is the noise perturbation term, is the terrain complexity threshold. The final action increases with the increase of terrain complexity.
[0158] Send the final action and the adaptive step size to the underlying controller to execute the action and verify the feasibility of the lunar rover dynamics model. If , , then enter step (6); otherwise, if the verification is infeasible, conversely, scale the , output by the Actor network to 70% of the feasible interval and clear the noise coefficient, and recalculate the final action.
[0159] (6) Calculate the multi-dimensional reward function for path planning.
[0160] Based on the current state and the target location, calculate the multi-dimensional reward function \(r\) (immediate reward), and dynamically adjust the weights to balance path safety and efficiency. When , the smoothness reward weight is increased to 0.5. When , the dynamic instantaneous progress reward coefficient is decreased to 2.0.
[0161] (7) Calculate the priority weight according to the terrain risk function, the multi-dimensional reward function, and the step loss , filter the experience samples into the experience pool; perform model training, and select the best path according to the path reachability rate.
[0162] The priority experience replay mechanism calculates the priority weight based on the terrain risk level and the absolute value of the reward in the experience as
[0163] ;
[0164] where, is the temporal difference error, which is used to measure the prediction deviation of the network Q value; is the step loss, represents the terrain risk function, is the state vector, represents the absolute value of the multi-dimensional reward of the i th experience, is the step decision error term coefficient, is the terrain risk weight coefficient, is the reward coefficient. Take , .
[0165] The temporal difference error is
[0166] ;
[0167] where, is the discount factor, is the target Q value of the target network for the next state and action, is the predicted Q value of the critic network for the current state and action.
[0168] The terrain risk function is
[0169] ;
[0170] where, 1 represents high-risk experience, and 0.2 represents low-risk experience.
[0171] Dynamically select the batch size according to the training stage :
[0172] ;
[0173] If the proportion of high - risk experience exceeds 30%, adjust the batch size to 1.5 Input the batch data ; Calculate the target Q - value and minimize the temporal - difference error Calculate the policy gradient , Update the policy network parameters , Update the target network .
[0174] When the average reward for 500 consecutive iterations and the standard deviation , Determine convergence, save the optimal policy network parameters , Output the training metrics
[0175] Output and analyze the training metrics according to the training results, such as path reachability rate, path length, path smoothness evaluation, training time, etc. Select the best path using the path reachability rate. The path reachability rate is defined as the probability of successfully reaching the target point within a given time. The path reachability evaluation is specifically
[0176] ;
[0177] Among them, is the Euclidean distance between the position at the end of the i - th training episode and the target, is the task success threshold, The function represents the indicator function, which takes 1 when the condition is satisfied and 0 otherwise. The path with the optimal path reachability value is the optimal path
Claims
1. A lunar rover path planning method based on deep deterministic gradient policy, characterized in that It includes the following steps: (1) Introduce the terrain slope angle , and based on the Ackermann steering principle, determine the differential equations of the heading angle and the steering angle to obtain the kinematic model of the lunar rover; (2)Based on the Bekker pressure-settlement theory, calculate the wheel settlement Z; according to the terrain slope angle and the wheel settlement Z, calculate the maximum steering angle of the vehicle ; according to the maximum allowable steering angle of the vehicle and the wheel settlement Z, calculate the minimum wheel turning radius ; based on the steering angle, minimum turning radius, maximum speed and slope angle, construct the constraint conditions of the kinematic model of the lunar rover based on the depth deterministic gradient strategy; (3) Based on the kinematic state vector of the lunar rover and the lunar terrain feature vector , define the state space S; based on the velocity increment and the heading angle increment , define the action space ; based on the distance reward, dynamic instantaneous progress reward, dynamic instantaneous direction reward, and smoothness reward, define the multi-dimensional reward function r for path planning; (4)Construct an adaptive step size based on terrain roughness, slope gradient, and terrain complexity ; Expand the state space to , introduce the adaptive step size into the action space , and obtain the basic action space . Input it into the critic network and output the step size evaluation value , which is used to update the parameters of the actor network; introduce Ornstein-Uhlenbeck (OU) noise into to obtain the final action , and execute the final action . If it satisfies and , then go to step (5); otherwise, recalculate the final action (5) Calculate the multi-dimensional reward function for path planning; (6) Calculate the priority weights based on the terrain risk function, multi-dimensional reward function, and step loss , and filter the experience samples to enter the experience pool; Perform model training and select the best path according to the path reachability rate; In step (3), the multi-dimensional reward function for path planning is ; Among them, , , and are weights, is the distance reward, is the dynamic instantaneous progress reward, is the dynamic instantaneous direction reward, is the smoothness reward; Distance Reward For ; Dynamic Instant Progress Reward For ; Dynamic instantaneous direction reward For ; Smoothness Reward For ; Weight , , and are as follows: ; ; Among them, represents the Euclidean distance from the current position to the target, is the initial distance, is the current speed, is the target speed, is the remaining distance from the current position to the target point, is the total task distance, is the distance from the position of the previous state to the target point, is the basic direction reward weight, 、 are adjustment coefficients, represents the angle between the current heading and the target direction, is the maximum tolerance threshold angle, is the target direction heading angle, is the current moment heading angle, is the basic smoothness reward weight, is the acceleration change rate of adjacent time steps, is the maximum allowable change rate, is the current moment heading angle, is the maximum allowable heading angle; In step (4), the adaptive step size is ; ; ; Among them, is the reference step size, is the terrain roughness, which is obtained by normalizing the elevation standard deviation ; is the slope gradient, which is obtained by calculating the slope change rate of adjacent grids; is the terrain complexity, which is a composite index based on the fusion of terrain roughness and terrain slope; , , are dynamic weight coefficients; is the maximum speed; When When it is not less than 0.7, that is, when encountering high-density terrain, the step length is shortened to 30% - 50% of the reference value; When When it is less than 0.3, that is, when encountering low-density terrain, the step size is expanded to 120% - 150% of the reference value.
2. The lunar rover path planning method based on the deep deterministic gradient strategy according to claim 1, wherein The kinematic model of the lunar rover is ; Among them, is the position information of the lunar rover, is the linear velocity of the vehicle body, is the wheelbase of the vehicle.
3. The lunar rover path planning method based on the deep deterministic gradient strategy according to claim 2, characterized in that The wheel sinkage Z is ; Among them, is the vertical soil pressure, is the soil cohesion modulus, is the vehicle wheel width, is the soil friction modulus, is the settlement index; Maximum steering angle of vehicle is ; Among them, represents the maximum turning angle of the flat ground when considering the subsidence amount and slope, is the vehicle radius; Minimum turning radius of the wheel is ; Maximum speed is ; Among them, is the friction coefficient, is the lunar gravity coefficient.
4. The lunar rover path planning method based on the deep deterministic gradient policy according to claim 3, wherein The state space S is ; Among them, the kinematic state vector of the lunar rover includes the position information of the lunar rover , the linear velocity of the vehicle body and the heading angle . The lunar terrain feature vector includes the elevation standard deviation , the terrain slope angle and the roughness R; The roughness R is ; where N represents the total number of sampling points, represents the elevation value of the i-th sampling point within the grid cell, represents the average elevation of all sampling points within the grid cell. If indicates that the lunar rover encounters rough terrain, the driving speed of the lunar rover is restricted ; Action space For ; Among them, and .
5. The lunar rover path planning method based on the deep deterministic gradient strategy according to claim 4, characterized in that, Basic action space For ; Step size evaluation value For ; Among them, is the weight coefficient, is the terrain feature embedding vector, is the dynamic state vector, is the bias term; Final action For ; Among them, represents the mean reversion rate, ; represents the noise mean, is the maximum slope threshold, represents the noise value generated by the OU process, is the noise value at time t, is the noise perturbation term, is the terrain complexity threshold.
6. The lunar rover path planning method based on the deep deterministic gradient policy according to claim 5, characterized in that Priority weight For ; in, is the time difference error, is the step loss, represents the terrain risk function, is the state vector, Indicates i The absolute value of the multi-dimensional rewards of experience. is the step size decision error term coefficient, is the terrain risk weight coefficient, is the reward coefficient.
7. The lunar rover path planning method based on the deep deterministic gradient policy according to claim 6, wherein Time difference error is ; wherein, is the discount factor, is the target Q-value of the target network for the next state and action, is the predicted Q-value of the critic network for the current state and action; Step loss is ; Among them, is the step size recommended value, is the actual executed step size value; Terrain risk function For ; Among them, 1 represents high-risk experience and 0.2 represents low-risk experience.
Citation Information
Patent Citations
Multi-patroller cooperative task planning method and system for lunar surface exploration time network
CN118584951A
Mobile robot local motion planning method and apparatus and computer storage medium
WO2019076044A1