An autonomous vehicle navigation control method and system

The integration of traditional navigation algorithms with Actor-Critic DRL networks using laser radar and inertial navigation enhances DRL training efficiency and accuracy for autonomous vehicle navigation.

CN115494849BActive Publication Date: 2025-07-15INST OF ELECTRICAL ENG CHINESE ACAD OF SCI
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211322372.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-10-27
Publication Date
2025-07-15
Estimated Expiration
2042-10-27

AI Technical Summary

Technical Problem

Traditional deep reinforcement learning methods are inefficient in navigation control of autonomous vehicles, and it is difficult to quickly converge to ideal state. It is difficult to improve accuracy by integrating traditional control theory with deep learning methods.

Method used

Combining lidar point cloud data and inertial navigation instrument data, navigation control algorithms and Actor-Critic type DRL decision networks are used to determine the vehicle control volume through DWA algorithm and path tracking algorithm, and the navigation model is optimized using the reward and punishment mechanism.

Benefits of technology

It improves the accuracy of navigation control of autonomous vehicles and the convergence speed of neural network training, reduces the difficulty of training, and enhances the actual performance of the DRL model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115494849B_ABST
    Figure CN115494849B_ABST
Patent Text Reader

Abstract

The present invention relates to a navigation control method and system for an autonomous vehicle. The method includes performing navigation using an optimized navigation model according to vehicle and environmental state data; the training process is to determine vehicle control quantities according to the point cloud data of obstacles using a navigation control algorithm; constructing a navigation model using a DRL decision network according to the vehicle control quantities and the vehicle and environmental state data; the DRL decision network includes a DRL Actor network that outputs a first vehicle control quantity according to the vehicle and environmental state data and the vehicle control quantity; a DRL Critic network that determines the expected rewards corresponding to two sets of vehicle control quantities according to the first vehicle control quantity, the vehicle control quantity, and the vehicle and environmental state data, determines the final control quantity, and outputs it; and optimizing the navigation model using a reward and punishment mechanism. The present invention can improve the accuracy of autonomous driving navigation control in the DRL mode, the convergence speed of neural network training, and reduce the difficulty of neural network training convergence.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of autonomous driving, and particularly to a navigation control method and system for an autonomous driving vehicle. Background Art

[0002] With the development of autonomous vehicles, researchers have found that the deep reinforcement learning (DRL) method can well convert road network data and various sensor information into decision values and control commands through unsupervised learning. This feature enables the DRL method to be widely applied to key scenarios such as path planning systems, behavior decision-making systems, and navigation control systems for autonomous driving. Among them, the vehicle navigation control system based on the DRL method requires the intelligent agent to directly output the control command of the vehicle according to the data of the perception system; at the same time, it is required that the decision value of the intelligent agent meets the vehicle dynamics requirements and can independently handle various situations in the running scenario. It belongs to an end-to-end vehicle control method. Among them, the deep neural network, as the "brain" of the intelligent vehicle, controls the vehicle movement according to the environmental state. Researchers expect that through a large number of trainings and learnings of the deep neural network, the intelligent vehicle can achieve the effect of human driving.

[0003] In the traditional DRL neural network training process, a reward and punishment mechanism is used to optimize the neural network. That is, the intelligent agent observes a multi-dimensional state S, calculates an action value A according to the weights and biases of the neural network. After the intelligent agent executes this action, it will enter the next state S'. If S' is a relatively ideal state, it will obtain a "reward", otherwise it will be "punished". According to the reward and punishment situation, it is judged whether the neural network parameters from S to A are appropriate. If not, the parameters are adjusted according to the reward and punishment values. In the initial stage of neural network training, since the mapping relationship from S to A is almost random, the probability of obtaining a "reward" during the operation of the intelligent agent is very small. This sparse reward largely leads to slow training of the neural network and difficulty in converging to an ideal state.

[0004] The intelligent agent uses sensors to observe the environment to obtain a multi-dimensional state S. Commonly used sensors are cameras and lidars. Among them, the camera obtains planar data and cannot observe depth information; using a binocular camera can calculate depth information, but this method has a large error; the lidar can directly observe the accurate three-dimensional information of the environment. Whether it is the image data obtained by the camera or the point cloud information obtained by the lidar, the data volume is too large, and it needs to be preprocessed to minimize the dimension of the state S.

[0005] Due to the complexity of the navigation control task of autonomous vehicles, there are many neuron parameters in the neural network for this task that need to be adjusted. Therefore, the training of the neural network is a very time-consuming task, usually requiring hundreds of thousands or even millions of training times. In response to the problems of low exploration efficiency and difficulty in quickly converging to an ideal state in traditional DRL methods, the current optimization solutions are mainly divided into two categories. The first category is to design different decision-making or control tasks by combining multiple reinforcement learning methods. For example, Y. Xiao et al. used a network with a discrete action space in the vehicle end-to-end navigation control task to make vehicle behavior decisions (turn left, turn right, follow the vehicle, etc.), and then switched different continuous action networks according to the selected action to output vehicle control instructions. L. Chen et al. also used a similar solution. This way of combining neural networks draws on hierarchical reinforcement learning. This processing method can often effectively improve the comprehensive performance of the DRL model, but the running speed will slow down. The second is to combine traditional control theory to correct the output of decision values during the DRL training process. For example, Xie L et al. used a random switch to randomly select and output actions of a PID, OA (obstacle avoidance) controller, and a DDPG decision maker, thereby solving the problem of high variance suffered by DDPG when applied to complex real-world environments. This method can often effectively improve the training speed of the neural network. However, the neural network is a black box system. How to effectively integrate traditional control methods into the DRL method to improve the accuracy of autonomous vehicle navigation control remains a key problem to be solved. Summary of the Invention

[0006] The object of the present invention is to provide a navigation control method and system for autonomous vehicles, which can improve the accuracy of autonomous vehicle navigation control.

[0007] To achieve the above object, the present invention provides the following solutions:

[0008] A navigation control method for autonomous vehicles, comprising:

[0009] Obtain vehicle and environmental state data; the vehicle and environmental state data includes: point cloud data of obstacles within a set distance of the vehicle obtained by a lidar and pose information of the vehicle obtained by an inertial navigator;

[0010] Navigate according to the vehicle and environmental state data using an optimized navigation model; the training process of the optimized navigation model is as follows:

[0011] Based on the point cloud data of the obstacle, a navigation control algorithm is used to determine the vehicle control quantity; the vehicle control quantity includes: the throttle action value, the steering action value, and the brake action value; the navigation control algorithm includes: the DWA algorithm.

[0012] Based on the vehicle control quantity determined by the navigation control algorithm and the vehicle and environmental state data, a navigation model is constructed using an Actor-Critic type DRL decision network; the Actor-Critic type DRL decision network includes: a DRL Actor network and a DRL Critic network; the DRL Actor network is used to output a first vehicle control quantity according to the vehicle and environmental state data and the vehicle control quantity of the optimal path; the DRL Critic network is used to output the expected rewards corresponding to the vehicle control quantity determined by the navigation control algorithm and the first vehicle control quantity according to the first vehicle control quantity, the vehicle control quantity determined by the navigation control algorithm, and the vehicle and environmental state data; the final control quantity is determined according to the expected rewards corresponding to the two sets of vehicle control quantities and output.

[0013] The navigation model is optimized using a reward and punishment mechanism to determine the optimized navigation model.

[0014] Optionally, the obtaining of the vehicle and environmental state data specifically includes:

[0015] Taking the vehicle as the center of the cylindrical surface, the point cloud data of the obstacle is reprojected and encoded to obtain a two-dimensional panoramic image.

[0016] Horizontal average pooling and vertical maximum pooling are performed on the two-dimensional panoramic image to obtain a 1*60 state matrix; the state matrix is used to represent the distance from each angle of the vehicle body to the obstacle.

[0017] Obtain the global path of the vehicle, and obtain the path point with the minimum distance value from the current position of the vehicle in the global path and the included angle with the vehicle heading.

[0018] Optionally, the using of the navigation control algorithm to determine the vehicle control quantity based on the point cloud data of the obstacle specifically includes:

[0019] Taking the current position of the vehicle as the origin, according to the current pose and target pose of the vehicle, a plurality of paths to be traveled are determined using a path planning algorithm.

[0020] Determine the optimal path according to the position of the target point and the point cloud data of the obstacle.

[0021] According to the optimal path, a vehicle control quantity is determined using a path tracking algorithm.

[0022] Optionally, the using of the reward and punishment mechanism to optimize the navigation model to determine the optimized navigation model specifically includes:

[0023] Use the formula to determine the reward function R;

[0024] where r done is the reward value obtained for completing the task. When the task is not completed, r done = 0, and r over is the penalty value when a collision occurs or the distance from the global path exceeds the set value. When no collision occurs or the distance from the global path does not exceed the set value, r over = 0, V angular_z is the angular velocity of the vehicle along the z-axis, λ1, λ2, λ3, λ4 are proportionality coefficients, speed is the forward linear velocity of the vehicle, and dis l is the Euclidean distance between the vehicle and the path point with the minimum distance from the vehicle's current position in the global path, and dis a is the angle between the global path and the vehicle's heading.

[0025] An autonomous vehicle navigation control system, comprising:

[0026] A data acquisition module for acquiring vehicle and environmental state data; the vehicle and environmental state data includes: point cloud data of obstacles within a set distance of the vehicle obtained using a lidar and pose information of the vehicle obtained using an inertial navigator;

[0027] A navigation module for navigating according to the vehicle and environmental state data using an optimized navigation model; the training process of the optimized navigation model is as follows:

[0028] According to the point cloud data of the obstacles, use a navigation control algorithm to determine the vehicle control amount; the navigation control algorithm includes: determining the route to be traveled using a local path planning algorithm and then using a path tracking algorithm to determine the vehicle control amount; the vehicle control amount includes: throttle action value, steering action value, and brake action value;

[0029] According to the vehicle control amount determined by the navigation control algorithm and the vehicle and environmental state data, use an Actor-Critic type DRL decision network to construct a navigation model; the Actor-Critic type DRL decision network includes: a DRL Actor network and a DRL Critic network; the DRL Actor network is used to output a first vehicle control amount according to the vehicle and environmental state data and the vehicle control amount of the optimal path; the DRL Critic network is used to output the expected return corresponding to the vehicle control amount determined by the navigation control algorithm and the first vehicle control amount according to the first vehicle control amount, the vehicle control amount determined by the navigation control algorithm, and the vehicle and environmental state data; determine the final control amount according to the expected returns corresponding to the two sets of vehicle control amounts and output it;

[0030] Optimize the navigation model using a reward and punishment mechanism to determine the optimized navigation model.

[0031] Optionally, the data acquisition module specifically includes:

[0032] A two-dimensional panoramic image determination unit, which takes the vehicle as the center of the cylindrical surface, reprojects and encodes the point cloud data of the obstacle to obtain a two-dimensional panoramic image;

[0033] A state matrix determination unit, which performs horizontal average pooling and vertical maximum pooling on the two-dimensional panoramic image to obtain a 1*60 state matrix; the state matrix is used to represent the distance from each angle of the vehicle body to the obstacle;

[0034] An included angle determination unit, which obtains the global path of the vehicle, and obtains the path point with the smallest distance value from the current position of the vehicle in the global path and the included angle with the vehicle's heading.

[0035] Optionally, the method for determining the vehicle control amount according to the point cloud data of the obstacle specifically includes:

[0036] Taking the current position of the vehicle as the origin, and using a path planning algorithm to determine multiple paths to be traveled according to the current pose and target pose of the vehicle;

[0037] Determine the optimal path according to the position of the target point and the point cloud data of the obstacle;

[0038] According to the optimal path, use a path tracking algorithm to determine the vehicle control amount.

[0039] Optionally, the method for optimizing the navigation model using a reward and punishment mechanism to determine the optimized navigation model specifically includes:

[0040] Use the formula to determine the reward function;

[0041] where r done is the reward value obtained for completing the task, when the task is not completed, r done = 0, r over is the penalty value when a collision occurs or the distance from the global path exceeds a set value, when no collision occurs or the distance from the global path does not exceed the set value, r over = 0, V angular_z is the angular velocity of the vehicle's z-axis, λ1, λ2, λ3, λ4 are proportionality coefficients, speed is the vehicle's forward linear velocity, dis l is the Euclidean distance between the vehicle and the path point with the smallest distance value from the current position of the vehicle in the global path, dis a is the included angle between the global path and the vehicle's heading.

[0042] According to the specific embodiments provided by the present invention, the following technical effects are disclosed by the present invention:

[0043] A navigation control method and system for an autonomous vehicle provided by the present invention integrates a traditional vehicle navigation control algorithm with a DRL decision network of the Actor-Critic type. During the process of guiding and training the DRL decision network of the Actor-Critic type by the traditional vehicle navigation control algorithm, the parameter adjustment method (algorithm) and network structure of the DRL algorithm are not involved. Therefore, it will not affect the characteristics and mechanisms of various Actor-Critic types of DRL algorithms, that is, it can greatly improve the convergence speed of the neural network while retaining the original characteristics of the DRL algorithm, reduce the convergence difficulty, improve the actual performance of the DRL model, and further improve the accuracy of the navigation control of the autonomous vehicle. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0045] Figure 1 It is a schematic flow chart of a navigation control method for an autonomous vehicle provided by the present invention;

[0046] Figure 2 It is a schematic overall flow chart of a navigation control method for an autonomous vehicle provided by the present invention;

[0047] Figure 3 It is a schematic diagram of the global path guiding the vehicle to travel;

[0048] Figure 4 It is a schematic diagram of the DRL Actor network structure;

[0049] Figure 5 It is a schematic diagram of the DRL Critic network structure. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0050] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.

[0051] The object of the present invention is to provide a navigation control method and system for an autonomous vehicle, which can improve the accuracy of the autonomous driving navigation control in the DRL (Actor-Critic type) manner, increase the neural network training convergence speed, and reduce the neural network training convergence difficulty.

[0052] In order to make the above objects, features and advantages of the present invention more obvious and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0053] Figure 1 It is a schematic flow chart of a navigation control method for an autonomous vehicle provided by the present invention. Figure 2 It is a schematic overall flow chart of a navigation control method for an autonomous vehicle provided by the present invention. As Figure 1 and Figure 2 shown, a navigation control method for an autonomous vehicle provided by the present invention includes:

[0054] S101, obtaining vehicle and environmental state data; the vehicle and environmental state data includes: point cloud data of obstacles within a set distance of the vehicle obtained by a lidar and pose information of the vehicle obtained by an inertial navigator; the point cloud data of obstacles within a set distance of the vehicle output by the lidar is filtered by the ground.

[0055] S101 specifically includes:

[0056] Taking the vehicle as the center of the cylindrical surface, re-projecting and encoding the point cloud data of the obstacles to obtain a two-dimensional panoramic image with depth information.

[0057] Performing horizontal average pooling and vertical maximum pooling on the two-dimensional panoramic image to obtain a 1*60 state matrix φ d ; the state matrix is used to represent the distance from each angle of the vehicle body to the obstacle.

[0058] Obtaining the global path of the vehicle, and obtaining the path point with the smallest distance value from the current position of the vehicle in the global path and the included angle with the vehicle heading.

[0059] As Figure 3As shown, the global path is pre-planned. The global path consists of multiple equidistant path points. Determine the position deviation between the vehicle and the global path based on the path point closest to the vehicle in the global path and the vehicle's current position. Determine the angle between the vehicle's current orientation and the direction of the global path based on the orientation of the line connecting the 2nd and 3rd path points after the path point closest to the vehicle in the global path. The global path is the route from the starting point to the ending point. The acquisition method is as follows: Use Unity or corresponding high-precision map drawing software to draw the route map of a certain area. After obtaining the area map, according to the positions of the starting point and the ending point, obtain an optimal route following the principle of the shortest distance or the shortest travel time. This route is the global path.

[0060] Among the global path points, the point closest to the vehicle's current position is c0, which is the Current waypoint. Here, define the Euclidean distance between c0 and the vehicle's coordinate origin as dis. i Define the path points 2m and 3m after c0 as c2 and c3. The vector direction from c2 to c3 The angle between the vehicle's heading is dis. a Among them, when the vehicle is on the left side of the global path, dis l Is a positive value. When the vehicle is on the right side of the global path, dis l Is a negative value. If yaw is relative to Left deviation, then dis a Is a positive value. Yaw is relative to Right deviation, then dis a Is a negative value. The advantage of this design is that when the vehicle is left-deviated relative to the global path, dis l And dis a The value will increase. Passing through the fully connected layer of the DRL network will more easily increase the value of steer (steering action), and then make the vehicle turn right.

[0061] The data volume of each frame of lidar is as high as 1 million to 10 million data points. Directly using it for the neural network will generate a huge amount of computation. First, use the ground removal algorithm to filter out unnecessary point clouds, and then remap the remaining point clouds (obstacle point clouds) into a one-dimensional array of 60 numbers, which can greatly reduce the amount of computation. At the same time, each number represents the distance to the obstacle in the direction of every 3° interval around the vehicle body, and can effectively retain the necessary information.

[0062] The mutual relationship between the vehicle and the global path is represented only by dis i And dis a These two data, and the changes in these two values can be effectively mapped to the action values of the vehicle. That is, when the vehicle is left-deviated relative to the global path, dis l And dis aThe value will increase, and it will be easier for the steer value to increase after passing through the fully connected layer of the drl neural network, which will then cause the vehicle to turn right, and vice versa. This can effectively reduce the number of input parameters, reduce the number of neurons, and thus reduce the training time.

[0063] S102, perform navigation using the optimized navigation model according to the vehicle and environmental state data; the training process of the optimized navigation model is as follows:

[0064] According to the point cloud data of the obstacle, use the navigation control algorithm to determine the vehicle control quantity O of the optimal path dwa ; the vehicle control quantity O dwa includes: throttle action value, steering action value, and brake action value; the navigation control algorithm is an algorithm that, in a non-self-learning manner, obtains the vehicle control quantity through a calculation method of local path planning and path tracking based on the vehicle's current position, target position, and the surrounding environment; the navigation control algorithm includes but is not limited to the DWA algorithm.

[0065] Construct a navigation model using the Actor-Critic type of DRL decision network according to the vehicle control quantity determined by the navigation control algorithm and the vehicle and environmental state data; the Actor-Critic type of DRL decision network includes: DRL Actor network and DRL Critic network; the DRL Actor network is used to output the first vehicle control quantity according to the vehicle and environmental state data and the vehicle control quantity of the optimal path; the DRL Critic network is used to output the vehicle control quantity determined by the traditional navigation control algorithm and the expected return corresponding to the first vehicle control quantity according to the first vehicle control quantity, the vehicle control quantity determined by the navigation control algorithm, and the vehicle and environmental state data; determine the final control quantity according to the expected returns corresponding to the two sets of vehicle control quantities and output it;

[0066] Optimize the navigation model using the reward and punishment mechanism to determine the optimized navigation model.

[0067] The determination of the vehicle control quantity according to the point cloud data of the obstacle using the navigation control algorithm, taking the improved DWA algorithm as an example, specifically includes:

[0068] Taking the vehicle's current position as the origin, use the path planning algorithm to determine multiple paths to be traveled according to the vehicle's current pose and target pose;

[0069] Use the formula to determine the turning angle δ i ;

[0070] where α is the maximum steering angle of the vehicle, and δ i being negative represents a left turn, and δ iPositive means turning right.

[0071] Determine the optimal path Path based on the position of the target point and the point cloud data of the obstacle best ;

[0072] Path best is selected in the following way:

[0073] Select the path point at a certain distance behind the current path point c0 of the vehicle as Target

[0074] Target = c0 + R + V × k target (2)

[0075] where R is the minimum turning radius of the vehicle, V is the forward speed of the vehicle, and k target is the proportionality coefficient.

[0076] For each driving trajectory, define the last trajectory point as P i , select the center point of the vehicle head as P0, then the included angle θ i = ∠P i P0T arget , and then select the optimal path Path best

[0077] Path best = Path j if θ j = min(θ i , i ∈ [1, n]) (3)

[0078] According to the optimal path, determine the vehicle control quantity O using the following formula dwa .

[0079] The vehicle control quantity O dwa is as follows:

[0080]

[0081] All states S required by the DWA algorithm dua , and all states S required by the Actor-Critic type DRL decision network drl . Where speed is the vehicle's forward linear speed.

[0082]

[0083] The Actor-Critic type DRL algorithm includes two types of neural networks. One is the DRL Actor network, and the input parameter is the observation value S used to characterize the current environment and state of the vehicle drl ; the output is the set of action values O drl, where the action values include throttle, steering, and braking, and the range of the action values is:

[0084]

[0085] Another type is the DRL Critic network. The input values of this network include the observation value s representing the current environment and state of the vehicle and the action value a corresponding to the observation value s; the output value is the expected return J of the corresponding relationship of [s, a]. The structures of the two types of neural networks are as Figure 4 and Figure 5 shown.

[0086] In the Actor-Critic type of DRL algorithm, the Actor network is used to calculate the action value a of vehicle control. Therefore, the training goal of the Actor network is to select appropriate parameters for each neuron in the Actor network to complete the navigation control task of autonomous driving. The Critic network is used to calculate the return that will be brought by executing a in the current environment s, that is, to evaluate this set of actions. Therefore, the training goal of the Critic network is to correctly evaluate the mapping of s and a. The basis for whether the evaluation is correct comes from the artificially set reward function, and the result of the evaluation will be used to optimize the parameters of the neural network.

[0087] The reward function is that after the vehicle executes the action a in the current environment s, people "reward" (meeting expectations) or "punish" (not meeting expectations or causing harm) this action.

[0088] Use the formula to determine the reward function;

[0089] where r done is the reward value obtained for completing the task. When the task is not completed, r done = 0, r over is the penalty value when a collision occurs or the distance from the global path exceeds the set value. When no collision occurs or the distance from the global path does not exceed the set value, r over = 0, V angular_z is the angular velocity of the vehicle along the z-axis, λ1, λ2, λ3, λ4 are proportionality coefficients, speed is the forward linear velocity of the vehicle, dis l is the Euclidean distance between the vehicle and the path point with the minimum distance from the vehicle's current position in the global path, dis a is the angle between the global path and the vehicle's heading.

[0090] As Figure 2 shown, the vehicle action values O dwa and O drl obtained by the DWA algorithm and the DRL module are respectively obtained. These two action values and the current state s are respectively input into the Critic network to obtain two expected values Jdwa With J drl 。Then the selector compares J dwa With J drl values, and takes the one with the larger expected value in O dwa With O drl as the vehicle action a output. After the action a is executed, the vehicle will enter a new state S′ drl 。Subsequently, calculate the reward value R obtained by this action, and store the (S td3 , a, S′ td3 , R) parameters before and after each action execution for adjusting the neural network parameters.

[0091] Perform multiple trainings until the vehicle performance converges.

[0092] In the initial stage of DRL neural network training, due to the initial parameters of the Actor network, the vehicle actions are nearly random, and the vehicle control effect is bound to be very poor. The actions output by the DWA algorithm are artificially set, and the control effect in the initial stage is much better than that of the DRL algorithm. And the initial parameters of the Critic network are random, so the action selection for the DWA algorithm and the DRL module is nearly random, so O dwa With O drl will be randomly selected for action output. Compared with the single DRL algorithm, the control effect of the whole set of algorithms is better. A better control effect means that the vehicle is more likely to obtain rewards during the training process, thus solving the problem that the DRL algorithm is difficult to converge and converges slowly due to sparse rewards in the initial stage.

[0093] In the later stage of DRL neural network training, due to the gradual improvement of the neural network parameters, the probability of the result of the DWA algorithm being selected will become smaller and smaller. In this mode, the DWA algorithm will not affect the final effect of the DRL module, but only provides help for DRL in the initial stage.

[0094] Since the parameter adjustment method (algorithm) and network structure of the DRL algorithm are not involved in the process of the DWA algorithm guiding the training of DRL, it will not affect the characteristics and mechanisms of various Actor-Critic DRL algorithms. That is, it can greatly improve the convergence speed of the neural network while retaining the characteristics of the original DRL algorithm, and at the same time reduce the convergence difficulty.

[0095] Corresponding to the above embodiments, the present invention also provides an autonomous driving vehicle navigation control system, including:

[0096] A data acquisition module for acquiring vehicle and environmental state data; the vehicle and environmental state data includes: point cloud data of obstacles within the set distance of the vehicle obtained by lidar and pose information of the vehicle obtained by an inertial navigator;

[0097] A navigation module, which is used to perform navigation according to the vehicle and environmental state data by using an optimized navigation model; the training process of the optimized navigation model is as follows:

[0098] According to the point cloud data of the obstacle, a navigation control algorithm is used to determine the vehicle control amount; the vehicle control amount includes: throttle action value, steering action value and brake action value; the navigation control algorithm includes: DWA algorithm.

[0099] According to the vehicle control amount determined by the navigation control algorithm and the vehicle and environmental state data, a navigation model is constructed by using an Actor-Critic type DRL decision network; the Actor-Critic type DRL decision network includes: DRL Actor network and DRL Critic network; the DRL Actor network is used to output the first vehicle control amount according to the vehicle and environmental state data and the vehicle control amount of the optimal path; the DRL Critic network is used to output the expected return corresponding to the vehicle control amount determined by the navigation control algorithm and the first vehicle control amount according to the first vehicle control amount, the vehicle control amount determined by the navigation control algorithm and the vehicle and environmental state data; the final control amount is determined according to the expected returns corresponding to the two groups of vehicle control amounts and output.

[0100] The navigation model is optimized by using a reward and punishment mechanism to determine the optimized navigation model.

[0101] The data acquisition module specifically includes:

[0102] A two-dimensional panoramic image determination unit, which is used to re-project and encode the point cloud data of the obstacle with the vehicle as the center of the cylindrical surface to obtain a two-dimensional panoramic image.

[0103] A state matrix determination unit, which is used to perform horizontal average pooling and vertical maximum pooling on the two-dimensional panoramic image to obtain a 1*60 state matrix; the state matrix is used to represent the distance from each angle of the vehicle body to the obstacle.

[0104] An included angle determination unit, which is used to obtain the global path of the vehicle, and obtain the path point with the minimum distance value from the current position of the vehicle in the global path and the included angle with the vehicle heading.

[0105] The specific process of determining the vehicle control amount of the optimal path according to the point cloud data of the obstacle by using the navigation control algorithm includes:

[0106] Taking the current position of the vehicle as the origin, according to the current pose and target pose of the vehicle, a path planning algorithm is used to determine multiple paths to be traveled.

[0107] Determine the optimal path based on the position of the target point and the point cloud data of the obstacle;

[0108] According to the optimal path, use the path tracking algorithm to determine the vehicle control quantity.

[0109] The use of a reward and punishment mechanism to optimize the navigation model to determine the optimized navigation model specifically includes:

[0110] Use the formula To determine the reward function;

[0111] Where r done Is the reward value obtained for completing the task. When the task is not completed, r done = 0, r over Is the penalty value when a collision occurs or the distance from the global path exceeds the set value. When no collision occurs or the distance from the global path does not exceed the set value, r over = 0, V angular_z Is the angular velocity of the vehicle along the z-axis, λ1, λ2, λ3, λ4 are proportionality coefficients, speed is the forward linear velocity of the vehicle, dis l Is the Euclidean distance between the vehicle and the path point with the minimum distance value from the vehicle's current position in the global path, dis a Is the angle between the global path and the vehicle's heading.

[0112] The various embodiments in this specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. The same or similar parts among the various embodiments can be referred to each other. For the system disclosed in the embodiments, since it corresponds to the method disclosed in the embodiments, the description is relatively simple. For the relevant parts, refer to the description in the method section.

[0113] Specific examples are used in this article to elaborate on the principles and implementation methods of the present invention. The descriptions of the above embodiments are only used to help understand the method of the present invention and its core idea; at the same time, for those of ordinary skill in the art, based on the idea of the present invention, there will be changes in the specific implementation methods and application scopes. In summary, the content of this specification should not be construed as a limitation to the present invention.

Claims

1. A navigation control method for an autonomous vehicle, characterized in that, Including: Obtain vehicle and environmental state data; The vehicle and environmental state data includes: point cloud data of obstacles within a set distance of the vehicle obtained by lidar and pose information of the vehicle obtained by an inertial navigator; Navigate according to the vehicle and environmental state data using an optimized navigation model; The training process of the optimized navigation model is as follows: According to the point cloud data of the obstacles, determine the vehicle control quantity using a navigation control algorithm; The vehicle control quantity includes: throttle action value, steering action value, and brake action value; According to the vehicle control quantity determined by the navigation control algorithm and the vehicle and environmental state data, construct a navigation model using an Actor-Critic type DRL decision network; The Actor-Critic type DRL decision network includes: a DRL Actor network and a DRL Critic network; The DRL Actor network is used to output a first vehicle control quantity according to the vehicle and environmental state data and the vehicle control quantity of the optimal path; The DRL Critic network is used to output the expected reward corresponding to the vehicle control quantity determined by the navigation control algorithm and the first vehicle control quantity according to the first vehicle control quantity, the vehicle control quantity determined by the navigation control algorithm, and the vehicle and environmental state data; Determine the final control quantity according to the expected rewards corresponding to the two sets of vehicle control quantities and output it; Optimize the navigation model using a reward and punishment mechanism to determine the optimized navigation model; The specific steps of determining the vehicle control quantity of the optimal path according to the point cloud data of the obstacles using a navigation control algorithm include: Taking the current position of the vehicle as the origin, determine multiple paths to be traveled using a path planning algorithm according to the current pose and target pose of the vehicle; Determine the optimal path according to the position of the target point and the point cloud data of the obstacles; According to the optimal path, determine the vehicle control quantity using a path tracking algorithm.

2. The method for navigating and controlling an autonomous vehicle according to claim 1, wherein The specific steps of obtaining the vehicle and environmental state data include: Taking the vehicle as the center of the cylindrical surface, perform reprojection and encoding on the point cloud data of the obstacles to obtain a two-dimensional panoramic image; Perform horizontal average pooling and vertical maximum pooling on the two-dimensional panoramic image to obtain a 1*60 state matrix; The state matrix is used to represent the distance from each angle of the vehicle body to the obstacles; Obtain the global path of the vehicle, and obtain the path point with the minimum distance value from the current position of the vehicle in the global path and the included angle with the vehicle heading.

3. The method for navigating and controlling an autonomous vehicle according to claim 2, wherein The specific steps of optimizing the navigation model using a reward and punishment mechanism to determine the optimized navigation model include: Using the formula Determine the reward function R; where r done is the reward value obtained for completing the task. When the task is not completed, r done = 0, r over is the penalty value when a collision occurs or the distance from the global path exceeds the set value. When no collision occurs or the distance from the global path does not exceed the set value, r over = 0, V angular_z is the angular velocity of the vehicle along the z-axis, λ1, λ2, λ3, λ4 are proportionality coefficients, speed is the forward linear velocity of the vehicle, dis l is the Euclidean distance between the vehicle and the path point with the minimum distance from the vehicle's current position in the global path, dis a is the angle between the global path and the vehicle's heading.

4. An autonomous vehicle navigation control system, characterized in that, Including: A data acquisition module for obtaining vehicle and environmental state data; The vehicle and environmental state data includes: point cloud data of obstacles within a set distance of the vehicle obtained by lidar and pose information of the vehicle obtained by an inertial navigator; A navigation module for navigating according to the vehicle and environmental state data using an optimized navigation model; The training process of the optimized navigation model is as follows: According to the point cloud data of the obstacles, determine the vehicle control quantity using a navigation control algorithm; The vehicle control quantity includes: throttle action value, steering action value, and brake action value; Based on the vehicle control quantity determined by the navigation control algorithm and the vehicle and environmental state data, a navigation model is constructed using a DRL decision network of the Actor-Critic type; the DRL decision network of the Actor-Critic type includes: a DRL Actor network and a DRL Critic network; the DRL Actor network is used to output a first vehicle control quantity according to the vehicle and environmental state data and the vehicle control quantity of the optimal path; the DRL Critic network is used to output the vehicle control quantity determined by the navigation control algorithm and the expected return corresponding to the first vehicle control quantity according to the first vehicle control quantity, the vehicle control quantity determined by the navigation control algorithm, and the vehicle and environmental state data; determine the final control quantity according to the expected returns corresponding to the two sets of vehicle control quantities, and output it; Optimize the navigation model using a reward and punishment mechanism to determine the optimized navigation model; The method for determining the vehicle control quantity according to the point cloud data of the obstacle by using the navigation control algorithm specifically includes: Taking the current position of the vehicle as the origin, and using a path planning algorithm to determine multiple paths to be traveled according to the current pose and the target pose of the vehicle; Determine the optimal path according to the position of the target point and the point cloud data of the obstacle; According to the optimal path, use a path tracking algorithm to determine the vehicle control quantity.

5. An automatic driving vehicle navigation control system according to claim 4, characterized in that, The data acquisition module specifically includes: A two-dimensional panoramic image determination unit, which is used to re-project and encode the point cloud data of the obstacle with the vehicle as the center of the cylindrical surface to obtain a two-dimensional panoramic image; A state matrix determination unit, which is used to perform horizontal average pooling and vertical maximum pooling on the two-dimensional panoramic image to obtain a 1*60 state matrix; the state matrix is used to represent the distance from each angle of the vehicle body to the obstacle; An angle determination unit, which obtains the global path of the vehicle, and obtains the path point with the minimum distance value from the current position of the vehicle in the global path and the included angle with the vehicle heading.

6. An automatic driving vehicle navigation control system according to claim 5, characterized in that, The method for optimizing the navigation model using a reward and punishment mechanism to determine the optimized navigation model specifically includes: Using the formula Determine the reward function; where r done is the reward value obtained for completing the task. When the task is not completed, r done = 0, r over is the penalty value when a collision occurs or the distance from the global path exceeds the set value. When no collision occurs or the distance from the global path does not exceed the set value, r over = 0, V angular_z is the angular velocity of the vehicle along the z-axis, λ1, λ2, λ3, λ4 are proportionality coefficients, speed is the forward linear velocity of the vehicle, dis l is the Euclidean distance between the vehicle and the path point with the minimum distance from the vehicle's current position in the global path, dis a is the angle between the global path and the vehicle's heading.

Citation Information

Patent Citations

  • System and method for automated driving

    CN111338333A

  • Unmanned vehicle path planning method based on improved A * algorithm and deep reinforcement learning

    CN111780777A