An automatic parking path planning method for multiple storage site scenarios
Patent Information
- Application Number
- CN202211398938.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-09
- Publication Date
- 2026-09-25
- Estimated Expiration
- 2042-11-09
AI Technical Summary
现有的混合A*算法、RRT算法、lattice算法、人工势场法均无法单独运用实现多段路径的规划.
[0055]1)统一性:本发明通过坐标转换并在新坐标系下进行规划,能够利用统一的算法完成在不同库位、不同初始位姿下的规划。
Smart Images

Figure CN115817455B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of driver assistance technology, and in particular to an automatic parking path planning method for various parking space scenarios. Background Technology
[0002] Automated parking technology, as a representative technology of automotive intelligence, has received high attention from universities and enterprises. High-precision automated parking path planning algorithms not only allow vehicles to park without excessive safety margins, thus improving parking space planning and increasing land use efficiency, but can also be organically combined with future technologies such as automated charging that require precise vehicle parking.
[0003] Compared to typical autonomous driving scenarios, path planning for automated parking has the following requirements: ① It should be able to plan segmented paths to handle multi-segment parking and parallel parking maneuvers; ② The planning should achieve high accuracy as much as possible while ensuring safety, so as to fully utilize space and cooperate with future charging pile technologies; ③ The planning should be completed using a unified planning algorithm as much as possible to facilitate algorithm simplification; ④ As vehicles become more intelligent in the future, the algorithm itself should have the potential for self-updating, so that it can continuously optimize itself like a human. Existing hybrid A* algorithm, RRT algorithm, lattice algorithm, and artificial potential field method cannot be used alone to plan multi-segment paths.
[0004] Chinese patent CN 113830079 A proposes a parking path planning algorithm that combines the hybrid A* algorithm with CC curve splicing. Although it can achieve planning functions under different parking locations, the design of the cost function of the hybrid A* algorithm requires a lot of preliminary calculation work, and different CC curve splicing forms still need to be designed for different parking locations. It does not fully realize the uniformity of the planning process under different parking locations.
[0005] Chinese patent CN 114906128 A uses reinforcement learning for parallel parking planning. Although it retains the ability to self-update, the neural network is limited to parallel parking scenarios for training, which restricts its versatility in other scenarios. At the same time, the training data used by the network has an expansion step size of 0.04m. This will cause the amount of data to surge when the number of initial poses increases, forcing the neural network to become more complex to handle more data. A more complex neural network corresponds to a longer training time and a greater amount of computation, which will also increase the computational burden on the vehicle system.
[0006] In summary, the current algorithm cannot fully meet the four requirements mentioned above, and developing an algorithm that can meet these requirements is of great significance. Summary of the Invention
[0007] The purpose of this invention is to overcome the shortcomings of the existing technology and provide an automatic parking path planning method for various parking space scenarios, thus meeting the above requirements.
[0008] The objective of this invention can be achieved through the following technical solutions:
[0009] An automated parking path planning method for various parking space scenarios, the method comprising the following steps:
[0010] Obtain the storage location type, storage location corner coordinates, and initial pose;
[0011] Coordinate transformation is performed based on parking space type to unify the coordinate system for vehicles in parallel, diagonal, and perpendicular parking scenarios;
[0012] Under a unified coordinate system, a sequence of scattered points for coarse-planned paths is obtained through a neural network;
[0013] The planned path is obtained from the scatter sequence of coarse-planned paths through post-processing including simulation tracking, DWA, and end smoothing.
[0014] As a preferred technical solution, the specific steps of the method for coordinate transformation and unifying the coordinate system include:
[0015] The vehicle's initial pose in the original coordinate system is Target pose is Parallel parking uses the parking position as the target position; angled and perpendicular parking use the final position as the target position.
[0016] The coordinate transformation is performed with the vehicle's target pose in the new coordinate system as (0,0,pi / 2). The new coordinate system is offset in the x and y directions compared to the original coordinate system. Rotation
[0017] The initial pose and storage location corners in the new coordinate system are obtained based on the coordinate system transformation relationship;
[0018] The initial pose (x0, y0, θ0) in the new coordinate system and the initial pose in the original coordinate system and target pose The coordinate transformation formula is:
[0019]
[0020] In the formula, and These are the horizontal and vertical coordinates and the angle of the initial pose in the original coordinate system, respectively. and x0, y0, and θ0 are the x and y coordinates and angles of the target pose in the original coordinate system, respectively; x0, y0, and θ0 are the x and y coordinates and angles of the initial pose in the new coordinate system, respectively.
[0021] As a preferred technical solution, the step of obtaining the scatter sequence of coarse-planned paths through a neural network includes:
[0022] Using the vehicle's current pose and the corner point of the parking space after coordinate transformation as input to the neural network, the output is a combination of motion with the extended direction π and the extended curvature k;
[0023] Based on the action combinations output by the neural network, the vehicle extends to the next state in large steps using geometric inference; until the vehicle extends from the initial pose to the target pose.
[0024] As a preferred technical solution, the training steps of the neural network include:
[0025] The neural network is used to learn by imitation using path data based on RS curve splicing;
[0026] By repeatedly setting different storage location categories and initial poses, the neural network is used to explore and optimize, and data with higher reward values is obtained to update the network again, thus completing reinforcement learning.
[0027] As a preferred technical solution, the path data based on RS curve splicing is generative data, and its generation process is as follows:
[0028] Based on human drivers' parking experience in parallel, oblique, and perpendicular parking spaces, various RS curve splicing forms are defined;
[0029] By utilizing the positional relationship between different initial poses and target poses under different storage locations, the length and curvature of the RS curve are obtained through geometric calculation;
[0030] The positional information and corresponding action categories needed to train the neural network are generated based on RS curves.
[0031] As a preferred technical solution, the return value R total The function is:
[0032] R total =R change_dir +R space +R lengt +R reach +R colli +R change_cur
[0033] In the formula:
[0034] R chang_dir The shift reward value decreases with more shifts.
[0035] R space The more space is utilized, the lower the return on path space utilization.
[0036] R lengt This is the path length reward value; the longer the path length, the lower the path length reward value.
[0037] R reac The final posture evaluation value represents whether the target pose has been achieved. A reward value can be obtained if the pose is within the allowable range. If the target pose can be matched with a smaller error on the basis of the previous value, an additional reward value can be obtained.
[0038] R colli The collision reward value is significantly reduced if a collision occurs during the expansion process.
[0039] R chang_cur The return value is the curvature change value; the greater the curvature change during the expansion process, the lower the return value.
[0040] As a preferred technical solution, the simulation tracking method is as follows:
[0041] Using the vehicle kinematics model, a lateral control state equation is established. The front wheel steering angle is designed through feedback and feedforward. Considering vehicle dynamics constraints, the vehicle's tracking process of coarse-planned path points is simulated. The kinematics model is expanded with small steps, and the path point sequence obtained by the simulation tracking is used as the planning data.
[0042] The kinematic model extension starts from the target pose and extends in reverse towards the initial pose.
[0043] As a preferred technical solution, the steps of the DWA include:
[0044] Sampling was performed on the front wheel steering angle of vehicles within the permitted range;
[0045] Set a fixed extension length and compare the reward values of the extension paths corresponding to different sampling results; select the value r. total The maximum value is the ideal front wheel steering angle at this moment;
[0046] Where the value r total The function is defined as:
[0047] r total =r follow_pat +r avoid_obstacle
[0048] In the formula:
[0049] r follow_path The score represents the degree of conformity between the sampled value and the angle value calculated by the simulation tracking. The smaller the deviation between the sampled value and the angle value calculated by the simulation tracking, the higher the degree of conformity.
[0050] r avoid_obstacleThis represents the vehicle collision report value after the expansion is complete. It is calculated using the distance between the vehicle and the obstacle. The smaller the distance, the lower the collision report value.
[0051] As a preferred technical solution, the end-effector smoothing involves performing two simulated tracking segments with the midline between the exploration termination position and the target pose as the boundary:
[0052] The first segment continues to expand in the opposite direction, extending a certain distance using simulated tracking to eliminate the lateral deviation between the vehicle and the centerline; the second segment is in the opposite direction to the first segment, eliminating the lateral deviation between the vehicle and the target pose.
[0053] As a preferred technical solution, when planning parallel parking paths, an additional scatter sequence of parking path points based on geometric circular arc curves is added after the planned path is obtained in post-processing.
[0054] Compared with the prior art, the present invention has the following beneficial effects:
[0055] 1) Uniformity: This invention, through coordinate transformation and planning in the new coordinate system, can complete planning under different storage locations and different initial poses using a unified algorithm.
[0056] 2) Accuracy: Because the present invention extends the path from the target pose in reverse during path post-processing, it achieves zero error of the path relative to the target pose.
[0057] 3) Self-updating ability: This invention uses neural networks, which have the ability to explore and have the opportunity to explore better data, which means that the algorithm has the potential to self-update.
[0058] 4) Low-cost implementation: This invention specifies that the neural network expansion step size is 1m, which allows for the use of less data to cover more working conditions during training without increasing the complexity of the neural network. Attached Figure Description
[0059] Figure 1 This is a schematic diagram of the overall principle of the present invention;
[0060] Figure 2 The effect of coordinate transformation under different storage locations;
[0061] Figure 3 This is a summary diagram of the effects of coordinate transformation on different storage locations;
[0062] Figure 4 A schematic diagram illustrating the method for calculating the pose of a parallel parking vehicle entering a parking space;
[0063] Figure 5 Schematic diagrams of four RS curve splicing methods;
[0064] Figure 6A schematic diagram for calculating the dimensions of a curve consisting of a straight line, an arc, and a straight line;
[0065] Figure 7 This is a schematic diagram of a dataset used for imitation learning based on RS curves.
[0066] Figure 8 A schematic diagram showing the corner points of obstacle vehicles in different parking spaces;
[0067] Figure 9 A schematic diagram illustrating the rolling exploration process for enhanced learning;
[0068] Figure 10 This is a schematic diagram illustrating the simulated tracking of vehicle anchor points by a neural network.
[0069] Figure 11 This is a schematic diagram of lateral deviation during vehicle simulation tracking.
[0070] Figure 12 A diagram illustrating the vehicle being aligned with the anchor point.
[0071] Figure 13 This is a schematic diagram illustrating the principle of DWA (Dynamic Window Method).
[0072] Figure 14 A schematic diagram illustrating the use of an envelope circle to cover the outline of an octagonal vehicle;
[0073] Figure 15 This is a schematic diagram illustrating the principle of end-smoothing.
[0074] Figure 16 A schematic diagram of the parallel parking planning results;
[0075] Figure 17 This is a schematic diagram of the angled parking planning results;
[0076] Figure 18 This is a schematic diagram of the vertical parking planning results. Detailed Implementation
[0077] The present invention will now be described in detail with reference to the accompanying drawings and specific embodiments. These embodiments are based on the technical solution of the present invention and provide detailed implementation methods and specific operating procedures. However, the scope of protection of the present invention is not limited to the following embodiments.
[0078] This invention proposes an automatic parking path planning method for various parking space scenarios, such as... Figure 1 As shown, the entire method is divided into a coordinate unification module and a path planning module.
[0079] As previously mentioned, future automatic parking path planning algorithms should meet four requirements: ① They should be able to plan segmented paths to handle multi-segment parking and parallel parking maneuvers; ② The planning should achieve high accuracy while ensuring safety, so as to fully utilize space and cooperate with future charging pile technologies; ③ The planning should be completed using a unified planning algorithm as much as possible to facilitate algorithm simplification; ④ As vehicles become more intelligent in the future, the algorithm itself should have the potential for self-updating, so that it can continuously optimize itself like a human. Based on requirement ①, methods that are not suitable for multi-segment path planning alone, such as hybrid A* algorithms, RRT algorithms, lattice algorithms, and artificial potential field methods, will not be used. The only methods that can be considered are traditional curve stitching methods, such as the CC curve stitching planning method proposed in Chinese patent CN 113830079 A, and machine learning methods, such as the method proposed in Chinese patent CN 114906128 A. Based on requirement ④, traditional curve stitching lacks self-updating capabilities, therefore, machine learning methods are the most suitable. However, this algorithm is limited to parallel parking scenarios; its network can only be applied to parallel parking. Secondly, the network is trained using trajectory point data at 5ms intervals, resulting in a large amount of data for a single scenario, and the limited network structure cannot handle training data for multiple scenarios. Furthermore, the trajectory points are entirely obtained through network exploration; to ensure the exploration ends, a relatively broad termination condition needs to be set, but such a broad termination condition cannot guarantee the accuracy of the planning results.
[0080] Based on the above analysis, the use of neural networks can satisfy requirements ① and ④, while reinforcement learning is an effective means to achieve superior network performance. In reinforcement learning, the value of data is determined by the reward function, and the degree of fit between the planned termination pose and the target pose is extremely important in the reward function. Since a single neural network should only be trained based on one reward function, and the termination condition should also be unique, this means that the target poses for parallel, oblique, and perpendicular parking must also be uniform. In this way, the neural network can explore data in different parking spaces according to a uniform standard and explore according to a uniform termination condition, thus being applicable to different parking spaces and achieving the algorithmic uniformity mentioned in requirement ③. Due to the randomness of neural network exploration, to ensure that its exploration process can end smoothly, a lenient termination condition needs to be set, that is, the exploration ends when the pose with an error of less than a certain range from the target pose is found. However, once such a lenient range exists, the neural network will stop exploring when there is an error, which cannot ensure the accuracy requirement mentioned in point ②. Since this requirement cannot be reliably achieved by neural networks alone, other algorithms need to be added to assist neural networks.
[0081] Therefore, it can be concluded that before applying neural networks, the target poses under different storage locations need to be unified through coordinate transformation before training. While this training method achieves algorithm consistency across different storage locations, it also means that the neural network's training targets will be arbitrary starting poses under all three storage locations, resulting in a massive amount of data. Therefore, strategies need to be developed to reduce the amount of data while ensuring training effectiveness, to avoid the increased network complexity caused by excessive data volume.
[0082] A. The first approach is to minimize the dimensionality of the network input, thereby reducing the number of corresponding scenarios and thus the amount of data. Chinese Patent CN 114906128 A uses vehicle coordinates, heading angle (x, y, θ), and the vehicle's current steering wheel angle as network inputs. Given the vehicle speed v, the output is the change in steering wheel angle. The acceleration 'a' ensures that the planning results conform to dynamic constraints, but this means that a single scenario is determined by five quantities, exponentially increasing the number of scenarios that need to be trained. Therefore, the new neural network will use vehicle coordinates and heading angles (x, y, θ), as well as the corner point of the storage location (x, y, θ) representing the storage location. k1 ,y k1 ), (x k2 ,y k2 The input is the forward direction π and the spread curvature k. The above input is the simplified case that can represent different poses in different scenarios.
[0083] B. The second approach is to increase the exploration step size. Chinese patent CN 114906128 A uses trajectory data at 5ms intervals, resulting in hundreds of data sets for a single working condition. As the number of working conditions increases, the data volume surges. Therefore, the new neural network explores with a 1m step size, ensuring that a single working condition only has about 20 data sets at most. Thus, the network can handle more data from different working conditions without increasing its complexity.
[0084] The above two measures can alleviate the training pressure on the neural network, but their shortcomings are also obvious: First, the neural network plans path points that simply represent pose, not trajectory points with real-time speed and front wheel angle, making it impossible to design speed and angle commands simultaneously; second, the exploration results do not consider dynamic constraints at all, and the curvature is discontinuous; third, the path point interval is 1m, and the lack of information between path points makes them unsuitable for direct tracking. These three points determine that the neural network's exploration can only be a coarse-grained plan, requiring post-processing to achieve fine-grained planning.
[0085] The fine-planning step size can be set to 0.04m, approximately the length of one encoder grid on the vehicle. To ensure the path reaches the target pose with zero error, the fine-planning expansion direction extends from the target pose to the starting pose. Since directly modifying the neural network's planning results is inconvenient, a simulated tracking method is used. Considering vehicle dynamics constraints, a kinematic model is used to establish the vehicle's lateral control state equation for the neural network-planned path, simulating the vehicle's actual path when following a path with discontinuous curvature. The simulated path result serves as the planning result. Simultaneously, to ensure the vehicle avoids collisions with obstacles or boundaries during the planning process, Dynamic Windowing (DWA) is used as an aid, aligning the vehicle's expansion direction with the calculated direction from the simulated tracking as much as possible while avoiding collisions. Finally, since the vehicle expands from the target pose to the starting pose, there may be a deviation from the starting pose when the expansion stops. An end-effector smoothing mechanism is added, which eliminates residual errors by setting a reference line and using two segments of simulated tracking, one before and one after. At this point, the path post-processing is complete and can work well with the neural network's coarse-planning.
[0086] Based on the above analysis, the two modules involved in this invention will now be described in detail.
[0087] 1. Coordinate unification module
[0088] The standardization of coordinates aims to unify the target pose for parking in various parking spaces, facilitating training with a single neural network applicable to all spaces. However, the selection of the target pose needs consideration. During planning, parallel parking spaces, as well as perpendicular and oblique parking spaces, all involve an entry process. The "parking merging" part is unique to parallel parking spaces. Therefore, the parking merging planning for parallel parking should not be considered in the unified algorithm design. The unified algorithm framework only plans up to the entry pose for parallel parking spaces. This means that the target pose for parallel parking should actually be the entry pose, not the final pose; the target pose for perpendicular and oblique parking can be set as the final pose.
[0089] The coordinate transformation effect of parking paths under the three parking spaces is as follows: Figure 2 As shown, the initial parking coordinate systems differ for the three parking positions, but once the target pose is determined, coordinate transformation can unify the target pose. And through... Figure 2 As can be seen, the parking process of a vehicle in different parking spaces is essentially just starting from different locations, avoiding obstacles with varying distributions, and finally arriving at the same location. This also means that the parking process of a vehicle in different parking spaces can be summarized as an obstacle avoidance problem. Figure 3 This is a summary of the above conclusions:
[0090] a. Vertical parking starts from point A, avoiding obstructing vehicles ① and ②, and proceeds towards point O;
[0091] b. Angled parking starts from point B, avoiding obstructing vehicles ③ and ④, and proceeds towards point O;
[0092] c. Parallel parking starts from point C, avoiding obstructing vehicles ⑤ and ⑥, and proceeds towards point O;
[0093] In addition, when planning to move a vehicle towards point O in different parking spaces, it is also necessary to avoid collisions with the planned boundary while avoiding obstructing vehicles.
[0094] If the initial pose of the vehicle in the original coordinate system is Target pose is To achieve a target pose of (0,0,pi / 2) for the vehicle in the new coordinate system, the new coordinate system should be offset in the x and y directions compared to the original coordinate system. Rotation Then, according to the coordinate transformation formula, the initial pose (x0, y0, θ0) in the new coordinate system can be established. and Relationship:
[0095]
[0096] Storage location corner point (x) in the new coordinate system k1 ,y k1 ), (x k2 ,y k2 It can be obtained in a similar way.
[0097] As specifically mentioned earlier, parallel parking requires calculating the parking position pose before coordinate transformation can be performed. The calculation of the parallel parking parking position pose is as follows: Figure 4 As shown. The rubbing-in-the-garage path is simply designed as a circular curve. The vehicle, starting from the ending position O within the parking space, rubs along the circular curve to point O'. If, at this point, the vehicle moves from its current position with the minimum radius of curvature, the most dangerous point K will not collide with the obstacle vehicle in front, then the vehicle's current posture is considered the parking posture. During the rubbing-in-the-garage process, the vehicle must ensure that it does not collide with the obstacle vehicles in front and behind, or with the right boundary, i.e., ensure... Figure 4 Δl1, Δl2, and Δl3 are always greater than the threshold, which can be set from 15cm to 20cm.
[0098] At this point, the coordinate unification is complete, which is an important foundation for the implementation of subsequent algorithms.
[0099] 2. Path planning module
[0100] As designed above, the path planning module consists of two parts: coarse-grained planning via neural network and fine-grained planning via post-processing. These will be described in detail below.
[0101] (1) Coarse programming of neural networks
[0102] Neural networks need to be trained before use to ensure good performance. To reduce the training process, neural networks can first acquire basic performance through imitation learning, and then further improve performance through reinforcement learning. Therefore, it is necessary to create datasets that can be applied to imitation learning.
[0103] a. Dataset Creation
[0104] Although the CC curve stitching method used in Chinese patent CN 113830079 A lacks self-updating capabilities, it relies on rigorous geometric calculations and exhibits relatively stable performance. Therefore, dataset creation can also be based on this method. Since CC curve calculations are complex, and as previously mentioned, neural networks do not require data points with continuous curvature, the requirements for dataset creation can be simplified; RS curves can suffice.
[0105] Before using RS curves for splicing, several splicing forms must first be defined for geometric calculations. Based on human drivers' parking experience in parallel, oblique, and perpendicular parking spaces, the following splicing forms are defined: Figure 5 The four curve forms shown are: straight line-arc-straight line, straight line-arc-arc-straight line, arc-straight line-arc-arc-straight line, and straight line-arc-straight line-arc-straight line. Figure 5 The crosses in the diagram represent the boundary points of segments with different shapes.
[0106] The path calculation is based on strict geometric relationships. The calculation process is illustrated using the simplest example: a straight line-arc-straight line. Figure 6 As shown, the vehicle starts from point S with initial coordinates (x... s ,y s ,θ s The vehicle needs to be parked at the destination O, whose coordinates are (x...). o ,y o ,θ o The point most likely to collide during the entire process is C(x). c ,t c The vehicle plans a path of straight line-circular arc-straight line, and will encounter an inflection point A(x). a ,y a ,θ a ) and B(x b ,y b ,θ b Furthermore, according to geometric relationships, the positional relationships between points A and B and points O and S are as follows:
[0107]
[0108] Here l BS l OALet BS and OA be the lengths of line segments. Given that the heading angles of A and B are determined, and the arc angle Δθ is also determined, once the arc radius r is given, the positional relationship between A and B can be determined based on the arc shape.
[0109]
[0110] Therefore, if the radius r of the arc can ensure that the vehicle does not collide with C, then l can be calculated from r. BS l OA .
[0111] The computational mechanism for other curve forms is similar. When choosing the arc radius r, i.e., the arc curvature κ, to ensure compatibility with subsequent neural network training, the given curvature κ is set between -0.2 and 0.2, with intervals of 0.05, resulting in nine selectable curvatures. The RS curve obtained after calculation according to the above rules is in implicit form as shown in equation (2.3).
[0112]
[0113] In the formula, l1, l2, and l3 are the segment boundaries along the arc length; ρ1, ρ2, and ρ3 are the curvature values within this segment; q1, q2, and q3 are the vehicle travel directions of this segment (forward is 1, reverse is -1); π1, π2, and π3 are the steering directions of this segment (left turn is 1, right turn is -1); s d The length along the path is given. Data points can be collected according to the shape and direction of expansion specified in equation (2.3), with one data point taken every 1m.
[0114] The data should include planning data for multiple starting poses under three storage locations: parallel, oblique, and vertical. All planned target poses should be (0, 0, pi / 2). Finally, data such as... Figure 7 The data results are shown. Figure 7 The system only displays the x and y coordinates of the data points, but in reality, the heading angle, direction of motion, curvature, and corner coordinates of the obstacle vehicle also need to be recorded. Figure 8 The corner points K1 and K2 of the obstacle vehicles that need to be recorded for different parking locations are marked. All the data recorded above should be coordinate-transformed data.
[0115] b. Imitation learning / reinforcement learning
[0116] The aforementioned RS curve-based dataset will be used for imitation learning. To reiterate, each data set in the dataset contains the vehicle's position information (x, y, θ), the coordinates of the corner point of the obstacle in the parking space where the vehicle is located, and the coordinates of the obstacle corner (x, y, θ). k1 ,y k1 ) and (x k2 ,y k2The inputs to the neural network are vehicle position information and corner points of vehicles facing obstacles in the parking space. The 18-classification output is composed of two combinations of different vehicle movement directions and different running curvatures. The movement direction is either forward or backward, and the running curvature is from -0.2 to 0.2, with an interval of 0.05.
[0117] Because data is extracted using a large step size, the amount of data remains moderate even with numerous operating conditions. In this invention, approximately 10,000 to 15,000 sets of data under about 500 different operating conditions were generated by changing the initial pose in three storage locations. This allows the neural network structure to remain relatively simple. The network used in this invention, excluding the input and classification layers, contains two hidden layers, each with 25 neurons. Of the aforementioned data, 70% is the training set, 15% is the validation set, and 15% is the test set.
[0118] Neural networks, after imitation learning, have developed preliminary planning capabilities, while reinforcement learning will further enhance their performance. The goal of reinforcement learning is to explore and obtain the best possible data sequences {(s)} based on existing policies. i ,a i The network is then categorized into groups (i=0,1,...n-1) and stored. Finally, the network is updated again using the stored data sequences under different operating conditions to further improve network performance. Here, s... i This represents the i-th state, containing seven data points: the vehicle's coordinates, heading angle, and the corner point of the obstacle vehicle under a certain working condition. i Represents state s i The optimal motion set, including the direction of motion and curvature, is given below. s0 when i=0 represents the initial state, where the vehicle is in state s. i In the state, take action group a. i Transition to the next state s i+1 When i = n-1, the vehicle takes action group a. n-1 Will transition to final state s n .
[0119] The vehicle is expanded according to a kinematic model, but due to the large step size (1m) expansion, the expansion result requires geometric calculations:
[0120] ① When curvature κ i =0
[0121]
[0122] ②When curvature κ i >0
[0123]
[0124] ③When curvature κ i <0 o'clock
[0125]
[0126] Reward function R in reinforcement learning total Designed as
[0127] R total =R change_dir +R space +R length +R reach +R colli +R chang_cur (2.7)
[0128] Each part is specifically defined as follows
[0129]
[0130] In the above formulas:
[0131] ①R change_dir This represents the shift report value, calculated based on the number of shifts (dir_change).
[0132] ②R space The space utilization return value is based on the leftmost x-coordinate value of the path. left The rightmost x-coordinate value x right The top y-coordinate value up The lowest coordinate value y down calculate;
[0133] ③R leng This represents the path length reward value, calculated based on the total path length_total.
[0134] ④R reach Represents the final posture evaluation value, derived from the base reward value R. reac_basic And additional reward value R reach_extra It consists of two parts. Because neural networks use large step sizes for exploration, and there is post-processing after the neural network is planned, the network's exploration does not need to reach the final pose with perfect precision; it only needs to reach a pose that is easy for post-processing. If the exploration result (x,y,θ) satisfies...
[0135]
[0136] You can get R rea_basic =40000. If the exploration results can get closer to the target pose, an additional reward value R can be obtained. reach_extra , can be calculated as
[0137] R rea_extr=20*(0.05-|x|) / 0.05+100*(0.25-|θ-pi / 2|) / 0.25(2.10)
[0138] ⑤R colli This represents the collision report value. If the path collides with an obstacle vehicle, is_colli is 1, otherwise it is 0.
[0139] ⑥R change_cur The value representing the curvature change reward is calculated from the maximum curvature change over two consecutive steps, max_curve_change.
[0140] The optimization process is as follows Figure 9 As shown. After selecting the storage location and starting pose, a single-step MCTS is used for exploration: before reaching the termination condition, the current state is expanded, traversing all 18 action combinations to obtain 18 corresponding sub-states. Then, starting from these 18 sub-states, a neural network combined with a roulette wheel rule is used for rolling exploration, resulting in 18 path point sequences. The action group corresponding to the sequence with the highest reward value is selected as the action group for this expansion, and the next state is entered. This sequence is compared with the sequence corresponding to the maximum reward value under the current storage location and starting pose, and the sequence with the larger reward value is retained. After reaching the next state, the above process is repeated. After reaching the termination condition, the sequence with the maximum reward value under the starting pose is stored in the data pool. The direct effect of this exploration is that during the expansion process, it continuously attempts to update the existing sequence with a better path point sequence, ultimately making the path point sequence under the current storage location and starting pose optimal. If the above process is performed multiple times under different storage locations and different starting poses, a large amount of optimized data will be stored in the data pool. Using this data to retrain the neural network can enhance the network performance.
[0141] (2) Path post-processing fine planning
[0142] Since the planning result of the neural network is a discrete path point with a spacing of 1m, there is a problem of discontinuous curvature of adjacent points, which cannot be directly used for tracking. Therefore, post-processing is required, and the step size can be set to the length of one code disk unit, 0.04m. Furthermore, the neural network planning only reaches a pose suitable for post-processing, not the target pose. To ensure the accuracy of the planning, the post-processing extension should be performed in reverse from the target pose to the starting pose.
[0143] a. Simulated tracking
[0144] Simulated tracking, as the name suggests, involves simulating a vehicle following a certain path. Although the vehicle cannot completely follow the path due to dynamic constraints, the actual path taken by the vehicle to follow the path must conform to the vehicle dynamic constraints and can be truly tracked. Therefore, the actual path taken by the vehicle during simulated tracking can be used as the final path planning result.
[0145] The path points planned by the neural network can actually be regarded as a number of "anchor points." During simulated tracking, the vehicle only needs to follow the expansion direction marked by these anchor points; precisely reaching the anchor points is not a requirement. A schematic diagram of the simulated tracking is shown below. Figure 10 .
[0146] Because the anchor points of a neural network contain the coordinates of the anchor point and the heading angle. The direction and curvature of expansion at this location This allows the path between any two points planned by the neural network to be viewed as a simple circular arc, while simulated tracking attempts to follow this path, which consists of several circular arcs connected end-to-end and has discontinuous curvature, under dynamic constraints. Because the path is an arc, the lateral offset and angular error between the vehicle and the path are easy to calculate; therefore, a simple feedforward + feedback design for the lateral controller is sufficient. The process of the vehicle following the circular arc is as follows: Figure 11 As shown, using the vehicle reference point at the midpoint of the rear axle, the controller design process is as follows:
[0147] ① Rate of change of lateral deviation e1 Related to vehicle speed V and angular deviation e2, and taking into account the small angle assumption, we have
[0148]
[0149] ② Rate of change of angular deviation e2 It is related to vehicle speed V, current front wheel steering angle ψ, vehicle wheelbase L, and path curvature κ, and combined with the small angle assumption, we have
[0150]
[0151] Let the control variable be u = tanψ, then we have the state equation.
[0152]
[0153] Establish the relationship between the control variable u and the deviations e1 and e2, and add a feedforward quantity to...
[0154] u=-k1e1-k2e2-Lκ (2.14)
[0155] Therefore, a lateral controller is established for simulated tracking, and during simulated tracking, the range of variation of the front wheel steering angle ψ needs to be set in conjunction with dynamic constraints.
[0156] As previously mentioned, the simulated tracking does not require the vehicle to precisely reach the anchor point; it only needs to extend in the direction indicated by the anchor point. In actual simulated tracking, when the vehicle reaches a position "aligned" with the anchor point, it is considered ready to proceed to the next arc of tracking. The state where the vehicle is "aligned" with the anchor point is as follows: Figure 12 As shown, when the vehicle state is (x, y, θ), the anchor point is... And the action group stored at the anchor point is When the angle between the vector formed by the anchor point and the vehicle and the direction of the anchor point itself is less than pi / 2, it is considered to be "aligned".
[0157] b. DWA (Dynamic Window Method)
[0158] One point to note is that the neural network does not consider the boundaries of the parking environment during planning. In order to further ensure that the vehicle does not collide with other vehicles in front or behind or other obstacles that may appear in the environment during the expansion process, it is necessary to use a method that can flexibly combine the current position for obstacle avoidance.
[0159] The method used in this invention is DWA (Dynamic Window Method). Its principle is to sample possible actions, advance a certain distance using the sampled actions, and evaluate the sampled actions based on factors such as whether they collide with obstacles and whether they conform to the intended forward direction. The action with the highest evaluation is selected as the next forward action. DWA is widely used in robot path planning, but this invention is geared towards vehicle path planning; therefore, the following modifications were made when using DWA:
[0160] ① Evaluation is performed using only the termination position reached after a single sampling action, rather than using the entire forward path. This is because the vehicle is much larger than a typical robot, and the distance extended by the sampling action is less than half the total length of the vehicle. If the vehicle collides with an obstacle during the extension process, the termination position will most likely also collide with an obstacle, and the state of the termination position is sufficient to represent the state of the entire path. Furthermore, calculating only the termination position saves considerable computation and avoids placing excessive pressure on the onboard system.
[0161] ② When sampling the front wheel turning angle, only the maximum range that the front wheel turning angle can reach is considered, without considering whether the sampling action at the current turning angle meets the speed requirement. This is because when parking, the vehicle is in a confined environment and always maintains a close distance to obstacles. If speed is considered, the exploration range will be reduced, which may lead to a situation where all sampling actions used for expansion result in collisions with obstacles. In this case, the vehicle needs to know whether there is a direction that can avoid the obstacle, rather than just the turning angle that meets the speed requirement.
[0162] A schematic diagram of DWA is shown below. Figure 13 During sampling, since the front wheel turning angle range of the vehicle is ±30°, in order to avoid excessive sampling actions and excessive calculation, sampling was performed at 3° intervals, for a total of 21 sampling actions. Based on the actual effect, an extension of 40*0.04=1.6m was selected to obtain 21 termination positions. These 21 termination positions were evaluated, and the action corresponding to the highest evaluation was selected as the next action.
[0163] The evaluation function is defined as follows:
[0164] r total =r follow_path +r avoid_obstacle (2.15)
[0165] in:
[0166] ①r follow_path Representative sample value ψ c The rotation angle value ψ calculated by simulation tracking p The better the conformity reward value, the smaller the deviation between the sampled value and the angle value calculated by the simulated tracking. Specifically, it is defined as follows:
[0167] r follow_path =-|ψ c -ψ p | (2.16)
[0168] ②r avoid_obstacle This represents the vehicle collision reward value after expansion, calculated using the distance *dis* between the vehicle and the obstacle. The smaller the distance, the lower the value of this item. Specifically, it is defined as follows:
[0169]
[0170] It's important to note that calculating the distance *dis* between the vehicle and obstacles is actually quite complex. The vehicle's original outline is octagonal, and calculating the distance to other obstacles involves calculating the distances between polygons, which is tedious. This invention uses four envelope circles to cover the vehicle's outline as shown... Figure 14 In this way, although it takes more space, the distance calculation between the vehicle and other obstacles can be simplified to the distance calculation between a point and a polygon.
[0171] c. End smoothing
[0172] As mentioned earlier, post-processing involves reverse exploration from the target pose back to the initial pose. However, due to the discontinuous curvature of the path points in the neural network, simulated tracking itself cannot perfectly follow the target pose. Secondly, the use of DWA (Directed Path Analysis) causes the vehicle to make slight adjustments when approaching obstacles. Both of these factors can lead to the vehicle failing to accurately return to the initial pose during reverse exploration. If the reverse exploration result differs significantly from the initial pose, the vehicle will have a large initial error, causing the subsequent path tracking module to calculate a large turning angle. This results in strong steering wheel vibration during initial vehicle operation, reducing passenger comfort. Therefore, an end-effector smoothing strategy is needed to avoid excessive initial error.
[0173] The steering wheel angle vibration is specifically caused by a significant lateral error between the termination position of the reverse exploration and the target pose. Therefore, the goal of end-of-pipe smoothing is to eliminate the lateral deviation by increasing the vehicle's movement through additional forward and backward movements. Since passengers will not intentionally activate the automatic parking function near obstacles in most cases, obstacle avoidance is not considered for the sake of design simplicity.
[0174] like Figure 15 As shown, there is a lateral deviation between the vehicle and the target pose at the end of the reverse exploration. The midline between the two can be used as the boundary for two-segment simulated tracking: the first segment maintains the direction of reverse exploration to eliminate the lateral deviation between the vehicle and the midline; the second segment is in the opposite direction to the first segment to eliminate the lateral deviation between the vehicle and the target pose.
[0175] The above process allows for the use of a unified framework to plan vehicle parking in different parking spaces and starting positions. Specifically, for parallel parking, the framework can only plan the entry and exit paths; the paths within the parking space need to be based on geometric arcs and merged with the entry path to form the final planned path. Figure 16 , 17 Figures 1 and 18 show the planned paths for vehicles in parallel, diagonal, and perpendicular parking spaces, respectively. The simulation results also show that the planned paths can be completely tracked.
[0176] The preferred embodiments of the present invention have been described in detail above. It should be understood that those skilled in the art can make numerous modifications and variations based on the concept of the present invention without creative effort. Therefore, all technical solutions that can be obtained by those skilled in the art based on the concept of the present invention through logical analysis, reasoning, or limited experimentation on the basis of existing technology should be within the scope of protection defined by the claims.
Claims
1. An automatic parking path planning method for various parking space scenarios, characterized in that, The method includes the following steps: Obtain the storage location type, storage location corner coordinates, and initial pose; Coordinate transformation is performed based on parking space type to unify the coordinate system for vehicles in parallel, diagonal, and perpendicular parking scenarios; In a unified coordinate system, the steps for obtaining the scattered point sequence of the coarse-planned path using a neural network include: Using the vehicle's current pose and the coordinate-transformed parking space corner points as inputs to the neural network, the output direction of expansion is... and extended curvature Action combinations; Based on the action combinations output by the neural network, the vehicle extends to the next state in large steps using geometric inference; until the vehicle extends from the initial pose to the target pose; the training steps of the neural network include: The neural network is used to learn by imitation using path data based on RS curve splicing; By repeatedly setting different storage location categories and different initial poses, the neural network is used to explore optimization, and data with higher reward values are obtained to update the network again, thus completing reinforcement learning. The return value The function is: In the formula: The shift reward value decreases with more shifts. The more space is utilized, the lower the return on path space utilization. This is the path length reward value; the longer the path length, the lower the path length reward value. The final posture evaluation value represents whether the target pose has been achieved. A reward value can be obtained if the pose is within the allowable range. If the target pose can be matched with a smaller error on this basis, an additional reward value can be obtained. The collision reward value is significantly reduced if a collision occurs during the expansion process. The return value is the curvature change reward; the greater the curvature change during the expansion process, the lower the reward value. The planned path is obtained from the scatter sequence of coarse-planned paths through post-processing including simulated tracking, DWA, and end-point smoothing; the DWA steps include: Sampling was performed on the front wheel steering angle of vehicles within the permitted range; Set a fixed extension length, compare the reward values of the extension paths corresponding to different sampling results; select the value. The maximum value is the ideal front wheel steering angle at this moment; Its value The function is defined as: In the formula: The score represents the degree of conformity between the sampled value and the angle value calculated by the simulation tracking. The smaller the deviation between the sampled value and the angle value calculated by the simulation tracking, the higher the degree of conformity. This represents the vehicle collision report value after the expansion is complete. It is calculated using the distance between the vehicle and the obstacle. The smaller the distance, the lower the collision report value.
2. The automatic parking path planning method for multiple parking space scenarios according to claim 1, characterized in that, The specific steps of the method for performing coordinate transformation and unifying the coordinate system include: The vehicle's initial pose in the original coordinate system is The target pose is For parallel parking, the target pose is the parking position; for diagonal and perpendicular parking, the target pose is the final position. Taking the target pose of the vehicle in the new coordinate system as Perform a coordinate transformation; the new coordinate system is different from the original coordinate system. , Direction offset , Rotation ; The initial pose and storage location corners in the new coordinate system are obtained based on the coordinate system transformation relationship; The initial pose in the new coordinate system relative to the initial pose in the original coordinate system and target pose The coordinate transformation formula is: In the formula, , and These are the horizontal and vertical coordinates and the angle of the initial pose in the original coordinate system, respectively. , and These are the horizontal and vertical coordinates and the angle of the target pose in the original coordinate system, respectively; , and These are the horizontal and vertical coordinates and the angle of the initial pose in the new coordinate system.
3. The automatic parking path planning method for multiple parking space scenarios according to claim 1, characterized in that, The path data based on RS curve splicing is generative data, and its generation process is as follows: Based on human drivers' parking experience in parallel, oblique, and perpendicular parking spaces, various RS curve splicing forms are defined; By utilizing the positional relationship between different initial poses and target poses under different storage locations, the length and curvature of the RS curve are obtained through geometric calculation; The positional information and corresponding action categories needed to train the neural network are generated based on RS curves.
4. The automatic parking path planning method for multiple parking space scenarios according to claim 1, characterized in that, The simulation tracking method is as follows: Using the vehicle kinematics model, a lateral control state equation is established. The front wheel steering angle is designed through feedback and feedforward. Considering vehicle dynamics constraints, the vehicle's tracking process of coarse-planned path points is simulated. The kinematics model is expanded with small steps, and the path point sequence obtained by the simulation tracking is used as the planning data. The kinematic model extension starts from the target pose and extends in the reverse direction towards the initial pose.
5. The automatic parking path planning method for multiple parking space scenarios according to claim 1, characterized in that, The end-effector smoothing involves performing two simulated tracking segments, with the midline between the exploration termination position and the target pose as the boundary: The first segment continues to expand in the opposite direction, extending a certain distance using simulated tracking to eliminate the lateral deviation between the vehicle and the centerline; the second segment is in the opposite direction to the first segment, eliminating the lateral deviation between the vehicle and the target pose.
6. The automatic parking path planning method for multiple parking space scenarios according to claim 1, characterized in that, When performing parallel parking path planning, an additional scatter sequence of parking path points based on geometric circular arc curves is added after the planned path is obtained in post-processing.
Citation Information
Patent Citations
Online planning method and system for continuous curvature parking path of any starting pose
CN113830079A
Automatic parking motion planning method based on MCTS algorithm
CN114906128A
Automatic vertical parking system and method based on multi-stage planning and machine learning
CN109131317A
Path planning method and device for automatic parking system
CN114030463A
Real-time trajectory planning method and system
CN114715193A