End-to-end automatic parking path planning method and system

CN122808709APending Publication Date: 2026-09-25TONGJI UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611071709.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-20
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

[0005]本发明的目的在于提供一种端到端自动泊车路径规划方法及系统,用于解决分层式泊车系统的信息损失及误差累计、复杂拥挤场景的规划能力和安全性不足、以及现有端到端方法路径可执行性差等问题,提高泊车系统对环境的适应性、路径规划的安全性和可执行性

Benefits of technology

1、端到端架构减少了信息损失与误差累计。本发明构建了从感知输入到控制指令输出的完整端到端系统,策略网络直接从原始观测状态映射到参数化动作,避免了传统分层式泊车系统中各模块之间的信息损失和误差累计,提高了路径规划策略对环境的适应性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122808709A_ABST
    Figure CN122808709A_ABST
Patent Text Reader

Abstract

The application provides an end-to-end automatic parking path planning method and system, the method comprising: acquiring vehicle parking demand, target parking space information and surrounding environment information, and constructing an observation state of a parking system; outputting a parameterized action sequence directly according to the observation state through a trained parking path planning strategy, constructing a safe parking path from the parameterized action sequence, and directly using the parameterized action sequence for vehicle control; and following the safe parking path by using a vehicle control and execution module to complete automatic parking. The parking path planning strategy is trained by using a constraint reinforcement learning algorithm, the action space of the strategy is a mixed action space, and the strategy network is updated by evaluating whether the strategy satisfies a safety cost constraint during the training process. The application reduces error accumulation caused by a hierarchical mechanism through an end-to-end architecture, and improves the adaptability of the path planning strategy to the environment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of intelligent driving technology, specifically relating to an end-to-end automatic parking path planning method and system, which is suitable for efficient and safe parking of intelligent vehicles in congested and complex environments. Background Technology

[0002] With the number of cars on the road increasing year by year, vehicle density in urban areas has increased significantly. Narrow and congested parking environments exacerbate parking difficulties, causing parking problems for novice drivers. Automated parking technology, by sensing the surrounding environment and planning routes to complete autonomous parking tasks, holds promise for solving these problems. Therefore, as a key technology for intelligent driving vehicles, it is crucial for improving the driving experience and safety of intelligent vehicles.

[0003] Existing path planning methods for automated parking systems mainly fall into two categories: The first category is traditional methods based on geometric curves or heuristic search algorithms. Although these methods can achieve certain results in simple or specific environments, they have significant limitations in complex and congested environments. Specifically, they are: (1) relying on prior knowledge, which leads to poor adaptability of the path planning strategy to the environment; (2) generally relying on structured perception information to complete path planning, and the separation of perception information from the planning layer leads to poor robustness of the path planning model to the environment; (3) the limitations of prior knowledge and the non-real-time nature of online search make it difficult to handle parking tasks in space-constrained or non-standard parking scenarios, and it is difficult to guarantee the safety and real-time performance of parking in complex scenarios.

[0004] The second category is end-to-end path planning methods based on deep reinforcement learning. These methods can directly generate parking paths from perception data, achieving seamless connection from perception to control, and showing significant advantages in improving parking efficiency and path smoothness. However, existing end-to-end parking methods generally have the following problems: (1) They directly output continuous lateral and longitudinal control actions with a Gaussian distribution, resulting in poor executability of the planned path and difficulty in ensuring safety; (2) Since they are directly mapped to the control layer, the trained end-to-end parking strategy has poor adaptability among different vehicle models. Summary of the Invention

[0005] The purpose of this invention is to provide an end-to-end automated parking path planning method and system to address issues such as information loss and error accumulation in hierarchical parking systems, insufficient planning capabilities and safety in complex congestion scenarios, and poor path executability of existing end-to-end methods. This improves the parking system's adaptability to the environment, the safety of path planning, and its executability. The technical solution adopted is as follows: An end-to-end automated parking path planning method includes the following steps: Step S1: Obtain vehicle parking demand, target parking space information, and surrounding environment information to construct the observation status of the parking system; Step S2: Based on the trained parking path planning strategy, output a parameterized action sequence directly according to the observed state, and construct a safe parking path from the parameterized action sequence; the parameterized action sequence is directly used for vehicle control. Step S3: Use the vehicle control and execution module to follow the safe parking path to complete automatic parking; The parking path planning strategy is trained using a constraint reinforcement learning algorithm, and its action space is a hybrid action space. During training, the strategy network is updated by evaluating whether the parking path planning strategy meets the safety cost constraint.

[0006] Preferably, each parameterized action in the parameterized action sequence includes a discrete action and a continuous parameter. The discrete actions include turning left, going straight, and turning right, and the continuous parameter is the arc length or straight distance of the vehicle moving forward or backward under the corresponding discrete action.

[0007] Preferably, the state transition is performed based on the parameterized action and the kinematic state transition relationship of the vehicle's rear axle center.

[0008] Preferably, the training process of the parking path planning strategy includes a constraint evaluation step and a strategy update step: The constraint evaluation steps are as follows: calculate the average cumulative cost of the parking path based on the collected samples, and determine whether the current strategy meets the safety constraints by comparing the relationship between the average cumulative cost and the cost threshold. The policy update steps are as follows: when the safety constraints are met, update the policy network by maximizing the cumulative reward; when the safety constraints are not met, update the policy network by minimizing the cumulative cost.

[0009] Preferably, during training, the reward value network and the cost value network are used to calculate the discrete action advantage function and continuous parameter advantage function of all action state pairs in the sample with respect to the reward value, as well as the discrete action advantage function and continuous parameter advantage function with respect to the cost value.

[0010] Preferably, the surrounding environment information is obtained through an in-vehicle surround-view camera and / or ultrasonic sensors, and the vehicle parking requirements are obtained through a human-machine interface.

[0011] An end-to-end automated parking path planning system includes: The perception module is used to acquire environmental information around the vehicle, the pose of the target parking space, and the pose of the vehicle. Human-computer interaction interface, used to obtain parking demand information; The path planning module is configured with a trained parking path planning strategy, which is used to directly output a safe parking path composed of parameterized action sequences based on the observation state obtained by the perception module and the human-machine interface; the parameterized action sequences are used for vehicle control. And a vehicle control and execution module, used to complete automatic parking by following the safe parking path.

[0012] Preferably, the parking path planning strategy is trained in the following way: Initialize the policy network parameters, reward value network parameters, and cost value network parameters; Sample data is collected in an automated parking simulation environment or a parking world model built based on real data using a policy network to calculate cumulative rewards and cumulative costs. The average cumulative cost of the trajectory is calculated based on the sample and compared with the cost threshold to determine whether the current strategy meets the safety constraints. Depending on whether the security constraints are met, the policy network should be updated by either maximizing the cumulative reward or minimizing the cumulative cost. Iterative training continues until the policy network converges and satisfies the cost constraint.

[0013] In summary, the key points of this invention are as follows: (1) End-to-end architecture solves the problems of information loss and error accumulation in layered systems.

[0014] The present invention constructs a complete end-to-end automatic parking system, that is, from perception input (step S1) directly to path output (step S2), and then directly used for vehicle control (step S3), without going through multiple intermediate processing links such as environmental perception and path planning in the traditional layered architecture.

[0015] The end-to-end architecture allows the policy network to directly map from the original observed states to parameterized actions, avoiding information loss and error accumulation between the perception and planning modules in traditional hierarchical parking systems. Simultaneously, the end-to-end training method enables the policy to better adapt to different parking environments and scenarios, improving the system's generalization ability and robustness.

[0016] (2) The hybrid action space solves the problem of path executability.

[0017] The parameterized action sequence of this invention consists of multiple parameterized actions, each of which includes a discrete action and a continuous parameter. The discrete actions include left turn, straight ahead, and right turn, and the continuous parameter is the arc length or straight-ahead distance of the vehicle moving forward or backward under the corresponding discrete action.

[0018] Compared to existing end-to-end methods that directly output continuous control actions, the hybrid action space design of this invention ensures that the planned path naturally satisfies vehicle kinematic characteristics (such as minimum turning radius constraints), significantly improving path executability. Each parameterized action corresponds to a clearly defined vehicle trajectory, facilitating direct tracking and execution by the vehicle control module.

[0019] (3) Constraint reinforcement learning solves the path safety problem.

[0020] This invention introduces a cost assessment value for collision safety and uses a constrained reinforcement learning algorithm to constrain the collision safety of the policy within a safety threshold range. During training, the policy network is updated by evaluating whether the policy meets the safety cost constraint: when the policy meets the safety constraint, the policy network is updated by maximizing the cumulative reward; when the policy does not meet the safety constraint, the policy network is updated by minimizing the cumulative cost.

[0021] By formalizing security into a quantifiable cost constraint and explicitly considering this constraint during policy updates, the security of the planned path is ensured, avoiding the problem of security being difficult to guarantee in existing end-to-end methods.

[0022] Compared with the prior art, the advantages of the present invention are: 1. End-to-end architecture reduces information loss and error accumulation. This invention constructs a complete end-to-end system from sensing input to control command output. The policy network directly maps from the original observed state to parameterized actions, avoiding information loss and error accumulation between modules in traditional hierarchical parking systems, and improving the adaptability of the path planning strategy to the environment.

[0023] 2. The hybrid motion space improves the executability of the path. This invention combines discrete motions with continuous parameters, so that the planned path naturally satisfies the vehicle's kinematic characteristics (minimum turning radius constraint), solving the problem of poor path executability caused by the direct output of continuous control motions in existing end-to-end methods.

[0024] 3. Constraint reinforcement learning ensures path safety. This invention formalizes collision safety into quantifiable cost constraints, and through a closed-loop design of constraint evaluation and conditional policy update, makes safety a rigid and insurmountable constraint during training, thus ensuring the safety of the planned path. Attached Figure Description

[0025] Figure 1 The execution process of end-to-end automated parking path planning methods, systems, and devices; Figure 2 A schematic diagram of parametric actions for parking path planning; Figure 3This is a schematic diagram of the rear axle center pose state transition during a vehicle left turn based on parametric motion. Figure 4 This is a schematic diagram of the rear axle center pose state transition during a vehicle right turn based on parametric motion. Figure 5 A diagram illustrating the training process for parking path planning strategies. Detailed Implementation

[0026] The following will describe in more detail an end-to-end automatic parking path planning method and system of the present invention with reference to the schematic diagrams, which illustrate preferred embodiments of the invention. It should be understood that those skilled in the art can modify the invention described herein while still achieving the advantageous effects of the invention. Therefore, the following description should be understood as being of general knowledge to those skilled in the art and is not intended to limit the invention.

[0027] like Figures 1-5 The end-to-end automated parking path planning method includes the following steps: Step S1: Obtain parking demand information through the human-computer interaction interface, and obtain environmental information near the parking space through vehicle surround view or ultrasonic (or both simultaneously). Position of the target parking space Vehicle position Based on the acquired information, construct the observation status of the vehicle parking system. .

[0028] Parking demand information is obtained through a human-computer interaction interface and is used to determine the target parking space and parking type for the parking task; the parking demand information itself is not directly used as a feature input for the policy network.

[0029] After identifying the target parking space, the perception system acquires the pose of the target parking space, which, together with the vehicle pose and environmental information, constitutes the observation state of the policy network.

[0030] Step S2, based on the information observed in step S1 Using a trained parking path planning strategy Output continuous parameterized actions The process continues until the deviation between the vehicle's position and the target position reaches the parking requirement.

[0031] The parking path is constructed based on a series of planned parametric actions. .

[0032] Among them, parameterized actions like Figure 2 As shown, this action includes offline actions. and continuous parameters .

[0033] Furthermore, based on parameterized actions, the state transition processes of the vehicle's rear axle center during left and right turns are as follows: Figure 3 and Figure 4 As shown, the state transition relationship of the rear axle center of the vehicle based on the geometric relationship is shown in formulas (1), (2), and (3).

[0034] When a vehicle turns left, based on parameterized actions The state transition relationship of the vehicle's rear axle center:

[0035] In the formula: This represents the radian angle corresponding to the arc length rotated; where the arc length is positive when moving forward and negative when moving backward. , and These respectively indicate that the vehicle is in The horizontal and vertical coordinates and heading of the step, , and These respectively indicate that the vehicle is in The horizontal and vertical coordinates and heading of the step; It is the vehicle's turning radius, which is specified in this patent as the vehicle's minimum turning radius. Therefore, the trajectory determined based on the vehicle's initial position and parametric actions is unique.

[0036] When a vehicle turns right, based on parameterized actions The state transition relationship of the vehicle's rear axle center:

[0037] When the vehicle is traveling straight, based on parameterized actions The state transition relationship of the vehicle's rear axle center:

[0038] Step S3, based on the safe parking path planned in step S2, utilize the vehicle control and execution module ( Figure 1 The vehicle controller in the system follows the parking trajectory or commands to complete autonomous parking.

[0039] like Figure 5 As shown, the training process of the parking path planning strategy based on constraint reinforcement learning in a hybrid action space includes the following steps: Step S21, initialize policy network parameters Reward value network parameters Cost value network parameters .

[0040] Step S22, through the policy network Collect data including environmental perception states in an automated parking simulation environment or a parking world model built based on real data. ,award ,cost and actions Sample data.

[0041] And based on the sample reward and cost Calculate the cumulative reward for the remaining trajectory in the current state. and cost .

[0042] In addition, using reward value networks and cost value network Calculate all action-state pairs in the sample respectively The advantage function, This includes a discrete action advantage function with respect to reward values. and continuous parameter advantage function And the discrete action advantage function regarding cost value and continuous parameter advantage function .

[0043] Step S23: Calculate the average cumulative cost of the trajectory based on the collected samples, and determine whether the current strategy meets the safety constraints by comparing the relationship between the average cumulative cost and the cost threshold.

[0044] If the average cumulative cost of the trajectory is greater than the cost threshold, it indicates that the current strategy does not meet the safety constraints; otherwise, it indicates that the current strategy meets the safety constraints.

[0045] Step S24: When the policy satisfies the cost constraint, update the policy network by maximizing the cumulative reward; Maximize cumulative rewards: ; in, This is the correction factor after cropping; ; This is the current strategy. With the Step strategy The ratio; It is the cutting factor.

[0046] When the policy does not meet the cost constraint, the policy network is updated by minimizing the cumulative cost. Minimize cumulative cost: .

[0047] in The same correction factor as in step S24 for maximizing cumulative rewards.

[0048] Step S25: Update the reward and cost based on the mean squared error loss function.

[0049] The mean squared error loss function used for rewards is:

[0050] The variance loss function used for update costs is:

[0051] Repeat steps S22, S23, S24, and S25 until the network policy converges and the cost constraint is met, thereby obtaining a real-time, efficient, and safe end-to-end parking path planning strategy. .

[0052] Furthermore, since the formulas involved in step S25 are all existing technologies, the meaning of the letters in them will no longer be explained.

[0053] The above are merely preferred embodiments of the present invention and do not constitute any limitation on the present invention. Any equivalent substitutions or modifications made by those skilled in the art to the technical solutions and content disclosed in the present invention without departing from the scope of the present invention shall be deemed to have remained within the protection scope of the present invention.

Claims

1. An end-to-end automated parking path planning method, characterized in that, Includes the following steps: Step S1: Obtain vehicle parking demand, target parking space information, and surrounding environment information to construct the observation status of the parking system; Step S2: Based on the trained parking path planning strategy, output a parameterized action sequence directly according to the observed state, and construct a safe parking path from the parameterized action sequence; the parameterized action sequence is directly used for vehicle control. Step S3: Use the vehicle control and execution module to follow the safe parking path to complete automatic parking; The parking path planning strategy is trained using a constraint reinforcement learning algorithm, and its action space is a hybrid action space. During training, the strategy network is updated by evaluating whether the parking path planning strategy meets the safety cost constraint.

2. The end-to-end automatic parking path planning method according to claim 1, characterized in that, Each parameterized action in the parameterized action sequence includes a discrete action and a continuous parameter. The discrete actions include turning left, going straight, and turning right. The continuous parameter is the arc length or straight distance of the vehicle moving forward or backward under the corresponding discrete action.

3. The end-to-end automatic parking path planning method according to claim 1, characterized in that, Based on the parameterized actions, the vehicle's rear axle center kinematic state transition relationship is used for state transition.

4. The end-to-end automatic parking path planning method according to claim 1, characterized in that, The training process of the parking path planning strategy includes a constraint evaluation step and a strategy update step: The constraint evaluation steps are as follows: calculate the average cumulative cost of the parking path based on the collected samples, and determine whether the current strategy meets the safety constraints by comparing the relationship between the average cumulative cost and the cost threshold. The policy update steps are as follows: when the safety constraints are met, update the policy network by maximizing the cumulative reward; when the safety constraints are not met, update the policy network by minimizing the cumulative cost.

5. The end-to-end automatic parking path planning method according to claim 1, characterized in that, During training, the reward value network and cost value network are used to calculate the discrete action advantage function and continuous parameter advantage function of all action state pairs in the sample with respect to the reward value, as well as the discrete action advantage function and continuous parameter advantage function with respect to the cost value.

6. The end-to-end automatic parking path planning method according to claim 1, characterized in that, The surrounding environment information is obtained through an in-vehicle surround-view camera and / or ultrasonic sensors, and the vehicle parking requirements are obtained through a human-machine interface.

7. An end-to-end automated parking path planning system, characterized in that, include: The perception module is used to acquire environmental information around the vehicle, the pose of the target parking space, and the pose of the vehicle. Human-computer interaction interface, used to obtain parking demand information; The path planning module is configured with a trained parking path planning strategy, which is used to directly output a safe parking path composed of parameterized action sequences based on the observation state obtained by the perception module and the human-machine interface; the parameterized action sequences are used for vehicle control. And a vehicle control and execution module, used to complete automatic parking by following the safe parking path.

8. The end-to-end automatic parking path planning system according to claim 7, characterized in that, The parking path planning strategy is trained in the following way: Initialize the policy network parameters, reward value network parameters, and cost value network parameters; Sample data is collected in an automated parking simulation environment or a parking world model built based on real data using a policy network to calculate cumulative rewards and cumulative costs. The average cumulative cost of the trajectory is calculated based on the sample and compared with the cost threshold to determine whether the current strategy meets the safety constraints. Depending on whether the security constraints are met, the policy network should be updated by either maximizing the cumulative reward or minimizing the cumulative cost. Iterative training continues until the policy network converges and satisfies the cost constraint.