Method, device and equipment for constructing automatic parking path planning model based on deep reinforcement learning, and vehicle
By building an automatic parking path planning model based on deep reinforcement learning, combining global and local state information, and optimizing strategy exploration, the error and environmental disturbance problems in the existing system are solved, achieving more efficient parking path planning.
Patent Information
- Application Number
- CN202510314159.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-17
- Publication Date
- 2025-10-21
- Estimated Expiration
- 2045-03-17
Smart Images

Figure CN120003469B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of automobile autonomous driving technology, and specifically relates to a method, device, equipment and vehicle for constructing an automatic parking path planning model based on deep reinforcement learning. Background Art
[0002] With the development of intelligent vehicles, people have higher expectations for the driving experience, and autonomous driving technology has become a hot topic in the automotive industry. Automatic parking, as a key feature of autonomous driving, is also attracting increasing attention. Traditional automatic parking systems that use trajectory planning and tracking can suffer from tracking errors, actuator control errors, and environmental disturbances, resulting in inconsistencies between the planned and actual trajectories and poor parking performance. As one of the hottest research and application areas in artificial intelligence, deep reinforcement learning algorithms inherit the advantages of deep learning algorithms in perception and feature extraction. They can map vehicle state information into a feature space to achieve end-to-end learning, theoretically minimizing the negative impact of errors and disturbances.
[0003] In existing technologies, a hybrid A* algorithm is first used to search for an initial parking path, and then reinforcement learning is used to optimize the parking path. This method has drawbacks: it fails to account for the presence of other moving vehicles, its application scenarios are relatively limited, and its scalability and generalization capabilities are limited. Furthermore, because this existing technology calculates the initial parking path using random exploration and does not optimize the strategy solution space, the number of training parameters is large and the training time is relatively long. Summary of the Invention
[0004] The present invention provides a method, device, equipment and vehicle for constructing an automatic parking path planning model based on deep reinforcement learning, which are used to improve generalization ability and reduce training workload.
[0005] The technical solution of the present invention is:
[0006] The present invention provides a method for constructing an automatic parking path planning model based on deep reinforcement learning, comprising:
[0007] Constructing a parking scenario, which includes a parking area formed by a wall, obstacles formed by the vehicle, multiple other moving vehicles, multiple stationary vehicles, and multiple parking spaces within the parking area;
[0008] Build deep reinforcement learning networks;
[0009] Obtain the vehicle's state observation information in the parking scenario, which includes global state observation information and local state observation information. Global state observation information includes the relative position coordinates of the vehicle and all available parking spaces, as well as the relative position coordinates of the vehicle and other moving vehicles. Local state observation information includes the distance between the vehicle and obstacles in multiple directions.
[0010] The global state observation information and the local state observation information are input into the constructed deep reinforcement learning network to obtain the parking execution action of the vehicle;
[0011] The set feedback reward function is used to guide the deep reinforcement learning network to iteratively update parameters; after the number of iterations reaches the preset number, the required automatic parking path planning model is obtained.
[0012] Preferably, the step of inputting the global state observation information and the local state observation information into the constructed deep reinforcement learning network to obtain the parking execution action of the vehicle includes:
[0013] The global state observation information is input into the target selection strategy subnetwork of the deep reinforcement learning network to obtain the target parking space;
[0014] The local state observation information is input into the collision avoidance strategy subnetwork of the deep reinforcement learning network to obtain the current action of the vehicle.
[0015] Preferably, the step of using a set feedback reward function to guide the deep reinforcement learning network to iteratively update parameters includes:
[0016] Obtain the new global state observation information and feedback reward value after the vehicle performs the current action;
[0017] Input the global state observation information, current action, feedback reward value and new global state observation information into the collision avoidance strategy sub-network of the deep reinforcement learning network to obtain the action value function value;
[0018] The collision avoidance policy subnetwork of the deep reinforcement learning network updates its parameters based on the action-value function value so that it can choose actions that maximize the action-value function value in the future.
[0019] Preferably, the step of obtaining the feedback reward value after the vehicle performs the current action includes:
[0020] After the vehicle performs the current action, it obtains a time-step reward for punishing the vehicle's stay in the environment, a target proximity reward for measuring how close the vehicle's current state is to the success state, a collision avoidance reward for measuring whether the distance between the vehicle and other moving vehicles is too close, and a collision avoidance reward for measuring whether the distance between the vehicle and an obstacle is too close.
[0021] The feedback reward value after the vehicle performs the current action is determined based on the sum of the time step reward, the target proximity reward, the collision avoidance reward between the vehicle and other moving vehicles, and the collision avoidance reward between the vehicle and obstacles.
[0022] Preferably, the step of obtaining a target proximity reward for measuring the degree of proximity between the current state of the vehicle and the success state includes:
[0023] Obtain the distance between the vehicle's state observation information at the current control step and the successful state observation information;
[0024] Obtain the distance between the vehicle's state observation information at the previous control step and the successful state observation information;
[0025] If the distance between the state observation information of the vehicle in the previous control step and the successful state observation information is smaller than the distance between the state observation information of the vehicle in the current control step and the successful state observation information, it means that the vehicle is closer to the successful state in the current control step, and a positive reward is given to the target proximity reward used to measure the degree of proximity between the current state of the vehicle and the successful state; otherwise, a negative reward is given to the target proximity reward used to measure the degree of proximity between the current state of the vehicle and the successful state.
[0026] Preferably, the step of obtaining a collision avoidance reward for measuring whether the distance between the host vehicle and other moving vehicles is too close comprises:
[0027] Calculate the Euclidean distance between the vehicle and other moving vehicles based on their relative position coordinates;
[0028] If the Euclidean distance between the vehicle and other moving vehicles is less than the set first safety distance threshold, a negative reward is given to the collision avoidance reward used to measure whether the distance between the vehicle and other moving vehicles is too close; otherwise, a positive reward is given to the collision avoidance reward used to measure whether the distance between the vehicle and other moving vehicles is too close.
[0029] Preferably, the step of obtaining a collision avoidance reward for measuring whether the distance between the vehicle and the obstacle is too close includes:
[0030] Determine whether the distance between the vehicle and the obstacle in the target observation direction is less than a set second safety distance threshold;
[0031] If the distance between the vehicle and the obstacle in the target observation direction is less than the set second safety distance threshold, a negative reward is given to the collision avoidance reward used to measure whether the distance between the vehicle and the obstacle is too close; otherwise, a positive reward is given to the collision avoidance reward used to measure whether the distance between the vehicle and the obstacle is too close.
[0032] The present invention also provides an automatic parking path model construction device based on deep reinforcement learning, comprising:
[0033] A parking scene construction module is used to construct a parking scene, which includes a parking area formed by a wall, obstacles formed by the vehicle, multiple other moving vehicles, multiple stationary vehicles, and multiple parking spaces within the parking area;
[0034] Network building blocks for building deep reinforcement learning networks;
[0035] The state observation information acquisition module is used to obtain the state observation information of the vehicle in the parking scene. The state observation information includes global state observation information and local state observation information. The global state observation information includes the relative position coordinates of the vehicle and all available parking spaces, as well as the relative position coordinates of the vehicle and other moving vehicles. The local state observation information includes the distance between the vehicle and obstacles in multiple directions.
[0036] The parking execution action output module is used to input global state observation information and local state observation information into the constructed deep reinforcement learning network to obtain the vehicle's parking execution action;
[0037] The iteration module is used to use the set feedback reward function to guide the deep reinforcement learning network to iteratively update the parameters; after the number of iterations reaches the preset number, the required automatic parking path planning model is obtained.
[0038] The present invention also provides a control device, including a processor, a memory, and a program or instruction stored in the memory and executable on the processor. When the program or instruction is executed by the processor, the steps of the method for constructing an automatic parking path planning model based on deep reinforcement learning as described above are implemented.
[0039] The present invention also provides a vehicle, on which is deployed an automatic parking path planning model constructed according to the above-mentioned method for constructing an automatic parking path planning model based on deep reinforcement learning.
[0040] The beneficial effects of the present invention are:
[0041] The relative position of the vehicle and other moving vehicles is added to the vehicle's state observation information, and a corresponding collision avoidance reward is set in the feedback reward function to guide the parameter update of the deep reinforcement learning network. This strategy can effectively avoid collisions between the vehicle and other moving vehicles, and the parking scenario is more realistic, thus the method has stronger generalization ability. The policy exploration method has been optimized, and pre-processing of the vehicle's observation information has been added to the processing of the collision avoidance strategy subnetwork. Only action strategies that point to the vehicle's selected target direction within a certain range are selected. When there are no obstacles in the selected target direction, the current action that deviates from the target direction will not be selected. Therefore, the vehicle's path will approach the target parking space, and path strategies that deviate from the target parking space direction will not be attempted. This reduces the policy solution space, accelerates network convergence, and reduces training time. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] Figure 1 This is a flowchart of a method for constructing an automatic parking path planning model based on deep reinforcement learning in an embodiment of the present application;
[0043] Figure 2 This is a flowchart of step S4 in the embodiment of the present application;
[0044] Figure 3 This is a flowchart of step S5 in the embodiment of the present application;
[0045] Figure 4 This is a flowchart of a device for constructing an automatic parking path planning model based on deep reinforcement learning in an embodiment of the present application. DETAILED DESCRIPTION
[0046] To facilitate understanding by those skilled in the art, the present invention is further described below with reference to the accompanying drawings. The detailed description is complete, but it should not be construed as limiting the scope of the present invention. Obvious variations and alternative forms of the following examples are all within the scope of protection of this patent.
[0047] Reference Figure 1 , an embodiment of the present invention proposes a method for constructing an automatic parking path planning model based on deep reinforcement learning, comprising the following steps:
[0048] S1, build parking scene.
[0049] During the research and development, the parking scene constructed was a simulation environment, which included obstacles such as walls, a stationary vehicle, the vehicle itself, other moving vehicles, and the target parking location. The simulation environment size was set to 36×36 square meters. The N moving vehicles and M parking spaces in the environment were randomly generated. Ten obstacles were also randomly placed in the simulation environment. Each obstacle was randomly formed into a circle or square, and its diameter or side length was generated by a uniform distribution U(1m, 4m). The vehicle moved at a constant speed v=1m / s control step and had a steering angle that varied within the range of [-(Π / 2), (Π / 2)]. The vehicle used ranging beams in K directions to evenly divide the steering range and detect surrounding obstacles. The effective detection distance was set to 4 meters.
[0050] S2, build and initialize the deep reinforcement learning network.
[0051] In an embodiment of the present application, the deep reinforcement learning network consists of a target selection strategy sub-network and a collision avoidance strategy sub-network.
[0052] The value function of the target selection sub-policy is estimated by a neural network with two hidden layers, containing 300 and 200 units, respectively. The parameters of this neural network are learned using the RMSProp optimizer with a learning rate of 0.01. The experience replay pool has a capacity of 1500, and the stochastic gradient descent mini-batch size is 32. The target selection policy sub-network dynamically selects the target parking space for the vehicle based on the global state observations of the vehicle.
[0053] The collision avoidance sub-strategy uses the vehicle's local state observations to determine whether there is an obstacle ahead. If so, the collision avoidance sub-strategy outputs an angle, at which the vehicle turns and moves forward one step. If not, the vehicle moves one step toward the selected target parking position, where forward refers to the angular direction in which the vehicle turns.
[0054] In the embodiment of the present application, the collision avoidance sub-strategy is constructed as two networks, Critic and Actor, both of which have two hidden layers, and each layer contains 100 units.
[0055] The actor's output layer contains an activation unit consisting of a tanh function multiplied by π / 2, which improves the network's nonlinearity fitting capabilities. The Adam optimizer is used to learn the critic and actor parameters, with learning rates of 0.002 and 0.001, respectively. The experience replay pool has a capacity of 3000, and the mini-batch size for stochastic gradient descent is 32.
[0056] The Actor network selects the current action based on the vehicle's local observations. The input is the local state observations of the parking environment, as described in the first section, and the output is the vehicle's current action. Because the vehicle's speed is assumed to be constant, the current action in this embodiment is represented by the steering angle.
[0057] The role of the Critic network is to judge the quality of the current action and calculate the action value function value; its input is the current global state observation information, the current action, the feedback reward value, and the global state observation information after executing the current action, and the output is the action value function value.
[0058] S3. Obtain state observation information of the vehicle in the parking scene, where the state observation information includes global state observation information and local state observation information.
[0059] In this embodiment, the vehicle will output two parts of state observation information in the parking environment, specifically including: global state observation information and local state observation information.
[0060] The global state observation information includes the relative position coordinates of vehicle i and all parking spaces at the t-th control step, as well as the relative position coordinates of vehicle i and other moving vehicles, which can be expressed as These locations can be obtained through positioning systems such as GPS or vehicle-mounted radar.
[0061] The local state observation information includes the distance between vehicle i and its surrounding obstacles in K directions at the tth control step, which is expressed as These distances can be measured using a rangefinder.
[0062] S4. Input the global state observation information and the local state observation information into the constructed deep reinforcement learning network to obtain the parking execution action of the vehicle.
[0063] Reference Figure 2 In the embodiment of the present application, step S4 of inputting global state observation information and local state observation information into the constructed deep reinforcement learning network to obtain the parking execution action of the vehicle includes:
[0064] S41. Input the global state observation information into the target selection strategy subnetwork of the deep reinforcement learning network to obtain the target parking space;
[0065] S42. Input the local state observation information into the collision avoidance strategy subnetwork of the deep reinforcement learning network to obtain the current action of the vehicle.
[0066] S5. Use the set feedback reward function to guide the deep reinforcement learning network to iteratively update parameters; after the number of iterations reaches the preset number, the required automatic parking path planning model is obtained.
[0067] In this embodiment of the present application, the feedback reward is the feedback from the environment on the state transition result after the vehicle makes the current action according to the policy network. The feedback reward function is expressed as
[0068] in, It is the time step reward given to vehicle i at each control step. The time step reward is a negative constant reward used to punish the vehicle for each time step it stays in the environment. Its function is to encourage the vehicle to complete the parking task as quickly as possible.
[0069] It is a target proximity reward used to measure the closeness between the vehicle's current state information and the successful state information. The target proximity reward is calculated by comparing the distance between the vehicle and the successful state in the current step and the previous step.
[0070] in,
[0071] represents the distance between the state observation information of vehicle i at the tth control step and the successful state observation information;
[0072] It represents the distance between the state observation information of vehicle i at the t-1th control step and the successful state observation information.
[0073] The state observation information of vehicle i at the tth control step includes the current global state observation information and local state observation information of the vehicle.
[0074] A successful state observation refers to the ideal state of the vehicle when completing the parking task. A successful state observation can be represented as an ideal state vector. In this state, the vehicle is accurately parked in the target parking space and meets all safety and positioning requirements. Specifically, a successful state observation may include the following: the relative position of the vehicle and the target parking space: the coordinate deviation between the vehicle's center point and the target parking space's center point is close to zero; the vehicle's orientation: the vehicle's orientation is consistent with the orientation of the target parking space; and the safe distance from other vehicles and obstacles: the distance between the vehicle and surrounding obstacles or other vehicles meets safety requirements.
[0075] In this embodiment of the present application, the distance between the state observation information of vehicle i at the tth control step and the successful state observation information is used to measure the closeness between the current state of the vehicle and the target state. This distance can be a comprehensive indicator, including the following:
[0076] Position deviation: The coordinate deviation between the vehicle's current position and the center point of the target parking space.
[0077] Orientation deviation: The deviation between the vehicle's current orientation and the orientation of the target parking space.
[0078] Safety distance deviation: Whether the distance between the vehicle and surrounding obstacles or vehicles meets safety requirements.
[0079] When the vehicle i is closer to the successful state (i.e. ),but is a positive value, indicating a reward; when the vehicle i is closer to the successful state (i.e. ) A negative value indicates a penalty.
[0080] Indicates the collision avoidance reward between the vehicle i and other moving vehicles. The two parts of the collision avoidance reward are used to penalize the situation where the distance between the vehicle i and other vehicles or obstacles is too close to ensure the safety of the parking process.
[0081] in, When the vehicle i and other moving vehicles car j The Euclidean distance between Less than the set first safety distance threshold d s1 , the value of the step function u is 1, otherwise it is 0. s1 The value of is a constant; C2 is the weight coefficient of the collision avoidance reward, which is a positive constant.
[0082] This car i and other sports vehicles car j The Euclidean distance between It is obtained by calculating the L2 norm of the difference vector of the two vehicle positions. For example, the position of vehicle i is (x i ,y i ), other sports vehicles j The position is (x j ,y j ), then the Euclidean distance between them is: This distance is used to calculate the collision avoidance reward To ensure that the car and other sports vehicles j Maintain a safe distance; if this distance is less than the set first safety distance threshold d s1 , a negative reward will be given to vehicle i to punish it for being too close to another moving vehicle, thereby avoiding potential collision risks.
[0083] When the distance between vehicle i and the obstacle in the kth target observation direction Less than the set second safety distance threshold d s2 , the value of the step function u is 1, otherwise it is 0. s2 The value of is a constant; C3 is the weight coefficient of the collision avoidance reward, which is a positive constant. The value of the step function u is 1, indicating that a negative reward will be given to vehicle i to punish it for being too close to another moving vehicle, thereby avoiding potential collision risks.
[0084] Reference Figure 3 , step S5 of using the set feedback reward function to guide the deep reinforcement learning network to iteratively update parameters includes:
[0085] S51, obtaining new global state observation information and feedback reward value after the vehicle performs the current action;
[0086] S52, inputting the global state observation information, the current action, the feedback reward value, and the new global state observation information into the collision avoidance strategy subnetwork of the deep reinforcement learning network to obtain the action value function value;
[0087] S53. The collision avoidance strategy subnetwork of the deep reinforcement learning network updates its parameters according to the action-value function value so that it can select actions that maximize the action-value function value in the future.
[0088] In conjunction with the above description, in the embodiment of the present application, the step of obtaining the feedback reward value after the vehicle performs the current action specifically includes:
[0089] After the car performs the current action, it obtains the time step reward used to punish the car for staying in the environment Target proximity reward used to measure how close the vehicle's current state is to the successful state Collision avoidance reward used to measure whether the distance between the vehicle and other moving vehicles is too close and a collision avoidance reward to measure whether the distance between the vehicle and the obstacle is too close.
[0090] Based on the time step reward Target proximity reward Reward for avoiding collision between this vehicle and other moving vehicles And the collision avoidance reward between the vehicle and the obstacle The sum of the two determines the feedback reward value after the vehicle performs the current action.
[0091] After learning is completed, the learned strategy is deployed to the vehicle, and the vehicle can be controlled to complete the parking task in an unknown environment.
[0092] The above-mentioned method of this embodiment adds the relative position between the vehicle and other moving vehicles to the state observation information of the vehicle, and sets the collision avoidance reward between the vehicle and other moving vehicles in the feedback reward function accordingly, thereby guiding the parameter update of the deep reinforcement learning network. This strategy can effectively avoid collisions between the vehicle and other moving vehicles, and the parking scenario is closer to the actual situation. Therefore, this method has stronger generalization ability. The strategy exploration method has been optimized, and pre-processing of the vehicle observation information has been added to the processing process of the collision avoidance strategy sub-network. Only action strategies pointing to a certain range of the selected target direction of the vehicle are selected. When there are no obstacles in the selected target direction, the current action that deviates from the target direction will not be selected. Therefore, the path of the vehicle will be close to the target parking space, and the path strategy that deviates from the target parking space direction will not be attempted. This reduces the strategy solution space, accelerates network convergence, and reduces training time.
[0093] Reference Figure 4 The present invention also provides an automatic parking path model construction device based on deep reinforcement learning, comprising:
[0094] A parking scene construction module 101 is used to construct a parking scene, which includes a parking area formed by a wall, obstacles formed by the vehicle, multiple other moving vehicles, multiple stationary vehicles, and multiple parking spaces within the parking area;
[0095] A network construction module 102 is used to construct a deep reinforcement learning network;
[0096] The state observation information acquisition module 103 is used to obtain the state observation information of the vehicle in the parking scene. The state observation information includes global state observation information and local state observation information. The global state observation information includes the relative position coordinates of the vehicle and all available parking spaces and the relative position coordinates of the vehicle and other moving vehicles. The local state observation information includes the distance between the vehicle and obstacles in multiple directions.
[0097] The parking execution action output module 104 is used to input the global state observation information and the local state observation information into the constructed deep reinforcement learning network to obtain the parking execution action of the vehicle;
[0098] The iteration module 105 is used to use the set feedback reward function to guide the deep reinforcement learning network to perform iterative parameter updates; after the number of iterations reaches a preset number, the required automatic parking path planning model is obtained.
[0099] The present invention also provides a control device, including a processor, a memory, and a program or instruction stored in the memory and executable on the processor. When the program or instruction is executed by the processor, the steps of the method for constructing an automatic parking path planning model based on deep reinforcement learning as described above are implemented.
[0100] The present invention also provides a vehicle, on which is deployed an automatic parking path planning model constructed according to the above-mentioned method for constructing an automatic parking path planning model based on deep reinforcement learning.
[0101] It should be noted that the various embodiments in this specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same and similar parts between the various embodiments can be referenced to each other.
[0102] Although the preferred embodiments of the present invention have been described, those skilled in the art may make additional changes and modifications to these embodiments once they become aware of the basic creative concepts. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the embodiments of the present invention.
[0103] It should also be noted that, in this document, the terms "center", "up", "down", "left", "right", "vertical", "horizontal", "inside", "outside" and the like indicate positions or positional relationships based on the positions or positional relationships shown in the accompanying drawings, and are for the purpose of facilitating the description of the present invention and simplifying the description, rather than indicating or implying that the devices or components referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be understood as limiting the present invention. In addition, relational terms such as "first" and "second" are used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply that there is any actual relationship or order between these entities or operations, nor should they be understood as indicating or implying relative importance. Moreover, the terms "comprises", "includes" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or terminal device comprising a series of elements does not include those elements, but also includes other elements not explicitly listed, or also includes elements inherent to such process, method, article or terminal device. In the absence of further restrictions, an element defined by the sentence "comprising a ..." does not exclude the presence of other identical elements in the process, method, article or terminal device comprising the element.
[0104] The technical solutions provided by the present invention have been described in detail above. Specific examples are used herein to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is intended only to facilitate understanding of the present invention, and the contents of this specification should not be construed as limiting the present invention. Furthermore, those skilled in the art will appreciate that various modifications may be made to the specific implementation methods and scope of application according to the present invention. It is not necessary and impossible to exhaustively enumerate all implementation methods herein, and any obvious variations or modifications derived therefrom remain within the scope of protection of the present invention.
Claims
1. A method for constructing an automatic parking path planning model based on deep reinforcement learning, characterized in that: include: Constructing a parking scenario, which includes a parking area formed by a wall, obstacles formed by the vehicle, multiple other moving vehicles, multiple stationary vehicles, and multiple parking spaces within the parking area; Build deep reinforcement learning networks; Obtaining state observation information of the vehicle in the parking scene, including global state observation information and local state observation information; Global state observation information includes the relative position coordinates of the vehicle and all available parking spaces, as well as the relative position coordinates of the vehicle and other moving vehicles; local state observation information includes the distance between the vehicle and obstacles in multiple directions; The global state observation information and the local state observation information are input into the constructed deep reinforcement learning network to obtain the parking execution action of the vehicle; The steps of using a set feedback reward function to guide the deep reinforcement learning network to iteratively update parameters; obtaining the required automatic parking path planning model after the number of iterations reaches a preset number; inputting global state observation information and local state observation information into the constructed deep reinforcement learning network to obtain the vehicle's parking execution action include: The global state observation information is input into the target selection strategy subnetwork of the deep reinforcement learning network to obtain the target parking space; The local state observation information is input into the collision avoidance strategy subnetwork of the deep reinforcement learning network to obtain the current action of the vehicle. The steps of using the set feedback reward function to guide the deep reinforcement learning network to iteratively update the parameters include: Obtain the new global state observation information and feedback reward value after the vehicle performs the current action; Input the global state observation information, current action, feedback reward value and new global state observation information into the collision avoidance strategy sub-network of the deep reinforcement learning network to obtain the action value function value; The collision avoidance policy subnetwork of the deep reinforcement learning network updates its parameters based on the action-value function value so that it can choose actions that maximize the action-value function value in the future.
2. The method for constructing an automatic parking path planning model based on deep reinforcement learning according to claim 1, characterized in that: The steps to obtain the feedback reward value after the vehicle performs the current action include: After the vehicle performs the current action, it obtains a time-step reward for punishing the vehicle's stay in the environment, a target proximity reward for measuring how close the vehicle's current state is to the success state, a collision avoidance reward for measuring whether the distance between the vehicle and other moving vehicles is too close, and a collision avoidance reward for measuring whether the distance between the vehicle and an obstacle is too close. The feedback reward value after the vehicle performs the current action is determined based on the sum of the time step reward, the target proximity reward, the collision avoidance reward between the vehicle and other moving vehicles, and the collision avoidance reward between the vehicle and obstacles.
3. The method for constructing an automatic parking path planning model based on deep reinforcement learning according to claim 2, characterized in that: The steps to obtain the target proximity reward used to measure how close the current state of the vehicle is to the successful state include: Obtain the distance between the vehicle's state observation information at the current control step and the successful state observation information; Obtain the distance between the vehicle's state observation information at the previous control step and the successful state observation information; If the distance between the state observation information of the vehicle in the previous control step and the successful state observation information is smaller than the distance between the state observation information of the vehicle in the current control step and the successful state observation information, it means that the vehicle is closer to the successful state in the current control step, and a positive reward is given to the target proximity reward used to measure the degree of proximity between the current state of the vehicle and the successful state; otherwise, a negative reward is given to the target proximity reward used to measure the degree of proximity between the current state of the vehicle and the successful state.
4. The method for constructing an automatic parking path planning model based on deep reinforcement learning according to claim 2, characterized in that: The steps to obtain the collision avoidance reward used to measure whether the distance between the own vehicle and other moving vehicles is too close include: Calculate the Euclidean distance between the vehicle and other moving vehicles based on their relative position coordinates; If the Euclidean distance between the vehicle and other moving vehicles is less than the set first safety distance threshold, a negative reward is given to the collision avoidance reward used to measure whether the distance between the vehicle and other moving vehicles is too close; otherwise, a positive reward is given to the collision avoidance reward used to measure whether the distance between the vehicle and other moving vehicles is too close.
5. The method for constructing an automatic parking path planning model based on deep reinforcement learning according to claim 2, characterized in that: The steps to obtain the collision avoidance reward used to measure whether the distance between the vehicle and the obstacle is too close include: Determine whether the distance between the vehicle and the obstacle in the target observation direction is less than a set second safety distance threshold; If the distance between the vehicle and the obstacle in the target observation direction is less than the set second safety distance threshold, a negative reward is given to the collision avoidance reward used to measure whether the distance between the vehicle and the obstacle is too close; otherwise, a positive reward is given to the collision avoidance reward used to measure whether the distance between the vehicle and the obstacle is too close.
6. A device for constructing an automatic parking path model based on deep reinforcement learning, characterized in that: include: A parking scene construction module is used to construct a parking scene, which includes a parking area formed by a wall, obstacles formed by the vehicle, multiple other moving vehicles, multiple stationary vehicles, and multiple parking spaces within the parking area; Network building blocks for building deep reinforcement learning networks; A state observation information acquisition module is used to obtain the state observation information of the vehicle in the parking scene. The state observation information includes global state observation information and local state observation information; Global state observation information includes the relative position coordinates of the vehicle and all available parking spaces, as well as the relative position coordinates of the vehicle and other moving vehicles; local state observation information includes the distance between the vehicle and obstacles in multiple directions; The parking execution action output module is used to input global state observation information and local state observation information into the constructed deep reinforcement learning network to obtain the vehicle's parking execution action; An iteration module is used to use a set feedback reward function to guide the deep reinforcement learning network to iteratively update parameters; after the number of iterations reaches a preset number, the required automatic parking path planning model is obtained; The steps of inputting global state observation information and local state observation information into the constructed deep reinforcement learning network to obtain the vehicle's parking execution action include: The global state observation information is input into the target selection strategy subnetwork of the deep reinforcement learning network to obtain the target parking space; The local state observation information is input into the collision avoidance strategy subnetwork of the deep reinforcement learning network to obtain the current action of the vehicle. The steps of using the set feedback reward function to guide the deep reinforcement learning network to iteratively update the parameters include: Obtain the new global state observation information and feedback reward value after the vehicle performs the current action; Input the global state observation information, current action, feedback reward value and new global state observation information into the collision avoidance strategy sub-network of the deep reinforcement learning network to obtain the action value function value; The collision avoidance policy subnetwork of the deep reinforcement learning network updates its parameters based on the action-value function value so that it can choose actions that maximize the action-value function value in the future.
7. A control device, characterized in that: The invention comprises a processor, a memory, and a program or instruction stored in the memory and executable on the processor, wherein when the program or instruction is executed by the processor, the steps of the method for constructing an automatic parking path planning model based on deep reinforcement learning are implemented as claimed in any one of claims 1 to 5.
8. A vehicle, characterized in that: The vehicle is deployed with an automatic parking path planning model constructed according to the automatic parking path planning model construction method based on deep reinforcement learning according to any one of claims 1-5.
Citation Information
Patent Citations
Parking task allocation and trajectory planning system based on multi-agent reinforcement learning
CN116620264A
Method for operating at least an assisted parking function for a vehicle
DE102023117282A1