A method and system for intelligent driving path planning based on reinforcement learning

Through the intelligent driving path planning method based on reinforcement learning, the traffic environment is analyzed in real time and the path strategy is dynamically adjusted, which solves the problem that traditional methods are difficult to take into account efficiency and safety in complex traffic environments, and achieves safe and efficient driving path planning.

CN119469192BActive Publication Date: 2025-05-06SHENZHEN SHENHANG HUACHUANG AUTOMOBILE TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510058234.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-14
Publication Date
2025-05-06
Estimated Expiration
2045-01-14

AI Technical Summary

Technical Problem

Traditional path planning methods are difficult to adapt to real-time changing traffic environments and dynamic road conditions information, especially in urban traffic congestion, emergencies and complex road conditions, and it is difficult to take into account efficiency and safety.

Method used

The intelligent driving path planning method based on reinforcement learning is adopted to analyze the traffic environment in real time, adjust the path planning strategy dynamically, and generate safe and efficient driving paths. The specific steps include dividing the detection area into multiple grids, acquiring detection data, analyzing and identifying the target's movement path and driving probability, calculating the intersection risk score, and adjusting the driving path.

Benefits of technology

It realizes the efficiency and safety of intelligent driving in complex traffic environments, can dynamically adjust paths to adapt to real-time traffic changes, and improves the safety and efficiency of the driving process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119469192B_ABST
    Figure CN119469192B_ABST
Patent Text Reader

Abstract

The present invention discloses a method and system for intelligent driving path planning based on reinforcement learning, the method comprising: dividing a detection area into multiple grids; acquiring detection data; inputting the detection data into a preset target recognition model for analysis, and outputting recognition data of the recognition target; analyzing the recognition data of the recognition target through a preset path planning model, determining the moving path of the recognition target, determining the second predicted driving probability and the second occupied time interval of each grid; calculating the intersection risk score of each grid, intercepting the preset driving path of the current vehicle, and determining the first driving path; adjusting the first driving path according to the intersection risk score of each grid on the first driving path. The present invention generates a safe and efficient driving path by analyzing the traffic environment in real time, thereby ensuring the safety of intelligent driving.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Please relate to the field of intelligent driving technology, and more specifically, to an intelligent driving path planning method and system based on reinforcement learning. Background Art

[0002] With the rapid development of intelligent driving technology, path planning, as a core component of intelligent driving systems, has become increasingly important. However, traditional path planning methods are mostly based on static maps and fixed traffic rules, which are difficult to adapt to real-time changing traffic environments and dynamic road conditions. Especially in the face of urban traffic congestion, emergencies and complex road conditions, traditional path planning methods often find it difficult to balance efficiency and safety.

[0003] Therefore, the prior art has defects and is in urgent need of improvement. Summary of the invention

[0004] In view of the above problems, the purpose of the present invention is to provide an intelligent driving path planning method and system based on reinforcement learning, which can generate a safe and efficient driving path by analyzing the traffic environment in real time, dynamically adjusting the path planning strategy, and ensuring the efficiency and safety of intelligent driving in complex traffic environments.

[0005] A first aspect of the present invention provides an intelligent driving path planning method based on reinforcement learning, comprising:

[0006] Divide the detection area into multiple grids;

[0007] Obtain test data;

[0008] The detection data is input into a preset target recognition model for analysis, and recognition data of the recognition target is output; the recognition data of the recognition target is analyzed by a preset path planning model to determine a moving path of the recognition target, and a first predicted travel probability and a first occupied time interval of the recognition target passing through each grid are determined; the moving path includes a first moving path and a second moving path;

[0009] Integrate the first predicted travel probability and the first occupied time interval of all identified targets passing through each grid respectively, and determine the second predicted travel probability and the second occupied time interval of each grid;

[0010] Analyze the preset driving path of the current vehicle through the preset path planning model, determine the third occupied time interval of the current vehicle passing through each grid, and calculate the intersection risk score of each grid in combination with the second predicted driving probability and the second occupied time interval of each grid;

[0011] Intercepting a preset driving path of the current vehicle based on a preset driving time and a driving speed of the current vehicle to determine a first driving path;

[0012] The first driving path is adjusted according to the intersection risk score of each grid on the first driving path.

[0013] In this solution, the identification data of the identified target is analyzed by a preset path planning model to determine the moving path of the identified target, including:

[0014] Filtering a plurality of identification data of the identified target within a first preset time interval from the historical detection data, and drawing a first moving path of the identified target;

[0015] The first movement data of the identified target is input into a preset path planning model, and a second movement path of the identified target in a second preset time interval is output.

[0016] In this solution, determining the first predicted travel probability and the first occupied time interval of the identified target passing through each grid includes:

[0017] Analyze the first movement data and road traffic data of the identified target through a preset path planning model, calculate the movement probability of the identified target in each direction, and predict the first predicted travel probability and the first predicted passing time of the identified target passing through each grid;

[0018] A first preset time length and a second preset time length are respectively added before and after the first predicted elapsed time to determine a first occupied time interval passing through the corresponding grid.

[0019] In this solution, the first predicted travel probability and the first occupied time interval of all identified targets passing through each grid are integrated to determine the second predicted travel probability and the second occupied time interval of each grid, including:

[0020] When there are predicted travel data of multiple identified targets in the grid, sorting the first occupied time intervals of the multiple identified targets in chronological order;

[0021] Calculating time intervals of adjacent first occupied time intervals, and merging adjacent first occupied time intervals whose time intervals are shorter than a preset time interval;

[0022] The predicted driving probability of the overlapping part of adjacent first occupied time intervals is determined as the larger first predicted driving probability among the corresponding identified targets, and the predicted driving probability of the interval part is determined as the smaller first predicted driving probability among the corresponding identified targets.

[0023] In this solution, the preset driving path of the current vehicle is analyzed by the preset path planning model to determine the third occupied time interval of each grid when the current vehicle passes through, and the intersection risk score of each grid is calculated by combining the second predicted driving probability and the second occupied time interval of each grid, including:

[0024] Get the preset driving path of the current vehicle;

[0025] Inputting the preset driving path of the current vehicle into a preset path planning model, analyzing it in combination with road traffic data, and outputting a second predicted passing time of the current vehicle passing through each grid;

[0026] Adding a first preset time length and a second preset time length before and after the second predicted passing time, respectively, to determine a third occupied time interval of the current vehicle passing through the corresponding grid;

[0027] Calculate the same occupied time interval of the second occupied time interval and the corresponding third occupied time interval of each grid;

[0028] Calculating a time length ratio of the same occupied time interval and a corresponding third occupied time interval, and determining a first influence weight of a second predicted travel probability in the corresponding grid;

[0029] Multiplying the second predicted travel probability corresponding to the same occupied time interval in the grid by the corresponding first impact weight to determine the intersection risk score of the grid;

[0030] The risk gradient curve is drawn according to the intersection risk scores of all grids.

[0031] In this solution, adjusting the first driving path according to the intersection risk score of each grid on the first driving path includes:

[0032] Marking grids in the first driving path whose intersection risk scores are greater than a first preset risk score threshold for avoidance;

[0033] determining a second impact weight of the intersection risk score of each grid according to a relative distance between the grid and the current position of the current vehicle;

[0034] Calculating a weighted average of the intersection risk score of each grid in the first driving path and the corresponding second impact weight to determine an average risk score of the first driving path;

[0035] When the average risk score is less than the second preset risk score threshold, no adjustment is made;

[0036] When the average risk score is greater than or equal to a second preset risk score threshold, marking the grids in the first driving path whose intersection risk scores are greater than a third preset risk score threshold for avoidance;

[0037] The first driving path is adjusted according to the grid where the avoidance mark exists.

[0038] In this solution, the first driving path is adjusted according to the grid with the avoidance mark, including:

[0039] Calculate the driving score of each grid based on the driving parameters of the current vehicle and the intersection risk score of each grid;

[0040] Determine an obstacle avoidance starting point and an obstacle avoidance end point according to the grids with avoidance marks in the first driving path;

[0041] Determine one or more obstacle avoidance paths according to the driving score, obstacle avoidance starting point, and obstacle avoidance end point of each grid;

[0042] Calculate the average driving score of each obstacle avoidance path, and filter the obstacle avoidance paths whose average driving score is less than the preset driving score threshold;

[0043] Calculate the obstacle avoidance score of the obstacle avoidance path according to the area of ​​the area enclosed by the obstacle avoidance path and the first driving path;

[0044] Determine the obstacle avoidance path with the smallest obstacle avoidance score as the second driving path;

[0045] The second driving path is replaced with a corresponding first driving path portion according to the obstacle avoidance starting point and the obstacle avoidance end point.

[0046] This plan also includes:

[0047] The grid specifications are dynamically adjusted according to the current vehicle speed, driving status, target type and target ratio in the monitoring area.

[0048] A second aspect of the present invention provides an intelligent driving path planning system based on reinforcement learning, comprising:

[0049] A region segmentation module is used to divide the detection area into multiple grids;

[0050] A data acquisition module, used for acquiring detection data;

[0051] A first analysis module is used to input the detection data into a preset target recognition model for analysis, and output recognition data of the recognition target; the recognition data of the recognition target is analyzed by a preset path planning model to determine a moving path of the recognition target, and to determine a first predicted travel probability and a first occupied time interval of the recognition target passing through each grid; the moving path includes a first moving path and a second moving path;

[0052] The second analysis module is used to integrate the first predicted travel probability and the first occupied time interval of all the identified targets passing through each grid, and determine the second predicted travel probability and the second occupied time interval of each grid;

[0053] A third analysis module is used to analyze the preset driving path of the current vehicle through a preset path planning model, determine the third occupied time interval of the current vehicle passing through each grid, and calculate the intersection risk score of each grid in combination with the second predicted driving probability and the second occupied time interval of each grid;

[0054] The path planning module is used to intercept the preset driving path of the current vehicle based on the preset driving time and the driving speed of the current vehicle to determine a first driving path; and adjust the first driving path according to the intersection risk score of each grid on the first driving path.

[0055] A third aspect of the present invention provides a computer-readable storage medium, which includes a program for an intelligent driving path planning method based on reinforcement learning. When the program for an intelligent driving path planning method based on reinforcement learning is executed by a processor, the steps of the above-mentioned intelligent driving path planning method based on reinforcement learning are implemented.

[0056] The present invention discloses a method and system for intelligent driving path planning based on reinforcement learning, the method comprising: dividing a detection area into multiple grids; acquiring detection data; inputting the detection data into a preset target recognition model for analysis, and outputting recognition data of the recognition target; analyzing the recognition data of the recognition target through a preset path planning model, determining the moving path of the recognition target, determining the second predicted driving probability and the second occupied time interval of each grid; calculating the intersection risk score of each grid, intercepting the preset driving path of the current vehicle, and determining the first driving path; adjusting the first driving path according to the intersection risk score of each grid on the first driving path. The present invention generates a safe and efficient driving path by analyzing the traffic environment in real time, thereby ensuring the safety of intelligent driving. BRIEF DESCRIPTION OF THE DRAWINGS

[0057] Figure 1 A flow chart of an intelligent driving path planning method based on reinforcement learning provided by the present invention is shown;

[0058] Figure 2 A flowchart of a method for determining a first predicted travel probability and a first occupied time interval of an identification target passing through a grid provided by the present invention is shown;

[0059] Figure 3 A flowchart showing a method for determining a second predicted travel probability and a second occupied time interval of a grid provided by the present invention is shown;

[0060] Figure 4 A block diagram of an intelligent driving path planning system based on reinforcement learning provided by the present invention is shown. DETAILED DESCRIPTION

[0061] In order to more clearly understand the above-mentioned purpose, features and advantages of the present invention, the present invention is further described in detail below in conjunction with the accompanying drawings and specific embodiments. It should be noted that the embodiments of the present application and the features in the embodiments can be combined with each other without conflict.

[0062] In the following description, many specific details are set forth to facilitate a full understanding of the present invention. However, the present invention may also be implemented in other ways different from those described herein. Therefore, the protection scope of the present invention is not limited to the specific embodiments disclosed below.

[0063] Figure 1 A flow chart of an intelligent driving path planning method based on reinforcement learning provided by the present invention is shown.

[0064] like Figure 1 As shown, the present invention discloses an intelligent driving path planning method based on reinforcement learning, comprising:

[0065] S102, dividing the detection area into multiple grids;

[0066] S104, acquiring detection data;

[0067] S106, inputting the detection data into a preset target recognition model for analysis, and outputting recognition data of the recognition target; analyzing the recognition data of the recognition target through a preset path planning model, determining a moving path of the recognition target, and determining a first predicted driving probability and a first occupied time interval of the recognition target passing through each grid; the moving path includes a first moving path and a second moving path;

[0068] S108, respectively integrating the first predicted travel probability and the first occupied time interval of all identified targets passing through each grid, and determining the second predicted travel probability and the second occupied time interval of each grid;

[0069] S110, analyzing the preset driving path of the current vehicle through a preset path planning model, determining the third occupied time interval of each grid when the current vehicle passes through, and calculating the intersection risk score of each grid in combination with the second predicted driving probability and the second occupied time interval of each grid;

[0070] S112, intercepting a preset driving path of the current vehicle based on the preset driving time and the current driving speed of the vehicle to determine a first driving path;

[0071] S114: Adjust the first driving path according to the intersection risk score of each grid on the first driving path.

[0072] According to an embodiment of the present invention, the system first divides the detection area into multiple grids of the same size according to a pre-set grid specification. It should be noted that the size of the grid specification is dynamically adjusted according to the current vehicle speed, driving status (such as straight driving, turning, etc.), target type (such as pedestrians, vehicles, etc.) and target ratio in the monitoring area.

[0073] The detection data at least includes the image data acquired by the vehicle-mounted camera and the signal data acquired by the vehicle-mounted radar. First, the acquired detection data is preprocessed by data cleaning, image denoising and other preprocessing steps, and the preprocessed detection data is input into the preset target recognition model. The preprocessed detection data is analyzed by the preset target recognition model, the target in the detection data is identified, and the recognition data of the identified target is determined. Among them, the recognition data of the identified target at least includes the target type (such as pedestrians, vehicles, etc.) and position coordinates of the identified target. The recognition data of the identified target within a certain period of time is retrieved from the historical detection data, and the first moving path of the identified target is drawn according to the position coordinates in the recognition data. The first moving path of the identified target is input into the preset path planning model. The preset path planning model performs path planning for the identified target based on the road traffic data, calculates the movement probability of the identified target in each direction, and determines the first predicted driving probability and the first occupied time interval passing through each grid in combination with the moving speed of the identified target, and determines the moving path formed by the grid with the highest first predicted driving probability as the second moving path of the identified target.

[0074] When there are multiple identification targets in the detection area, the first predicted travel probability and the first occupied time interval of each identification target passing through the grid are integrated according to the grid to determine the second predicted travel probability and the second occupied time interval of the grid. The third occupied time interval passing through each grid is calculated according to the preset driving path of the current vehicle through the preset path planning model, and the same occupied time interval as the second occupied time interval of each grid is calculated. According to the time length ratio of the same occupied time interval and the corresponding third occupied time interval, the first influence weight of the second predicted travel probability in the corresponding grid is determined, and the corresponding second predicted travel probability in the grid is weightedly calculated to determine the intersection risk score of each grid, which is used to indicate the possibility that the current vehicle and the identification target enter the same grid at the same time. The larger the intersection risk score, the higher the possibility that the current vehicle and the identification target enter the same grid at the same time.

[0075] Based on the preset driving time and the driving speed of the current vehicle, a driving path of a certain distance is intercepted within the preset driving path of the current vehicle, and determined as the first driving path. The grids in the first driving path whose intersection risk scores are greater than the first preset risk score threshold are marked for avoidance. The intersection risk scores of the remaining grids in the first driving path are weighted according to the relative distance between the grid and the current position of the current vehicle to determine the average risk score of the first driving path. When the average risk score is greater than or equal to the second preset risk score threshold, the grids in the first driving path whose intersection risk scores are greater than the third preset risk score threshold are marked for avoidance. The obstacle avoidance starting point and obstacle avoidance end point are determined according to the grids with avoidance marks for analysis to determine the second driving path, and the corresponding first driving path part is replaced to complete the adjustment of the first driving path. In addition, during the driving process of the current vehicle, the driving path of the current vehicle is adjusted in real time according to the distribution of the identified targets and the road traffic status.

[0076] Among them, the preset target recognition model and the preset path planning model are both trained by historical detection data collected during the vehicle driving process.

[0077] Among them, the preset driving time and the preset driving route are set by technical personnel in this field according to actual needs.

[0078] According to an embodiment of the present invention, the identification data of the identification target is analyzed by a preset path planning model to determine the moving path of the identification target, including:

[0079] Filtering a plurality of identification data of an identified target within a first preset time interval from historical detection data, and drawing a first moving path of the identified target;

[0080] The first movement data of the identified target is input into a preset path planning model, and a second movement path of the identified target in a second preset time interval is output.

[0081] It should be noted that during the current vehicle driving process, the road conditions are monitored in real time, and each target in the acquired detection data is monitored and identified. After the identification data of the identified target in the current detection data is determined, multiple identification data of the identified target are retrieved from the historical detection data within a first preset time interval based on the identification data of the identified target, and multiple coordinate data of the identified target during the movement are extracted, and the multiple coordinate data are connected in chronological order to determine the first moving path of the identified target, that is, the moving path of the identified target within the preset time interval.

[0082] Then, a preset path planning model is used to perform movement path prediction and extension according to the first movement path of the identified target, so as to determine a second movement path of the identified target in a second preset time interval in the future.

[0083] Among them, the end time of the first preset time interval is the current time, and the start time of the second preset time interval is the current time. The interval sizes of the first preset time interval and the second preset time interval are set by technical personnel in this field according to actual needs.

[0084] Figure 2 A flow chart of a method for determining a first predicted travel probability and a first occupied time interval of an identification target passing through a grid provided by the present invention is shown.

[0085] like Figure 2 As shown, according to an embodiment of the present invention, determining a first predicted travel probability and a first occupied time interval of an identified target passing through each grid includes:

[0086] S202, analyzing the first movement data and road traffic data of the identified target through a preset path planning model, calculating the movement probability of the identified target in each direction, and predicting the first predicted travel probability and the first predicted passing time of the identified target passing through each grid;

[0087] S204: Add a first preset time length and a second preset time length before and after the first predicted passing time, respectively, to determine a first occupied time interval passing through the corresponding grid.

[0088] It should be noted that the preset path planning model determines whether the identified target has changed its moving path based on the second moving path of the identified target according to the road traffic data, and determines the first predicted driving probability of each grid on each other alternative path by predicting the moving probability of the identified target to other alternative paths in each direction. And the first predicted passing time of each grid is predicted based on the road traffic data and the moving speed of the identified target. Among them, the road traffic data is obtained by analyzing the detection data, and the road traffic data includes the coordinate data of each identified target within a certain range where the current vehicle is located, the target state (such as stationary or moving), etc.

[0089] In order to avoid the influence of the time error between the actual elapsed time and the first predicted elapsed time on the final path planning effect, the first predicted elapsed time is used as a time node, a first preset time length is added before the time node, and a second preset time length is added after the time node to determine the first occupied time interval passing through the corresponding grid. The first preset time length and the second preset time length are set by those skilled in the art according to actual needs.

[0090] In addition, the first predicted driving probability and the first occupied time interval of the identified target passing through each grid can be integrated to establish a preset driving path change area of ​​the identified target, so as to analyze and view the predicted driving path of each identified target.

[0091] Figure 3 A flowchart showing a method for determining a second predicted travel probability and a second occupied time interval of a grid provided by the present invention is shown;

[0092] like Figure 3 As shown, according to an embodiment of the present invention, the first predicted travel probability and the first occupied time interval of all identified targets passing through each grid are integrated respectively to determine the second predicted travel probability and the second occupied time interval of each grid, including:

[0093] S302, when there are predicted travel data of multiple identified targets in the grid, sorting the first occupied time intervals of the multiple identified targets in chronological order;

[0094] S304, calculating time intervals of adjacent first occupied time intervals, and merging adjacent first occupied time intervals whose time intervals are less than a preset time interval;

[0095] S306: Determine the predicted driving probability of the overlapping part of adjacent first occupied time intervals as the larger first predicted driving probability among the corresponding identified targets, and determine the predicted driving probability of the interval part as the smaller first predicted driving probability among the corresponding identified targets.

[0096] It should be noted that multiple identified targets may pass through the same grid at different times. By analyzing the first predicted travel probability and the first occupied time interval of all identified targets passing through each grid, the first predicted travel probability and the first occupied time interval of the identified targets passing through the same grid are integrated based on the grid, and the first predicted travel probability and the first occupied time interval of all identified targets corresponding to each grid are determined, and expressed as the second predicted travel probability and the second occupied time interval of the grid.

[0097] When the time interval between adjacent first occupied time intervals is less than the preset time interval, it means that the time interval between the adjacent first occupied time intervals is small, and the time interval does not meet the path insertion condition of the current vehicle, and the adjacent first occupied time intervals are merged. Since the predicted time for different identified targets to pass through the same grid is represented by a time interval, there may be an overlap of the occupied time intervals, and the predicted travel probability of the overlapping part of the adjacent first occupied time intervals is determined as the larger first predicted travel probability of the corresponding identified target.

[0098] The preset time interval is set by those skilled in the art according to actual needs.

[0099] According to an embodiment of the present invention, a preset driving path of the current vehicle is analyzed by a preset path planning model to determine the third occupied time interval of each grid through which the current vehicle passes, and an intersection risk score of each grid is calculated in combination with the second predicted driving probability and the second occupied time interval of each grid, including:

[0100] Get the preset driving path of the current vehicle;

[0101] Input the preset driving path of the current vehicle into the preset path planning model, analyze it in combination with the road traffic data, and output the second predicted passing time of the current vehicle passing through each grid;

[0102] Adding the first preset time length and the second preset time length before and after the second predicted passing time respectively, to determine a third occupied time interval of the current vehicle passing through the corresponding grid;

[0103] Calculate the same occupied time interval of the second occupied time interval and the corresponding third occupied time interval of each grid;

[0104] Calculating a time length ratio of the same occupied time interval and a corresponding third occupied time interval, and determining a first impact weight of a second predicted travel probability in the corresponding grid;

[0105] The second predicted travel probability corresponding to the same occupied time interval in the grid is multiplied by the corresponding first impact weight to determine the intersection risk score of the grid;

[0106] The risk gradient curve is drawn according to the intersection risk scores of all grids.

[0107] It should be noted that the preset driving path of the current vehicle can be obtained through the destination set by the driver, the in-vehicle navigation, etc. The preset driving path of the current vehicle is analyzed through the preset path planning model, and the second predicted passing time of the current vehicle passing through each grid based on the preset driving path is calculated in combination with the road traffic data. The second predicted passing time is used as the time node, and the first preset time length is added before the time node, and the second preset time length is added after the time node to determine the third occupied time interval passing through the corresponding grid. The second occupied time interval and the third occupied time interval in each grid are compared in turn, and the same occupied time interval with overlapping time is determined as the time interval with intersection risk. The first influence weight of the second predicted driving probability in the corresponding grid is determined by the time length ratio of the same occupied time interval and the corresponding third occupied time interval, and the second predicted driving probability is multiplied by the corresponding first influence weight to determine the intersection risk score of the grid. In addition, when the second occupied time interval corresponding to the same occupied time interval in the grid is obtained by combining multiple first occupied time intervals, the length ratio of each first occupied time interval and the corresponding third occupied time interval in the same occupied time is calculated to determine the first impact weight of the first predicted travel probability corresponding to each first occupied time interval. The first predicted travel probability corresponding to each first occupied time interval is multiplied by the corresponding first impact weight, and the calculation results are accumulated to determine the intersection risk score of the grid.

[0108] At the same time, a risk gradient curve of the intersection risk can be drawn according to the intersection risk score within each grid, so as to facilitate the analysis and viewing of the intersection risk of the current vehicle and the identified target.

[0109] According to an embodiment of the present invention, adjusting the first driving path according to the intersection risk score of each grid on the first driving path includes:

[0110] Marking grids in the first driving path whose intersection risk scores are greater than a first preset risk score threshold for avoidance;

[0111] determining a second impact weight of the intersection risk score of each grid according to a relative distance between the grid and the current position of the current vehicle;

[0112] Calculate the weighted average of the intersection risk score of each grid in the first driving path and the corresponding second impact weight to determine the average risk score of the first driving path;

[0113] When the average risk score is less than the second preset risk score threshold, no adjustment is made;

[0114] When the average risk score is greater than or equal to the second preset risk score threshold, marking the grids in the first driving path whose intersection risk scores are greater than the third preset risk score threshold for avoidance;

[0115] The first driving path is adjusted according to the grid where the avoidance mark exists.

[0116] It should be noted that, in the process of planning and adjusting the driving path of the current vehicle, it is preferred to determine whether there are grids with an intersection risk score greater than the first preset risk score threshold on the first driving path of the current vehicle, and such grids are marked for avoidance. Then, the second impact weight of the intersection risk score of each grid is determined by the relative distance between the grid and the current position of the current vehicle. The closer the relative distance between the grid and the current position of the current vehicle, the smaller the second impact weight of the intersection risk score of the grid. The intersection risk score of each grid in the first driving path is multiplied by the corresponding second impact weight, the calculation results are accumulated, and finally divided by the number of grids to determine the average risk score of the first driving path. Among them, the first preset risk score threshold, the second preset risk score threshold and the third preset risk score threshold are all set by those skilled in the art according to actual needs, and the first preset risk score threshold> the third preset risk score threshold> the second preset risk score threshold.

[0117] According to an embodiment of the present invention, adjusting the first driving path according to the grid having the avoidance mark includes:

[0118] Calculate the driving score of each grid based on the driving parameters of the current vehicle and the intersection risk score of each grid;

[0119] Determine an obstacle avoidance starting point and an obstacle avoidance end point according to the grids with avoidance marks in the first driving path;

[0120] Determine one or more obstacle avoidance paths according to the driving score, obstacle avoidance starting point and obstacle avoidance end point of each grid;

[0121] Calculate the average driving score of each obstacle avoidance path, and filter the obstacle avoidance paths whose average driving score is less than the preset driving score threshold;

[0122] Calculate the obstacle avoidance score of the obstacle avoidance path according to the area of ​​the area enclosed by the obstacle avoidance path and the first driving path;

[0123] Determine the obstacle avoidance path with the smallest obstacle avoidance score as the second driving path;

[0124] The second driving path is replaced with a corresponding first driving path portion according to the obstacle avoidance starting point and the obstacle avoidance end point.

[0125] It should be noted that the preset path planning model is used to analyze the road traffic data and the current vehicle's driving parameters (including driving direction, speed, and turning angle, etc.), starting from the current position of the current vehicle, the grids that meet the current vehicle's movement conditions (i.e., the grids that the current vehicle can pass through) are marked as movable, and each movable marked grid is used as the starting position to move again, and the grids that meet the current vehicle's movement conditions are marked as movable, and the above process is repeated until there are no grids that meet the current vehicle's movement conditions. The intersection risk score of the grid with movable marks is standardized, and the intersection risk score of the grid is converted into a value between 0 and 100, and the driving score of the grid is calculated based on the intersection risk score after the grid is standardized.

[0126] ;

[0127] Among them, P a is the driving score of the grid, P b Intersection risk score after normalization for the grid.

[0128] Then, the preset path planning model is used to analyze the road traffic data, the driving parameters of the current vehicle (including driving direction, speed, turning angle, etc.) and the grids with avoidance marks to determine the obstacle avoidance starting point and obstacle avoidance end point of the current vehicle without passing through the avoidance grid. One or more obstacle avoidance paths are generated based on the obstacle avoidance starting point, obstacle avoidance end point and the grids with movable marks. The driving scores of all grids on the obstacle avoidance path are accumulated and divided by the number of grids to determine the average driving score of the obstacle avoidance path. When the average driving score of the obstacle avoidance path is less than the preset driving score threshold, the obstacle avoidance path does not meet the obstacle avoidance condition and the obstacle avoidance path is filtered. The path deviation between the obstacle avoidance path and the original first driving path can be determined by calculating the area enclosed by the obstacle avoidance path and the first driving path. The obstacle avoidance score is expressed by selecting the obstacle avoidance path with the smallest obstacle avoidance score to replace the corresponding part of the driving path in the first driving path, and the obstacle avoidance adjustment of the first driving path is completed.

[0129] The preset driving score threshold is set by those skilled in the art according to actual needs.

[0130] According to an embodiment of the present invention, it also includes:

[0131] The grid specifications are dynamically adjusted according to the current vehicle speed, driving status, target type and target ratio in the monitoring area.

[0132] It should be noted that the size of the grid specification is dynamically adjusted according to the current vehicle speed, driving status (such as straight, turning, etc.), target type (such as pedestrians, vehicles, etc.) and target proportion in the monitoring area. For example, when the current vehicle is driving on the same road, the faster the speed, the smaller the corresponding grid specification; the grid specification when turning is smaller than the grid specification when driving straight; when there are only pedestrians and vehicles in the detection area, the greater the proportion of pedestrians, the smaller the grid specification. By adjusting the grid size, the vehicle's obstacle avoidance accuracy requirements under different driving conditions and environmental conditions can be met while ensuring analysis efficiency.

[0133] Figure 4 A block diagram of an intelligent driving path planning system based on reinforcement learning provided by the present invention is shown.

[0134] like Figure 4 As shown, the second aspect of the present invention provides an intelligent driving path planning system based on reinforcement learning, comprising:

[0135] A region segmentation module is used to divide the detection area into multiple grids;

[0136] A data acquisition module, used for acquiring detection data;

[0137] The first analysis module is used to input the detection data into a preset target recognition model for analysis, and output recognition data of the recognition target; the recognition data of the recognition target is analyzed by a preset path planning model to determine the movement path of the recognition target, and determine the first predicted travel probability and the first occupied time interval of the recognition target passing through each grid; the movement path includes a first movement path and a second movement path;

[0138] The second analysis module is used to integrate the first predicted travel probability and the first occupied time interval of all the identified targets passing through each grid, and determine the second predicted travel probability and the second occupied time interval of each grid;

[0139] A third analysis module is used to analyze the preset driving path of the current vehicle through a preset path planning model, determine the third occupied time interval of the current vehicle passing through each grid, and calculate the intersection risk score of each grid in combination with the second predicted driving probability and the second occupied time interval of each grid;

[0140] The path planning module is used to intercept the preset driving path of the current vehicle based on the preset driving time and the current driving speed of the vehicle to determine the first driving path; and adjust the first driving path according to the intersection risk score of each grid on the first driving path.

[0141] The third aspect of the present invention provides a computer-readable storage medium, which includes a program for an intelligent driving path planning method based on reinforcement learning. When the program for an intelligent driving path planning method based on reinforcement learning is executed by a processor, the steps of the above-mentioned method for intelligent driving path planning based on reinforcement learning are implemented.

[0142] The information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data used for analysis, stored data, displayed data, etc.) and signals (including but not limited to signals transmitted between user terminals and other devices, etc.) involved in this application are all authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with relevant laws, regulations and standards of relevant countries and regions. For example, the "detection data" and "preset driving path of the current vehicle" involved in this disclosure are all obtained with full authorization.

[0143] The present invention discloses a method and system for intelligent driving path planning based on reinforcement learning, the method comprising: dividing a detection area into multiple grids; acquiring detection data; inputting the detection data into a preset target recognition model for analysis, and outputting recognition data of the recognition target; analyzing the recognition data of the recognition target through a preset path planning model, determining the moving path of the recognition target, determining the second predicted driving probability and the second occupied time interval of each grid; calculating the intersection risk score of each grid, intercepting the preset driving path of the current vehicle, and determining the first driving path; adjusting the first driving path according to the intersection risk score of each grid on the first driving path. The present invention generates a safe and efficient driving path by analyzing the traffic environment in real time to ensure the safety of intelligent driving. In several embodiments provided in this application, it should be understood that the disclosed device and method can be implemented in other ways. The device embodiments described above are only schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation, such as: multiple units or components can be combined, or can be integrated into another system, or some features can be ignored or not executed. In addition, the coupling, direct coupling, or communication connection between the components shown or discussed may be through some interfaces, indirect coupling or communication connection of devices or units, which may be electrical, mechanical or other forms.

[0144] The units described above as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units; they may be located in one place or distributed on multiple network units; some or all of the units may be selected according to actual needs to achieve the purpose of the present embodiment.

[0145] In addition, all functional units in the embodiments of the present invention may be integrated into one processing unit, or each unit may be separately used as a unit, or two or more units may be integrated into one unit; the above-mentioned integrated units may be implemented in the form of hardware or in the form of hardware plus software functional units.

[0146] A person of ordinary skill in the art can understand that: all or part of the steps of implementing the above method embodiment can be completed by hardware related to program instructions, and the aforementioned program can be stored in a computer-readable storage medium. When the program is executed, it executes the steps of the above method embodiment; and the aforementioned storage medium includes: a mobile storage device, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, and other media that can store program codes.

[0147] Alternatively, if the above-mentioned integrated unit of the present invention is implemented in the form of a software function module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the embodiment of the present invention can be essentially or partly reflected in the form of a software product that contributes to the prior art. The computer software product is stored in a storage medium and includes several instructions for a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the methods described in each embodiment of the present invention. The aforementioned storage medium includes: various media that can store program codes, such as mobile storage devices, ROM, RAM, magnetic disks or optical disks.

Claims

1. An intelligent driving path planning method based on reinforcement learning, characterized in that: include: Divide the detection area into multiple grids; Obtain test data; Input the detection data into a preset target recognition model for analysis, and output recognition data of the recognized target; Analyzing the identification data of the identified target through a preset path planning model, determining a moving path of the identified target, and determining a first predicted travel probability and a first occupied time interval of the identified target passing through each grid; the moving path includes a first moving path and a second moving path; Integrate the first predicted travel probability and the first occupied time interval of all identified targets passing through each grid respectively, and determine the second predicted travel probability and the second occupied time interval of each grid; Analyze the preset driving path of the current vehicle through the preset path planning model, determine the third occupied time interval of the current vehicle passing through each grid, and calculate the intersection risk score of each grid in combination with the second predicted driving probability and the second occupied time interval of each grid; Intercepting a preset driving path of the current vehicle based on a preset driving time and a driving speed of the current vehicle to determine a first driving path; Adjusting the first driving path according to the intersection risk score of each grid on the first driving path; The preset path planning model is used to analyze the preset driving path of the current vehicle, determine the third occupied time interval of each grid through which the current vehicle passes, and calculate the intersection risk score of each grid in combination with the second predicted driving probability and the second occupied time interval of each grid, including: Get the preset driving path of the current vehicle; Inputting the preset driving path of the current vehicle into a preset path planning model, analyzing it in combination with road traffic data, and outputting a second predicted passing time of the current vehicle passing through each grid; Adding a first preset time length and a second preset time length before and after the second predicted passing time, respectively, to determine a third occupied time interval of the current vehicle passing through the corresponding grid; Calculate the same occupied time interval of the second occupied time interval and the corresponding third occupied time interval of each grid; Calculating a time length ratio of the same occupied time interval and a corresponding third occupied time interval, and determining a first influence weight of a second predicted travel probability in the corresponding grid; Multiplying the second predicted travel probability corresponding to the same occupied time interval in the grid by the corresponding first impact weight to determine the intersection risk score of the grid; Draw a risk gradient curve according to the intersection risk scores of all grids; When the second occupied time interval corresponding to the same occupied time interval in the grid is obtained by merging multiple first occupied time intervals, the ratio of the length of each part of the first occupied time interval and the corresponding third occupied time interval in the same occupied time is calculated to determine the first influence weight of the first predicted driving probability corresponding to each part of the first occupied time interval; the first predicted driving probability corresponding to each part of the first occupied time interval is multiplied by the corresponding first influence weight, the calculation results are accumulated, and the intersection risk score of the grid is determined.

2. The intelligent driving path planning method based on reinforcement learning according to claim 1, characterized in that: The analyzing the identification data of the identification target by using a preset path planning model to determine the moving path of the identification target includes: Filtering a plurality of identification data of the identified target within a first preset time interval from the historical detection data, and drawing a first moving path of the identified target; The first movement data of the identified target is input into a preset path planning model, and a second movement path of the identified target in a second preset time interval is output.

3. The intelligent driving path planning method based on reinforcement learning according to claim 1 is characterized in that: The determining of a first predicted travel probability and a first occupied time interval of the identified target passing through each grid includes: Analyze the first movement data and road traffic data of the identified target through a preset path planning model, calculate the movement probability of the identified target in each direction, and predict the first predicted travel probability and the first predicted passing time of the identified target passing through each grid; A first preset time length and a second preset time length are respectively added before and after the first predicted elapsed time to determine a first occupied time interval passing through the corresponding grid.

4. The intelligent driving path planning method based on reinforcement learning according to claim 1 is characterized in that: The step of integrating the first predicted travel probability and the first occupied time interval of all identified targets passing through each grid to determine the second predicted travel probability and the second occupied time interval of each grid comprises: When there are predicted travel data of multiple identified targets in the grid, sorting the first occupied time intervals of the multiple identified targets in chronological order; Calculating time intervals of adjacent first occupied time intervals, and merging adjacent first occupied time intervals whose time intervals are shorter than a preset time interval; The predicted driving probability of the overlapping part of adjacent first occupied time intervals is determined as the larger first predicted driving probability among the corresponding identified targets, and the predicted driving probability of the interval part is determined as the smaller first predicted driving probability among the corresponding identified targets.

5. The intelligent driving path planning method based on reinforcement learning according to claim 1, characterized in that: The adjusting the first driving path according to the intersection risk score of each grid on the first driving path includes: Marking grids in the first driving path whose intersection risk scores are greater than a first preset risk score threshold for avoidance; determining a second impact weight of the intersection risk score of each grid according to a relative distance between the grid and the current position of the current vehicle; Calculating a weighted average of the intersection risk score of each grid in the first driving path and the corresponding second impact weight to determine an average risk score of the first driving path; When the average risk score is less than the second preset risk score threshold, no adjustment is made; When the average risk score is greater than or equal to a second preset risk score threshold, marking the grids in the first driving path whose intersection risk scores are greater than a third preset risk score threshold for avoidance; The first driving path is adjusted according to the grid where the avoidance mark exists.

6. The intelligent driving path planning method based on reinforcement learning according to claim 5 is characterized in that: The adjusting the first driving path according to the grid having the avoidance mark includes: Calculate the driving score of each grid based on the driving parameters of the current vehicle and the intersection risk score of each grid; Determine an obstacle avoidance starting point and an obstacle avoidance end point according to the grids with avoidance marks in the first driving path; Determine one or more obstacle avoidance paths according to the driving score, obstacle avoidance starting point, and obstacle avoidance end point of each grid; Calculate the average driving score of each obstacle avoidance path, and filter the obstacle avoidance paths whose average driving score is less than the preset driving score threshold; Calculate the obstacle avoidance score of the obstacle avoidance path according to the area of ​​the area enclosed by the obstacle avoidance path and the first driving path; Determine the obstacle avoidance path with the smallest obstacle avoidance score as the second driving path; The second driving path is used to replace the corresponding first driving path portion according to the obstacle avoidance starting point and the obstacle avoidance end point.

7. The intelligent driving path planning method based on reinforcement learning according to claim 1 is characterized in that: Also includes: The grid specifications are dynamically adjusted according to the current vehicle speed, driving status, target type and target ratio in the monitoring area.

8. An intelligent driving path planning system based on reinforcement learning, used to implement the intelligent driving path planning method based on reinforcement learning as claimed in any one of claims 1 to 7, characterized in that: include: A region segmentation module is used to divide the detection area into multiple grids; A data acquisition module, used for acquiring detection data; A first analysis module, used to input the detection data into a preset target recognition model for analysis, and output recognition data of the recognized target; Analyzing the identification data of the identified target through a preset path planning model, determining a moving path of the identified target, and determining a first predicted travel probability and a first occupied time interval of the identified target passing through each grid; the moving path includes a first moving path and a second moving path; The second analysis module is used to integrate the first predicted travel probability and the first occupied time interval of all the identified targets passing through each grid, and determine the second predicted travel probability and the second occupied time interval of each grid; A third analysis module is used to analyze the preset driving path of the current vehicle through a preset path planning model, determine the third occupied time interval of the current vehicle passing through each grid, and calculate the intersection risk score of each grid in combination with the second predicted driving probability and the second occupied time interval of each grid; A path planning module, used to intercept a preset driving path of the current vehicle based on a preset driving time and a driving speed of the current vehicle to determine a first driving path; The first driving path is adjusted according to the intersection risk score of each grid on the first driving path.

9. The intelligent driving path planning system based on reinforcement learning according to claim 8, characterized in that: The preset path planning model is used to analyze the preset driving path of the current vehicle, determine the third occupied time interval of each grid through which the current vehicle passes, and calculate the intersection risk score of each grid in combination with the second predicted driving probability and the second occupied time interval of each grid, including: Get the preset driving path of the current vehicle; Inputting the preset driving path of the current vehicle into a preset path planning model, analyzing it in combination with road traffic data, and outputting a second predicted passing time of the current vehicle passing through each grid; Adding a first preset time length and a second preset time length before and after the second predicted passing time, respectively, to determine a third occupied time interval of the current vehicle passing through the corresponding grid; Calculate the same occupied time interval of the second occupied time interval and the corresponding third occupied time interval of each grid; Calculating a time length ratio of the same occupied time interval and a corresponding third occupied time interval, and determining a first influence weight of a second predicted travel probability in the corresponding grid; Multiplying the second predicted travel probability corresponding to the same occupied time interval in the grid by the corresponding first impact weight to determine the intersection risk score of the grid; The risk gradient curve is drawn according to the intersection risk scores of all grids.

Citation Information

Patent Citations

  • Autonomous Driving Control System and Collision Avoidance Control Method Therewith

    US20240157976A1

  • Method and system for determining sub-path in a navigational path for a vehicle

    WO2024184918A1