Vehicle trajectory planning method and device, electronic equipment and storage medium
By extracting and fusing the features of the vehicle's surrounding environment information to generate a reward map, and combining it with an autoregressive trajectory generation network and a kinematic model, the real-time and stability issues of vehicle trajectory planning in complex scenarios are solved, accurate future trajectory prediction is achieved, and the safety of autonomous driving is improved.
Patent Information
- Application Number
- CN202510901709.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-30
- Publication Date
- 2025-09-30
AI Technical Summary
Existing vehicle trajectory planning methods need to adapt to various dynamic and static targets in complex scenarios, which leads to increased resource consumption, reduced search efficiency, and challenges in real-time performance and stability.
By determining the static road network information, dynamic obstacle information and vehicle navigation information of the vehicle's surrounding environment, feature extraction and fusion are performed to generate a reward map, and the target value expectation grid map is iteratively generated. Combined with the autoregressive trajectory generation network and kinematic deduction model, the vehicle's motion trajectory in the future time period is predicted.
It improves the real-time and stability of vehicle trajectory planning, achieves accurate prediction of complex driving scenarios, provides reliable prior reference information for the decision-making and control of autonomous driving vehicles, and improves driving safety.
Smart Images

Figure CN120721086A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to intelligent assisted driving technology, and in particular to a vehicle trajectory planning method, device, electronic device, and storage medium. Background Art
[0002] During the driving process of an intelligent assisted driving vehicle, generally, the vehicle's movement trajectory in a future time period (such as the next 6 seconds) is planned based on navigation information and perception information around the vehicle to control the vehicle to drive safely.
[0003] Existing vehicle trajectory planning methods typically use a state-space search method to represent the vehicle state as a state vector, divide the specified state space into multiple discrete grids, and sample the grids to search for the optimal state sequence to generate the optimal driving path. However, in complex scenarios, such as urban areas, where roads contain a variety of dynamic targets such as vehicles, pedestrians, bicycles, motorcycles, and pets, as well as static targets such as traffic signs, traffic lights, lane height restrictions, and obstacles, state-space search methods require additional search conditions adapted to each target to search for the optimal path. This results in exponentially increasing resource consumption and poses significant challenges to the real-time and stability of trajectory planning. Summary of the Invention
[0004] Embodiments of the present disclosure provide a vehicle trajectory planning method, apparatus, electronic device, and storage medium to improve the real-time performance and stability of vehicle trajectory planning.
[0005] According to a first aspect of an embodiment of the present disclosure, a vehicle trajectory planning method is provided, comprising:
[0006] Determine static road network information, dynamic obstacle information and vehicle navigation information of the vehicle's surrounding environment;
[0007] Extracting and fusing features of the static road network information, the dynamic obstacle information, and the vehicle navigation information to obtain a reward map;
[0008] Iteratively generating a target value expectation grid map based on the reward map;
[0009] Sampling the target value expectation grid map to obtain a path map, wherein the path map includes at least one state path;
[0010] Based on the scene features and vehicle state information corresponding to any state path, the trajectory points of the vehicle at any time step in a future time period are predicted to obtain the motion trajectory of the vehicle in the future time period.
[0011] According to a second aspect of an embodiment of the present disclosure, a vehicle driving trajectory planning device is provided, comprising:
[0012] An information determination module, used to determine static road network information, dynamic obstacle information and vehicle navigation information of the vehicle's surrounding environment;
[0013] A first encoding module is used to extract and fuse features of the static road network information, the dynamic obstacle information, and the vehicle navigation information to obtain a reward map;
[0014] An iteration module, configured to iteratively generate a target value expectation grid map based on the reward map;
[0015] A sampling module is used to sample the target value expectation grid map to obtain a path map, wherein the path map includes at least one state path;
[0016] The trajectory prediction module is used to predict the trajectory point of the vehicle at any time step in a preset future time period based on the scene features and vehicle state information corresponding to any state path, and obtain the vehicle trajectory of the vehicle in the preset future time period.
[0017] According to a third aspect of an embodiment of the present disclosure, a computer-readable storage medium is provided, wherein the storage medium stores a computer program, and when the computer program is executed by a processor, it is used to implement the above-mentioned vehicle driving trajectory planning method.
[0018] According to a fourth aspect of the embodiments of the present disclosure, there is provided an electronic device, comprising:
[0019] processor;
[0020] a memory for storing instructions executable by the processor;
[0021] The processor is used to read the executable instructions from the memory and execute the instructions to implement the above-mentioned vehicle driving trajectory planning method.
[0022] Based on the vehicle trajectory planning method provided by the above-mentioned embodiment of the present disclosure, the static road network information, dynamic obstacle information and vehicle navigation information of the vehicle's surrounding environment are determined; the static road network information, dynamic obstacle information and vehicle navigation information are feature extracted and integrated to obtain a reward map; based on the reward map, a target value expectation grid map is iteratively generated; the target value expectation grid map is sampled to obtain a path map; finally, based on the scene features and vehicle state information corresponding to any state path, the trajectory points of the vehicle at any time step in the future time period are predicted to obtain the vehicle's motion trajectory in the future time period. Therefore, the embodiment of the present disclosure avoids the difficulties of cost design and time-consuming operation in traditional search methods through a learnable reward map, and combines the autoregressive trajectory generation network and the kinematic deduction model to achieve accurate prediction of the vehicle's driving trajectory in a specified future time period, providing reliable prior reference information for downstream tasks such as decision-making and control of autonomous driving vehicles, thereby improving the safety of vehicle driving.
[0023] The technical solution of the present disclosure is further described in detail below through the accompanying drawings and examples. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] Figure 1 An electronic device to which the vehicle trajectory planning method according to an embodiment of the present disclosure can be applied is shown;
[0025] Figure 2 is a flowchart of a vehicle driving trajectory planning method provided by an exemplary embodiment of the present disclosure;
[0026] Figure 3 202 is a flow chart of step 202 of a vehicle driving trajectory planning method provided by an exemplary embodiment of the present disclosure;
[0027] Figure 4 This is a schematic diagram of a process for determining a path map in a vehicle driving trajectory planning method provided by an exemplary embodiment of the present disclosure;
[0028] Figure 5 This is a schematic diagram of the output result of behavior planning in the planning of a vehicle driving trajectory provided by an exemplary embodiment of the present disclosure;
[0029] Figure 6 205 is a flow chart of step 205 of a vehicle driving trajectory planning method provided by an exemplary embodiment of the present disclosure;
[0030] Figure 7 This is a schematic diagram of trajectory planning in the planning of a vehicle driving trajectory provided by an exemplary embodiment of the present disclosure;
[0031] Figure 8 This is a schematic diagram of a trajectory planning output result in vehicle driving trajectory planning provided by an exemplary embodiment of the present disclosure;
[0032] Figure 9 1 is a schematic diagram of a training process of a trajectory planning model in a vehicle driving trajectory planning method provided by an exemplary embodiment of the present disclosure;
[0033] Figure 10 is an exemplary system framework diagram of an embodiment of the present disclosure;
[0034] Figure 11 1 is a schematic structural diagram of a vehicle driving trajectory planning device provided by an exemplary embodiment of the present disclosure;
[0035] Figure 12 It is a structural diagram of a vehicle driving trajectory planning device provided by another exemplary embodiment of the present disclosure. DETAILED DESCRIPTION
[0036] To explain the present disclosure, example embodiments of the present disclosure will be described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present disclosure, rather than all the embodiments. It should be understood that the present disclosure is not limited to the example embodiments.
[0037] It should be noted that the relative arrangement of components and steps, the numerical expressions and numerical values set forth in these embodiments do not limit the scope of the present disclosure unless specifically stated otherwise.
[0038] Application Overview
[0039] In the process of implementing the present disclosure, the inventors discovered through research that there are at least the following problems: when planning vehicle trajectories through the state space search method, it is necessary to add search conditions to adapt to various dynamic targets and static targets that may appear in complex driving scenarios. As the search conditions increase, the number of states in the state space also increases accordingly, the search process becomes more complicated, the computational complexity of the search is increased, and the search time is prolonged, resulting in reduced search efficiency and increased resource consumption, which poses a relatively large challenge to the real-time and stability of trajectory planning.
[0040] In order to determine the future driving trajectory of a vehicle in real time under various complex driving scenarios, the inventors proposed the technical solution disclosed herein.
[0041] Exemplary devices
[0042] Figure 1 An electronic device to which the vehicle trajectory planning method according to an embodiment of the present disclosure can be applied is shown.
[0043] like Figure 1 As shown, the electronic device includes at least one processor 11 and a memory 12 .
[0044] The processor 11 may be a central processing unit (CPU) or other forms of processing units having data processing capabilities and / or instruction execution capabilities, and may control other components in the electronic device to perform desired functions.
[0045] The memory 12 may include one or more computer program products, which may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. Volatile memory may include, for example, random access memory (RAM) and / or cache memory. Non-volatile memory may include, for example, read-only memory (ROM), a hard disk, flash memory, etc. One or more computer program instructions may be stored on the computer-readable storage medium, and the processor 11 may execute one or more computer program instructions to implement the vehicle trajectory planning method and / or other desired functions provided in the present disclosure. The vehicle trajectory planning method includes: determining static road network information, dynamic obstacle information, and vehicle navigation information of the vehicle's surrounding environment; extracting and fusing features of the static road network information, dynamic obstacle information, and vehicle navigation information to obtain a reward map; iteratively generating a target value expectation grid map based on the reward map; sampling the target value expectation grid map to obtain a path map, the path map containing at least one state path; and predicting the vehicle's trajectory points at any time step in a future time period based on the scene features and vehicle state information corresponding to any state path, thereby obtaining the vehicle's motion trajectory in the future time period.
[0046] In one example, the electronic device may further include an input device 13 and an output device 14 , and these components are interconnected via a bus system and / or other forms of connection mechanisms (not shown).
[0047] The input device 13 may also include, for example, a keyboard, a mouse, etc.
[0048] The output device 14 can output various information to the outside, and may include, for example, a display, a speaker, a printer, a communication network and a remote output device connected thereto, and the like.
[0049] Of course, to simplify, Figure 1 Only some of the components related to the present disclosure in the electronic device are shown, and components such as a bus, an input / output interface, etc. are omitted. In addition, the electronic device may further include any other appropriate components according to specific application scenarios.
[0050] It should be noted that the electronic devices in the technical solution disclosed in the present invention can be various electronic devices, including but not limited to mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), vehicle-mounted terminals (such as vehicle-mounted navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc.
[0051] Exemplary Methods
[0052] Figure 2 This is a flow chart of a vehicle driving trajectory planning method provided by an exemplary embodiment of the present disclosure. Figure 1 The electronic device shown. The disclosed embodiment can be applied to any electronic device with data processing capabilities, such as but not limited to terminal devices with data processing capabilities (such as vehicle-mounted terminals, mobile phone terminals, tablet computers, PCs, etc.), cloud servers, computing platforms in vehicle driving control systems, etc., by obtaining the data of upstream perception tasks as input information and processing it, and obtaining the motion trajectory in the future time period, so as to perform downstream tasks such as driving control accordingly. Figure 2 As shown, the vehicle driving trajectory planning method of the embodiment of the present disclosure includes:
[0053] Step 201: Determine static road network information, dynamic obstacle information, and vehicle navigation information of the vehicle's surrounding environment.
[0054] In the embodiment of the present disclosure, the static road network information and dynamic obstacle information of the vehicle's surrounding environment can be obtained by sensing the surrounding environment through the target detection model after the sensors deployed on the vehicle collect images or point clouds. The target detection model is a model used to identify and locate target objects in images or videos. Common target detection models include R-CNN (Region-based Convolutional Neural Networks), Fast R-CNN, Faster R-CNN and YOLO. The sensors deployed on the vehicle may include but are not limited to any number of sensors of the same type or different types deployed at different locations on the vehicle and / or in different orientations. Among them, the sensors include but are not limited to visual sensors, lidar, millimeter wave radar, ultrasonic radar, etc. In a specific implementation, the number of sensors deployed on the vehicle can be one or more. When there are multiple sensors deployed on a vehicle, each sensor can collect data according to its preset frame rate. For a frame of data collected by any sensor at any data collection moment (as the first moment) (or further combined with several adjacent historical frame data), target detection or perception processing can be performed through a deep learning model or other algorithm model to obtain static road network information and dynamic obstacle information of the vehicle's surrounding environment.
[0055] Static road network information can include road information and static object information around the vehicle. Road information includes, but is not limited to, curbs, lane markings, stop signs, zebra crossings, road arrows, virtual lane markings, and other information. Static object information includes, but is not limited to, traffic signs, traffic lights, lane height limits, sewer openings, obstacles, and other road details. Other road details include, but are not limited to, overhead objects, guardrails, road edge types, roadside landmarks, buildings, flowers, trees, and other infrastructure information.
[0056] Among them, dynamic obstacle information may include pedestrians, animals, vehicles, bicycles, motorcycles, etc. in the vehicle's surrounding environment.
[0057] In the disclosed embodiments, static road network information and dynamic obstacle information can be obtained from images of the surrounding environment captured by the vehicle. Dynamic obstacle coordinates and static road network element coordinates from the BEV's perspective are then obtained based on an object detection model. The dynamic obstacle coordinates and static road network element coordinates are then rendered into two-dimensional images, respectively, to obtain a two-dimensional image corresponding to the dynamic obstacle information and a two-dimensional image corresponding to the static road network information. For example, the two-dimensional image corresponding to the static road network information is a 512-pixel by 512-pixel two-dimensional image from a bird's-eye view perspective. The image illustrates the local field of view, including the location, shape, and status information of static objects such as curbs, lane markings, stop signs, zebra crossings, road arrows, and virtual lane markings. The two-dimensional image corresponding to the dynamic obstacle information can also be a 512-pixel by 512-pixel two-dimensional image from a bird's-eye view perspective. To accurately describe the changes in the motion state of dynamic obstacles over a historical time period, the dynamic obstacle information can be obtained by continuously capturing multiple frames of two-dimensional images from the historical time period. Each frame of the two-dimensional image can be used to obtain information such as the position, velocity, acceleration, and direction of the dynamic objects at different points in time. The historical time period is a shorter time period before the current moment of the predicted vehicle driving trajectory, such as within two seconds before the current moment.
[0058] In the disclosed embodiment, the vehicle navigation information includes the navigation target points visible within the field of view of the BEV in the current frame and the vehicle's current location. The vehicle navigation information is also rendered as a two-dimensional image. For example, the vehicle navigation point is rendered as a 65 pixel by 65 pixel two-dimensional image from a bird's-eye view.
[0059] Step 202 : extract and fuse the static road network information, dynamic obstacle information, and vehicle navigation information to obtain a reward map.
[0060] In the embodiment of the present disclosure, a convolutional neural network (CNN) can be used to extract features from the two-dimensional image corresponding to the static road network information, the two-dimensional image corresponding to the dynamic obstacle information, and the two-dimensional image corresponding to the vehicle navigation information, and then the extracted features are fused to obtain scene features, which are scene representation vectors used to represent scene features. The fused scene features are then downsampled and feature extracted to obtain a reward map. For details, see Figure 3 The embodiment shown.
[0061] Feature fusion methods include: adding the features of static road network information and dynamic obstacle information; or concatenating the features of static road network information and dynamic obstacle information. Addition can be achieved by element-by-element addition of the features corresponding to the static road network information and the features corresponding to the dynamic obstacle information, or by element-by-element weighted addition of the features corresponding to the static road network information and the features corresponding to the dynamic obstacle information. Feature concatenation can be achieved by channel dimension, using the concatenate function to concatenate the features corresponding to the static road network information and the features corresponding to the dynamic obstacle information along the channel dimension.
[0062] Step 203: Iteratively generate a target value expectation grid map based on the reward map.
[0063] In the disclosed embodiment, the reward map is a deterministic function defined on a grid environment map. It represents the immediate reward value a vehicle receives for taking an action to enter any grid cell. Higher reward values indicate better and more reasonable returns for that action. The reward map is generated by the model at each time step through real-time predictions based on input information such as the static road network and dynamic obstacles. Each grid cell corresponds to a reward value, which can be used in a value iteration algorithm. Through multiple iterative updates, the expected value of each grid cell is ultimately obtained.
[0064] In the disclosed embodiment, the expected value of any grid in the expected value grid diagram is used to represent an estimate of the weighted sum of the reward values corresponding to all actions when the vehicle travels from the current grid position to the navigation target point.
[0065] In the embodiment of the present disclosure, the implementation method of iteratively generating the target value grid map can be found in Figure 4 The embodiment shown will not be described in detail here.
[0066] Step 204: Sampling the target value expectation grid map to obtain a path map, wherein the path map includes at least one state path.
[0067] In the disclosed embodiment, according to the target value expectation grid diagram, a sampling process can be repeatedly performed through a Markov Decision Process (MDP) to obtain multiple state paths, and the multiple state paths form a path map.
[0068] The path map includes an optimal state path from the starting point to the navigation destination, or includes multiple drivable state paths from the starting point to the navigation destination. The state path is used to indicate the grid sequence connecting the starting point and the navigation target point, and the starting point represents the grid position of the vehicle at the current moment (the moment when the vehicle driving trajectory planning is executed). In the embodiment of the present disclosure, the specific implementation method of sampling the reward map to obtain the path map can be found in Figure 4 The embodiment shown will not be described in detail here.
[0069] Step 205 : Based on the scene features and vehicle state information corresponding to any state path, the trajectory point of the vehicle at any time step in the future time period is predicted to obtain the motion trajectory of the vehicle in the future time period.
[0070] The vehicle status information may include, but is not limited to, vehicle position, vehicle direction, vehicle speed, vehicle acceleration, vehicle angular acceleration, etc. The various speeds and / or positions in the vehicle status information may be obtained in a specified coordinate system, which may include, but is not limited to, any one or more of the following coordinate systems: the vehicle coordinate system, the world coordinate system, and the image coordinate system. Coordinate systems may also be transformed into each other based on the relationship between the coordinate systems, thereby obtaining vehicle status information in the same coordinate system.
[0071] In the disclosed embodiment, the scene features corresponding to the state path are used to indicate a vector composed of feature values at the corresponding index of each grid on the state path on the scene feature map. The scene feature map is a feature map obtained by extracting features from the two-dimensional image corresponding to the static road network information and the two-dimensional image corresponding to the dynamic obstacle information, and then fusing the extracted features. The feature map contains the scene features corresponding to each grid (the above-mentioned feature values). The future time period is defined as a continuous time interval after the current moment, such as within 6 seconds after the current moment. The time step is a discrete time point with equal spacing in the future time period. For example, if the future time period is 6 seconds and a trajectory point is predicted every half second, the time step is 12 time steps such as 0.5 seconds, 1 second, 1.5 seconds, 2 seconds, 2.5 seconds, 3 seconds, 3.5 seconds, 4 seconds, 4.5 seconds, 5 seconds, 5.5 seconds, and 6 seconds after the current moment. The trajectory points separated by 0.5 seconds in the 6-second time period are directly connected to obtain the vehicle's motion trajectory in the next 6 seconds.
[0072] Among them, the trajectory points of each time step in the future time period can be fitted by a variety of fitting methods, including but not limited to ordinary least squares method, stepwise regression algorithm, polynomial fitting method, logarithmic fitting method, etc. The embodiment of the present disclosure does not limit the fitting method.
[0073] Based on the embodiment of the present disclosure, by determining the static road network information, dynamic obstacle information and vehicle navigation information of the vehicle's surrounding environment; extracting and fusing the static road network information, dynamic obstacle information and vehicle navigation information, a reward map is obtained; based on the reward map, a target value expectation grid map is iteratively generated; the target value expectation grid map is sampled to obtain a path map; finally, based on the scene features and vehicle state information corresponding to any state path, the trajectory points of the vehicle at any time step in the future time period are predicted to obtain the vehicle's motion trajectory in the future time period. Therefore, the embodiment of the present disclosure avoids the difficulties of cost design and time-consuming operation in traditional search methods through a learnable reward map, and combines the autoregressive trajectory generation network and the kinematic deduction model to achieve accurate prediction of the vehicle's driving trajectory in a specified future time period, providing reliable prior reference information for downstream tasks such as decision-making and control of autonomous driving vehicles, thereby improving the safety of vehicle driving.
[0074] Figure 3 FIG. 2 is a flow chart of step 202 of a vehicle driving trajectory planning method provided by an exemplary embodiment of the present disclosure. Figure 3 As shown in the above Figure 2 Based on the illustrated embodiment, the method for determining the reward map in step 202 may include the following steps:
[0075] Step 221 : extract and fuse the static road network information, dynamic obstacle information, and vehicle navigation information to obtain a first fusion feature.
[0076] In the embodiment of the present disclosure, CNN can be used to extract features from two-dimensional images corresponding to static road network information, two-dimensional images corresponding to dynamic obstacle information, and two-dimensional images corresponding to vehicle navigation information, and then the extracted features are fused to obtain a first fused feature.
[0077] In the embodiments of the present disclosure, various CNN models, such as LeNet-5, AlexNet, VGGNet, and ResNet, can be used to perform feature extraction on two-dimensional images corresponding to static road network information, two-dimensional images corresponding to dynamic obstacle information, and two-dimensional images corresponding to vehicle navigation information. The embodiments of the present disclosure do not limit the CNN models used.
[0078] Step 222: downsample the first fused feature to obtain a second fused feature.
[0079] In the embodiment of the present disclosure, the first fusion feature may be downsampled, and the downsampling multiple may be a preset multiple, for example, downsampling by 8 times.
[0080] In the disclosed embodiment, the amount of computation can be reduced through downsampling processing. In autonomous driving scenarios with high real-time requirements, a larger downsampling multiple, such as 8 times or 16 times, is usually chosen to reduce the computational burden.
[0081] Step 223: Fuse the second fused feature with the feature corresponding to the vehicle state information to obtain a reward map.
[0082] The vehicle state information includes linear velocity and rotational angular velocity. In other possible implementations, the vehicle state information may include more information, such as acceleration, angular acceleration, etc., according to actual needs.
[0083] In the disclosed embodiment, a reward map generation model can be used to obtain a reward map based on the second fusion feature and the feature corresponding to the vehicle state information. The reward map generation model is a model used to generate a reward map and is a commonly used reward map generation model.
[0084] Based on the disclosed embodiment, a first fused feature is obtained by extracting and fusing features from static road network information, dynamic obstacle information, and vehicle navigation information. The first fused feature is then downsampled to obtain a second fused feature. The second fused feature is then fused with features corresponding to vehicle status information to create a reward map. Thus, the reward value for each grid in the reward map in the disclosed embodiment is derived by comprehensively considering static road network information, dynamic obstacle information, vehicle navigation information, and vehicle status information, providing data support for planning vehicle trajectories in complex driving scenarios such as urban areas.
[0085] Figure 4 This is a flow chart of determining a path map of a vehicle driving trajectory planning method provided by an exemplary embodiment of the present disclosure. Figure 4 As shown in the above Figure 2 Based on the illustrated embodiment, the method for determining a path map may include the following steps:
[0086] Step 231 : Initialize the expected value of any grid in the reward map to obtain an expected value grid map for the initial round.
[0087] The expected value of any grid in the expected value grid map is used to represent the estimated weighted sum of the reward values corresponding to all actions when the vehicle travels from the current grid position to the navigation target point.
[0088] In the disclosed embodiment, before performing iterative value update, the expected value of the grid where the vehicle navigation target point is located can be set to be constant at 1, and the initial expected value values of the remaining grids can be set to 0.
[0089] Step 232 , based on the known state transition equation and the expected value of any grid in the expected value grid diagram of the previous iteration round, determine the expected action value corresponding to different actions taken at any grid.
[0090] Among them, the state transfer equation is s * =T(s,a), where s * Indicates the next grid to jump to after executing action a from the current grid s. Action a represents the actions that can be performed at grid s, including up, down, left, right, upper left, lower left, upper right, and lower right. T(s,a) indicates the transition to the new state s * Probabilistic or deterministic rules. In the case of a determined reward map, for any grid s, the grid s reached after taking different actions * The location can be determined.
[0091] In the disclosed embodiment, the reward value of any grid in the reward map is obtained by real-time prediction based on input information such as static road network information, dynamic obstacle information, and navigation target points at the current time step, and is used to represent the immediate benefit obtained by the vehicle after reaching a certain grid.
[0092] Among them, there are 8 actions that can be taken at any grid, namely 8 actions around the grid: up, down, left, right, upper left, lower left, upper right, and lower right.
[0093] For example, based on the above state transition equation and the reward value of any grid in the reward map, the expected value of taking different actions can be obtained using formula (1).
[0094] Q k (s,a)=r θ (s)+V k (s * ) Formula (1)
[0095] In formula (1), Q k (s,a) represents the expected value (action expected value) of taking action a in grid s in the kth iteration round, and the initial value of k is 0; s represents the grid, V k is the expected value of grid s, r is the learnable reward function output by the network, r θ (s) represents the reward value of grid s, s * is the state transfer equation s * =T(s,a).
[0096] Step 233 , based on the expected action value of taking different actions in any grid, the expected value value of any grid in the value expectation grid diagram of the previous round is updated to obtain the expected value grid diagram of the current round.
[0097] Among them, the expected value of taking different actions at any grid is the above Q k , the value expectation grid map of the previous round (kth round) is the value expectation grid map obtained in the previous round. At the initial iteration, the value expectation grid map of the previous round is the value expectation grid map of the initial round obtained in step 231.
[0098] Among them, the expected value of taking different actions in any grid represents the estimate of the weighted sum of multiple future rewards that can be obtained after taking a certain action in the current state.
[0099] In the embodiment of the present disclosure, the expected value of any grid in the value expectation grid map obtained in the previous round can be updated based on the expected value of taking different actions in any grid to obtain the value expectation grid map of the current round. This can be achieved using formula (2).
[0100]
[0101] In formula (2), V k+1 (s) represents the expected value of the grid s obtained in the k+1th iteration round; Q k (s,a) represents the expected value (action expected value) of taking action a in grid s in the kth iteration round, and the initial value of k is 0.
[0102] Through the above steps 232 and 233, an iterative update of the value expectation grid map is achieved. Through this iterative update, the value expectation value of each grid in the value expectation grid map is changed accordingly.
[0103] After executing step 232 and step 233 , an operation of determining expected values of actions taken at any grid may be performed.
[0104] In the embodiment of the present disclosure, the value expectation grid map of the initial round is the value expectation grid map initialized in step 231, that is, the value expectation grid map of the 0th iteration round. The action expectation value Q of taking different actions at any grid in the 0th iteration round can be determined based on the value expectation grid map of the 0th iteration round. 0 (s, a), and then based on the expected value Q of taking different actions at any grid in the 0th iteration round 0 (s, a), update the expected value of any grid in the expected value grid map of the 0th iteration round to obtain the expected value grid map of the 1st iteration round; then, based on the expected value grid map of the 1st iteration round, determine the expected action value Q of taking different actions at any grid in the 1st iteration round. 1 (s, a), and then take different actions based on the expected value Q of any grid in the first iteration round 1(s, a), update the expected value of any grid in the expected value grid map of the first iteration round to obtain the expected value grid map of the second iteration round; and so on, the iterative update of the expected value grid map of the initial round can be achieved.
[0105] Step 234 , in response to reaching the iteration termination condition, obtaining a target value expectation grid map of the final iteration round.
[0106] In the disclosed embodiment, the iteration termination condition indicates the conditions for stopping the iterative update of the expected value grid. The iteration termination condition may include any of the following: the difference between the expected value of any grid and the expected value of the previous round is less than a given threshold; or the iteration is terminated when a specified maximum number of iterations is reached, for example, 10 iterations.
[0107] Among them, the target value expectation grid map is the value expectation grid map generated by the last iteration round.
[0108] Step 235 , based on the action expectation values of taking different actions in any grid and the value expectation value of any grid in the target value expectation grid diagram, obtain the action probability of taking different actions in any grid.
[0109] In the embodiment of the present disclosure, after determining the expected action value of taking different actions in any grid and the expected value of any grid in the target value expectation grid diagram through the above steps, the action probability of taking different actions in any grid can be further determined.
[0110] For example, the action probability of taking different actions in any grid can be determined by formula (3).
[0111]
[0112] In formula (3), represents the probability of taking different actions in any grid, Q k (s,a) represents the expected value (action expected value) of taking action a at any grid s in the kth iteration round. The initial value of k is 0. V k (s) represents the expected value of any grid s in the kth iteration round.
[0113] The above formulas (1), (2), and (3) provide implementation methods for determining the expected value of action and the expected value, and the probability of action at grid s. Based on the description of the embodiments of the present disclosure, those skilled in the art can know that a similar implementation method can be used to determine the expected value of action and the expected value at grid s, which will not be repeated here.
[0114] Step 236 , based on the action probabilities of taking different actions in any grid, probability sampling is performed on the target value expectation grid map to obtain a path map.
[0115] In the embodiment of the present disclosure, based on the target value expectation grid diagram that has been iteratively converged, the action expectation value corresponding to any action taken in any grid state can be obtained starting from the starting grid of the vehicle, that is, the above Q k The corresponding selection probability is assigned according to the expected value of each action, and the actions are sampled step by step to complete the deduction, forming a grid path from the navigation starting point to the navigation target point. Multiple samplings can be performed to obtain multiple such paths to form a path map.
[0116] For example, Figure 5 The reward map shown in the left figure is sampled probabilistically, and we can get Figure 5 The path map is shown on the right.
[0117] Based on the embodiment of the present disclosure, the expected value of any grid in the reward map is initialized based on vehicle navigation information to obtain the value expectation grid map of the initial round; based on the known state transition equation and the reward value of any grid in the reward map, the action expectation value of taking different actions at any grid is determined; based on the action expectation value of taking different actions at any grid, the expected value of any grid in the value expectation grid map of the previous round is updated to obtain the value expectation grid map of the current round; iteratively executes the operation of determining the action expectation value of taking different actions at any grid; in response to reaching the iteration termination condition, the target value expectation grid map of the final iteration round is obtained; based on the action expectation value of taking different actions at any grid, the action probability of taking different actions at any grid is obtained, and further action sampling is performed to finally obtain a path map, which provides scene prior information support for the subsequent generation of continuous trajectory points.
[0118] Figure 6 FIG. 2 is a flow chart of step 205 of a vehicle driving trajectory planning method provided by an exemplary embodiment of the present disclosure. Figure 6 As shown in the above Figure 2 Based on the illustrated embodiment, the method for determining the vehicle trajectory in the future time period in step 205 may include the following steps:
[0119] Step 251 : performing feature fusion on scene features corresponding to any state path of at least one state path to obtain a third fused feature.
[0120] The state path is used to indicate the grid sequence connecting the starting point and the navigation target point, while the starting point represents the grid position of the vehicle at the current moment (the moment when the vehicle's driving trajectory is planned). The scene feature of the state path is used to indicate the vector composed of the feature values of each grid in the state path at the corresponding index on the scene feature map. The scene feature map is a feature map obtained by extracting features from the two-dimensional image corresponding to the static road network information and the two-dimensional image corresponding to the dynamic obstacle information, and then fusing the extracted features. This feature map contains the scene features (the aforementioned feature values) corresponding to each grid.
[0121] In the disclosed embodiment, feature fusion can be performed on the scene features corresponding to the multiple state paths to obtain a third fused feature that combines the scene features of the multiple state paths. The fusion method can be the feature addition or feature concatenation method described above, which will not be detailed here.
[0122] The third fusion feature is used to indicate a vector composed of feature values of corresponding indexes of each grid in the path map on the scene feature map.
[0123] Step 252: extract features from the vehicle status information to obtain vehicle status features.
[0124] In the embodiment of the present disclosure, the vehicle state information includes linear velocity and rotational angular velocity. In other possible implementations, the vehicle state information may include more information, such as acceleration, angular acceleration, etc., according to actual needs.
[0125] In the embodiment of the present disclosure, a convolutional neural network can be used to extract features of vehicle status information to obtain vehicle status features corresponding to the vehicle status information. The embodiment of the present disclosure does not limit the method of extracting features of vehicle status information.
[0126] Step 253 : Generate the vehicle acceleration and curvature at any time step based on the vehicle state feature and the third fusion feature.
[0127] Among them, the time step is a time point in the future time period. For example, if the future time period is 6 seconds and a trajectory point is predicted every half second, the time step is 12 time steps such as 0.5 seconds, 1 second, 1.5 seconds, 2 seconds, 2.5 seconds, 3 seconds, 3.5 seconds, 4 seconds, 4.5 seconds, 5 seconds, 5.5 seconds, and 6 seconds after the current moment.
[0128] In the disclosed embodiment, a recurrent neural network, such as a long short-term memory network (LSTM), may be used to cyclically output acceleration and curvature information for each time step based on vehicle state features and the third fusion feature.
[0129] In step 254 , a kinematic model is used to generate a trajectory point of the next time step adjacent to any time step based on the vehicle acceleration and curvature at any time step.
[0130] A kinematic model is a mathematical model that describes the motion of an object through geometric methods, primarily studying the relationship between an object's position, velocity, and time. In the disclosed embodiments, the kinematic model can be used to generate trajectory points for the next time step adjacent to a given time step, i.e., the position of the next time step, based on the vehicle's acceleration and curvature at any time step.
[0131] In the disclosed embodiment, the speed and heading angle of the vehicle at any time step can be determined based on the acceleration and curvature of the vehicle at any time step; the trajectory point of the vehicle at the next time step adjacent to any time step can be determined based on the speed and heading angle corresponding to the vehicle at any time step, as well as the position of the vehicle at any time step.
[0132] For example, Figure 7 As shown, after encoding the scene features of each state path to obtain the third fused feature and extracting the vehicle state information (state), the acceleration (acc) and curvature (cur) information, as well as the model's intermediate layer features (h, such as h1, h2, h3...hn-1), are output through the LSTM. The model's intermediate layer features h and the third fused feature are then input into the LSTM to execute the acceleration and curvature information for the next time step, until the acceleration and curvature information for all time steps in the future time period is output. The output acceleration and curvature information for each time step is input into the kinematic model (model), which generates the trajectory points for the next time step adjacent to any time step based on the vehicle acceleration and curvature at any time step.
[0133] In specific implementation, the kinematic model is based on the acceleration and curvature information of each time step, and the speed and heading angle of the vehicle at the current time step can be determined by equations (4) and (5):
[0134]
[0135] v t+1 =v t +a t Δt formula (5)
[0136] In formula (4), represents the vehicle's heading angle at the previous time step, v t represents the speed of the vehicle at the previous time step, κ trepresents the curvature information of the vehicle in the previous time step, and Δt represents the time interval between two adjacent time steps.
[0137] In formula (5), v t represents the speed of the vehicle in the previous time step, a t represents the acceleration information of the vehicle in the previous time step, and Δt represents the time interval between two adjacent time steps.
[0138] After determining the vehicle's speed and heading angle at the current time step through equations (4) and (5), the vehicle's position information at the current time step can be further determined through equations (6) and (7):
[0139]
[0140] In formula (6), x t Indicates the horizontal coordinate of the vehicle's trajectory point at the previous time step, v t represents the speed of the vehicle at the previous time step, represents the vehicle's heading angle at the previous time step, and Δt represents the time interval between two adjacent time steps.
[0141] In formula (7), y t Indicates the vertical coordinate of the vehicle's trajectory point at the previous time step, v t represents the speed of the vehicle at the previous time step, represents the vehicle's heading angle at the previous time step, and Δt represents the time interval between two adjacent time steps.
[0142] Step 255 : Based on the trajectory points at any time step, the vehicle trajectory of the vehicle in the future time period is obtained.
[0143] In the disclosed embodiment, the vehicle trajectory of the vehicle in a preset future time period can be obtained by performing curve fitting processing on the trajectory points at any time step in the future time period.
[0144] In specific implementation, after obtaining the trajectory points of each time step in the future time period, the trajectory points can be connected to obtain the motion trajectory of the vehicle in the future time period, such as Figure 8 The vehicle trajectory indicated by the middle number 81.
[0145] Based on the embodiment of the present disclosure, a third fusion feature is obtained by performing feature fusion on the scene features corresponding to any state path of at least one state path; feature extraction is performed on the vehicle state information to obtain vehicle state features; based on the vehicle state features and the third fusion features, the vehicle acceleration and curvature at any time step are generated; based on the vehicle acceleration and curvature at any time step, a kinematic model is used to generate the trajectory points of the next time step adjacent to any time step; based on the trajectory points at any time step, the vehicle trajectory of the vehicle in the future time period is obtained. Therefore, the embodiment of the present disclosure discloses an implementation method for obtaining the vehicle trajectory of the vehicle in the future time period based on the scene features corresponding to each state path in the path map. The recurrent neural network can process and learn temporal dependencies in the middle of the sequence number, can capture long-term dependencies in the trajectory data, and thus achieve accurate prediction of the vehicle driving trajectory in various scenarios, with strong generalization ability.
[0146] In some optional implementations, the behavior planning module and trajectory planning module ( Figure 2 The functions shown in the system framework can be realized by a trajectory planning model. Figure 9 FIG. 1 is a schematic diagram of a training process of a trajectory planning model in a vehicle driving trajectory planning method provided by an exemplary embodiment of the present disclosure. Figure 9 As shown, the training operation of the trajectory planning model includes the following steps:
[0147] Step 901: Acquire a trajectory data sample set. The trajectory data samples in the trajectory data sample set carry trajectory tags. The trajectory data include static road network information samples, dynamic obstacle information samples, and navigation map information samples.
[0148] The trajectory data in the trajectory data sample may be trajectory data sampled during a real road test, and may include static road network information samples, dynamic obstacle information samples, and navigation map information samples. The trajectory label is the actual driving trajectory of the vehicle during the road test.
[0149] In the disclosed embodiment, the trajectory data samples collected during the road test can be fed back into the system as the data type and format required by the trajectory planning model, such as the BEV static road network, dynamic obstacle rendering, and navigation point rendering.
[0150] In some optional implementations, new training samples can be obtained through offline optimization methods based on the deviation from the trajectory data sampled from the actual road test process, and aggregated into the above-mentioned trajectory data samples for model training, which helps to obtain richer samples and improve the generalization performance of the trained model.
[0151] In step 902 , for each trajectory data sample, the initial trajectory planning model is used to perform trajectory prediction based on the trajectory data sample to obtain a predicted trajectory.
[0152] In the disclosed embodiments, the initial trajectory planning model may be a model for predicting the vehicle's driving trajectory based on trajectory data samples to obtain a predicted trajectory. The initial trajectory planning model may utilize various existing network types, such as Faster RCNN (Faster Region-CNN), Retinanet, YoloX, FCOS (Fully Convolutional One-Stage), and AutoAssign. The initial trajectory planning model may determine a predicted trajectory from the trajectory data.
[0153] In step 903 , the initial trajectory planning model is iteratively trained based on the predicted trajectory and trajectory label corresponding to the trajectory data sample until a preset training completion condition is met, and a trajectory planning model is obtained corresponding to the initial trajectory planning model.
[0154] In this embodiment, the electronic device can determine a loss value for the error between the predicted trajectory and the trajectory label based on a preset loss function. The preset loss function can employ an existing loss function used to train a trajectory planning model. In this embodiment, the electronic device can adjust parameters of the initial trajectory planning model based on the loss value.
[0155] Among them, the training process of the initial trajectory planning model is a process of finding the optimal solution. The process of fitting the model to the optimal solution is mainly carried out iteratively by minimizing the error. For an input trajectory data sample, the above-mentioned preset loss function can be used to calculate the difference between the actual output of the initial trajectory planning model (i.e., the above-mentioned predicted trajectory) and the expected output (i.e., the above-mentioned trajectory label). The backpropagation algorithm then transmits this difference to the connection between each neuron in the initial trajectory planning model. The difference signal transmitted to each connection represents the contribution rate of the connection to the overall error. The model parameters of the initial trajectory planning model are then updated and modified using the gradient descent algorithm, so that the loss value calculated during the iterative training process gradually decreases.
[0156] By repeatedly executing steps 902 to 903, i.e., using multiple sets of training samples, the model is iteratively trained, and the model after each iterative training is the initial trajectory planning model for the next training. When the initial trajectory data sample after adjusting the parameters meets the training end condition, the current initial trajectory data sample model is the trained trajectory data sample. The training end condition may include but is not limited to at least one of the following: the loss value of the above-mentioned loss function converges, the training time exceeds the preset time length, and the number of training times exceeds the preset number.
[0157] In some optional implementations, after completing the model training through the above operations, the trained trajectory planning model can be packaged and adapted to the real vehicle interface, so that the system can support the operation test of the real vehicle and the trajectory planning model simulator, and the trajectory planning sample data can be loaded into the trajectory planning model simulator in the form of re-injection to perform closed-loop test evaluation and improve the quality of the model.
[0158] Based on the embodiment of the present disclosure, by obtaining a trajectory data sample set, the trajectory data in the trajectory data sample set carries a trajectory label, and the trajectory data includes a static road network information sample, a dynamic obstacle information sample, and a navigation map information sample; for each trajectory data sample, respectively, using the initial trajectory planning model, trajectory prediction is performed based on the trajectory data sample to obtain a predicted trajectory; based on the predicted trajectory and trajectory label corresponding to the trajectory data sample, the initial trajectory planning model is iteratively trained until the preset training completion conditions are met, and a trajectory planning model is obtained corresponding to the initial trajectory planning model. The embodiment of the present disclosure trains a trajectory planning model through real trajectory data from road tests, which helps to reduce the impact of uncertainty in the vehicle driving environment and make more robust decisions; in addition, by testing the model with a real vehicle, it helps to reduce the number of road tests, shorten development time and reduce costs, and improve the planning effect of the vehicle's driving trajectory.
[0159] The electronic device (also referred to as an intelligent agent) of the present disclosure can be Figure 10 The trajectory planning system shown implements the vehicle trajectory planning method and / or other desired functions provided by the present disclosure.
[0160] Figure 10 2 is an exemplary system framework diagram of an embodiment of the present disclosure, including: input information 21, a behavior planning module 22 and a trajectory planning module 23.
[0161] The input information 21 can be obtained through upstream perception tasks and includes static road network information, dynamic obstacle information, vehicle navigation information, and vehicle status information. Static road network information includes, but is not limited to, information about static objects such as buildings, trees, plants, traffic signs, intersections, roads, stop lines, road arrows, and curbs; dynamic obstacle information includes, but is not limited to, information about dynamic objects such as vehicles, pedestrians, bicycles, motorcycles, and pets.
[0162] The behavior planning module 22 can extract features of static road network information, dynamic obstacle information, and vehicle navigation information respectively, and then fuse the extracted features to obtain scene features ( Figure 10 221 in ); then the fused scene features ( Figure 10221) in the downsampling process, and feature fusion with the vehicle status information to obtain the reward map ( Figure 10 222 in ); the probability of taking different actions under any grid ( Figure 10 223 in), sample the reward map and get the path map ( Figure 10 224).
[0163] The trajectory planning module 23 can integrate the scene features 231 of each state path in the path map 224, and combine it with the vehicle state information to use the kinematic model to predict the vehicle's movement trajectory in the future time period to obtain the vehicle trajectory 24.
[0164] The behavior planning module 22 may be configured to execute the above Figure 2 In the embodiment shown, the operations of step 202 to step 204 are to extract and fuse the features of the static road network information, dynamic obstacle information, vehicle navigation information and vehicle status information to obtain a reward map, and based on the reward map, the value iteration update algorithm is used to iteratively converge to obtain a target value expectation grid map, and the target value expectation grid map is sampled to obtain a path map; the trajectory planning module 23 can be configured to perform the above Figure 2 The operation of step 205 in the illustrated embodiment is to autoregressively predict the trajectory point sequence of the vehicle in the future time period based on the scene features of each state path in the path map and the vehicle state information, and obtain the movement trajectory of the vehicle in the future time period.
[0165] Figure 10 This is only an exemplary system framework of the embodiment of the present disclosure. Those skilled in the art can know based on the description of the embodiment of the present disclosure that the embodiment of the present disclosure can also adopt any other feasible implementation method. For example, the behavior planning module and the trajectory planning module can be a trajectory planning model or a sub-model in the model. By connecting with the upstream perception task, it receives the task perception information of the upstream perception task and processes it using the vehicle trajectory planning method provided by the embodiment of the present disclosure to obtain the vehicle's motion trajectory in the future time period and feedback to the vehicle driving control system.
[0166] Exemplary devices
[0167] Figure 11 FIG. 1 is a schematic diagram of a vehicle trajectory planning device provided by an exemplary embodiment of the present disclosure. Figure 11 As shown, the vehicle trajectory planning device includes:
[0168] An information determination module 111 is used to determine static road network information, dynamic obstacle information, and vehicle navigation information of the vehicle's surrounding environment;
[0169] A first encoding module 112 is configured to extract and fuse features of the static road network information, the dynamic obstacle information, and the vehicle navigation information to obtain a reward map;
[0170] An iteration module 113 is configured to iteratively generate a target value expectation grid map based on the reward map;
[0171] A sampling module 114 is configured to sample the target value expectation grid map to obtain a path map, wherein the path map includes at least one state path;
[0172] The first prediction module 115 is used to predict the trajectory point of the vehicle at any time step in a preset future time period based on the scene features and vehicle state information corresponding to any state path, and obtain the vehicle trajectory of the vehicle in the preset future time period.
[0173] Figure 12 FIG. 1 is a schematic diagram of a vehicle driving trajectory planning device provided by another exemplary embodiment of the present disclosure. Figure 12 As shown, in Figure 11 Based on the illustrated embodiment, in some implementations, the first encoding module 112 includes:
[0174] A first encoding submodule 1121 is configured to extract and fuse features of the static road network information, the dynamic obstacle information, and the vehicle navigation information to obtain a first fusion feature;
[0175] A downsampling submodule 1122 is configured to perform downsampling processing on the first fused feature to obtain a second fused feature;
[0176] The first fusion submodule 1123 is configured to fuse the second fusion feature with the feature corresponding to the vehicle state information to obtain a reward map.
[0177] In some implementations, the iteration module 113 may include:
[0178] Initialization submodule 1131 is used to initialize the expected value of any grid in the reward map based on the vehicle navigation information to obtain the expected value grid map of the initial round;
[0179] A first determination submodule 1132 is configured to determine the expected value of taking different actions at any grid based on a known state transition equation and a reward value at any grid in the reward map;
[0180] An updating submodule 1133 is configured to update the expected value of any grid in the value expectation grid map of the previous round based on the expected value of taking different actions in any grid, to obtain the expected value grid map of the current round;
[0181] The iterative submodule 1134 is configured to iteratively execute an operation of determining expected action values of different actions taken at any grid, and obtain a target value expectation grid diagram of a final iteration round in response to reaching an iteration termination condition.
[0182] In some implementations, the sampling module 114 may include:
[0183] The probability determination submodule 1141 is configured to obtain the action probability of taking different actions at any grid based on the action expected value of taking different actions at any grid and the value expected value of any grid in the target value expected grid graph;
[0184] The sampling submodule 1142 is used to perform probability sampling on the target value expectation grid map based on the action probability of taking different actions in any grid to obtain a path map.
[0185] In some embodiments, the first prediction module 115 may include:
[0186] The second fusion submodule 1151 is configured to perform feature fusion on scene features corresponding to any state path of at least one state path to obtain a third fused feature;
[0187] The feature extraction submodule 1152 is used to extract features from the vehicle state information to obtain vehicle state features;
[0188] The state generation submodule 1153 is used to generate the vehicle acceleration and curvature at any time step based on the vehicle state feature and the third fusion feature;
[0189] The trajectory point generation submodule 1154 is used to generate a trajectory point of the next time step adjacent to any time step based on the vehicle acceleration and curvature at any time step using a kinematic model;
[0190] The trajectory fitting submodule 1155 is used to obtain the vehicle trajectory in the future time period based on the trajectory points at any time step.
[0191] In some embodiments, the trajectory point generation submodule 1154 is specifically used to: determine the speed and heading angle of the vehicle at any time step based on the acceleration and curvature of the vehicle at any time step; and determine the trajectory point of the vehicle at the next time step adjacent to any time step based on the speed and heading angle corresponding to the vehicle at any time step, as well as the position of the vehicle at any time step.
[0192] In some embodiments, the trajectory fitting submodule 1155 is specifically configured to perform curve fitting on the trajectory points at any time step in the future time period to obtain the vehicle trajectory of the vehicle in the preset future time period.
[0193] In some embodiments, the vehicle trajectory planning device may further include:
[0194] A sample acquisition module 116 is configured to acquire a trajectory data sample set, wherein the trajectory data samples in the trajectory data sample set carry trajectory tags, and the trajectory data include static road network information samples, dynamic obstacle information samples, and navigation map information samples;
[0195] The second prediction module 117 is used to use the initial trajectory planning model to perform trajectory prediction based on the trajectory data sample for each trajectory data sample to obtain a predicted trajectory;
[0196] The model training module 118 is used to iteratively train the initial trajectory planning model based on the predicted trajectory and trajectory label corresponding to the trajectory data sample until the preset training completion condition is met, and the trajectory planning model corresponding to the initial trajectory planning model is obtained.
[0197] It should be noted that the specific implementation of the vehicle driving trajectory planning device in the embodiment of the present disclosure is similar to the specific implementation of the vehicle driving trajectory planning method in the embodiment of the present disclosure. Please refer to the vehicle driving trajectory planning method part for details. In order to reduce redundancy, it will not be described in detail.
[0198] Exemplary computer program products and computer-readable storage media
[0199] In addition to the above-mentioned methods and devices, embodiments of the present disclosure may also provide a computer program product, including computer program instructions, which, when executed by a processor, enable the processor to execute the steps of the vehicle driving trajectory planning method of various embodiments of the present disclosure described in the above-mentioned "Exemplary Method" section.
[0200] The computer program product may be written in any combination of one or more programming languages to implement the operations of the disclosed embodiments, including object-oriented programming languages such as Java, C++, and conventional procedural programming languages such as C or similar programming languages. The program code may be executed entirely on the user's computing device, partially on the user's computing device, as a stand-alone software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server.
[0201] In addition, an embodiment of the present disclosure may also be a computer-readable storage medium having computer program instructions stored thereon. When the computer program instructions are executed by a processor, the processor executes the steps in the vehicle driving trajectory planning method of various embodiments of the present disclosure described in the above-mentioned "Exemplary Method" section.
[0202] Computer readable storage media can adopt any combination of one or more readable media. The readable medium can be a readable signal medium or a readable storage medium. The readable storage medium is, for example, but not limited to, a system, device or component comprising electricity, magnetism, light, electromagnetic, infrared, or semiconductor, or any combination thereof. More specific examples (non-exhaustive list) of readable storage media include: an electrical connection with one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof.
[0203] The basic principles of the present disclosure have been described above in conjunction with specific embodiments. However, the advantages, strengths, and effects mentioned in this disclosure are merely illustrative and not restrictive, and should not be considered as essential to each embodiment of the present disclosure. Furthermore, the specific details disclosed above are provided for illustrative purposes and to facilitate understanding, rather than as limitations. These details do not limit the present disclosure to necessarily being implemented using these specific details.
[0204] Those skilled in the art may make various changes and modifications to the present disclosure without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present disclosure and their equivalents, the present disclosure is intended to include these modifications and variations.
Claims
1. A vehicle trajectory planning method, comprising: Determine static road network information, dynamic obstacle information and vehicle navigation information of the vehicle's surrounding environment; Extracting and fusing features of the static road network information, the dynamic obstacle information, and the vehicle navigation information to obtain a reward map; Iteratively generating a target value expectation grid map based on the reward map; Sampling the target value expectation grid map to obtain a path map, wherein the path map includes at least one state path; Based on the scene features and vehicle state information corresponding to any state path, the trajectory points of the vehicle at any time step in a future time period are predicted to obtain the motion trajectory of the vehicle in the future time period.
2. The method according to claim 1, wherein The feature extraction and fusion of the static road network information, the dynamic obstacle information, and the vehicle navigation information to obtain a reward map includes: Extracting and fusing features of the static road network information, the dynamic obstacle information, and the vehicle navigation information to obtain a first fusion feature; Downsampling the first fused features to obtain second fused features; The second fused feature is fused with the feature corresponding to the vehicle state information to obtain the reward map.
3. The method according to claim 1, wherein The iterative generation of a target value expectation grid map based on the reward map includes: Initializing the expected value of any grid in the reward map to obtain an expected value grid map for the initial round; Determine the action expectation value corresponding to different actions taken at any grid based on the known state transition equation and the value expectation value of any grid in the value expectation grid diagram of the previous round, where the value expectation grid diagram of the previous round is the value expectation grid diagram of the initial round in the initial round; Based on the expected action values corresponding to different actions taken in any of the grids, the expected value value of any grid in the value expectation grid map of the previous round is updated to obtain the expected value grid map of the current round; Iteratively performing the operation of determining expected action values corresponding to taking different actions at any grid; In response to reaching the iteration termination condition, the target value expectation grid diagram of the final iteration round is obtained.
4. The method according to claim 3, wherein: The sampling of the target value expectation grid map to obtain a path map includes: Based on the expected value of taking different actions at any grid and the expected value of any grid in the target value expectation grid graph, the action probability of taking different actions at any grid is obtained; Based on the action probabilities of taking different actions in any grid, probability sampling is performed on the target value expectation grid map to obtain the path map.
5. The method according to any one of claims 1 to 4, wherein: The step of predicting a trajectory point of the vehicle at any time step in a future time period based on scene features and vehicle state information corresponding to any state path, and obtaining a motion trajectory of the vehicle in the future time period, includes: Performing feature fusion on scene features corresponding to any state path of the at least one state path to obtain a third fused feature; Performing feature extraction on the vehicle state information to obtain vehicle state features; generating the vehicle acceleration and curvature at any time step based on the vehicle state feature and the third fused feature; Using a kinematic model, based on the acceleration and curvature of the vehicle at any time step, a trajectory point of a next time step adjacent to the any time step is generated; Based on the trajectory points at any time step, the vehicle trajectory of the vehicle in the future time period is obtained.
6. The method according to claim 5, wherein: The generating of a trajectory point of a next time step adjacent to any time step based on the vehicle acceleration and curvature at any time step by using a kinematic model includes: Determining the speed and heading angle of the vehicle at any time step based on the acceleration and curvature of the vehicle at any time step; Based on the speed and orientation angle of the vehicle at any time step, and the position of the vehicle at any time step, a trajectory point of the vehicle at a next time step adjacent to the any time step is determined.
7. The method according to claim 5, wherein: The obtaining of the vehicle trajectory of the vehicle in the future time period based on the trajectory point at any time step includes: Curve fitting is performed on the trajectory points at any time step in the future time period to obtain the vehicle trajectory of the vehicle in the preset future time period.
8. The method according to any one of claims 1 to 7, further comprising: Acquire a trajectory data sample set, wherein the trajectory data samples in the trajectory data sample set carry trajectory tags, and the trajectory data include static road network information samples, dynamic obstacle information samples, and navigation map information samples; For each trajectory data sample, using the initial trajectory planning model, perform trajectory prediction based on the trajectory data sample to obtain a predicted trajectory; Based on the predicted trajectory and trajectory label corresponding to the trajectory data sample, the initial trajectory planning model is iteratively trained until a preset training completion condition is met, and a trajectory planning model corresponding to the initial trajectory planning model is obtained.
9. A vehicle trajectory planning device, comprising: An information determination module, used to determine static road network information, dynamic obstacle information and vehicle navigation information of the vehicle's surrounding environment; A first encoding module is used to extract and fuse features of the static road network information, the dynamic obstacle information, and the vehicle navigation information to obtain a reward map; An iteration module, configured to iteratively generate a target value expectation grid map based on the reward map; A sampling module, configured to sample the target value expectation grid map to obtain a path map, wherein the path map includes at least one state path; The trajectory prediction module is used to predict the trajectory point of the vehicle at any time step in a preset future time period based on the scene features and vehicle state information corresponding to any state path, and obtain the vehicle trajectory of the vehicle in the preset future time period.
10. A computer-readable storage medium, wherein the storage medium stores computer program instructions, and when the computer program instructions are executed by a processor, the computer program instructions are used to implement the method according to any one of claims 1 to 8.
11. An electronic device, comprising: processor; a memory for storing instructions executable by the processor; The processor is configured to read the executable instructions from the memory and execute the instructions to implement the method according to any one of claims 1 to 8.