Intelligent driving trajectory prediction method and device, equipment and storage medium
By integrating data from LiDAR, radar, and cameras to generate multimodal perception features and combining them with a target cost function to optimize trajectory selection, the problem of understanding complex environments and predicting trajectories in intelligent driving is solved, achieving more accurate environmental perception and trajectory prediction.
Patent Information
- Application Number
- CN202411903359.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-23
- Publication Date
- 2025-10-21
- Estimated Expiration
- 2044-12-23
AI Technical Summary
In existing intelligent driving technologies, trajectory prediction schemes fail to fully consider real-time changes in the surrounding environment, resulting in an inability to accurately understand complex driving environments and predict the optimal driving trajectory.
Integrate lidar, radar, and camera data, convert them into bird's-eye view features, align them, generate multimodal perception features, and optimize the selection of the optimal trajectory based on the target cost function.
It provides richer and more accurate environmental perception, enabling a better understanding of complex driving environments, including traffic flow, pedestrians, obstacles, and the behavior of other vehicles, and predicting the optimal driving trajectory.
Smart Images

Figure CN119682788B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of intelligent driving technology, and in particular to intelligent driving trajectory prediction methods, devices, equipment and storage media. Background Art
[0002] With the development of intelligent driving technology, higher requirements are being placed on the accuracy and real-time performance of trajectory prediction. A trajectory prediction method with good generalization capabilities is needed, which can accurately reflect vehicle dynamics, environmental changes, and inter-vehicle interactions to adapt to different driving scenarios and conditions.
[0003] Currently, some solutions use long short-term memory (LSTM) networks for trajectory prediction and behavioral intent classification. This information is then fed into a multimodal LSTM network to generate the final predicted trajectory. However, these predictions fail to fully account for real-time environmental changes. Therefore, in the process of intelligent driving, accurately understanding the complex driving environment to predict the optimal driving trajectory remains an unresolved issue.
[0004] The above content is only used to assist in understanding the technical solution of this application and does not constitute an admission that the above content is prior art. Summary of the Invention
[0005] The main purpose of this application is to provide an intelligent driving trajectory prediction method, device, equipment and storage medium, aiming to solve the technical problem of how to accurately understand the complex driving environment and predict the optimal driving trajectory during the intelligent driving process.
[0006] To achieve the above objectives, the present application proposes a method for intelligent driving trajectory prediction, which includes:
[0007] Acquire environmental lidar data, environmental radar data, and environmental image data;
[0008] converting the environmental lidar data, the environmental radar data, and the environmental image data into bird's-eye view features;
[0009] Performing feature alignment on the bird's-eye view features based on the sensor field of view angle identifier to obtain a multimodal perception feature;
[0010] determining a candidate trajectory based on the multimodal perception features;
[0011] A target trajectory is determined from the candidate trajectories according to a target cost function to complete intelligent driving trajectory prediction.
[0012] In one embodiment, the step of determining candidate trajectories based on the multimodal perception features includes:
[0013] Determining path planning language information based on the multimodal perception features and the environmental image data;
[0014] Obtain vehicle driving status information, road environment information, obstacle status information, and traffic rules information;
[0015] A candidate trajectory is generated according to the path planning language information, the environmental image data, the vehicle driving state information, the road environment information, the obstacle dynamic information and the traffic rule information.
[0016] In one embodiment, the step of determining the path planning language information based on the multimodal perception features and the environmental image data includes:
[0017] Decoding the multimodal perception features to obtain a bird's-eye view visual marker;
[0018] Encoding the environmental image data to obtain a text label;
[0019] Path planning language information is determined based on the bird's-eye view visual mark and the text mark.
[0020] In one embodiment, the step of generating candidate trajectories based on the path planning language information, the environmental image data, the vehicle driving state information, the road environment information, the obstacle dynamic information, and the traffic rule information includes:
[0021] Encoding the trajectory planning language information to obtain a trajectory planning vector;
[0022] A candidate trajectory is generated according to the trajectory planning vector, the environmental image data, the vehicle driving state information, the road environment information, the obstacle dynamic information and the traffic rule information.
[0023] In one embodiment, the sensor field of view angle identifier includes a laser radar field of view angle identifier, a radar field of view angle identifier, and an image field of view angle identifier;
[0024] The step of aligning the features of the bird's-eye view based on the sensor field of view angle identifier to obtain multimodal perception features includes:
[0025] Obtaining sensor field of view angle identifiers corresponding to the environmental lidar data, the environmental radar data, and the environmental image data, and obtaining the lidar field of view angle identifier, the radar field of view angle identifier, and the image field of view angle identifier;
[0026] The bird's-eye view features are aligned based on the laser radar field of view angle identifier, the radar field of view angle identifier, and the image field of view angle identifier to obtain multimodal perception features.
[0027] In one embodiment, the step of aligning the bird's-eye view features based on the laser radar field of view angle identifier, the radar field of view angle identifier, and the image field of view angle identifier to obtain multimodal perception features includes:
[0028] Assigning a modal weight to the bird's-eye view feature based on the laser radar field of view angle identifier, the radar field of view angle identifier, and the image field of view angle identifier;
[0029] The bird's-eye view features are aligned according to the modal weights to obtain multimodal perception features.
[0030] In one embodiment, the step of determining a target trajectory from the candidate trajectories according to a target cost function to complete intelligent driving trajectory prediction includes:
[0031] Constructing the target cost function based on safety cost, comfort cost, compliance cost, prediction uncertainty cost and path continuity cost;
[0032] A target trajectory is determined from the candidate trajectories according to the target cost function to complete intelligent driving trajectory prediction.
[0033] In addition, to achieve the above objectives, the present application also proposes an intelligent driving trajectory prediction device, which includes:
[0034] A data acquisition module is used to acquire environmental lidar data, environmental radar data, and environmental image data;
[0035] a feature extraction module, configured to convert the environmental lidar data, the environmental radar data, and the environmental image data into bird's-eye view features;
[0036] A feature alignment module, configured to align the features of the bird's-eye view based on the sensor field of view angle identifier to obtain a multimodal perception feature;
[0037] a trajectory generation module, configured to determine candidate trajectories based on the multimodal perception features;
[0038] The trajectory prediction module is used to determine the target trajectory from the candidate trajectories according to the target cost function to complete the intelligent driving trajectory prediction.
[0039] In addition, to achieve the above-mentioned purpose, the present application also proposes an intelligent driving trajectory prediction device, which includes: a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein the computer program is configured to implement the steps of the intelligent driving trajectory prediction method as described above.
[0040] In addition, to achieve the above-mentioned purpose, the present application also proposes a storage medium, which is a computer-readable storage medium. A computer program is stored on the storage medium, and when the computer program is executed by the processor, the steps of the intelligent driving trajectory prediction method as described above are implemented.
[0041] In addition, to achieve the above-mentioned purpose, the present application also provides a computer program product, which includes a computer program, and when the computer program is executed by a processor, it implements the steps of the intelligent driving trajectory prediction method as described above.
[0042] One or more technical solutions proposed in this application have at least the following technical effects:
[0043] Acquire environmental lidar data, environmental radar data, and environmental image data; convert the environmental lidar data, environmental radar data, and environmental image data into bird's-eye view features; perform feature alignment on the bird's-eye view features based on sensor field of view angle identifiers to obtain multimodal perception features; determine candidate trajectories based on the multimodal perception features; and determine a target trajectory from the candidate trajectories based on a target cost function to complete intelligent driving trajectory prediction. By integrating data from different sensors, such as lidar, radar, and cameras, a richer and more accurate environmental perception is provided. By converting the data from each sensor into bird's-eye view features and aligning the features based on the sensor's field of view, a multimodal perception feature containing perspective information from different sensors is obtained. This can comprehensively describe the behavior of traffic flow, pedestrians, obstacles, and other vehicles in the vehicle's driving environment. Based on these multimodal perception features, potential trajectories, namely candidate trajectories, are identified and predicted. The optimal trajectory is selected through optimization using a target cost function, enabling the intelligent driving system to accurately understand and predict complex driving environments, providing richer and more accurate environmental perception. The solution of this application enables the intelligent driving system to better understand the complex driving environment, including traffic flow, pedestrians, obstacles and the behavior of other vehicles, in order to predict the optimal driving trajectory. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.
[0045] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0046] Figure 1A flowchart illustrating the first embodiment of the intelligent driving trajectory prediction method of this application;
[0047] Figure 2 A flowchart illustrating the second embodiment of the intelligent driving trajectory prediction method of this application;
[0048] Figure 3 A schematic diagram of a simplified process of the intelligent driving trajectory prediction method provided in Example 2 of this application;
[0049] Figure 4 This is a schematic diagram of the module structure of the intelligent driving trajectory prediction device according to an embodiment of the present application;
[0050] Figure 5 Schematic diagram of the device structure of the hardware operating environment involved in the intelligent driving trajectory prediction method in the embodiment of the present application.
[0051] The purpose, features and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. DETAILED DESCRIPTION
[0052] It should be understood that the specific embodiments described herein are merely used to explain the technical solutions of the present application and are not intended to limit the present application.
[0053] In order to better understand the technical solution of the present application, a detailed description will be given below in conjunction with the accompanying drawings and specific implementation methods.
[0054] The main solutions of the embodiments of this application are:
[0055] Acquire environmental lidar data, environmental radar data, and environmental image data;
[0056] converting the environmental lidar data, the environmental radar data, and the environmental image data into bird's-eye view features;
[0057] Performing feature alignment on the bird's-eye view features based on the sensor field of view angle identifier to obtain a multimodal perception feature;
[0058] determining a candidate trajectory based on the multimodal perception features;
[0059] A target trajectory is determined from the candidate trajectories according to a target cost function to complete intelligent driving trajectory prediction.
[0060] In this embodiment, for ease of description, the following description is made with the identification of the intelligent driving system as the execution subject.
[0061] With the development of intelligent driving technology, higher requirements are being placed on the accuracy and real-time performance of trajectory prediction. A trajectory prediction method with good generalization capabilities is needed, which can accurately reflect vehicle dynamics, environmental changes, and inter-vehicle interactions to adapt to different driving scenarios and conditions.
[0062] Currently, some solutions use long short-term memory (LSTM) networks for trajectory prediction and behavioral intent classification. This information is then fed into a multimodal LSTM network to generate the final predicted trajectory. However, these predictions fail to fully account for real-time environmental changes. Therefore, in the process of intelligent driving, accurately understanding the complex driving environment to predict the optimal driving trajectory remains an unresolved issue.
[0063] The present application provides a solution that provides richer and more accurate environmental perception by integrating data from different sensors such as lidar, radar, and cameras. The data from each sensor is converted into bird's-eye view features, and the features are aligned using the sensor's field of view to obtain multimodal perception features containing perspective information from different sensors. This can comprehensively describe the traffic flow, pedestrians, obstacles, and other vehicle behaviors in the vehicle's driving environment. Based on these multimodal perception features, potential trajectories, i.e., candidate trajectories, are identified and predicted. The optimal trajectory is selected through target cost function optimization, enabling the intelligent driving system to accurately understand and predict complex driving environments, providing richer and more accurate environmental perception. The solution of the present application enables the intelligent driving system to better understand complex driving environments, including traffic flow, pedestrians, obstacles, and other vehicle behaviors, in order to predict the optimal driving trajectory.
[0064] It should be noted that the execution subject of this embodiment can be a computing service device with data processing, network communication, and program execution functions, such as a tablet computer, personal computer, mobile phone, etc., or an electronic device capable of implementing the above functions, such as an intelligent driving system. The following uses the intelligent driving system as an example to illustrate this embodiment and the following embodiments.
[0065] Based on this, the embodiment of the present application provides an intelligent driving trajectory prediction method, referring to Figure 1 , Figure 1 This is a flow chart of the first embodiment of the intelligent driving trajectory prediction method of this application.
[0066] In this embodiment, the intelligent driving trajectory prediction method includes steps S10 to S40:
[0067] Step S10, acquiring environmental lidar data, environmental radar data, and environmental image data;
[0068] It should be understood that environmental lidar data is the three-dimensional point cloud data of the vehicle's surroundings collected by lidar while the vehicle is driving, reflecting the three-dimensional spatial coordinates (x, y, z) and reflection intensity information of each point. Environmental radar data is the distance, speed, and angle information collected by radar sensors (radar), such as millimeter-wave radar, while the vehicle is driving. Environmental image data is image data of the surrounding environment collected by the vehicle's onboard camera while the vehicle is driving.
[0069] Step S20, converting the environmental lidar data, the environmental radar data, and the environmental image data into bird's-eye view features;
[0070] It should be understood that the bird's-eye view is a two-dimensional representation of the vehicle's surroundings, which provides a perspective of looking down at the vehicle and its surrounding objects from above, and projects all target objects in the view space (such as vehicles, pedestrians, obstacles, etc.) onto the ground to form a simplified view. The Bird's Eye View (BEV) feature includes a first bird's-eye view feature, a second bird's-eye view feature, and a third bird's-eye view feature. Among them, the first bird's-eye view feature is a 3D feature extracted from the environmental lidar data through a 3D encoder (such as PointPillars); the second bird's-eye view feature is a 3D feature extracted from the environmental radar data through a 3D encoder (such as PointPillars); the third bird's-eye view feature is a visual feature extracted from the environmental image data through a visual encoder (such as ResNet) and a multi-view feature fusion module (such as a Transformer model).
[0071] Specifically, PointPillars divides the three-dimensional space into a grid, generating multiple pillars. Each pillar represents a small spatial region, and a fixed size (e.g., 0.1m × 0.1m × 3m) can be selected to discretize the ambient space. For 3D features from LiDAR and Radar sensors, PointPillars maps the information of each point to a corresponding Pillar on the BEV plane, with each Pillar containing a certain number of features. Pillar features are extracted and concatenated to form a unified BEV feature map, namely the first or second bird's-eye view features corresponding to the ambient LiDAR or radar data. ResNet is used to extract features from the ambient image data, generating multi-view visual features containing key information about the ambient image (such as road signs, traffic lights, pedestrians, and vehicles). Transformers can be used to integrate multi-view visual features extracted from different cameras to generate third bird's-eye view features with image feature representation.
[0072] Step S30, aligning the bird's-eye view features based on the sensor field of view angle identifier to obtain multimodal perception features;
[0073] It should be noted that the information obtained by extracting features from different sources (i.e., lidar, radar, and camera) or different perspectives (such as front, back, left, and right) based on the bird's-eye view features requires feature alignment through the sensor fusion model. Sensor field of view angles can be used to assign a specific label to each sensor, helping the sensor fusion model identify and distinguish sensors in different orientations, thereby obtaining multimodal perception features.
[0074] In a feasible implementation, step S30 may include steps S31 to S32:
[0075] Step S31, assigning a modal weight to the bird's-eye view feature based on the laser radar field of view angle identifier, the radar field of view angle identifier, and the image field of view angle identifier;
[0076] It should be noted that the sensor field of view angle identification includes the laser radar field of view angle identification, the radar field of view angle identification and the image field of view angle identification. Among them, the laser radar field of view angle identification is the label assigned to the laser radar with different viewing angles. For example, the label corresponding to the laser radar on the right can be<RIGHT_LIDAR> ; Radar field of view angle is the label assigned to radars with different viewing angles. For example, the label corresponding to the radar on the left can be<LEFT_RADAR> ; The image field of view is identified by the labels assigned to cameras with different viewing angles. For example, the label corresponding to the front-view camera can be<FRONT_CAM> . By comparing sensor signals with different fields of view, the sensor fusion model can infer the relative position of objects, the direction of movement of the vehicle, the geometric structure of the road, etc. Specifically, it can identify vehicles in the adjacent lane ahead based on the forward visual image of the environmental image data, and verify these vehicles based on the side radar information of the environmental radar data, thereby unifying the data from different perspectives into a complete BEV view space, so that the network of the sensor fusion model can learn efficiently without the need to retrieve data from different perspectives separately. In the process of feature alignment, a weight, namely the modal weight, is assigned to the feature set corresponding to each bird's-eye view feature. The modal weight can be learned through the training process of the preset sensor fusion model using a multimodal dataset.
[0077] Step S32: aligning the bird's-eye view features according to the modal weights to obtain multimodal perception features.
[0078] It should be understood that after the bird's-eye view features and the corresponding sensor field of view angle identifiers are input into the sensor fusion model, the bird's-eye view features of each sensor (i.e., lidar, radar, and camera) will be properly projected and aligned according to the sensor field of view angle identifier, and the data of different sensors will be unified into the BEV view space to obtain multimodal perception features. The weighted bird's-eye view features are added according to the modal weights to obtain the fused feature representation F fused , that is, multimodal perception features. F fused The expression formula is as follows:
[0079] F fused =F1·w1+F2·w2+...+F n w n
[0080] Among them, F1, F2, ..., F n are bird’s-eye view features from different sources, w1, w2, ..., w n is the corresponding weight, and n depends on the number of sensors of different types and orientations. For example, if a vehicle is equipped with one lidar in one viewing orientation (e.g., the front), four radars in four viewing orientations (e.g., the front, rear, left, and right), and four cameras in four viewing orientations (e.g., the front, rear, left, and right), then n is 9.
[0081] Step S40, determining candidate trajectories based on the multimodal perception features;
[0082] It should be noted that candidate trajectories are multiple possible vehicle trajectories. Candidate trajectories can be obtained based on a kinematic model or trajectory planning algorithm, combined with multimodal perception features, vehicle driving state information, and other information.
[0083] In a feasible implementation, step S40 may include steps S41 to S43:
[0084] Step S41, determining path planning language information based on the multimodal perception features and the environmental image data;
[0085] It should be noted that path planning language information is high-level planning decisions expressed in natural language, such as "go straight," "turn left," or "yield to pedestrians." After feature processing of multimodal perception features and environmental image data, corresponding visual tokens and textual tokens are generated. These visual tokens and textual tokens are then input into a pre-trained Large Language Model (LLM) to obtain the corresponding path planning language information.
[0086] In a feasible implementation, step S41 may include steps S411 to S413:
[0087] Step S411, decoding the multimodal perception features to obtain a bird's-eye view visual mark;
[0088] It's important to note that the BEV decoder decodes multimodal perception features to produce corresponding bird's-eye-view visual markers, or visual tokens. Visual tokens are images that contain information such as road boundaries, traffic signs, obstacles, and vehicle positions. While preserving depth information, these images highlight the spatial distribution of each object or obstacle, providing the intelligent driving system with a vehicle-centric bird's-eye view.
[0089] Step S412, encoding the environment image data to obtain a text label;
[0090] It should be noted that the environmental image data can be encoded through a text encoder to obtain corresponding text tags, namely text tokens. Specifically, the environmental image data can be processed through a convolutional neural network to obtain a high-dimensional feature vector. The high-dimensional feature vector contains object type information (such as lane lines, obstacles, traffic lights, etc.) in the environmental image data, object location information (such as the coordinates of obstacles, etc.) and object state information (such as stationary or moving). The high-dimensional feature vector is input into the text encoder, and the high-dimensional feature vector is mapped to text tokens to obtain a natural language description of the image, such as "There is a parked vehicle 50 meters ahead."
[0091] Step S413: determining path planning language information according to the bird's-eye view visual mark and the text mark.
[0092] It should be understood that by inputting the bird's-eye view visual markers (i.e., visual tokens) and text markers (i.e., text tokens) into the pre-trained large language model (Large Language Model, LLM), the corresponding path planning language information will be obtained. The large language model includes a self-attention mechanism and a multi-layer perceptron (Multi-Layer Perceptron, MLP). Specifically, the LLM can process the input bird's-eye view visual markers and text markers through the self-attention mechanism, identify the semantic association between the bird's-eye view visual markers and the text markers, and assign corresponding weights to perform feature fusion to obtain fused features. The fused features are successively subjected to nonlinear transformation and feature mapping by the multi-layer perceptron, and path planning language information that is easy to understand and interpret is generated at the end of the large language model. The large language model based on vision and semantics can process and interpret natural language instructions, improve the generalization ability of the intelligent driving system, and while enabling the intelligent driving system to work in different driving scenarios and conditions, its interpretability enhances the transparency of the intelligent driving system, allowing developers and users to understand the decision-making process of the intelligent driving system and increase trust.
[0093] Step S42, obtaining vehicle driving status information, road environment information, obstacle status information, and traffic regulation information;
[0094] It should be noted that vehicle driving state information refers to dynamic information about the vehicle during driving (such as speed, steering angle, acceleration, etc.). Road environment information refers to the road type (urban road, highway, etc.) and road surface type (such as wet, dry, icy, etc.) of the road section on which the vehicle is traveling. Obstacle state information includes static obstacles (such as parked vehicles, roadblocks, etc.) and dynamic obstacles (such as pedestrians, other moving vehicles, etc.) on the current road on which the vehicle is traveling. Traffic rule information refers to traffic signs, signal light status, speed limit regulations, etc. on the current road on which the vehicle is traveling.
[0095] Step S43 , generating a candidate trajectory based on the path planning language information, the environmental image data, the vehicle driving state information, the road environment information, the obstacle dynamic information, and the traffic rule information.
[0096] It should be noted that after data processing of path planning language information, environmental image data, vehicle driving status information, road environment information, obstacle dynamic information and traffic rule information, and inputting the processed data into a pre-trained end-to-end trajectory planning model (such as a Transformer model), several possible trajectories, namely candidate trajectories, can be obtained. The end-to-end trajectory planning model can realize the complete trajectory planning process from the initial state of the vehicle to the target state, and can be obtained by training the preset Transformer model. Each candidate trajectory is a possible driving path of the vehicle in the future, represented by a series of location information (such as longitude and latitude or coordinate points in a local coordinate system).
[0097] In a feasible implementation, step S43 may include steps S431 to S432:
[0098] Step S431, encoding the trajectory planning language information to obtain a trajectory planning vector;
[0099] It should be noted that the trajectory planning vector is a high-dimensional feature vector, and each high-dimensional feature vector represents a specific meta-action. The trajectory planning language information can be encoded into a trajectory planning vector through the meta-action encoder for subsequent calculations. Specifically, the meta-action encoder maps the output of the language model (i.e., the trajectory planning language information) to a set of predefined meta-actions, encodes the meta-actions into high-dimensional feature vectors, and realizes the conversion of meta-actions to trajectory planning feature vectors through a set of learnable embedding vectors (Embeddings). Each meta-action corresponds to an embedding vector, and the embedding vector can be learned in the process of training the meta-action encoder. Predefined meta-actions can include lateral meta-actions and longitudinal meta-actions. Horizontal meta-actions include at least turning left, going straight, and turning right, and longitudinal meta-actions include at least accelerating, holding, decelerating, and stopping.
[0100] For example, the mapping process of meta-actions is explained by taking the trajectory planning language information "go straight", "turn left" or "avoid pedestrians" as an example: "go straight" is mapped to the horizontal meta-action "go straight" and the vertical meta-action "keep the speed unchanged"; "turn left" is mapped to the horizontal meta-action "turn left" and the vertical meta-action "slow down"; "avoid pedestrians" is mapped to the successive execution of the horizontal meta-action "turn right" and the vertical meta-action "slow down", the horizontal meta-action "turn left" and the vertical meta-action "keep the speed unchanged", the horizontal meta-action "turn left" and the vertical meta-action "keep the speed unchanged", the horizontal meta-action "turn right" and the vertical meta-action "keep the speed unchanged", and the horizontal meta-action "go straight" and the vertical meta-action "accelerate".
[0101] Step S432 : generating a candidate trajectory based on the trajectory planning vector, the environmental image data, the vehicle driving state information, the road environment information, the obstacle dynamic information, and the traffic rule information.
[0102] It should be understood that after obtaining the trajectory planning vector corresponding to the meta-action mapped by the trajectory planning vector, the trajectory planning vector, environmental image data, vehicle driving status information, road environment information, obstacle dynamic information and traffic rule information are input into the pre-trained end-to-end trajectory planning model (such as the Transformer model), and several possible trajectories, namely candidate trajectories, can be obtained.
[0103] Step S50 , determining a target trajectory from the candidate trajectories according to a target cost function to complete intelligent driving trajectory prediction.
[0104] It should be understood that the target cost function is a mathematical function designed based on multiple factors (such as safety, driving comfort, traffic rule compliance, obstacle avoidance, etc.). It can calculate the cost of each candidate trajectory and then select the optimal trajectory. After obtaining the target cost function score for each candidate trajectory, an optimization algorithm (gradient descent) can be used to find the trajectory that minimizes the cost function, thus filtering the optimal trajectory from the predicted trajectories.
[0105] This embodiment provides an intelligent driving trajectory prediction method, which obtains environmental lidar data, environmental radar data, and environmental image data; converts the environmental lidar data, environmental radar data, and environmental image data into bird's-eye view features; performs feature alignment on the bird's-eye view features based on sensor field of view angle identifiers to obtain multimodal perception features; determines candidate trajectories based on the multimodal perception features; and determines a target trajectory from the candidate trajectories based on a target cost function to complete intelligent driving trajectory prediction. By integrating data from different sensors, such as lidar, radar, and cameras, a richer and more accurate environmental perception is provided. The data from each sensor is converted into bird's-eye view features, and the features are aligned using the sensor field of view angle to obtain multimodal perception features containing perspective information from different sensors. This can comprehensively describe the behavior of traffic flow, pedestrians, obstacles, and other vehicles in the vehicle's driving environment. Based on these multimodal perception features, potential trajectories, namely candidate trajectories, are identified and predicted. The optimal trajectory is selected through optimization using a target cost function, enabling the intelligent driving system to accurately understand and predict complex driving environments, providing richer and more accurate environmental perception. The solution of this application enables the intelligent driving system to better understand the complex driving environment, including traffic flow, pedestrians, obstacles and the behavior of other vehicles, in order to predict the optimal driving trajectory.
[0106] Based on the first embodiment of the present application, in the second embodiment of the present application, the same or similar contents as those in the above embodiment 1 can be referred to the above introduction and will not be described in detail later. Figure 2 , step S50 may include steps S51 to S52:
[0107] Step S51, constructing a target cost function based on safety cost, comfort cost, compliance cost, prediction uncertainty cost, and path continuity cost;
[0108] It should be understood that safety costs are used to assess the potential risks and collision probability of candidate trajectories, including the risk of collision between the vehicle and obstacles, pedestrians, other vehicles, etc. in the candidate trajectory. Comfort costs are used to assess the comfort experienced by passengers during driving, including the assessment of sudden braking, acceleration, or sharp turns in the candidate trajectory. Compliance costs are used to assess the degree to which the vehicle's driving behavior in the candidate trajectory complies with traffic regulations (such as speed limits, traffic lights, traffic signs, lane keeping, etc.), including the assessment of violations of traffic signals, speeding, and irregular lane changes in the candidate trajectory. Prediction uncertainty costs are used to assess the errors introduced in the trajectory prediction process due to environmental factors and the behavior of other traffic participants, including the assessment of prediction errors in the behavior of other vehicles and pedestrians in the candidate trajectory and the uncertainty of changing road conditions. Path continuity costs are used to assess the smoothness and continuity of the vehicle's driving trajectory, including the assessment of sudden trajectory changes, sharp turns, and excessive path curvature in the candidate trajectory.
[0109] It should be noted that the safety cost, comfort cost, compliance cost, prediction uncertainty cost, and path continuity cost can be calculated for candidate trajectories using vehicle body information (such as speed, steering angle, acceleration, etc.). Based on the safety cost, comfort cost, compliance cost, prediction uncertainty cost, and path continuity cost, a cost function C, namely the target cost function, can be constructed and defined to quantify the pros and cons of each candidate trajectory. The construction formula of the target cost function C is as follows:
[0110] C=λ1·C safety +λ2·C comfort +λ3·C compliance +λ4·C uncertainty +λ5·C continuity
[0111] Among them, λ i is the weight of the i-th cost factor, reflecting the importance of this factor in the total cost. safety Used to represent security cost; C comfort Used to express comfort cost; C complianceUsed to represent compliance costs; C uncertainty Used to represent the forecast uncertainty cost; C continuity Used to represent path continuity cost.
[0112] Step S52 : determining a target trajectory from the candidate trajectories according to the target cost function to complete intelligent driving trajectory prediction.
[0113] It should be understood that after obtaining all candidate trajectories and their corresponding safety cost, comfort cost, compliance cost, prediction uncertainty cost, and path continuity cost scores, an optimization algorithm can be used to select the optimal trajectory from the candidate trajectories. Specifically, the optimization algorithm can be gradient descent. After setting the initial trajectory and the target cost function C, the trajectory is updated in the direction of the decreasing cost function by calculating the gradient. Each updated trajectory is adjusted in the direction of decreasing cost. Among all the generated candidate trajectories, the trajectory that minimizes the cost function after this optimization process is the optimal trajectory, that is, the target trajectory.
[0114] This embodiment provides a method for intelligent driving trajectory prediction. A target cost function is constructed based on safety cost, comfort cost, compliance cost, prediction uncertainty cost, and path continuity cost. A target trajectory is determined from the candidate trajectories based on the target cost function to complete intelligent driving trajectory prediction. By constructing an efficient trajectory screening strategy, the intelligent system can quickly identify and select the optimal trajectory, i.e., the target trajectory, from multiple possible trajectories, i.e., candidate trajectories. This enables more efficient vehicle navigation during intelligent driving, reduces vehicle energy consumption, and improves driving smoothness, safety, and comfort.
[0115] For example, in order to help understand the implementation process of the intelligent driving trajectory prediction method obtained by combining this embodiment with the above embodiment 1, please refer to Figure 3 , Figure 3 A brief flowchart of an intelligent driving trajectory prediction method is provided. Specifically:
[0116] Data on the vehicle's surroundings is collected through three sensors: lidar, radar, and camera. This data is then converted into lidar data, radar data, and image data. The 3D encoder PointPillars extracts features from the lidar and radar data to obtain corresponding bird's-eye view features. A visual encoder (such as ResNet) and a multi-view feature fusion module (such as the Transformer model) extract features from the image data to obtain corresponding bird's-eye view features. The bird's-eye view features are then aligned to obtain multimodal perception features. The BEV decoder decodes the multimodal perception features to obtain corresponding BEV features (i.e., bird's-eye view visual markers), i.e., visual tokens. The text encoder converts the image data into text tokens. The visual and text tokens are then fed into the large language model (LLM) to obtain high-level planning decisions (i.e., path planning language information). The high-level planning decisions are encoded through the meta-action encoder to obtain the trajectory planning vector. The trajectory planning vector is combined with vehicle body information (such as speed, steering angle, acceleration, etc.) to perform end-to-end trajectory planning and generate vehicle trajectories (i.e., candidate trajectories). The vehicle trajectory is then verified based on the vehicle body information to achieve optimal trajectory screening and output the final trajectory planning result, i.e., the target trajectory.
[0117] It should be noted that the above examples are only used to understand the present application and do not constitute a limitation on the intelligent driving trajectory prediction method of the present application. More simple transformations based on this technical concept are all within the scope of protection of the present application.
[0118] This application also provides an intelligent driving trajectory prediction device, please refer to Figure 4 , the intelligent driving trajectory prediction device includes:
[0119] The data acquisition module 10 is used to acquire environmental lidar data, environmental radar data and environmental image data;
[0120] a feature extraction module 20 for converting the environmental lidar data, the environmental radar data, and the environmental image data into bird's-eye view features;
[0121] A feature alignment module 30 is configured to perform feature alignment on the bird's-eye view features based on the sensor field of view angle identifier to obtain a multimodal perception feature;
[0122] a trajectory generation module 40, configured to determine candidate trajectories based on the multimodal perception features;
[0123] The trajectory prediction module 50 is used to determine the target trajectory from the candidate trajectories according to the target cost function to complete the intelligent driving trajectory prediction.
[0124] In one embodiment, the trajectory generation module 40 is further used to determine path planning language information based on the multimodal perception features and the environmental image data; obtain vehicle driving state information, road environment information, obstacle state information, and traffic rule information; and generate a candidate trajectory based on the path planning language information, the environmental image data, the vehicle driving state information, the road environment information, obstacle dynamic information, and the traffic rule information.
[0125] In one embodiment, the trajectory generation module 40 is further used to decode the multimodal perception features to obtain a bird's-eye view visual mark; encode the environmental image data to obtain a text mark; and determine the path planning language information based on the bird's-eye view visual mark and the text mark.
[0126] In one embodiment, the trajectory generation module 40 is further used to encode the trajectory planning language information to obtain a trajectory planning vector; and generate a candidate trajectory based on the trajectory planning vector, the environmental image data, the vehicle driving state information, the road environment information, the obstacle dynamic information and the traffic rule information.
[0127] In one embodiment, the feature alignment module 30 is further used to obtain the sensor field of view angle identifiers corresponding to the environmental lidar data, the environmental radar data, and the environmental image data, and obtain the lidar field of view angle identifier, the radar field of view angle identifier, and the image field of view angle identifier; based on the lidar field of view angle identifier, the radar field of view angle identifier, and the image field of view angle identifier, feature alignment is performed on the bird's-eye view features to obtain multimodal perception features.
[0128] In one embodiment, the feature alignment module 30 is further used to assign modal weights to the bird's-eye view features based on the lidar field of view angle identifier, the radar field of view angle identifier, and the image field of view angle identifier; and perform feature alignment on the bird's-eye view features according to the modal weights to obtain multimodal perception features.
[0129] In one embodiment, the trajectory prediction module 50 is further used to construct a target cost function based on safety cost, comfort cost, compliance cost, prediction uncertainty cost, and path continuity cost; and determine a target trajectory from the candidate trajectories according to the target cost function to complete intelligent driving trajectory prediction.
[0130] The intelligent driving trajectory prediction device provided in this application, which employs the intelligent driving trajectory prediction method described in the aforementioned embodiments, can address the technical problem of accurately understanding complex driving environments to predict the optimal driving trajectory during intelligent driving. Compared to the prior art, the beneficial effects of the intelligent driving trajectory prediction device provided in this application are the same as those of the intelligent driving trajectory prediction method described in the aforementioned embodiments. Other technical features of the intelligent driving trajectory prediction device are the same as those disclosed in the aforementioned embodiments and are not further elaborated here.
[0131] The present application provides an intelligent driving trajectory prediction device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the intelligent driving trajectory prediction method in the above-mentioned embodiment one.
[0132] Reference below Figure 5 , which shows a schematic diagram of the structure of an intelligent driving trajectory prediction device suitable for implementing the embodiments of the present application. The intelligent driving trajectory prediction device in the embodiments of the present application may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Portable Application Descriptions), PMPs (Portable Media Players), in-vehicle terminals (such as in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. Figure 5 The intelligent driving trajectory prediction device shown is merely an example and should not impose any limitations on the functions and scope of use of the embodiments of the present application.
[0133] like Figure 5As shown, the intelligent driving trajectory prediction device may include a processing device 1001 (e.g., a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes based on programs stored in a read-only memory (ROM) 1002 or programs loaded from a storage device 1003 into a random access memory (RAM) 1004. RAM 1004 also stores various programs and data required for the operation of the intelligent driving trajectory prediction device. The processing device 1001, ROM 1002, and RAM 1004 are connected to each other via a bus 1005. An input / output (I / O) interface 1006 is also connected to the bus. Typically, the following systems can be connected to the I / O interface 1006: an input device 1007 including, for example, a touch screen, touchpad, keyboard, mouse, image sensor, microphone, accelerometer, gyroscope, etc.; an output device 1008 including, for example, a liquid crystal display (LCD), speaker, vibrator, etc.; a storage device 1003 including, for example, a magnetic tape, hard disk, etc.; and a communication device 1009. The communication device 1009 can allow the intelligent driving trajectory prediction device to communicate wirelessly or wired with other devices to exchange data. Although the figure shows an intelligent driving trajectory prediction device with various systems, it should be understood that it is not required to implement or include all of the systems shown. More or fewer systems may be implemented or included instead.
[0134] In particular, according to the embodiments disclosed in the present application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, the embodiments disclosed in the present application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program comprising program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network via a communication device, or installed from a storage device 1003, or installed from a ROM 1002. When the computer program is executed by the processing device 1001, the above-mentioned functions defined in the method of the embodiment disclosed in the present application are executed.
[0135] The intelligent driving trajectory prediction device provided in this application, which utilizes the intelligent driving trajectory prediction method described in the aforementioned embodiment, can address the technical problem of accurately understanding complex driving environments to predict the optimal driving trajectory during intelligent driving. Compared to the prior art, the beneficial effects of the intelligent driving trajectory prediction device provided in this application are the same as those of the intelligent driving trajectory prediction method described in the aforementioned embodiment. Other technical features of this intelligent driving trajectory prediction device are the same as those disclosed in the aforementioned embodiment and are not further elaborated here.
[0136] It should be understood that the various parts disclosed in this application can be implemented using hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in any one or more embodiments or examples in a suitable manner.
[0137] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.
[0138] The present application provides a computer-readable storage medium having computer-readable program instructions (i.e., computer programs) stored thereon, and the computer-readable program instructions are used to execute the intelligent driving trajectory prediction method in the above-mentioned embodiment.
[0139] The computer-readable storage medium provided in this application may be, for example, a USB flash drive, but is not limited to electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, systems or devices, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In this embodiment, the computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, system or device. The program code contained on the computer-readable storage medium may be transmitted using any appropriate medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination thereof.
[0140] The above-mentioned computer-readable storage medium may be included in the intelligent driving trajectory prediction device; or it may exist independently without being assembled into the intelligent driving trajectory prediction device.
[0141] The computer-readable storage medium carries one or more programs. When the one or more programs are executed by the intelligent driving trajectory prediction device, the intelligent driving trajectory prediction device:
[0142] Acquire environmental lidar data, environmental radar data, and environmental image data;
[0143] converting the environmental lidar data, the environmental radar data, and the environmental image data into bird's-eye view features;
[0144] Performing feature alignment on the bird's-eye view features based on the sensor field of view angle identifier to obtain a multimodal perception feature;
[0145] determining a candidate trajectory based on the multimodal perception features;
[0146] A target trajectory is determined from the candidate trajectories according to a target cost function to complete intelligent driving trajectory prediction.
[0147] Computer program code for performing the operations of the present application may be written in one or more programming languages, or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, C++, and conventional procedural programming languages such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on the remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider).
[0148] The flow charts and block diagrams in the accompanying drawings illustrate the possible architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present application. In this regard, each box in the flow chart or block diagram can represent a module, program segment or a part of code, and the module, program segment or a part of code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order than that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flow chart, and the combination of the boxes in the block diagram and / or flow chart can be implemented by a dedicated hardware-based system that performs the specified function or operation, or can be implemented by a combination of dedicated hardware and computer instructions.
[0149] The modules described in the embodiments of the present application may be implemented in software or hardware, wherein the name of a module does not necessarily limit the unit itself.
[0150] The computer-readable storage medium provided in this application stores computer-readable program instructions (i.e., a computer program) for executing the aforementioned intelligent driving trajectory prediction method. This computer-readable storage medium can address the technical problem of accurately understanding complex driving environments to predict optimal driving trajectories during intelligent driving. Compared to the prior art, the beneficial effects of the computer-readable storage medium provided in this application are the same as those of the intelligent driving trajectory prediction method provided in the aforementioned embodiments, and are not further elaborated here.
[0151] The present application also provides a computer program product, including a computer program, which, when executed by a processor, implements the steps of the above-mentioned intelligent driving trajectory prediction method.
[0152] The computer program product provided in this application can solve the technical problem of accurately understanding complex driving environments and predicting optimal driving trajectories during intelligent driving. Compared to the prior art, the beneficial effects of the computer program product provided in this application are the same as those of the intelligent driving trajectory prediction method provided in the above-mentioned embodiments, and will not be elaborated here.
[0153] The above description is only part of the embodiments of the present application and does not limit the patent scope of the present application. All equivalent structural transformations made by using the contents of the present application specification and drawings under the technical concept of the present application, or direct / indirect application in other related technical fields are included in the patent protection scope of the present application.
Claims
1. An intelligent driving trajectory prediction method, characterized in that: The intelligent driving trajectory prediction method includes: Acquire environmental lidar data, environmental radar data, and environmental image data; converting the environmental lidar data, the environmental radar data, and the environmental image data into bird's-eye view features; Performing feature alignment on the bird's-eye view features based on the sensor field of view angle identifier to obtain a multimodal perception feature; determining a candidate trajectory based on the multimodal perception features; Determining a target trajectory from the candidate trajectories according to a target cost function to complete intelligent driving trajectory prediction; The step of determining candidate trajectories based on the multimodal perception features includes: Determining path planning language information based on the multimodal perception features and the environmental image data; Obtain vehicle driving status information, road environment information, obstacle status information, and traffic rules information; generating a candidate trajectory based on the path planning language information, the environmental image data, the vehicle driving state information, the road environment information, the obstacle dynamic information, and the traffic rule information; The step of determining a target trajectory from the candidate trajectories according to a target cost function to complete intelligent driving trajectory prediction includes: Constructing the target cost function based on safety cost, comfort cost, compliance cost, prediction uncertainty cost and path continuity cost; A target trajectory is determined from the candidate trajectories according to the target cost function to complete intelligent driving trajectory prediction.
2. The method according to claim 1, wherein The step of determining the path planning language information based on the multimodal perception features and the environmental image data includes: Decoding the multimodal perception features to obtain a bird's-eye view visual marker; Encoding the environmental image data to obtain a text label; Path planning language information is determined based on the bird's-eye view visual mark and the text mark.
3. The method according to claim 1, wherein The step of generating candidate trajectories based on the path planning language information, the environmental image data, the vehicle driving state information, the road environment information, the obstacle dynamic information, and the traffic rule information includes: Encoding the path planning language information to obtain a trajectory planning vector; A candidate trajectory is generated according to the trajectory planning vector, the environmental image data, the vehicle driving state information, the road environment information, the obstacle dynamic information and the traffic rule information.
4. The method according to claim 1, wherein The sensor field of view angle identifier includes a laser radar field of view angle identifier, a radar field of view angle identifier, and an image field of view angle identifier; The step of aligning the features of the bird's-eye view based on the sensor field of view angle identifier to obtain multimodal perception features includes: Obtaining sensor field of view angle identifiers corresponding to the environmental lidar data, the environmental radar data, and the environmental image data, and obtaining the lidar field of view angle identifier, the radar field of view angle identifier, and the image field of view angle identifier; The bird's-eye view features are aligned based on the laser radar field of view angle identifier, the radar field of view angle identifier, and the image field of view angle identifier to obtain multimodal perception features.
5. The method according to claim 4, wherein The feature alignment of the bird's-eye view features is performed based on the laser radar field of view angle identifier, the radar field of view angle identifier, and the image field of view angle identifier. The steps of obtaining multimodal perception features include: Assigning a modal weight to the bird's-eye view feature based on the laser radar field of view angle identifier, the radar field of view angle identifier, and the image field of view angle identifier; The bird's-eye view features are aligned according to the modal weights to obtain multimodal perception features.
6. An intelligent driving trajectory prediction device, characterized in that: The device comprises: A data acquisition module is used to acquire environmental lidar data, environmental radar data, and environmental image data; a feature extraction module, configured to convert the environmental lidar data, the environmental radar data, and the environmental image data into bird's-eye view features; A feature alignment module, configured to align the features of the bird's-eye view based on the sensor field of view angle identifier to obtain a multimodal perception feature; a trajectory generation module, configured to determine candidate trajectories based on the multimodal perception features; a trajectory prediction module, configured to determine a target trajectory from the candidate trajectories according to a target cost function to complete intelligent driving trajectory prediction; The trajectory generation module is further configured to determine path planning language information based on the multimodal perception features and the environmental image data; obtain vehicle driving state information, road environment information, obstacle state information, and traffic rule information; and generate a candidate trajectory based on the path planning language information, the environmental image data, the vehicle driving state information, the road environment information, obstacle dynamic information, and the traffic rule information; The trajectory prediction module is further used to construct a target cost function based on safety cost, comfort cost, compliance cost, prediction uncertainty cost and path continuity cost; and determine the target trajectory from the candidate trajectories according to the target cost function to complete intelligent driving trajectory prediction.
7. An intelligent driving trajectory prediction device, characterized in that: The device includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program is configured to implement the steps of the intelligent driving trajectory prediction method according to any one of claims 1 to 5.
8. A storage medium, characterized in that: The storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by a processor, the steps of the intelligent driving trajectory prediction method according to any one of claims 1 to 5 are implemented.
Citation Information
Patent Citations
Trajectory prediction method and device
CN114708723A
Target detection method and device, computer equipment and storage medium
CN117911824A