An autonomous driving trajectory prediction method and device based on long-term and short-term prediction
By adopting a long-term prediction method in the autonomous driving system, combining long-term target prediction, short-term trajectory prediction and future trajectory prediction modules, dynamically fusion of future information is solved, and the problem of long-term prediction error accumulation is improved, and the accuracy and robustness of prediction are improved.
Patent Information
- Application Number
- CN202410664898.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-05-27
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2044-05-27
AI Technical Summary
The existing autonomous driving trajectory prediction methods are prone to error accumulation during long-term prediction, difficult to capture complex traffic scenarios and behavior patterns, and lack effective integration of future dynamic scenario information.
The method based on long-term prediction is adopted, and the combination of long-term target prediction Gru, short-term trajectory prediction Gru and future trajectory prediction Gru is used to dynamically integrate information from future moments, and the method of cyclic step-by-step prediction is adopted to enhance the prediction ability of future trajectories.
It effectively solves the problem of long-term prediction error accumulation, improves the modeling ability of complex traffic scenarios and behavioral patterns, enhances the ability to integrate future dynamic scenario information, and significantly improves the accuracy and robustness of trajectory prediction.
Smart Images

Figure CN118521003B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of autonomous driving trajectory prediction, and particularly to an autonomous driving trajectory prediction method and device based on long short-term prediction. Background Art
[0002] In order to ensure the safe and effective operation of autonomous vehicles on the road, it is not only necessary to understand the motion states of surrounding traffic participants in real time, but also to predict their future motion trajectories. This prediction can be regarded as simulating the prediction ability of human drivers, which helps to improve the awareness of driving risks of autonomous vehicles and estimate potential dangers. The autonomous driving system is mainly divided into three modules: perception, prediction, and planning. Among them, the trajectory prediction module plays a crucial role in connecting the above and the below. This module predicts the action trajectories of other traffic participants by receiving the surrounding environment information provided by the perception module, and provides accurate and timely trajectory information of other vehicles for the planning module, so as to help the planning module better formulate appropriate action strategies. Therefore, accurately predicting the trajectories of surrounding vehicles is crucial for the safety of autonomous driving. The goal of the trajectory prediction task is to predict multiple possible future two-dimensional coordinate trajectories given the historical information of the target vehicle, the high-definition map information within a certain observation radius, and the interaction information with surrounding vehicles.
[0003] With the development and popularization of deep learning technology, significant progress has been made in trajectory prediction methods based on deep learning. Typical methods include those based on recurrent neural networks (RNNs), convolutional neural networks (CNNs), and attention mechanisms. RNNs are used to model sequential data, such as vehicle trajectory data. By iteratively processing time series data, RNNs can capture the temporal information in the trajectory data, thus realizing the prediction of future trajectories. CNNs are mainly used to extract spatial features, such as the road and obstacle conditions around the vehicle. It can effectively capture local spatial information and maintain the translational invariance of spatial relationships. The attention mechanism is widely used to fuse the historical trajectories of the target vehicle and the interaction effects of surrounding other vehicles, helping the model to focus on the important aspects of the scene during the prediction process and improve the prediction accuracy. In addition, some methods based on graph neural networks (GNNs) have gradually received attention, which can better capture the spatial relationships and interaction information between vehicles, improving the accuracy and robustness of trajectory prediction.
[0004] Recurrent Neural Networks (RNNs) play an important role in sequence modeling and prediction tasks, and their development has gone through multiple stages and improvements: The initial RNNs were proposed by Elman et al. They have a recursive structure and can capture temporal dependencies in sequence data. However, basic RNNs are prone to the problems of vanishing or exploding gradients when dealing with long sequences, making it difficult to capture long-term dependencies. To address the problems of vanishing or exploding gradients, Hochreiter and Schmidhuber proposed the Long Short-Term Memory (LSTM) network, introducing a gating mechanism to control the flow of information. LSTM effectively captures long-term dependencies through three gates (input gate, forget gate, and output gate) and a memory cell, becoming an important tool for processing sequence data. Cho et al. proposed the Gate Recurrent Unit (GRU), which is a simplified gated recurrent unit with a gating mechanism similar to LSTM but with fewer parameters and a simpler structure. Compared with LSTM, GRU is faster in training speed, has fewer parameters, and performs better in some tasks. Bi-GRU is a bidirectional recurrent neural network structure that combines forward and backward GRU layers and can consider both the historical information and future information of the input sequence. This bidirectional structure enables Bi-GRU to better capture the context information in sequence data, thereby improving the modeling ability and prediction accuracy for sequence data. Here, the idea of Bi-Gru is adopted as the main module of the long-short-term prediction decoder. In order to be able to consider both the historical information and future information of the sequence, thus more comprehensively understand and model the dependencies between data, the generalization ability and prediction accuracy of the model are improved.
[0005] The long-term prediction problem refers to the task of predicting vehicle trajectories in the long term in autonomous driving trajectory prediction, usually exceeding several seconds or even dozens of seconds. This problem is challenging because as the prediction time increases, the uncertainty of the prediction results also increases. In the long term, vehicles may encounter various complex traffic conditions and road condition changes, such as intersections, vehicle lane changes, traffic signal changes at intersections, etc., which makes it more difficult to accurately predict the future trajectories of vehicles. Even if the encoding of the scene is rich, existing decoding strategies are difficult to capture the inherent multimodality in the future behavior of the agent, especially when the prediction range is very long.
[0006] Solving the long-term prediction problem requires the model to have the ability to effectively model complex traffic scenarios and various behavior patterns. In addition, the model also needs to be able to effectively process trajectory information over a long period of time and make accurate predictions in a changing environment. Therefore, solving the long-term prediction problem involves comprehensive consideration and optimization of aspects such as model structure, data representation, and training strategies. Inspired by human hierarchical decision-making, where the motion intention determines the specific trajectory, most existing methods adopt prior methods based on target points. Conditional prediction based on the target first predicts or predefined target candidates, and then predicts the trajectory under its conditions. This approach has been proven to be effective and has been widely adopted in state-of-the-art methods. Multipath predefines a set of anchor trajectories through clustering and predicts the trajectory prediction offset. TNT predicts the target point offset from the center line of the lane. Goal-Net uses lane segments as trajectory anchors and predicts the lane that the vehicle will pass through in the future. FRM predicts the occupancy of each road punctuation on each lane and then predicts the fine-grained trajectory. GANet proposes a framework based on the target area for predicting the target area and fusing key long-range map features. MTR predefines a set of target points as queries by clustering the trajectory data of each agent type and uses an attention layer to aggregate context information based on these queries. Prophnet uses trajectory proposals as learnable anchors, enabling the attention layer to encode the target-oriented scene context. However, although these methods utilize target-based conditional prediction, they only utilize the target-based context once.
[0007] Scene context fusion provides rich information for trajectory prediction. Early research rasterized the world state into a multi-channel image and used classical convolutional neural networks for learning. Due to the problems of information loss, limited receptive field, and high cost in the rasterization method, the research community has turned to vector-based encoding schemes. By using permutation-invariant set operators such as pooling, graph convolution, and attention mechanisms, vector-based methods can effectively aggregate sparse information in traffic scenarios. For the vast majority of methods, only the historical environment information of a fixed length before the current time step is fused, and the environmental information that the vehicle may travel in the short term in the future is not fused. However, as the vehicle moves, the surrounding scene will also change dynamically. If only the original historical scene information is statically utilized, some scene information observed during the vehicle's dynamic driving will be lost. Although previous research has explored the interaction between vehicles and the environment, the interaction between the vehicle and the map at future moments and the interaction between future vehicles have received less attention. This results in the loss of future information and difficulty in handling some sudden dangerous situations, leading to a problem of reduced prediction performance.
[0008] Although existing trajectory prediction models can achieve good prediction accuracy, most of them regard trajectory prediction as a static task. By using historical frames of fixed length to predict future trajectories, they predict all trajectory points at future time steps at once. This static paradigm of trajectory prediction may lead to instability and temporal inconsistency in consecutive predictions, which is not conducive to the autonomous driving system making safe and reliable decisions. Moreover, as the prediction time increases, the prediction results deteriorate. GANet and mmTransformer use the final point of the predicted trajectory as a future prior. Such a long-term target is difficult to capture the short-term dynamics of the trajectory, and the span is too large when generating the trajectory, which will also cause the loss of future information. There is also a method that clusters multiple representative trajectories at the dataset level, and then selects a trajectory as a coarse-grained prior during prediction and then calculates the bias for optimization. This method is highly dependent on the selection of the prior trajectory. If the selection is not good, the effect will be very poor, and it is difficult to have good generalization ability for variable driving scenarios. To achieve more stable prediction, DCMS proposes to model trajectory prediction as a dynamic problem. It explicitly considers the correlation between consecutive predictions and imposes a temporal consistency constraint, requiring the overlapping parts of the predicted trajectories at adjacent time steps to be the same. In addition, QCNet uses a two-stage decoding strategy to output future trajectories through iterative loops in the proposed trajectory generation module. However, this method also does not introduce the scene information at future moments, lacks the guidance of long-term goals, and due to the high complexity, it cannot achieve good real-time performance. These preliminary studies demonstrate the effectiveness and rationality of modeling trajectory prediction as a dynamic task. Summary of the Invention
[0009] To solve the technical problem of the cumulative long-term prediction error in autonomous driving trajectory prediction in the prior art, an embodiment of the present invention provides an autonomous driving trajectory prediction method and device based on long-term and short-term prediction. The technical solution is as follows:
[0010] On the one hand, an autonomous driving trajectory prediction method based on long-term and short-term prediction is provided. This method is implemented by an autonomous driving trajectory prediction device, and the method includes:
[0011] S1. Obtain the information data of the target vehicle to be predicted for the trajectory.
[0012] S2. Input the information data into the constructed long-term and short-term recurrent prediction decoder.
[0013] Among them, the long-term and short-term recurrent prediction decoder includes a long-term target prediction Gru, a short-term trajectory prediction Gru, and a future trajectory prediction Gru.
[0014] S3. Based on the information data, the long-term target prediction Gru, the short-term trajectory prediction Gru, and the future trajectory prediction Gru, obtain the automatic driving trajectory prediction result.
[0015] Optionally, in S3, based on the information data, the long-term target prediction Gru, the short-term trajectory prediction Gru, and the future trajectory prediction Gru, obtaining the automatic driving trajectory prediction result includes:
[0016] S31. Input the information data into the long-term target prediction Gru to obtain the long-term target features corresponding to each future time step.
[0017] S32. Input the hidden features of the future trajectory prediction Gru into the short-term trajectory prediction Gru to obtain the future short-term target features.
[0018] S33. Input the information data, the long-term target features, and the future short-term target features into the future trajectory prediction Gru to obtain the hidden features of the future trajectory prediction Gru, and then obtain the automatic driving trajectory prediction result.
[0019] Optionally, in S31, inputting the information data into the long-term target prediction Gru to obtain the long-term target features corresponding to each future time step includes:
[0020] S311. Obtain the scene data and trajectory data in the information data.
[0021] S312. According to the scene data and the trajectory data, divide the scene into multiple regions through the K-means clustering model, and each region in the multiple regions corresponds to multiple trajectories.
[0022] S313. Extract the features of the multiple trajectories corresponding to each region respectively and perform feature aggregation to obtain the trajectory prototype of each region, and obtain the initial hidden features of the long-term target prediction Gru according to the trajectory prototype.
[0023] S314. Use a learnable feature encoding initialized by a normal distribution as the input of the long-term target prediction Gru. According to the input of the long-term target prediction Gru, the initial hidden features of the long-term target prediction Gru, and the long-term target prediction Gru, obtain the long-term target features corresponding to each future time step.
[0024] Optionally, in S32, inputting the hidden features of the future trajectory prediction Gru into the short-term trajectory prediction Gru to obtain the future short-term target features includes:
[0025] S321. Obtain the hidden features of the future trajectory prediction Gru at time step as the input of the short-term trajectory prediction Gru at time step ; use an empty vector with a value of 0 as the time step The initial hidden features of the short-term trajectory prediction Gru, according to the time step The input of the short-term trajectory prediction Gru, the time step The initial hidden features of the short-term trajectory prediction Gru and the short-term trajectory prediction Gru, to obtain the time step The hidden state of the short-term trajectory prediction Gru, according to the time step The hidden state of the short-term trajectory prediction Gru is predicted to obtain the time step The short-term predicted trajectory points of the time step The short-term predicted trajectory points of the time step are transformed to obtain the time step The feature encoding of the short-term trajectory prediction Gru, and use the feature encoding as the time step The hidden features are passed on.
[0026] S322. Obtain the hidden features of the short-term trajectory prediction Gru at the time step, and obtain the time step The hidden features of the short-term trajectory prediction Gru, and obtain the time step The hidden features of the short-term trajectory prediction Gru.
[0027] S323. Add the hidden state of the short-term trajectory prediction Gru at the time step, the hidden features of the short-term trajectory prediction Gru at the time step, and the hidden features of the short-term trajectory prediction Gru at the time step to obtain the future short-term features. The hidden state of the short-term trajectory prediction Gru at the time step, the hidden features of the short-term trajectory prediction Gru at the time step, and the hidden features of the short-term trajectory prediction Gru at the time step The hidden features of the short-term trajectory prediction Gru at the time step, and the time step The hidden features of the short-term trajectory prediction Gru at the time step are added to obtain the future short-term features.
[0028] Optionally, in S33, input the information data, the long-term target features, and the future short-term target features into the future trajectory prediction Gru to obtain the hidden features of the future trajectory prediction Gru, and further obtain the autonomous driving trajectory prediction result, including:
[0029] S331. Use a learnable feature encoding initialized by a normal distribution as the input of the future trajectory prediction Gru, and obtain the feature encoding of the dynamic scene information at the location of the current target vehicle according to the information data and the adaptive scene information encoder.
[0030] S332. Use the cross-attention module to fuse the input of the future trajectory prediction Gru and the feature encoding of the dynamic scene information at the location of the current target vehicle to obtain the updated feature encoding of the dynamic scene information at the location of the current target vehicle.
[0031] S333. Concatenate the updated feature encoding of the dynamic scene information at the location of the current target vehicle, the long-term target features, and the future short-term target features to obtain the fused input of the future trajectory prediction Gru.
[0032] S334. Based on the input of Gru predicted by the fused future trajectory and the future trajectory prediction Gru, obtain the hidden features of the trajectory prediction Gru at each future time step.
[0033] S335. Predict the trajectory points at each future time step based on the hidden features of the trajectory prediction Gru at each future time step, and obtain the automatic driving trajectory prediction result according to the trajectory points at each future time step.
[0034] Optionally, obtaining the feature encoding of the dynamic scene information at the location of the current target vehicle according to the information data and the adaptive scene information encoder includes:
[0035] The trajectory points at each future time step predicted by the future trajectory prediction Gru and the vehicle speeds at each time step of the target vehicle are used to set an observation radius according to the vehicle speed, and the feature encoding of the dynamic scene information at the location of the current target vehicle is obtained according to the observation radius and the information data.
[0036] Optionally, the loss of the long short-term recurrent prediction decoder is shown by the following formula (1):
[0037] (1)
[0038] In the formula, represents the loss of the long short-term recurrent prediction decoder, represents the total number of time steps, represents the number of vehicles, represents the short-term prediction task loss, represents the future trajectory prediction task loss, represents the multi-modal trajectory confidence classification task loss.
[0039] On the other hand, an automatic driving trajectory prediction device based on long short-term prediction is provided. The device is applied to the automatic driving trajectory prediction method based on long short-term prediction. The device includes:
[0040] An acquisition module for acquiring information data of a target vehicle to be predicted for its trajectory.
[0041] An input module for inputting the information data into the constructed long short-term recurrent prediction decoder.
[0042] Among them, the long short-term recurrent prediction decoder includes a long-term target prediction Gru, a short-term trajectory prediction Gru, and a future trajectory prediction Gru.
[0043] An output module for obtaining the automatic driving trajectory prediction result according to the information data, the long-term target prediction Gru, the short-term trajectory prediction Gru, and the future trajectory prediction Gru.
[0044] On the other hand, there is provided an automatic driving trajectory prediction device, which includes: a processor; a memory storing computer-readable instructions, and when the computer-readable instructions are executed by the processor, any one of the methods in the above-mentioned automatic driving trajectory prediction method based on long short-term prediction is implemented.
[0045] On the other hand, there is provided a computer-readable storage medium storing at least one instruction, and the at least one instruction is loaded and executed by a processor to implement any one of the methods in the above-mentioned automatic driving trajectory prediction method based on long short-term prediction.
[0046] The beneficial effects brought by the technical solutions provided in the embodiments of the present invention at least include:
[0047] In the embodiments of the present invention, a long short-term feedback method is proposed: by predicting future long-term information and future short-term information to assist in predicting the future trajectory at the current moment, dynamically fusing the information of future moments, and adopting a cyclic step-by-step prediction method to effectively solve the problem of long-term prediction error accumulation.
[0048] An adaptive future scenario information fusion encoder is proposed: by predicting the vehicle speed of the current vehicle to set the size of the observation radius, entities with different motion modes will benefit from multiple different granularity scenario information.
[0049] Experimental results show that the model of the present invention has achieved excellent results on the Argoverse1 and Argoverse2 data sets, and there is a great improvement compared with the baseline model. BRIEF DESCRIPTION OF THE DRAWINGS
[0050] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0051] Figure 1 is a flowchart of an automatic driving trajectory prediction method based on long short-term prediction provided by an embodiment of the present invention;
[0052] Figure 2 is a flowchart of long short-term feedback provided by an embodiment of the present invention;
[0053] Figure 3 is a model flowchart provided by an embodiment of the present invention;
[0054] Figure 4 is an unfolded view of a short-term prediction module provided by an embodiment of the present invention
[0055] Figure 5 It is the flow chart of regional feature prototype extraction provided by the embodiments of the present invention;
[0056] Figure 6 It is the model flow chart provided by the embodiments of the present invention;
[0057] Figure 7 It is the visualization comparison chart provided by the embodiments of the present invention;
[0058] Figure 8 It is the block diagram of an automatic driving trajectory prediction device based on long - short - term prediction provided by the embodiments of the present invention;
[0059] Figure 9 It is the structural schematic diagram of an automatic driving trajectory prediction device provided by the embodiments of the present invention. Detailed implementation manners
[0060] Next, the technical solutions in the present invention will be described with reference to the accompanying drawings.
[0061] In the embodiments of the present invention, words such as "exemplarily" and "for example" are used to represent examples, illustrations or explanations. Any embodiment or design solution described as an "example" in the present invention should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Exactly speaking, the use of the word "example" is intended to present concepts in a specific way. In addition, in the embodiments of the present invention, the meaning expressed by "and / or" can be both, or either one of the two.
[0062] In the embodiments of the present invention, "image" and "picture" can sometimes be used interchangeably. It should be noted that when not emphasizing their differences, the meanings they express are the same. "(of)", "corresponding", and "corresponding" can sometimes be used interchangeably. It should be noted that when not emphasizing their differences, the meanings they express are the same.
[0063] In the embodiments of the present invention, sometimes subscripts such as W1 may be written in a non - subscript form such as W1. When not emphasizing their differences, the meanings they express are the same.
[0064] To make the technical problems, technical solutions and advantages to be solved by the present invention clearer, the following will be described in detail with reference to the accompanying drawings and specific embodiments.
[0065] The embodiments of the present invention provide an automatic driving trajectory prediction method based on long - short - term prediction. This method can be implemented by an automatic driving trajectory prediction device, and the automatic driving trajectory prediction device can be a terminal or a server. As Figure 1Flowchart of an autonomous driving trajectory prediction method based on long - short - term prediction. The processing flow of this method may include the following steps:
[0066] S1. Obtain the information data of the target vehicle for which trajectory prediction is to be performed.
[0067] In a feasible implementation, the information data may include the historical information of the target vehicle, scene data, interaction information of surrounding vehicles, etc.
[0068] S2. Input the information data into the constructed long - short - term recurrent prediction decoder.
[0069] Among them, the long - short - term recurrent prediction decoder includes a long - term target prediction Gru, a short - term trajectory prediction Gru, and a future trajectory prediction Gru.
[0070] In a feasible implementation, in the prior art, it is difficult to capture the dynamic changes in the scene with a one - time prediction. During the driving process of the present invention, the information of the surrounding scene and the vehicle are both dynamically changing. Due to the variability and complexity of the driving environment, it is very difficult to estimate and capture the dynamic changes at future moments only based on the current moment and historical information. Therefore, the dynamic changes in the scene may have an important impact on trajectory prediction, such as Figure 2As shown in the figure, the upper part is the strategy for predicting the trajectory of the existing model, and the lower part is the long-short-term cyclic prediction strategy proposed by the present invention. For example, if the goal of the present invention is to turn left, but pedestrians will be encountered during driving. If the future information is not predicted in the short term at this time, a very serious traffic accident will occur. Based on this problem, the present invention designs a long-short-term cyclic prediction decoder, which mainly includes three core modules: the long-term target prediction Gru is responsible for predicting the long-term target, the short-term trajectory prediction Gru is responsible for predicting the short-term target, and the future trajectory prediction Gru is responsible for outputting all future trajectory points. If only the long-term target, that is, the final position in the future, is considered and the short-term target is ignored, the middle process will be missing, and it is difficult to capture the short-term dynamic changes of the trajectory. Similarly, the future information will be missing when generating the trajectory. If only the short-term dynamic changes in the future are considered and the long-term target is ignored, the purpose will be lost and the goal-driven will be lacking. Therefore, the present invention takes the long-short-term target prediction as the future feature and gradually backtracks it, increasing the dynamic prediction of future information, so as to avoid the error accumulation problem of predicting all future time-long trajectories at one time. By first predicting a long-term target code that does not change with time, using the long-term target and the scene features at the current moment as the guidance for short-term prediction, the trajectory of short-term driving in the future is predicted. Then, these short-term targets that change with time are converted into feature codes and backtracked into the input of the trajectory prediction decoding module at the current moment to decode the trajectory points at the current moment. Then, they are spliced with the long-term target code and jointly used as the input for short-term prediction at the next time step to obtain the short-term target at the next moment, so as to perform cyclic decoding. And this exactly conforms to the driving scenario in real life. There is always a final destination when driving in the present invention, which does not change with time, while the short-term driving route may change at any time. Therefore, the present invention combines the two to jointly assist the prediction of the present invention.
[0071] Further, in order to make the model of the present invention easier to be inserted into the existing method, the present invention takes the long-short-term prediction module as a plug-and-play decoder module, and the model process is as Figure 3 、 Figure 4 shown. (1) First, the long-term prediction Gru takes a learnable embedding as the input and uses the trajectory prototype features in the area where the final point of the trajectory is located as to obtain the long-term target features corresponding to each time step; (2) The short-term prediction Gru takes the hidden features output by the trajectory prediction Gru at the current moment as the input, thereby predicting the trajectory points in the short-term future time, and then converts them into features and backtracks them into the trajectory prediction Gru at the next moment; (3) The input of the future trajectory prediction Gru is the learnable embedding that is the same as the long-term prediction, spliced with the output of the long-term prediction at the corresponding moment and the output of the short-term prediction at the previous moment, The scene information obtained by the corresponding encoder; (4) Finally, an attention layer is used to selectively fuse the hidden features output by the trajectory prediction Gru at all moments; thereby obtaining the final multi-modal trajectory coordinates at all time steps and the probabilities of each modality.
[0072] S3. According to the information data, the long-term target prediction Gru, the short-term trajectory prediction Gru, and the future trajectory prediction Gru, obtain the autonomous driving trajectory prediction result.
[0073] Optionally, the above step S3 may include the following steps S31 - S33:
[0074] S31. Input the information data into the long-term target prediction Gru to obtain the long-term target features corresponding to each future time step.
[0075] Optionally, the above step S31 may include the following steps S311 - S314:
[0076] S311. Obtain the scene data and trajectory data in the information data.
[0077] S312. According to the scene data and the trajectory data, divide the scene into multiple regions through a K-means clustering model, and each region in the multiple regions corresponds to multiple trajectories.
[0078] S313. Extract the features of the multiple trajectories corresponding to each region and perform feature aggregation to obtain the trajectory prototype of each region, and obtain the initial hidden feature of the long-term target prediction Gru according to the trajectory prototype.
[0079] S314. Use a learnable feature encoding initialized by a normal distribution as the input of the long-term target prediction Gru, and according to the input of the long-term target prediction Gru, the initial hidden feature of the long-term target prediction Gru, and the long-term target prediction Gru, obtain the long-term target features corresponding to each future time step.
[0080] In a feasible implementation manner, during the driving process, at the beginning, the present invention has a final destination position, which is used as a drive to guide the driving target at the current moment, thereby avoiding the lack of a target and causing the trajectory points to be too scattered. Therefore, the present invention adopts a method similar to that of mmTransformer to divide the scene into regions, and then extracts a feature prototype of a real trajectory in each region as the long-term target.
[0081] Specifically, first, at the dataset level, collect the endpoints of all real trajectories, and then divide the scene into N regions through the K-means clustering model, so that each trajectory corresponds to a region. After that, extract and aggregate the features of all trajectories in the region by means of LSTM and 1DCNN according to the region to obtain the trajectory prototypes of each region. As Figure 5 shown, in the long-term prediction stage, the present invention uses a learnable feature encoding as the input of the Gru, and predicts the distribution of the region where the final point of the future trajectory is located by inputting the output of the encoder into a region prediction module represented by an MLP. Multiply this distribution by the feature prototypes of each region as the initial hidden feature long term region goal of the long-term Gru for transmission, so as to obtain the long-term goals of each future time step, and it will not change with time, so as to provide a goal orientation for short-term prediction and future trajectory prediction later; among them, the encoder is the baseline model HiVT, and the output of the encoder is the fused scene information encoding.
[0082] Furthermore, the present invention sets the region prediction module as an offline model and uses the cross-entropy loss function to calculate the loss of this classification task for optimization. The present invention takes each region as a category, obtains the output distribution probability by predicting the region where the final point of the trajectory is located, and selects the region with the largest probability as the prediction label and calculates the loss with the correct label. Set it as untrainable during inference, so as to reduce the learning difficulty of the model and accelerate convergence.
[0083] In a feasible implementation manner, the long-term goal prediction Gru, the short-term trajectory prediction Gru, and the future trajectory prediction Gru obtain the output h of each time step according to an input and a hidden feature; specifically, at the initial time step , according to the input at the time step and the set hidden feature to obtain the output at the time step ; at the time step , take the output as the hidden feature at the time step, and according to the input at the time step and the hidden feature at the time step to obtain the output at the time step, and so on.
[0084] S32. Input the hidden feature of the future trajectory prediction Gru into the short-term trajectory prediction Gru to obtain the future short-term target feature.
[0085] Optionally, the above step S32 may include the following steps S321 - S323:
[0086] S321. Obtain the time step The hidden features of the future trajectory prediction Gru as the time step The input of the short-term trajectory prediction Gru for the time step; use the empty vector with a value of 0 as the initial hidden feature of the short-term trajectory prediction Gru for the time step According to the input of the short-term trajectory prediction Gru for the time step, the initial hidden feature of the short-term trajectory prediction Gru for the time step And the short-term trajectory prediction Gru for the time step To obtain the hidden state of the short-term trajectory prediction Gru for the time step According to the hidden state of the short-term trajectory prediction Gru for the time step Predict the short-term predicted trajectory points for the time step Perform transformation on the short-term predicted trajectory points for the time step To obtain the feature encoding of the short-term trajectory prediction Gru for the time step And use the feature encoding as the hidden feature for the time step For transmission.
[0087] S322. Obtain the hidden features of the short-term trajectory prediction Gru for the time step And obtain the hidden features of the short-term trajectory prediction Gru for the time step Of the short-term trajectory prediction Gru.
[0088] S323. Add the hidden state of the short-term trajectory prediction Gru for the time step The hidden features of the short-term trajectory prediction Gru for the time step And the hidden features of the short-term trajectory prediction Gru for the time step To obtain the future short-term features.
[0089] In a feasible implementation, statically fusing the historical scene features with a fixed length will result in the lack of future dynamic features. Short-term prediction can expand the visible window length instead of fixing the visible window to the historical time step length. As the vehicle moves, the surrounding scene will also change. Generally, the method is to set an observation radius to observe the surrounding environment. However, when the vehicle moves forward, the environment changes, and the predictions made will also be different. For example, if the vehicle is currently driving on a straight road but there is an intersection ahead, and if the prediction is still made according to the current straight road scene, it will be difficult to predict accurately, and directly predicting the complete sequence will inevitably lead to a decrease in smoothness or noise accumulation. Therefore, the present invention proposes that the short-term prediction branch enhances the smoothness and rationality of the predicted trajectory by adding the prediction of the short-term future trajectory and the feature feedback during the trajectory loop decoding.
[0090] Such as Figure 5As shown, the short-term prediction Gru takes the hidden features output by the future trajectory prediction Gru at the previous moment as input, thereby outputting the hidden state at this moment, and predicting the short-term prediction trajectory point Trajs1 based on this. Then, this trajectory point is converted into a feature encoding as the hidden feature for the next moment, while the output hidden feature of the previous moment is passed on as input.
[0091] As Figure 6 shown, below is the short-term prediction Gru. The output of the middle future trajectory Gru at the moment is input into the short-term prediction Gru, and then the trajectory points for the next 3 time steps are predicted backward. The hidden features obtained from these 3 time steps of the short Gru are added and fused, and then passed back to the future trajectory Gru at the moment as input.
[0092] Specifically, the short-term prediction Gru requires an input, which is the output of the future trajectory prediction Gru at the moment . The short-term prediction Gru also requires an initial hidden feature, which can be set as an empty vector of all 0s. In this way, an output is obtained. Through , a moment trajectory point is predicted, and this trajectory point is called . Then, at the next time step , the short-term prediction Gru also requires an input and a hidden feature. In the present invention, is converted into a feature encoding as the hidden feature, and the input is represented by . At the moment, an output is also obtained. Through , the trajectory point at is predicted. By analogy, the trajectory points at three moments are obtained, as well as the hidden features at three moments . The three hidden features are added, and the result is passed back to the future trajectory prediction Gru at the moment, spliced into the input of the future trajectory prediction Gru, enriching the future short-term context information, thereby effectively capturing and foreseeing some future behaviors.
[0093] And in the training stage, the present invention calculates the loss between the multi-modal future trajectory points obtained by short-term prediction and the real trajectory points. Adopting the winner-takes-all training strategy, among the multi-modal prediction trajectories, the trajectory with the smallest sum of the Euclidean distances between the final point and the real trajectory is selected as the optimal trajectory to calculate the loss, as shown in the following formula:
[0094] (1)
[0095] where k represents the number of modes of the trajectory, represents all the trajectory points of the future F time steps obtained by prediction, represents all the true trajectory points of the future F time steps, and t + F represents the final moment of future prediction.
[0096] S33. Input the information data, long-term target features, and future short-term target features into the future trajectory prediction Gru to obtain the hidden features of the future trajectory prediction Gru, and then obtain the automatic driving trajectory prediction result.
[0097] Optionally, the above step S33 may include the following steps S331 - S335:
[0098] S331. Use a scientific feature encoding initialized by a normal distribution as the input of the future trajectory prediction Gru, and obtain the feature encoding of the dynamic scene information at the location of the current target vehicle according to the information data and the adaptive scene information encoder.
[0099] S332. Use the cross-attention module to fuse the input of the future trajectory prediction Gru and the feature encoding of the dynamic scene information at the location of the current target vehicle to obtain the updated feature encoding of the dynamic scene information at the location of the current target vehicle.
[0100] S333. Concatenate the updated feature encoding of the dynamic scene information at the location of the current target vehicle, the long-term target features, and the future short-term target features to obtain the fused input of the future trajectory prediction Gru.
[0101] S334. According to the fused input of the future trajectory prediction Gru and the future trajectory prediction Gru, obtain the hidden features of the trajectory prediction Gru at each future time step.
[0102] S335. Predict the trajectory points at each future time step according to the hidden features of the trajectory prediction Gru at each future time step, and obtain the automatic driving trajectory prediction result according to the trajectory points at each future time step.
[0103] In a feasible implementation manner, in the final future trajectory prediction decoding module, the present invention combines long- and short-term features and the scene information at the location of the future moment to generate the trajectory. The traditional recurrent neural network unit only models the previous time dependence by accumulating the historical evidence of the input sequence, while by introducing the long- and short-term prediction module of the present invention, the time correlation between the current and future actions can also be utilized. By predicting the upcoming actions and explicitly using these estimated values to help identify the current action.
[0104] Specifically, the present invention still uses the scientifically feature-encoded long-term prediction input initialized by a normal distribution as the input of this module. However, here the present invention obtains the feature encoding of the dynamic scene information at the current vehicle position as the vehicle travels through an adaptive scene information encoder, and then fuses it with the input of the present invention through a cross-attention module, so as to further update the scene information at the current position. Then, the long-term features and short-term features at this moment are concatenated as the fused input. The selects the output of the model encoder for transmission, so that the hidden feature at this moment can be obtained, and it is used as the input of the short-term prediction Gru at the next moment for predicting future short-term targets, and continues to loop and back-reference until the hidden features of all future time steps are predicted. Finally, the future multi-modal trajectory and the corresponding probability are output through the mlp; among them, the encoder is the baseline model HiVT, and the output of the encoder is the fused scene information encoding.
[0105] Optionally, obtaining the feature encoding of the dynamic scene information at the current target vehicle position according to the information data and the adaptive scene information encoder in S331 includes:
[0106] The trajectory points at each future time step predicted by the future trajectory prediction Gru are used to obtain the vehicle speed at each time step of the target vehicle. An observation radius is set according to the vehicle speed. According to the observation radius and the information data, the feature encoding of the dynamic scene information at the current target vehicle position is obtained.
[0107] In a feasible implementation, as the vehicle travels, the surrounding scene will also change dynamically. Therefore, the present invention introduces a scene information fusion encoder here to update and fuse the scene information around the current vehicle position in real time. And considering that entities with different motion modes will benefit from multiple different granularities of scene information, the present invention also predicts the driving speed of the current vehicle to set different observation radii according to the speed to fuse the surrounding scene information, which also conforms to the driving scenario in real life. When the vehicle speed is very fast, it is necessary to observe the scene information at a relatively long distance, and when it is slower, the attention should be focused on the surrounding environment.
[0108] Specifically, the present invention uses the trajectory point at the current moment predicted by the future trajectory prediction module as the coordinate of the current position, sets the observation radius according to the predicted vehicle speed, and then converts the surrounding scene into a vehicle-centered representation method to obtain the features of all objects within the observation radius and fuse them with the input at the current moment.
[0109] In the task of the present invention, the Huber loss is adopted to calculate the loss of the regression task, and the regression task includes two sub-tasks, namely the short-term prediction task reg1 and the future trajectory prediction task reg2. The formulas are as shown below:
[0110] (2)
[0111] (3)
[0112] In the formula, k is the number of multi-modal trajectories in each region, n is the number of vehicles, t is the time step, and G represents the correct trajectory coordinates.
[0113] For predicting the probability of each modal trajectory, which is expressed as a classification task, the loss is calculated through the cross-entropy loss function. The formula is as shown below:
[0114] (4)
[0115] The loss formula of the entire model can be expressed as:
[0116] (5)
[0117] In the formula, represents the loss of the long-short-term cyclic prediction decoder, represents the total number of time steps, represents the number of vehicles, represents the loss of the short-term prediction task, represents the loss of the future trajectory prediction task.
[0118] To demonstrate the effectiveness of the proposed method, the present invention verifies the effectiveness of the method on two datasets, Argoverse1 and Argoverse2 (denoted as ours in the table). Table 1 shows the results of comparing with the Hivt as the baseline model and after adding the method of the present invention. The method of the present invention performs better in various metrics. Table 2 shows the results of experiments on the Argoverse2 dataset with the QCNet as the baseline model. It can be seen that the method of the present invention has improved in most metrics, and the complexity of the model has decreased significantly, effectively improving the real-time performance of model prediction and achieving better prediction results.
[0119] Table 1
[0120]
[0121] Table 2
[0122]
[0123] As shown in Table 3, the present invention conducted ablation experiments on each module of the method of the present invention using Hivt as the baseline model on the Argoverse1 dataset to better verify the effectiveness of each module, where ST represents the short-term prediction module, LT represents the long-term prediction module, and ASE represents the adaptive scenario information fusion module. It can be seen from the results that each module plays a certain positive role.
[0124] Table 3
[0125]
[0126] Figure 7 It shows a comparison chart of the trajectory prediction visualization results of the method of the present invention with Hivt as the baseline model. Among them, the yellow represents the true trajectory, the pink represents the predicted trajectory. The first row is the prediction result of the baseline model without adding the module of the present invention, and the second row is the result after adding it. It can be seen that the trajectory predicted by the method of the present invention is significantly better than the baseline model.
[0127] In the embodiment of the present invention, a long and short-term feedback method is proposed: by predicting future long-term information and future short-term information to assist the prediction of the future trajectory at the current moment, dynamically fusing the information of future moments, and adopting a cyclic step-by-step prediction method to effectively solve the problem of long-term prediction error accumulation.
[0128] An adaptive future scenario information fusion encoder is proposed: by predicting the vehicle speed of the current vehicle to set the size of the observation radius, so that entities with different motion patterns will benefit from multiple different granularity scenario information.
[0129] The experimental results show that the model of the present invention has achieved excellent results on the Argoverse1 and Argoverse2 datasets, and has been greatly improved compared with the baseline model.
[0130] Figure 8 It is a block diagram of an automatic driving trajectory prediction device based on long and short-term prediction shown according to an exemplary embodiment. This device is used for the automatic driving trajectory prediction method based on long and short-term prediction. Refer to Figure 8 , this device includes an acquisition module 310, an input module 320, and an output module 330. Among them:
[0131] The acquisition module 310 is used to acquire the information data of the target vehicle to be predicted for the trajectory.
[0132] The input module 320 is used to input the information data into the constructed long and short-term cyclic prediction decoder.
[0133] Among them, the long and short-term cyclic prediction decoder includes a long-term target prediction Gru, a short-term trajectory prediction Gru, and a future trajectory prediction Gru.
[0134] An output module 330 is configured to obtain an autonomous driving trajectory prediction result based on information data, a long-term target prediction Gru, a short-term trajectory prediction Gru, and a future trajectory prediction Gru.
[0135] In an embodiment of the present invention, a long- and short-term feedback method is proposed: by predicting future long-term information and future short-term information to assist in predicting the future trajectory at the current moment, dynamically fusing the information at future moments, and adopting a cyclic step-by-step prediction method to effectively solve the problem of long-term prediction error accumulation.
[0136] An adaptive future scenario information fusion encoder is proposed: by predicting the vehicle speed of the current vehicle to set the size of the observation radius, entities in different motion modes will benefit from multiple scene information with different granularities.
[0137] Experimental results show that the model of the present invention has achieved excellent results on the Argoverse1 and Argoverse2 datasets, with a significant improvement compared to the baseline model.
[0138] Figure 9 is a schematic structural diagram of an autonomous driving trajectory prediction device provided by an embodiment of the present invention, as Figure 9 shown, the autonomous driving trajectory prediction device may include the above-mentioned Figure 8 autonomous driving trajectory prediction device based on long- and short-term prediction. Optionally, the autonomous driving trajectory prediction device 410 may include a first processor 2001.
[0139] Optionally, the autonomous driving trajectory prediction device 410 may further include a memory 2002 and a transceiver 2003.
[0140] Wherein, the first processor 2001 is connected to the memory 2002 and the transceiver 2003, such as through a communication bus.
[0141] Next, in conjunction with Figure 9 each component of the autonomous driving trajectory prediction device 410 will be specifically introduced:
[0142] Among them, the first processor 2001 is the control center of the autonomous driving trajectory prediction device 410, which can be a single processor or a collective term for multiple processing elements. For example, the first processor 2001 is one or more central processing units (CPUs), or can be an application specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present invention, such as: one or more digital signal processors (DSPs), or one or more field programmable gate arrays (FPGAs).
[0143] Optionally, the first processor 2001 can execute various functions of the autonomous driving trajectory prediction device 410 by running or executing software programs stored in the memory 2002 and calling data stored in the memory 2002.
[0144] In a specific implementation, as an embodiment, the first processor 2001 can include one or more CPUs, such as Figure 9 CPU0 and CPU1 shown in
[0145] In a specific implementation, as an embodiment, the autonomous driving trajectory prediction device 410 can also include multiple processors, such as Figure 9 the first processor 2001 and the second processor 2004 shown in
[0146] Each of these processors can be a single-core processor (single-CPU) or a multi-core processor (multi-CPU). Here, the processor can refer to one or more devices, circuits, and / or processing cores for processing data (such as computer program instructions).
[0147] Optionally, the memory 2002 can be a read-only memory (ROM) or other type of static storage device that can store static information and instructions, a random access memory (RAM) or other type of dynamic storage device that can store information and instructions, or can also be an electrically erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM), or other optical disc storage, optical disc storage (including compact discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), magnetic disk storage media, or other magnetic storage devices, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto. The memory 2002 can be integrated with the first processor 2001 or can exist independently and be coupled to the first processor 2001 through an interface circuit ( Figure 9 not shown) of the autonomous driving trajectory prediction device 410. The embodiments of the present invention do not make specific limitations in this regard.
[0148] The transceiver 2003 is used to communicate with a network device or with a terminal device.
[0149] Optionally, the transceiver 2003 can include a receiver and a transmitter ( Figure 9 not shown separately). Among them, the receiver is used to implement the receiving function, and the transmitter is used to implement the sending function.
[0150] Optionally, the transceiver 2003 can be integrated with the first processor 2001 or can exist independently and be coupled to the first processor 2001 through an interface circuit ( Figure 9 not shown) of the autonomous driving trajectory prediction device 410. The embodiments of the present invention do not make specific limitations in this regard.
[0151] It should be noted that Figure 9 the structure of the autonomous driving trajectory prediction device 410 shown in does not constitute a limitation on the router. The actual knowledge structure recognition device can include more or fewer components than shown in the figure, or combine some components, or have different component arrangements.
[0152] In addition, the technical effects of the autonomous driving trajectory prediction device 410 can refer to the technical effects of the autonomous driving trajectory prediction method based on long short-term prediction described in the above method embodiments, and will not be elaborated here.
[0153] It should be understood that the first processor 2001 in the embodiments of the present invention may be a central processing unit (CPU), and the processor may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0154] It should also be understood that the memory in the embodiments of the present invention may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable ROM (PROM), an erasable PROM (EPROM), an electrically erasable PROM (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of random access memory (RAM) are available, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchlink DRAM (SLDRAM), and direct rambus RAM (DR RAM).
[0155] The above embodiments can be implemented in whole or in part by software, hardware (such as circuits), firmware, or any other combination. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, the processes or functions described in the embodiments of the present invention are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center by wired (such as infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that contains one or more collections of available media. The available medium can be a magnetic medium (such as a floppy disk, hard disk, or magnetic tape), an optical medium (such as a DVD), or a semiconductor medium. The semiconductor medium can be a solid-state drive.
[0156] It should be understood that the term "and / or" in this document is merely a description of the association relationship between associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. Here, A and B can be singular or plural. In addition, the character " / " in this document generally represents an "or" relationship between the associated objects before and after, but it may also represent an "and / or" relationship, which can be specifically understood with reference to the context before and after.
[0157] In the present invention, "at least one" means one or more, and "a plurality" means two or more. "At least one of the following" or its similar expressions refer to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b, or c can represent: a, b, c, a - b, a - c, b - c, or a - b - c, where a, b, and c can be single or multiple.
[0158] It should be understood that in various embodiments of the present invention, the magnitudes of the sequence numbers of the above processes do not mean the order of execution. The order of execution of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present invention.
[0159] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.
[0160] Those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working processes of the devices, apparatuses, and units described above can refer to the corresponding processes in the foregoing method embodiments, and will not be repeated herein.
[0161] In several embodiments provided by the present invention, it should be understood that the disclosed devices, apparatuses, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division, and there can be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces, and the indirect couplings or communication connections of the devices or units can be in electrical, mechanical, or other forms.
[0162] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place, or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0163] In addition, the functional units in each embodiment of the present invention can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit.
[0164] When the above-mentioned function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art or a part of this technical solution can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store program codes.
[0165] As described above, the above is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of changes or substitutions, which should all be covered by the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.
Claims
1. An automatic driving trajectory prediction method based on long-term and short-term prediction, characterized in that: The method comprises: S1. Obtain information data of the target vehicle to be tracked; S2, inputting the information data into the constructed long-short time cyclic prediction decoder; Wherein, the long-short cycle prediction decoder includes long-term target prediction Gru, short-term trajectory prediction Gru and future trajectory prediction Gru; S3. Obtain an automatic driving trajectory prediction result according to the information data, the long-term target prediction Gru, the short-term trajectory prediction Gru and the future trajectory prediction Gru; The step S3 obtains the automatic driving trajectory prediction result according to the information data, the long-term target prediction Gru, the short-term trajectory prediction Gru and the future trajectory prediction Gru, including: S31, inputting the information data into the long-term target prediction Gru to obtain the long-term target features corresponding to each future time step; S32, inputting the hidden features of the future trajectory prediction Gru into the short-term trajectory prediction Gru to obtain the future short-term target features; S33, inputting the information data, long-term target features, and future short-term target features into the future trajectory prediction Gru, obtaining hidden features of the future trajectory prediction Gru, and then obtaining an autonomous driving trajectory prediction result; The step S33 inputs the information data, long-term target features, and future short-term target features into the future trajectory prediction Gru to obtain hidden features of the future trajectory prediction Gru, and then obtains the automatic driving trajectory prediction result, including: S331, using a learnable feature code obtained by initializing a normal distribution as an input of the future trajectory prediction Gru, and obtaining a feature code of the dynamic scene information of the current target vehicle location according to the information data and the adaptive scene information encoder; S332, fusing the input of the future trajectory prediction Gru and the feature coding of the dynamic scene information at the current target vehicle location through a cross attention module to obtain an updated feature coding of the dynamic scene information at the current target vehicle location; S333, combining the updated feature coding of the dynamic scene information of the current target vehicle location, the long-term target features, and the future short-term target features to obtain the input of the fused future trajectory prediction Gru; S334, obtaining hidden features of the trajectory prediction Gru at each future time step according to the fused input of the future trajectory prediction Gru and the future trajectory prediction Gru; S335. Predict the hidden features of Gru according to the trajectory prediction of each future time step to obtain the trajectory points of each future time step, and obtain the automatic driving trajectory prediction result according to the trajectory points of each future time step.
2. The automatic driving trajectory prediction method based on long-term and short-term prediction according to claim 1 is characterized in that: The step S31 of inputting the information data into the long-term target prediction Gru to obtain the long-term target features corresponding to each future time step includes: S311, acquiring scene data and trajectory data in the information data; S312, dividing the scene into multiple regions by using a K-means clustering model according to the scene data and the trajectory data, each of the multiple regions corresponding to multiple trajectories; S313, extracting features of multiple trajectories corresponding to each region respectively and performing feature aggregation to obtain a trajectory prototype of each region, and obtaining an initial hidden feature of long-term target prediction Gru according to the trajectory prototype; S314. A learnable feature encoding obtained by initializing a normal distribution is used as the input of the long-term target prediction Gru. According to the input of the long-term target prediction Gru, the initial hidden features of the long-term target prediction Gru and the long-term target prediction Gru, the long-term target features corresponding to each future time step are obtained.
3. The automatic driving trajectory prediction method based on long-term and short-term prediction according to claim 1 is characterized in that: The step S32 of inputting the hidden features of the future trajectory prediction Gru into the short-term trajectory prediction Gru to obtain the future short-term target features includes: S321. Get time step t i The hidden features of the future trajectory prediction Gru are used as the time step t i The short-term trajectory prediction Gru input; the empty vector with a value of 0 is used as the time step t i The short-term trajectory predicts the initial hidden features of Gru, according to the time step t i The short-term trajectory prediction Gru input, time step t i The initial hidden features of the short-term trajectory prediction Gru and the short-term trajectory prediction Gru are obtained at time step t i The short-term trajectory predicts the hidden state of Gru according to the time step t i The short-term trajectory prediction of Gru is obtained by predicting the hidden state of time step t i The short-term prediction trajectory point for the time step t i The short-term prediction trajectory points are converted to obtain the time step t i The feature encoding of the short-term trajectory prediction Gru is used as the feature encoding at time step t i+1 The hidden features of S322. Get time step t i+1 The short-term trajectory predicts the hidden features of Gru, and obtains the time step t i+2 The short-term trajectory predicts the hidden features of Gru; S323, for the time step t i The short-term trajectory predicts the hidden state of Gru, time step t i+1 The short-term trajectory prediction Gru's hidden features and time step t i+2 The short-term trajectory prediction of Gru is added to the hidden features of Gru to obtain the future short-term features.
4. The automatic driving trajectory prediction method based on long-term and short-term prediction according to claim 1 is characterized in that: The step S331 of obtaining feature coding of dynamic scene information of the current target vehicle location according to the information data and the adaptive scene information encoder includes: The future trajectory prediction Gru predicts the trajectory points of each future time step and the speed of the target vehicle at each time step, establishes an observation radius according to the speed, and obtains the feature coding of the dynamic scene information of the current target vehicle location based on the observation radius and information data.
5. The automatic driving trajectory prediction method based on long-term and short-term prediction according to claim 1, characterized in that: The loss of the long short-term cyclic prediction decoder is expressed by the following formula (1): In the formula, represents the loss of the long-short cycle prediction decoder, T represents the total time step, N represents the number of vehicles, represents the short-term prediction task loss, represents the future trajectory prediction task loss, Represents the multimodal trajectory confidence classification task loss.
6. An automatic driving trajectory prediction device based on long-term and short-term prediction, the automatic driving trajectory prediction device based on long-term and short-term prediction is used to implement the automatic driving trajectory prediction method based on long-term and short-term prediction according to any one of claims 1 to 5, characterized in that: The device comprises: An acquisition module is used to acquire information data of a target vehicle for trajectory prediction; An input module, used for inputting the information data into the constructed long-short time cyclic prediction decoder; Wherein, the long-short cycle prediction decoder includes long-term target prediction Gru, short-term trajectory prediction Gru and future trajectory prediction Gru; The output module is used to obtain the automatic driving trajectory prediction result according to the information data, the long-term target prediction Gru, the short-term trajectory prediction Gru and the future trajectory prediction Gru.
7. An automatic driving trajectory prediction device, characterized in that: The automatic driving trajectory prediction device comprises: processor; A memory having computer-readable instructions stored thereon, wherein when the computer-readable instructions are executed by the processor, the method according to any one of claims 1 to 5 is implemented.
8. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores program codes, which can be called by a processor to execute the method according to any one of claims 1 to 5.
Citation Information
Patent Citations
Intelligent vehicle track prediction system and method fusing peripheral vehicle interaction information
CN113954864A
Peripheral vehicle track prediction method and system based on long and short time motion track fusion
CN115465296A