Automatic driving vehicle motion prediction method based on scene understanding enhancement
Through a variety of trajectory masking strategies and a lightweight encoder-decoder architecture with hybrid pre-training, combined with GPU deployment, the problem of computing burden and catastrophic forgetting in deep learning motion prediction is solved, and the real-time performance of motion prediction of autonomous vehicles is improved.
Patent Information
- Application Number
- CN202510814107.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-18
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-06-18
AI Technical Summary
The existing deep learning motion prediction methods have a high computing burden in autonomous driving and are difficult to meet real-time performance requirements. In addition, there are catastrophic forgetting problems in multiple occlusion strategies, which affects prediction performance.
A variety of trajectory masking strategies and hybrid pre-training strategies are adopted, combined with a lightweight encoder-decoder architecture, and scenario understanding is enhanced through self-supervised learning and model deployment is used to improve prediction performance.
It reduces the complexity of the model, enhances the semantic feature learning ability, solves the catastrophic forgetting problem, and improves the real-time prediction performance of the model.
Smart Images

Figure CN120339990A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of autonomous driving, and particularly relates to a method for predicting the motion of an autonomous driving vehicle enhanced by scene understanding. Background Art
[0002] Autonomous driving technology is an important part of intelligent transportation systems, aiming to improve traffic safety, reduce traffic accidents, and enhance road utilization efficiency. With the progress of artificial intelligence, sensor, and communication technologies, autonomous driving has gradually moved from theoretical research to practical applications. Autonomous driving technology can not only reduce human driving errors but also optimize traffic flow, reduce energy consumption, and environmental pollution, which is of great significance for the construction of future smart cities. An autonomous driving system includes modules such as perception, decision-making, planning, and control. Among them, motion prediction is a key link, which receives perception data and provides effective reference trajectories for downstream modules. Therefore, improving the accuracy of motion prediction is an important direction in autonomous driving technology.
[0003] Deep learning has shown good long-term prediction performance, generalization performance, and computational cost in motion prediction, and has received extensive attention from the automotive industry and academia. The main deep network models include sequence networks (such as convolutional neural networks, recurrent neural networks, attention mechanisms, etc.), generative networks, and graph networks. The framework based on encoder-social module-decoder has become the mainstream. However, while improving the prediction performance by continuously increasing the number of network layers, it also brings a huge computational burden, limiting the practical application of the above methods.
[0004] Self-supervised learning, as a deep learning paradigm that uses unlabeled data to discover its inherent features, has achieved remarkable results in the fields of computer vision and natural language processing, and has also been introduced into the field of motion prediction. It adopts a simple encoder-decoder design, understands scene information through pre-training, and can improve the prediction performance while simplifying the network architecture. Early self-supervised learning mostly adopted a single masking strategy, which could only learn fixed scene features and had limited gain in prediction performance. To address this problem, researchers have tried to adopt multiple masking strategies corresponding to multiple pre-training tasks, which involves the problem of multi-task learning. The main methods are divided into parallel training and sequential training. Parallel training performs multiple pre-training tasks at once, but the loss is difficult to converge. Sequential training performs pre-training tasks sequentially, but faces the problem of catastrophic forgetting, that is, the model will forget the knowledge learned in the previous stage. In addition, most deep learning trajectory prediction methods only verify the prediction accuracy through simulation experiments, but the inference speed of the model cannot meet the actual requirements, affecting its practical value.
[0005] In view of the above problems, the present invention proposes a method for predicting the motion of an autonomous driving vehicle enhanced by scene understanding. Summary of the Invention
[0006] The object of the present invention is to provide a motion prediction method for an autonomous driving vehicle enhanced by scene understanding, aiming to solve the problems raised in the above-mentioned background technology.
[0007] The object of the present invention is achieved through the following technical solutions: A motion prediction method for an autonomous driving vehicle enhanced by scene understanding, comprising the following steps: Step 1, design of multiple trajectory masking strategies and hybrid pre-training strategies; define a traffic scene with N vehicles, including a target vehicle and N- adjacent vehicles of 1 target vehicle, the historical trajectory length of each vehicle is T h and the number of feature dimensions of each trajectory point is D ; denote the set containing all vehicle historical trajectory data as ; store all vehicle historical trajectory data in the traffic scene in the form of discrete points and construct a data structure to reflect the multi-level characteristics of vehicles, trajectories and features; Step 2, construction of a pre-training network enhanced by scene understanding; adopt an encoder-decoder architecture to construct a pre-training model, the encoder includes a spatio-temporal module, an aggregation module and a social module, and the decoder is a multi-layer perceptron; Step 3, construction of a motion prediction network based on end-to-end fine-tuning; use the trained pre-training encoder for parameter initialization to obtain a motion prediction encoder; redesign a motion prediction decoder based on a multi-layer perceptron, that is, obtain a motion prediction network model based on end-to-end fine-tuning; Step 4, GPU deployment and inference acceleration of the prediction model.
[0008] Furthermore, the design of the multiple trajectory masking strategies includes: Design of a pre-training task based on a random masking strategy: randomly mask the historical trajectories of all vehicles in the scene, generate a random index sequence according to the masking ratio α, set the mask tokens at the corresponding positions according to the random index sequence, and the randomly masked part of the trajectory is regarded as unknown during the training process, and the network completes the complete historical trajectory through the unmasked part of the trajectory; Design of a pre-training task based on a social masking strategy: mask the consecutive paragraphs at the end of the historical trajectories of all vehicles in the scene according to the ratio β and use the historical state to predict the future state; The design of the hybrid pre-training strategy includes: dividing the entire training process into multiple stages, each stage corresponding to different training objectives and the number of subtasks; in the i th stage, train i subtasks simultaneously, and use the network parameters of the current stage at the same time and Initialize the model for the next stage.
[0009] Furthermore, the spatio-temporal module sequentially includes a linear mapping layer and multiple Transformer encoders; First, perform a unified dimensionality transformation through the linear mapping layer, and then calculate spatio-temporal attention and complete feature encoding through the Transformer encoder; specifically expressed as: ; Among them, is the spatio-temporal feature vector; is the output of the linear mapping layer, is the masking ratio (with different values for different tasks); D e is the unified encoding dimension; represents the mask tensor of the trajectory data, containing boolean data, 1 represents masked, 0 represents unmasked; TE is the Transformer encoder; Linear is the linear mapping layer; and are the weight and bias of the linear mapping layer respectively; For the i th layer of the multiple Transformer encoders, denote its input and output as and respectively; First, perform a learnable position encoding operation, expressed as: ; Among them, is the output of the learnable position encoding; is the encoding matrix, and its elements are optimized and iterated during backpropagation; is the loss; Secondly, perform a layer normalization. For a single feature in the input , its layer normalization formula is: ; Among them, is the n th feature of the l th trajectory point of the th vehicle, is the normalized output of the feature; is the scaling parameter; is a positive number; is the offset parameter; and are the n th vehicle's lThe mean and variance of each trajectory point; D is the number of features for each trajectory point; After that, the result of the first layer normalization is input into the multi-head self-attention layer, and the output is: ; where, is the attention vector; MSA is the multi-head self-attention layer; Then, a residual connection and another layer normalization operation are performed, and the formula is: ; where, is the output of the second layer normalization; LN is the layer normalization operation; DropPath is the DropPath operation; Finally, through the feed-forward layer, residual connection, and another DropPath operation, the i final output of the layer is obtained: ; where, FF is the feed-forward layer; The output of the i th layer of the multi-layer Transformer encoder is used as the input of the i +1th layer, and finally the final output of the spatio-temporal module is obtained ; The aggregation module first scales the input feature dimension and then compresses the trajectory dimension multiple times through max-pooling operations to increase the information density; the process is expressed as: ; where, is the aggregated feature vector; is the mask vector of the aggregated features, obtained by compressing the last dimension of the mask tensor before aggregation; The input of the social module is the aggregated feature , with a dimension of , represents the number of vehicles, is the unified encoding dimension; the output is the social feature ; the process is expressed as: ; The decoder decodes the social feature and generates the reconstructed trajectory: ; where, is the reconstructed historical trajectory; MLP is the multi-layer perceptron.
[0010] Furthermore, the specific process of step 3 is as follows: First, fine-tune the pre-trained model. The specific operation is as follows: Remove the masking operation in the pre-trained model, and only use the part corresponding to the target vehicle in the social features as the input of the decoder, with a dimension of 1 ×D e ; Second, the redesigned decoder is initialized with random parameters to decode the social features of the target vehicle and generate its future trajectory: ; Among them, is the predicted trajectory of the target vehicle, is the length of the predicted trajectory; represents the social features of the target vehicle; MLP is a multi-layer perceptron.
[0011] Furthermore, the specific process of step 4 is as follows: First, build an environment suitable for GPU development. After the model training is completed, convert it to the ONNX format; second, use TensorRT as a tool for inference acceleration to optimize the ONNX model; finally, generate a TensorRT engine based on the optimized model and deploy it to the Tianzhun GEACX1 GPU development board for inference acceleration.
[0012] Compared with the prior art, the beneficial effects of the present invention are: 1. Compared with the existing deep learning motion prediction methods, the motion prediction method proposed by the present invention adopts a lightweight network architecture design, reduces the complexity of the model, and performs hardware deployment to enhance the real-time performance of the model, thus improving the practical application value.
[0013] 2. Compared with the existing single masking strategy pre-training methods, the scene understanding pre-training method proposed by the present invention designs a variety of trajectory masking strategies, allowing the model to learn richer semantic features and enhancing the pre-training effect.
[0014] 3. Compared with the existing multi-masking strategy pre-training methods, the scene understanding pre-training method proposed by the present invention adopts a hybrid pre-training strategy to arrange multiple sub-tasks, solves the problem of catastrophic forgetting in the pre-training process, and enhances the learning ability of the model. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] Figure 1 is a flowchart of the method of the present invention.
[0016] Figure 2 are schematic diagrams of a variety of trajectory masking strategies; among them, (a) historical trajectory, (b) random masking trajectory, (c) social masking trajectory.
[0017] Figure 3 It is a schematic diagram of a hybrid pre-training strategy.
[0018] Figure 4 It is a structural diagram of a pre-training network enhanced by scene understanding.
[0019] Figure 5 It is a structural diagram of a motion prediction network based on end-to-end fine-tuning. Detailed implementation manners
[0020] For a clearer understanding of the technical features, objectives, and beneficial effects of the present invention, the technical solutions of the present invention will be described in detail below, but it should not be construed as a limitation on the scope of implementation of the present invention. The experimental methods described in the following embodiments are all conventional methods unless otherwise specified; the reagents and materials are all commercially available unless otherwise specified.
[0021] The following describes the specific implementation of the present invention in detail in conjunction with specific embodiments.
[0022] As Figure 1 shown, a motion prediction method for an autonomous driving vehicle based on enhanced scene understanding provided by an embodiment of the present invention. First, a variety of trajectory masking strategies are adopted to mask at different data levels respectively, and different features in the scene are learned specifically. A hybrid pre-training strategy is adopted to prevent potential catastrophic forgetting problems and enhance the learning ability of the model. Secondly, according to the requirements of the pre-training task, a lightweight encoder-decoder structure network model is designed, and scene understanding pre-training is carried out according to the above trajectory masking and pre-training strategies. In addition, the trained pre-training model is fine-tuned end-to-end and applied to the motion prediction task to verify the effectiveness of the multi-strategy. Finally, the motion prediction network model is optimized using the TensorRT optimizer, a TensorRT engine is generated and deployed to the GPU platform for inference acceleration to improve the real-time prediction performance.
[0023] The motion prediction method includes the following steps: Step 1: Design of a variety of trajectory masking strategies and a hybrid pre-training strategy; Self-supervised pre-training aims to learn some inherent but general knowledge in unlabeled data, and what the model can specifically learn depends on the design of the pre-training task. Consider a traffic scene with N vehicles, including a target vehicle (i.e., the subject of motion prediction) and N- adjacent vehicles of 1 target vehicle. The historical trajectory length of each vehicle is T h and the number of feature dimensions of each trajectory point is D . Denote the set containing all vehicle historical trajectory data as ;Store all the historical trajectory data of vehicles in the traffic scenario in the form of discrete points, and construct a data structure to reflect the multi-level characteristics of the data structure, that is, from top to bottom are vehicles, trajectories, and features respectively.
[0024] (1) For the input multi-level data structure, the present invention designs a variety of masking strategies corresponding to pre-training subtasks with different focuses to comprehensively and deeply understand the internal structure and interrelationships of the scenario data, as Figure 2 shown. The pre-training process needs to complete two subtasks, namely the pre-training task based on the random masking strategy and the pre-training task based on the social masking strategy, which will be introduced in detail below.
[0025] Design of the pre-training task based on the random masking strategy: The purpose is to help the model understand the temporal and spatial connections within the trajectory. Implement a random masking strategy for the historical trajectories of all vehicles in the scenario, generate a random index sequence according to the masking ratio α, and set the mask tokens at the corresponding positions according to this sequence. During the training process, the randomly masked part of the trajectory is unknown, and the network completes the complete historical trajectory through the remaining part of the trajectory.
[0026] Design of the pre-training task based on the social masking strategy: There are common social relationships among vehicles in the traffic scenario. Driving behaviors such as overtaking, lane changing, and lane keeping between vehicles will undoubtedly affect the future trajectories of adjacent vehicles. The pre-training task based on the social masking strategy aims to learn the above social relationships. For all vehicles in the scenario, according to the ratio β mask the continuous paragraphs at the ends of their historical trajectories. Such a masking strategy design follows the basic pattern of motion prediction, that is, using the historical state to predict the future state, capturing social relationships while preventing information leakage.
[0027] Design of the hybrid pre-training strategy: A variety of trajectory masking strategies are designed to help the model learn different information in the input with emphasis. Therefore, in practice, different pre-training tasks need to be used for implementation, and how to arrange the training processes of multiple subtasks is related to the learning effect of the model. The present invention fully considers the characteristics of parallel training and sequential training and proposes a hybrid pre-training strategy. Specifically, as the training stage progresses, increase the corresponding number of pre-training tasks. At the th stage, train subtasks simultaneously, similar to parallel training; at the same time, use the network parameters and of the current stage to initialize the model of the next stage, similar to sequential training. As Figure 3 shown, in stage 1, first train task 1, initialize a new pre-training network with the weight and bias parameters of the trained model, and then train task 1 and task 2 simultaneously in stage 2.
[0028] Step 2: Building a pre-trained network enhanced by scene understanding; To help the model learn the semantic features in the scene information, the present invention uses a lightweight encoder-decoder architecture to build a pre-trained model, as Figure 4 shown. Among them, the encoder includes a spatio-temporal module, an aggregation module, and a social module, and the decoder is a simple multi-layer perceptron.
[0029] (1) Encoder; a. The role of the spatio-temporal module is to perform spatio-temporal hybrid feature encoding on the input data at the trajectory level, which can be understood as a local social module, and extracts the temporal and spatial relationships between trajectory points. The spatio-temporal module sequentially includes a linear mapping layer (Linear) and multiple Transformer encoders (Transformer Encoder, TE). As described above, the input scene is denoted as , first passes through the linear mapping layer for unified dimensionality transformation, and then calculates spatio-temporal attention through the Transformer encoder to complete feature encoding. These two steps can be expressed as: ; Among them, is the spatio-temporal feature vector; is the output of the linear mapping layer, is the masking ratio (with different values in different tasks, i.e., α or β ); D e is the unified encoding dimension; represents the mask tensor of the trajectory data, containing boolean data, 1 represents masked, and 0 represents unmasked; TE is the Transformer encoder; Linear is the linear mapping layer; and are the weights and biases of the linear mapping layer respectively.
[0030] The standard Transformer encoder includes parts such as layer normalization operation (Layer Norm, LN), multi-head self-attention layer (Multi-head Self Attention, MSA), residual connection (Add), and feed-forward network (Feed Forward, FF). Different from the standard Transformer encoder, first, the present invention introduces learnable position encoding (Learnable Position Encoding, LPE) to mark the time order of trajectory points. Compared with the absolute position encoding used in the standard Transformer, the learnable position encoding can self-optimize during backpropagation, with better encoding effect and stronger adaptability. Second, the Transformer encoder module of the present invention uses the DropPath operation, which randomly and completely masks some network paths from the inside out with increasing probability in a multi-layer network, resulting in a higher degree of regularization. Finally, considering the characteristics of the trajectory occlusion reconstruction task, the present invention adopts Pre-Norm, which performs layer normalization on the features before the attention layer. Compared with the traditional Post-Norm, Pre-Norm can enhance the robustness of the network and accelerate convergence.
[0031] For the i th layer of the multi-layer Transformer encoder, denote its input and output as and ; First, perform the learnable position encoding operation, expressed as: ; where is the output of the learnable position encoding; is the encoding matrix, and its elements are optimized and iterated during backpropagation; is the loss; as the loss decreases, the module will iterate to the optimal encoding matrix, helping the model gradually learn the temporal relationship of trajectory points.
[0032] Second, perform layer normalization once. For a single feature in the input , its layer normalization formula is: ; where is the n th feature of the l th trajectory point of the th vehicle, is the normalized output of the feature; is the scaling parameter; is a very small positive number used to prevent division by 0 and ensure numerical stability; is the offset parameter; and are the mean and variance of the n -th trajectory point of the l -th vehicle respectively; D is the number of features of each trajectory point; After that, the result of the first layer normalization is input into the multi-head self-attention layer, and the output is: ; where is the attention vector; MSA is the multi-head self-attention layer; Then, a residual connection and another layer normalization operation are performed, and the formula is: ; where is the output of the second layer normalization; LN is the layer normalization operation; the specific process is the same as above; DropPath is the DropPath operation; Finally, through the feed-forward layer, residual connection, and another DropPath operation, the i -th layer's final output is obtained: ; where FF is the feed-forward layer; The output of the i -th layer of the multi-layer Transformer encoder is used as the input of the i +1-th layer, and so on to obtain the final output of the spatio-temporal module .
[0033] b. Different from the spatio-temporal module, the social module extends the social scope to the global, that is, capturing dependencies at the vehicle level. Therefore, as an intermediate module, the role of the aggregation module is to perform feature compression at the trajectory level and enrich the spatio-temporal features of a single vehicle into a matrix with a dimension of . The aggregation module in the present invention adopts a simplified version of the PointNet structure, first scales the input feature dimension, and then compresses the trajectory dimension multiple times through the max pooling operation (MaxPooling, MP) to increase the information density; the above process can be expressed as: ; where is the aggregated feature vector; is the mask vector of the aggregated feature, which is obtained by compressing the last dimension of the mask tensor before aggregation.
[0034] c. The current and future states of the vehicle are directly affected by its historical state. In addition, social interaction with surrounding vehicles is another major influencing factor. To learn this social relationship, the present invention designs a social module. The social module adopts the design of a multi-layer Transformer encoder, but the specific number of layers and input / output dimensions are different from those of the spatio-temporal module. The input of the social module is the aggregated feature , with a dimension of , representing the number of vehicles, being the unified encoding dimension; the output is the social feature , because the attention calculation does not change the dimension of the data. This process can be expressed as: .
[0035] (2) Decoder; The encoder is used to extract semantic features in the scene data, such as time, space, and social relationships, while the role of the decoder is to infer the unknown part based on the above-known information. To simplify the model structure as much as possible while meeting the performance requirements, the present invention designs the decoder based on a multi-layer perceptron (MLP) to decode the social features and generate the reconstructed trajectory: ; wherein, is the reconstructed historical trajectory; MLP is the multi-layer perceptron.
[0036] Step 3, Construction of a motion prediction network based on end-to-end fine-tuning; (1) Construction of the fine-tuning model; Fine-tuning refers to using a pre-trained model for downstream tasks. The present invention initializes the parameters with the trained pre-trained encoder to obtain a motion prediction encoder; considering the differences between the motion prediction and trajectory masking reconstruction tasks, a motion prediction decoder based on a multi-layer perceptron is redesigned, and thus a motion prediction network model based on end-to-end fine-tuning is obtained, as shown in Figure 5 . The encoder and decoder of the motion prediction network model based on end-to-end fine-tuning are specifically introduced below: First, fine-tune the pre-trained model. The specific operation is as follows: Remove the masking operation in the pre-trained model, that is, do not mask the trajectory data; only use the part of the social feature corresponding to the target vehicle as the input of the decoder, with a dimension of 1 ×D e ; Secondly, the redesigned decoder is initialized with random parameters to decode the social features of the target vehicle and generate its future trajectory: ; Among them, is the predicted trajectory of the target vehicle, and is the length of the predicted trajectory; represents the social characteristics of the target vehicle; MLP is a multi-layer perceptron.
[0037] (2)Training strategy design; To explore the influence of different combinations of various trajectory occlusion strategies and hybrid pre-training strategies on the motion prediction results, the present invention summarizes the training strategies as follows: Strategy 1: Training from scratch, without pre-training, directly performing motion prediction training; Strategy 2: Only completing the pre-training of the random occlusion reconstruction task, and performing motion prediction training after fine-tuning; Strategy 3: Only completing the pre-training of the social occlusion reconstruction task, and performing motion prediction training after fine-tuning; Strategy 4: Hybrid pre-training based on the random occlusion reconstruction task, that is, initializing the model with the parameters of the random task, then training the random task and the social task in parallel, and performing motion prediction training after fine-tuning; Strategy 5: Hybrid pre-training based on the social occlusion reconstruction task, that is, initializing the model with the parameters of the social task, then training the random task and the social task in parallel, and performing motion prediction training after fine-tuning.
[0038] Step 4: GPU deployment and inference acceleration of the prediction model; In actual autonomous driving applications, it is generally required that the inference speed of the motion prediction module reaches the millisecond level. However, with the iteration of the prediction network architecture, the real-time prediction performance under the CPU platform cannot meet the requirements. To address this issue, the present invention deploys the prediction model to a vehicle-grade GPU platform for inference acceleration to improve real-time performance.
[0039] First, set up an environment suitable for GPU development to ensure effective model conversion and optimization. After the model training is completed, convert it to the ONNX format, which not only supports cross-platform model conversion but also provides a basis for subsequent optimization processes. Secondly, use TensorRT as a key tool for inference acceleration and optimize the ONNX model through a series of technical means. First, predict data accuracy to ensure computational accuracy and efficiency during the inference process; second, perform layer and tensor fusion to reduce unnecessary intermediate steps in the computational graph; third, perform kernel auto-tuning to optimize the execution efficiency of network layers and other computational operations through automated tools; fourth, set dynamic tensor memory to dynamically manage memory usage at runtime; fifth, perform data stream parallel optimization to utilize the parallel computing power of the GPU to accelerate the inference process. Finally, generate a TensorRT engine based on the optimized model and deploy it to the Tianzhun GEACX1 GPU development board for inference acceleration. Verify the enhancement effect of GPU inference acceleration on the real-time prediction performance of the model by comparing the prediction metrics, inference speed, and model size of the same motion prediction model on CPU and GPU platforms.
[0040] Example 1: Simulation experiment verification; The present invention uses the Average Displacement Error (ADE) to measure the overall prediction performance of the model and the Final Displacement Error (FDE) to evaluate the final prediction performance of the model. In terms of the regression performance evaluation of trajectory prediction, most existing methods only focus on the ADE index because it can reflect the overall prediction performance of the model. However, in high-speed scenarios, the final prediction error of the vehicle's future trajectory (especially the longitudinal error) will be amplified, thus affecting the overall performance index. Therefore, it is necessary to also include the FDE in the evaluation system.
[0041] The present invention evaluates the prediction performance indicators of the model under training strategies 1 to 5, and the results are shown in Table 1: Table 1 Comparison of motion prediction ADE and FDE indicators (unit: m) under different training strategies
[0042] As can be seen from Table 1, compared with Strategy 1, the ADE and FDE indicators of the model under Strategies 2-5 show a downward trend with the addition of various trajectory masking strategies and hybrid pre-training strategies. In particular, the comparison of the indicators between Strategy 1 and Strategies 2 and 3 illustrates the effectiveness of various trajectory masking strategies, and the comparison of the indicators between Strategies 2, 3 and Strategies 4, 5 illustrates the effectiveness of the hybrid pre-training strategy. In addition, the ADE indicator of the model under Strategy 3 is higher than that of the model under Strategy 2, but the FDE indicator is lower. This may be because the random masking reconstruction task has a finer granularity and is more inclined to learn the local spatio-temporal relationship between trajectory points, so the ADE is smaller; while the social masking reconstruction task has a coarser granularity and learns more global social relationships, so the FDE is smaller.
[0043] The above are only the preferred embodiments of the present invention. It should be noted that for those skilled in the art, without departing from the concept of the present invention, several modifications and improvements can still be made, which should also be regarded as the protection scope of the present invention, and these will not affect the implementation effect of the present invention and the practicality of the patent.
Claims
1. A method for predicting the motion of an autonomous vehicle enhanced by scene understanding, characterized in that, It includes the following steps: Step 1: Design of multiple trajectory masking strategies and hybrid pre-training strategies; Define a traffic scenario with N vehicles, including a target vehicle and N- adjacent vehicles of 1 target vehicle. The historical trajectory length of each vehicle is T h , and the number of feature for each trajectory point is D . Denote the set containing all vehicles' historical trajectory data as ; Store all vehicles' historical trajectory data in the traffic scenario in the form of discrete points and construct a data structure to reflect the multi-level characteristics of vehicles, trajectories, and features. Step 2: Build a pre-training network enhanced by scene understanding; Adopt an encoder-decoder architecture to build a pre-training model. The encoder includes a spatio-temporal module, an aggregation module, and a social module, and the decoder is a multi-layer perceptron; Step 3: Build a motion prediction network based on end-to-end fine-tuning; use the trained pre-training encoder for parameter initialization to obtain a motion prediction encoder; redesign the motion prediction decoder based on a multi-layer perceptron, that is, obtain a motion prediction network model based on end-to-end fine-tuning; Step 4: GPU deployment and inference acceleration of the prediction model.
2. The method for predicting the movement of an autonomous driving vehicle enhanced based on scenario understanding according to claim 1, wherein The design of the multiple trajectory masking strategies includes: Pre-training task design based on a random masking strategy: Randomly mask the historical trajectories of all vehicles in the scene, generate a random index sequence according to the masking ratio α, set the mask tokens at the corresponding positions according to the random index sequence, and the randomly masked partial trajectories are regarded as unknown during the training process. The network completes the complete historical trajectory through the unmasked partial trajectories; Pre-training task design based on social masking strategy: according to a ratio β Mask consecutive paragraphs at the ends of the historical trajectories of all vehicles in the scenario, and use historical states to predict future states; The design of the hybrid pre-training strategy includes: dividing the entire training process into multiple stages, where each stage corresponds to different training objectives and the number of subtasks; training i multiple i subtasks simultaneously in the th stage, and initializing the model of the next stage with the network parameters of the current stage.
3. The method for predicting the movement of an autonomous driving vehicle enhanced by scene understanding according to claim 1, wherein The spatio-temporal module sequentially includes a linear mapping layer and multiple Transformer encoders; First, a unified dimensionality transformation is performed through the linear mapping layer, and then spatio-temporal attention is calculated through the Transformer encoder to complete feature encoding; specifically expressed as: ; Among them, is the spatio-temporal feature vector; is the output of the linear mapping layer, is the masking ratio; D e is the unified coding dimension; represents the mask tensor of the trajectory data, containing boolean data, where 1 represents masked and 0 represents unmasked; TE is the Transformer encoder; Linear is the linear mapping layer; and are the weight and bias of the linear mapping layer respectively; For the i -th layer of the multi-layer Transformer encoder, denote its input and output as and ; First, perform a learnable position encoding operation, expressed as: ; Among them, is the output of learnable position encoding; is the encoding matrix, and its elements are optimized and iterated during backpropagation; is the loss; Next, perform layer normalization once. For a single feature in the input , its layer normalization formula is as follows: ; Wherein, is the n th l feature of the th trajectory point of the th vehicle; is the normalized output of the said feature; is a positive number; is the offset parameter; and are respectively the mean and variance of the n th trajectory point of the l th vehicle; D is the number of features of each trajectory point; The result of normalizing the first layer is input into the multi-head self-attention layer, and the output is: ; Among them, is the attention vector; MSA is the multi-head self-attention layer; Then perform a residual connection and another layer normalization operation, with the formula: ; Among them, is the output of the second-layer normalization; LN is the layer normalization operation; DropPath is the DropPath operation; Finally, through the feedforward layer, residual connections, and another DropPath operation, we obtain i the final output of the layer: ; where FF is the feed-forward layer; The output of the i -th layer of the multi-layer Transformer encoder serves as the input to the i +1-th layer, and finally the final output of the spatio-temporal module is obtained. ; The aggregation module first scales the input feature dimension, and then compresses the trajectory dimension multiple times through max pooling operations to increase the information density; the process is expressed as: ; Among them, is the aggregated feature vector; is the mask vector of the aggregated feature, which is obtained by performing a compression operation on the last dimension of the mask tensor before aggregation; The input of the social module is the aggregated feature , with the dimension of , representing the number of vehicles, being the unified coding dimension; the output is the social feature ; the process is expressed as: ; The decoder decodes the social features and generates a reconstructed trajectory: ; Among them, is the reconstructed historical trajectory; MLP is the multi-layer perceptron.
4. The method for predicting the movement of an autonomous driving vehicle based on enhanced scene understanding according to claim 1, wherein The specific process of Step 3 is as follows: First, fine-tune the pre-trained model. The specific operation is as follows: Remove the masking operation in the pre-trained model, and only use the part corresponding to the target vehicle in the social features as the input of the decoder, with a dimension of 1 ×D e ; Secondly, the redesigned decoder uses random parameter initialization to decode the social features of the target vehicle and generate its future trajectory: ; Among them, is the predicted trajectory of the target vehicle, is the length of the predicted trajectory; represents the social characteristics of the target vehicle; MLP is a multi-layer perceptron.
5. The method for predicting the motion of an autonomous driving vehicle based on enhanced scene understanding according to claim 1, wherein The specific process of Step 4 is as follows: First, build an environment suitable for GPU development. After the model training is completed, convert it to the ONNX format; secondly, use TensorRT as a tool for inference acceleration to optimize the ONNX model; finally, generate a TensorRT engine based on the optimized model and deploy it to the Tianzhun GEACX1 GPU development board for inference acceleration.
Citation Information
Patent Citations
Pedestrian trajectory prediction method based on conditional variational autoencoder and social converter
CN113870309A
Information cascade prediction method and device based on spatio-temporal characteristics and content preferences
CN119622114A
Implementation method for vehicle track prediction
CN120011845A
Track prediction method and device for surrounding vehicles in automatic driving scene and medium
CN120071281A
Vehicle trajectory prediction method based on physical social soft attention Transformer
CN120145000A